An Open Training Set For AI Goes Global

peter.suber's bookmarks 2026-03-24

Summary:

"However, there is an alternative to this “grab it all” approach. It involves using materials that are either in the public domain or released under a “permissive” license that allows LLMs to be trained on them without any problems. There’s plenty of such material online, but its scattered nature puts it at a serious disadvantage compared to downloading everything without worrying about licensing issues. To address that, the Common Corpus was created and released just over a year ago by the French startup Pleias. A press release from the AI Alliance explains the key characteristics of the Common Corpus:

Truly Open: contains only data that is permissively licensed and provenance is documented

Multilingual: mostly representing English and French data, but contains at least 1[billion] tokens for over 30 languages

Diverse: consisting of scientific articles, government and legal documents, code, and cultural heritage data, including books and newspapers

Extensively Curated: spelling and formatting has been corrected from digitized texts, harmful and toxic content has been removed, and content with low educational content has also been removed."

Link:

https://www.techdirt.com/2026/03/24/an-open-training-set-for-ai-goes-global/

From feeds:

Open Access Tracking Project (OATP) » peter.suber's bookmarks
Music and Digital Media » Techdirt.

Tags:

oa.new oa.ai oa.pleias oa.licensing oa.common_corpus oa.multilingualism

Authors:

Glyn Moody

Date tagged:

03/24/2026, 20:05

Date published:

03/24/2026, 06:23