AI training text · Wikimedia Foundation
Wikipedia (full dumps)
Every article in every language edition, as published in the regular Wikimedia database dumps.
- Updated
- twice a month
- Licence
- CC BY-SA 4.0 and GFDL
An open archive of the web, crawled every month since 2008. The raw material behind most large language models.
We haven't connected this source yet. Requests decide what we add next, so tell us you need it.
More in AI training text
AI training text · Wikimedia Foundation
Every article in every language edition, as published in the regular Wikimedia database dumps.
AI training text · Hugging Face
Cleaned and de-duplicated English web text from 96 Common Crawl snapshots, built for LLM pre-training.
AI training text · EleutherAI
A diverse English text corpus from 22 sources, including papers, code, books and web text.
AI training text · Google and Allen Institute for AI
Cleaned Common Crawl text created for Google's T5 model.
AI training text · Technology Innovation Institute (TII)
Filtered, de-duplicated web text used to train the Falcon models.
AI training text · Together AI
Open web text with quality signals, built to reproduce LLaMA-style training data.