Images and video · ImageNet (Stanford and Princeton)
ImageNet
Millions of photos labelled with WordNet categories; the ILSVRC 1,000-class subset is the classic image-classification benchmark.
- Licence
- Non-commercial research and education only
The 20 datasets AI start-ups want most right now: web text, code, images, speech, maps, medicine and Earth observation.
20 datasets
Images and video · ImageNet (Stanford and Princeton)
Millions of photos labelled with WordNet categories; the ILSVRC 1,000-class subset is the classic image-classification benchmark.
Images and video · COCO Consortium
Everyday scenes with object boxes, segmentation masks, keypoints and captions.
AI training text · Common Crawl Foundation
An open archive of the web, crawled every month since 2008. The raw material behind most large language models.
Places and geography · OpenStreetMap Foundation
The free, editable map of the world: roads, buildings, places and land use, updated by millions of mappers.
AI training text · Wikimedia Foundation
Every article in every language edition, as published in the regular Wikimedia database dumps.
Images and video · LAION e.V.
Image-text pairs gathered from Common Crawl, used to train open image and multimodal models.
Companies and business · McAuley Lab, UC San Diego (Amazon Reviews 2023)
Product reviews, ratings and item metadata across 33 categories, widely used for recommendation and sentiment models.
Language and AI benchmarks · Stanford NLP
Questions about Wikipedia passages with answers marked in the text; v2 adds unanswerable questions.
Biology and medicine · MIT Laboratory for Computational Physiology (PhysioNet)
De-identified hospital and intensive-care records from Beth Israel Deaconess Medical Center: vitals, labs, medications and notes.
AI training text · Hugging Face
Cleaned and de-duplicated English web text from 96 Common Crawl snapshots, built for LLM pre-training.
Code · GitHub (via GH Archive and public repositories)
Public source code and activity on GitHub. Licences vary by repository, so filter before training.
Images and video · Google
Image-level labels, object boxes, segmentation masks and relationships across 600 classes.
Images and video · Waymo
Camera and lidar recordings from Waymo's self-driving cars, with 3D labels and motion data.
Images and video · Meta AI
The dataset behind Meta's Segment Anything model: high-resolution photos with automatic segmentation masks.
Audio and speech · Mozilla Foundation
Crowd-sourced recordings of people reading sentences aloud, with age, gender and accent where given.
Language and AI benchmarks · LMSYS
Real user conversations with chatbots, with pairwise human votes from the Arena.
Language and AI benchmarks · Tsinghua University (THUNLP)
Multi-turn synthetic dialogues used to fine-tune chat models.
Places and geography · Google Research
Building outlines detected from satellite imagery across Africa, Asia, Latin America and the Caribbean.
Satellite and Earth observation · European Space Agency (Copernicus)
10-metre multispectral images of Earth's land and coasts, used for crops, forests and change detection.
Weather and hazards · Copernicus Climate Change Service
Hourly global weather history since 1940 on a 31 km grid.