Images and video · ImageNet (Stanford and Princeton)
ImageNet
Millions of photos labelled with WordNet categories; the ILSVRC 1,000-class subset is the classic image-classification benchmark.
- Licence
- Non-commercial research and education only
24 datasets
Images and video · ImageNet (Stanford and Princeton)
Millions of photos labelled with WordNet categories; the ILSVRC 1,000-class subset is the classic image-classification benchmark.
Images and video · COCO Consortium
Everyday scenes with object boxes, segmentation masks, keypoints and captions.
Images and video · LAION e.V.
Image-text pairs gathered from Common Crawl, used to train open image and multimodal models.
Images and video · LeCun, Cortes and Burges
Handwritten digits 0 to 9, the 'hello world' of computer vision.
Images and video · University of Toronto (Krizhevsky)
Small colour images in 10 classes such as aeroplane, cat and truck.
Images and video · Google
Image-level labels, object boxes, segmentation masks and relationships across 600 classes.
Images and video · Waymo
Camera and lidar recordings from Waymo's self-driving cars, with 3D labels and motion data.
Images and video · Meta AI
The dataset behind Meta's Segment Anything model: high-resolution photos with automatic segmentation masks.
Images and video · Stanford University
Images annotated with objects, attributes, relationships and question-answer pairs.
Images and video · MIT CSAIL
Scene parsing: every pixel labelled across 150+ object and stuff classes.
Images and video · PASCAL network (University of Oxford)
The classic object detection and segmentation benchmark (2005-2012).
Images and video · MMLab, Chinese University of Hong Kong
Celebrity face images with 40 attribute labels and landmarks.
Images and video · Cityscapes team (Daimler, MPI, TU Darmstadt)
Urban street scenes with pixel-level labels for driving research.
Images and video · Karlsruhe Institute of Technology and Toyota Technological Institute
Driving recordings with stereo cameras, lidar and GPS for depth, odometry and 3D detection.
Images and video · Mapillary (Meta)
Street-level images from six continents with detailed segmentation.
Images and video · UC Berkeley
Diverse driving video with labels for detection, lanes, segmentation and tracking.
Images and video · Motional
Full sensor suite recordings from Boston and Singapore with 3D boxes.
Images and video · Ego4D consortium (Meta and universities)
Daily-life video recorded from head-mounted cameras in 9 countries.
Images and video · Institute of Mathematics of the Romanian Academy
Actors performing everyday activities, captured with motion capture and four cameras.
Images and video · MIT CSAIL
Scene recognition: kitchens, beaches, stadiums and hundreds more.
Images and video · Qualcomm
People doing basic actions with everyday objects, for action recognition.