Language and AI benchmarks · Stanford NLP
SQuAD (Stanford Question Answering)
Questions about Wikipedia passages with answers marked in the text; v2 adds unanswerable questions.
- Licence
- CC BY-SA 4.0
The general language understanding benchmark that shaped BERT-era models.
We haven't connected this source yet. Requests decide what we add next, so tell us you need it.
More in Language and AI benchmarks
Language and AI benchmarks · Stanford NLP
Questions about Wikipedia passages with answers marked in the text; v2 adds unanswerable questions.
Language and AI benchmarks · LMSYS
Real user conversations with chatbots, with pairwise human votes from the Arena.
Language and AI benchmarks · Tsinghua University (THUNLP)
Multi-turn synthetic dialogues used to fine-tune chat models.
Language and AI benchmarks · Google Research
Real Google search questions answered from Wikipedia pages.
Language and AI benchmarks · NYU, Facebook AI, University of Washington and DeepMind
A harder successor to GLUE for reading comprehension and reasoning.
Language and AI benchmarks · Hendrycks et al. (UC Berkeley)
Multiple-choice exam questions from law to physics, the most cited LLM knowledge benchmark.