Language and AI benchmarks · Stanford NLP
SQuAD (Stanford Question Answering)
Questions about Wikipedia passages with answers marked in the text; v2 adds unanswerable questions.
- Licence
- CC BY-SA 4.0
Multi-turn synthetic dialogues used to fine-tune chat models.
We haven't connected this source yet. Requests decide what we add next, so tell us you need it.
More in Language and AI benchmarks
Language and AI benchmarks · Stanford NLP
Questions about Wikipedia passages with answers marked in the text; v2 adds unanswerable questions.
Language and AI benchmarks · LMSYS
Real user conversations with chatbots, with pairwise human votes from the Arena.
Language and AI benchmarks · Google Research
Real Google search questions answered from Wikipedia pages.
Language and AI benchmarks · NYU, University of Washington and DeepMind
The general language understanding benchmark that shaped BERT-era models.
Language and AI benchmarks · NYU, Facebook AI, University of Washington and DeepMind
A harder successor to GLUE for reading comprehension and reasoning.
Language and AI benchmarks · Hendrycks et al. (UC Berkeley)
Multiple-choice exam questions from law to physics, the most cited LLM knowledge benchmark.