Language and AI benchmarks · Stanford NLP
SQuAD (Stanford Question Answering)
Questions about Wikipedia passages with answers marked in the text; v2 adds unanswerable questions.
- Licence
- CC BY-SA 4.0
Good and featured Wikipedia articles for language-model evaluation.
We haven't connected this source yet. Requests decide what we add next, so tell us you need it.
More in Language and AI benchmarks
Language and AI benchmarks · Stanford NLP
Questions about Wikipedia passages with answers marked in the text; v2 adds unanswerable questions.
Language and AI benchmarks · LMSYS
Real user conversations with chatbots, with pairwise human votes from the Arena.
Language and AI benchmarks · Tsinghua University (THUNLP)
Multi-turn synthetic dialogues used to fine-tune chat models.
Language and AI benchmarks · Google Research
Real Google search questions answered from Wikipedia pages.
Language and AI benchmarks · NYU, University of Washington and DeepMind
The general language understanding benchmark that shaped BERT-era models.
Language and AI benchmarks · NYU, Facebook AI, University of Washington and DeepMind
A harder successor to GLUE for reading comprehension and reasoning.