Language and AI benchmarks · Stanford NLP
SQuAD (Stanford Question Answering)
Questions about Wikipedia passages with answers marked in the text; v2 adds unanswerable questions.
- Licence
- CC BY-SA 4.0
13 datasets
Language and AI benchmarks · Stanford NLP
Questions about Wikipedia passages with answers marked in the text; v2 adds unanswerable questions.
Language and AI benchmarks · LMSYS
Real user conversations with chatbots, with pairwise human votes from the Arena.
Language and AI benchmarks · Tsinghua University (THUNLP)
Multi-turn synthetic dialogues used to fine-tune chat models.
Language and AI benchmarks · Google Research
Real Google search questions answered from Wikipedia pages.
Language and AI benchmarks · NYU, University of Washington and DeepMind
The general language understanding benchmark that shaped BERT-era models.
Language and AI benchmarks · NYU, Facebook AI, University of Washington and DeepMind
A harder successor to GLUE for reading comprehension and reasoning.
Language and AI benchmarks · Hendrycks et al. (UC Berkeley)
Multiple-choice exam questions from law to physics, the most cited LLM knowledge benchmark.
Language and AI benchmarks · Salesforce Research
Good and featured Wikipedia articles for language-model evaluation.
Language and AI benchmarks · Zhu et al. (University of Toronto and MIT)
Self-published books used to train early models such as BERT and GPT.
Language and AI benchmarks · Stanford CRFM
Instruction-following examples generated with OpenAI's text-davinci-003.
Language and AI benchmarks · OpenAssistant (LAION)
Human-written, human-ranked assistant conversations.
Language and AI benchmarks · Google Research
Templated instruction-tuning tasks used for Flan-T5 and Flan-PaLM.