SQuAD (Stanford Question Answering Dataset)
SQuAD is a benchmark dataset for evaluating question answering systems, featuring questions based on a set of Wikipedia articles.
Free Information Center
SQuAD is a benchmark dataset for evaluating question answering systems, featuring questions based on a set of Wikipedia articles.
TruthfulQA is a benchmark designed to evaluate the truthfulness of language models by testing their ability to provide accurate and truthful answers to questions that may induce false or misleading responses.