ARC (AI2 Reasoning Challenge)

The ARC (AI2 Reasoning Challenge) is a benchmark dataset and challenge designed to evaluate the reasoning abilities of artificial intelligence systems on grade-school level science questions. Developed by the Allen Institute for Artificial Intelligence, it aims to advance AI research in natural language understanding and reasoning.

Read More →

HellaSwag

HellaSwag is a benchmark dataset designed to evaluate commonsense reasoning and natural language understanding in artificial intelligence models. It presents multiple-choice questions requiring contextual inference and grounded reasoning.

Read More →

OpenBookQA

OpenBookQA is a benchmark dataset designed for evaluating artificial intelligence systems’ ability to answer elementary science questions using a provided set of facts. It challenges models to perform reasoning beyond simple retrieval by leveraging both a curated ‘open book’ of knowledge and commonsense reasoning.

Read More →

WinoGrande

WinoGrande is a large-scale dataset designed for evaluating commonsense reasoning in natural language processing. It extends the Winograd Schema Challenge by providing thousands of carefully constructed sentence pairs that test an AI’s ability to resolve ambiguous pronouns.

Read More →

CommonsenseQA

CommonsenseQA is a benchmark dataset designed to evaluate the ability of artificial intelligence systems to perform commonsense reasoning through multiple-choice questions. It consists of questions that require understanding and applying everyday knowledge beyond factual recall.

Read More →