The Pile (dataset)
The Pile is a large-scale dataset designed for training language models. It consists of diverse text sources, enhancing the capabilities of AI in natural language understanding.
Free Information Center
The Pile is a large-scale dataset designed for training language models. It consists of diverse text sources, enhancing the capabilities of AI in natural language understanding.
COCO WholeBody is a significant dataset used in computer vision, particularly for human pose estimation and related research.
SciQ is a dataset designed for evaluating question answering systems in the domain of science education. It consists of multiple-choice science questions paired with supporting facts, intended to aid research in natural language processing and machine learning.
OCHuman is a dataset used in computer vision for human detection and segmentation, particularly focusing on occlusion scenarios.
AudioSet is a large-scale dataset for sound classification created by Google. It contains over 2 million human-labeled sound clips from various sources.
MultiNLI is a large-scale natural language inference dataset designed for evaluating machine learning models in the field of natural language processing.
ActivityNet is a large-scale dataset for video understanding, focusing on complex human activities and events.
WikiText-2 is a dataset designed for training language models, particularly in understanding and generating text.
CommonsenseQA is a benchmark dataset designed to evaluate the ability of artificial intelligence systems to perform commonsense reasoning through multiple-choice questions. It consists of questions that require understanding and applying everyday knowledge beyond factual recall.
The Charades dataset is a large-scale dataset for human activity recognition in video, widely used for training and evaluating machine learning models.