EPIC-KITCHENS (egocentric video dataset)
EPIC-KITCHENS is a large-scale egocentric video dataset designed for action recognition and video understanding in kitchen environments.
Free Information Center
EPIC-KITCHENS is a large-scale egocentric video dataset designed for action recognition and video understanding in kitchen environments.
HMDB51 is a comprehensive database designed for human motion analysis, containing videos of various human activities categorized for research purposes.
Tencent AI Lab is a research division of Tencent focused on advancing artificial intelligence technologies. Established to explore AI applications across various domains, the lab conducts research in areas such as machine learning, natural language processing, and computer vision.
COCO WholeBody is a significant dataset used in computer vision, particularly for human pose estimation and related research.
Neuralangelo is a 3D reconstruction technology that utilizes neural networks to generate detailed three-dimensional models from images or video data. It leverages advances in artificial intelligence to create accurate, textured models useful for various applications including cultural heritage preservation and virtual reality.
OCHuman is a dataset used in computer vision for human detection and segmentation, particularly focusing on occlusion scenarios.
PC-AVS (pose-controllable audio-visual system) is a technology that integrates pose estimation with audio-visual synthesis, enabling interactive and controllable multimedia experiences. It allows users to manipulate audio and visual outputs based on detected human poses or movements.
PaLM-E is an advanced embodied language model that integrates vision and language processing, enhancing the interaction between AI and the physical world.
Third-person imitation learning is a machine learning technique where an autonomous agent learns to perform tasks by observing demonstrations from a third-person perspective. It enables learning from videos or observations where the demonstrator’s viewpoint differs from the learner’s, facilitating broader applications in robotics and AI.
A neural occupancy field is a representation used in computer vision and graphics to encode 3D shapes or scenes by learning a continuous function that predicts the occupancy status of any point in space. This technique leverages neural networks to model detailed geometric information efficiently.