Learn
Teaching Machines How the World Works with Video Intelligence
Shannon Hong, Yoon Kim
TwelveLabs Bridges Language and Physical Intelligence
TwelveLabs Bridges Language and Physical Intelligence

Sep 30, 2026
2 Minutes
Copy link to article
Introduction
With cameras everywhere, machines can already see. But seeing alone is not enough. Machines need to understand and reason over what they see to act upon our world. At TwelveLabs, our video intelligence models already understand human-curated video content including news, movies, sports plays, actions, and more. Now, TwelveLabs is expanding its model capabilities to help machines learn about the physical world: its people, objects, events, tasks and environments.
Physical AI is the embodiment of AI into hardware that can perceive, interpret, reason about, and act within the physical 3D world, and it will be yet another tool of human creation and creativity: What more will we discover? What more can we know about our world? What can we do now to advance the state of the art? Bringing about this capability requires video intelligence, which includes the broad abilities of 1) understanding novel stimuli across time and space and 2) interpreting and reasoning to support the best actions.
Currently, language intelligence exists primarily in text and code, acting in the digital world, blind to what physically happened, and physical intelligence operates moment to moment, with no memory of what it has seen and no true comprehension of its actions. LLMs think but cannot see well. Robots can see and react but cannot judge or reason well. Video Intelligence bridges language and physical intelligence, by understanding the physical world through video and making judgments through reasoning.
Video Intelligence Gives Physical AI Temporal Context, Spatial Reasoning, and Judgment
The first test of video intelligence’s ability to power Physical AI is in robotics, through the understanding of first-person perspective (egocentric) video. Over the past five years, TwelveLabs has built general video intelligence infrastructure and models that are used by the largest organizations in media, entertainment, and government sectors. TwelveLabs’ Pegasus 1.5 targets temporal segmentation and structured metadata—identifying where events begin and end, then describing them in a form downstream systems can use. These capabilities provide a foundation for the detailed annotations Physical AI workflows require. Advancing our models lets us improve their representations, training, and outputs together as those requirements evolve rapidly. The technical foundation and know-how of building video-native intelligence enables TwelveLabs to build an effective data intelligence layer for egocentric video: turning raw footage of human work into structured datasets with dense, time-based metadata.
With TwelveLabs, Physical AI will gain a novel understanding of time, space, and the judgment to know whether a task succeeded in the physical world: Our models help Physical AI understand both archive-level quantities of video as well as the most dense and granular segment of an action. Marengo processes raw video footage into structured embeddings to enable search and retrieval; Pegasus takes in video and describes the people, objects, events, and actions at a dense level of detail.
Temporal context: understand how actions become outcomes. Pegasus enables you to go beyond isolated moments to follow a task as it unfolds, revealing what happened before, what is happening now, and what comes next with dense captioning. Connect actions across time, and recognize progress and interruptions over time.
Spatial reasoning: understand how the physical world fits together. Identify which hand is acting, what it is holding, what surface the work sits on, and which objects are involved. Using Pegasus, turn video into accounts of hands, tools, surfaces, and objects that a downstream system can index and filter, ahead of the geometric work that recovers pose and contact points.
Judgment: determine whether the job was done. We aim to extend Pegasus from describing actions to assessing progress, completion, and failure against explicit criteria. Calibrated against human raters, these judgments could support evaluation, production QA, and reward signals.
Egocentric Video Intelligence Unlocks Physical AI Training at Scale
Video has become the predominant means by which humanity will “show and tell” its history and document the unfolding of everyday events. For machines that move through the world, that record is the raw material of intelligence, and it is already piling up in vehicles, factories, warehouses, and homes. But raw footage does not provide the structure a training pipeline requires. Robotics data can’t be scraped from the web like text; it has to be captured, structured, and made accessible before a model can learn from it.¹ A generalist robot policy such as π0 rests on roughly 10,000 hours of teleoperated robot data,² and the human-video corpora behind recent results are orders of magnitude larger.³ Across public egocentric datasets, hand labeling costs 70⁴ to 155⁵ hours of annotator work for every hour of footage. At that scale, no team can review, label, or search footage by hand.
TwelveLabs alleviates this bottleneck: with video intelligence, the economics and sustainability of training data production for Physical AI finally make sense. Our models ingest raw footage at archive scale and turn it into curated training datasets at a fraction of the cost of manual processing, with human review prioritized for uncertain cases. That collapses the time to understand all of that video, so Physical AI researchers can spend the next few years on breakthroughs instead of the next decade organizing data. TwelveLabs will be the intelligence layer that gives sight and insight to machines, shaping machines’ ability to reason and act with precision, speed and scale.
References
¹ https://www.ibm.com/think/news/the-data-gap-holding-back-robotics
² https://www.pi.website/blog/pi0
³ https://topicqueue.substack.com/p/a-million-hours-of-human-video-zero
⁴ Grauman et al., Ego4D: Around the World in 3,000 Hours of Egocentric Video, arXiv:2110.07058, CVPR 2022.
⁵ Grauman et al., Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives, arXiv:2311.18259, CVPR 2024.
Introduction
With cameras everywhere, machines can already see. But seeing alone is not enough. Machines need to understand and reason over what they see to act upon our world. At TwelveLabs, our video intelligence models already understand human-curated video content including news, movies, sports plays, actions, and more. Now, TwelveLabs is expanding its model capabilities to help machines learn about the physical world: its people, objects, events, tasks and environments.
Physical AI is the embodiment of AI into hardware that can perceive, interpret, reason about, and act within the physical 3D world, and it will be yet another tool of human creation and creativity: What more will we discover? What more can we know about our world? What can we do now to advance the state of the art? Bringing about this capability requires video intelligence, which includes the broad abilities of 1) understanding novel stimuli across time and space and 2) interpreting and reasoning to support the best actions.
Currently, language intelligence exists primarily in text and code, acting in the digital world, blind to what physically happened, and physical intelligence operates moment to moment, with no memory of what it has seen and no true comprehension of its actions. LLMs think but cannot see well. Robots can see and react but cannot judge or reason well. Video Intelligence bridges language and physical intelligence, by understanding the physical world through video and making judgments through reasoning.
Video Intelligence Gives Physical AI Temporal Context, Spatial Reasoning, and Judgment
The first test of video intelligence’s ability to power Physical AI is in robotics, through the understanding of first-person perspective (egocentric) video. Over the past five years, TwelveLabs has built general video intelligence infrastructure and models that are used by the largest organizations in media, entertainment, and government sectors. TwelveLabs’ Pegasus 1.5 targets temporal segmentation and structured metadata—identifying where events begin and end, then describing them in a form downstream systems can use. These capabilities provide a foundation for the detailed annotations Physical AI workflows require. Advancing our models lets us improve their representations, training, and outputs together as those requirements evolve rapidly. The technical foundation and know-how of building video-native intelligence enables TwelveLabs to build an effective data intelligence layer for egocentric video: turning raw footage of human work into structured datasets with dense, time-based metadata.
With TwelveLabs, Physical AI will gain a novel understanding of time, space, and the judgment to know whether a task succeeded in the physical world: Our models help Physical AI understand both archive-level quantities of video as well as the most dense and granular segment of an action. Marengo processes raw video footage into structured embeddings to enable search and retrieval; Pegasus takes in video and describes the people, objects, events, and actions at a dense level of detail.
Temporal context: understand how actions become outcomes. Pegasus enables you to go beyond isolated moments to follow a task as it unfolds, revealing what happened before, what is happening now, and what comes next with dense captioning. Connect actions across time, and recognize progress and interruptions over time.
Spatial reasoning: understand how the physical world fits together. Identify which hand is acting, what it is holding, what surface the work sits on, and which objects are involved. Using Pegasus, turn video into accounts of hands, tools, surfaces, and objects that a downstream system can index and filter, ahead of the geometric work that recovers pose and contact points.
Judgment: determine whether the job was done. We aim to extend Pegasus from describing actions to assessing progress, completion, and failure against explicit criteria. Calibrated against human raters, these judgments could support evaluation, production QA, and reward signals.
Egocentric Video Intelligence Unlocks Physical AI Training at Scale
Video has become the predominant means by which humanity will “show and tell” its history and document the unfolding of everyday events. For machines that move through the world, that record is the raw material of intelligence, and it is already piling up in vehicles, factories, warehouses, and homes. But raw footage does not provide the structure a training pipeline requires. Robotics data can’t be scraped from the web like text; it has to be captured, structured, and made accessible before a model can learn from it.¹ A generalist robot policy such as π0 rests on roughly 10,000 hours of teleoperated robot data,² and the human-video corpora behind recent results are orders of magnitude larger.³ Across public egocentric datasets, hand labeling costs 70⁴ to 155⁵ hours of annotator work for every hour of footage. At that scale, no team can review, label, or search footage by hand.
TwelveLabs alleviates this bottleneck: with video intelligence, the economics and sustainability of training data production for Physical AI finally make sense. Our models ingest raw footage at archive scale and turn it into curated training datasets at a fraction of the cost of manual processing, with human review prioritized for uncertain cases. That collapses the time to understand all of that video, so Physical AI researchers can spend the next few years on breakthroughs instead of the next decade organizing data. TwelveLabs will be the intelligence layer that gives sight and insight to machines, shaping machines’ ability to reason and act with precision, speed and scale.
References
¹ https://www.ibm.com/think/news/the-data-gap-holding-back-robotics
² https://www.pi.website/blog/pi0
³ https://topicqueue.substack.com/p/a-million-hours-of-human-video-zero
⁴ Grauman et al., Ego4D: Around the World in 3,000 Hours of Egocentric Video, arXiv:2110.07058, CVPR 2022.
⁵ Grauman et al., Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives, arXiv:2311.18259, CVPR 2024.
Platform
Enterprise
©
2026
TwelveLabs, Inc. All Rights Reserved
Platform
Enterprise
©
2026
TwelveLabs, Inc. All Rights Reserved
Platform
Enterprise
©
2026
TwelveLabs, Inc. All Rights Reserved





