Product

What’s New: Pegasus 1.6

Christine Lim

Pegasus 1.6 adds first-person video understanding, stronger entity recognition, native image analysis, and sharper visual comprehension.

Pegasus 1.6 adds first-person video understanding, stronger entity recognition, native image analysis, and sharper visual comprehension.

In this article

No headings found on page

Search, analyze, and explore your videos with AI.

Join our newsletter

Receive the latest advancements, tutorials, and industry insights in video understanding

Oct 6, 2026

2 minutes

Copy link to article

If you've used Pegasus before, you know the foundational idea: it watches a video and turns it into structured, searchable text, such as summaries, tags, timestamps, whatever shape your app needs. Pegasus 1.6 is the newest version of that model, and the upgrade comes down to three things: it understands video shot from a first-person point of view, it's noticeably better at recognizing who and what is actually in a scene, and it now handles still images directly instead of treating them as a video's afterthought.

Here's what changed, and why each piece is useful in practice.

It finally understands first-person video

Almost all the video Pegasus has analyzed so far was shot the normal way — a camera pointed at a scene from the outside, like a security camera or a film crew. Pegasus 1.6 adds real support for egocentric video: footage captured from the point of view of whoever (or whatever) is doing the thing, like a body camera or a robot's onboard camera.

Why it matters: this is exactly the kind of footage that robotics and embodied-AI teams need to label and train on — teleoperation recordings, first-person action data — and it's a category TwelveLabs simply didn't support well before this release.

It's better at knowing who's who

Pegasus 1.6 is meaningfully better at identifying people, characters, and objects in a video and getting their names right. That sounds like a small thing until you're relying on it: a model that confidently mislabels a person isn't just unhelpful, it's actively wrong in a way someone has to catch.

Why it matters: this is the backbone of any content-ID or character-recognition workflow — media cataloging, security review, anything where "who is this" needs to be a reliable answer rather than a best guess.

Images don't need a separate pipeline anymore

Up until now, if a team wanted to analyze still images in addition to video, they had to stand up a second tool just for that. Pegasus 1.6 reads images natively through the same API used for video, using the same model and the same prompts.

Why it matters: one less system to build and maintain. Teams working across both video and image content — think product photos alongside product videos, or screenshots alongside screen recordings — get one integration instead of two.

Why this is different from a general-purpose AI model

General multimodal models, the kind built primarily for text and adapted to handle video, tend to treat video as a stack of still images to caption, one frame at a time. Pegasus was built for video from the start, which is why it can reason about an entire two-hour video in one call, return structured, timestamped output natively, and keep getting sharper at recognizing entities release over release instead of that being a bolted-on afterthought.

Where this actually helps

  • Media, sports, and broadcasting teams get more complete summaries — on-screen graphics and jersey numbers included, not just the commentary.

  • Security teams get entity recognition that holds up across long stretches of footage instead of losing track of who’s who.

  • Robotics and physical AI teams finally have a model that understands first-person footage for training and labeling pipelines.

  • Developers get one API for both video and images, with structured output shaped by whatever schema their app needs.

Try it

Pegasus 1.6 works with the same SDKs, prompts, and API you’re already using if you’re on Pegasus today. Start Building or Talk to Sales if you want help scoping a use case.

If you've used Pegasus before, you know the foundational idea: it watches a video and turns it into structured, searchable text, such as summaries, tags, timestamps, whatever shape your app needs. Pegasus 1.6 is the newest version of that model, and the upgrade comes down to three things: it understands video shot from a first-person point of view, it's noticeably better at recognizing who and what is actually in a scene, and it now handles still images directly instead of treating them as a video's afterthought.

Here's what changed, and why each piece is useful in practice.

It finally understands first-person video

Almost all the video Pegasus has analyzed so far was shot the normal way — a camera pointed at a scene from the outside, like a security camera or a film crew. Pegasus 1.6 adds real support for egocentric video: footage captured from the point of view of whoever (or whatever) is doing the thing, like a body camera or a robot's onboard camera.

Why it matters: this is exactly the kind of footage that robotics and embodied-AI teams need to label and train on — teleoperation recordings, first-person action data — and it's a category TwelveLabs simply didn't support well before this release.

It's better at knowing who's who

Pegasus 1.6 is meaningfully better at identifying people, characters, and objects in a video and getting their names right. That sounds like a small thing until you're relying on it: a model that confidently mislabels a person isn't just unhelpful, it's actively wrong in a way someone has to catch.

Why it matters: this is the backbone of any content-ID or character-recognition workflow — media cataloging, security review, anything where "who is this" needs to be a reliable answer rather than a best guess.

Images don't need a separate pipeline anymore

Up until now, if a team wanted to analyze still images in addition to video, they had to stand up a second tool just for that. Pegasus 1.6 reads images natively through the same API used for video, using the same model and the same prompts.

Why it matters: one less system to build and maintain. Teams working across both video and image content — think product photos alongside product videos, or screenshots alongside screen recordings — get one integration instead of two.

Why this is different from a general-purpose AI model

General multimodal models, the kind built primarily for text and adapted to handle video, tend to treat video as a stack of still images to caption, one frame at a time. Pegasus was built for video from the start, which is why it can reason about an entire two-hour video in one call, return structured, timestamped output natively, and keep getting sharper at recognizing entities release over release instead of that being a bolted-on afterthought.

Where this actually helps

  • Media, sports, and broadcasting teams get more complete summaries — on-screen graphics and jersey numbers included, not just the commentary.

  • Security teams get entity recognition that holds up across long stretches of footage instead of losing track of who’s who.

  • Robotics and physical AI teams finally have a model that understands first-person footage for training and labeling pipelines.

  • Developers get one API for both video and images, with structured output shaped by whatever schema their app needs.

Try it

Pegasus 1.6 works with the same SDKs, prompts, and API you’re already using if you’re on Pegasus today. Start Building or Talk to Sales if you want help scoping a use case.