PEGASUS 1.5
Turn raw video into structured data.
2hrs
Max video duration with full temporal context retained end-to-end.
12
Languages supported across input prompts and generated output.
0
Pre-indexing steps. Send a URL, an asset, or base64, get text back.
JSON
Structured segmentation output ready for your editor, or pipeline.
Things only Pegasus does.

The answer comes with a timestamp.
Pegasus answers with exact timestamp, not “somewhere in the middle.” Granular temporal reasoning is built into the model.

Reads the fine print.
On-screen text. Jersey numbers. Whiteboards. Receipts. Pegasus parses the frame alongside the speech track, so summaries include the slides too.

From a raw video to structured JSON.
Define a segment such as a speaker change, brand appearance, or scene cut, choose your fields, and Pegasus returns timestamped JSON.

Reference an image. Ask about it.
Drop in a reference image such as a logo, face, or product, and Pegasus uses it as visual context.
From signup to first result in 5 minutes.
Built for video. Not for frames.
7,200s
Continuous video Pegasus handles in a single prompt — two full hours.

CAPABILITY
PEGASUS 1.5
Gemini 3.1 PRO
GPT-5.5
Max single-call duration
120 min
90 min
Not specified (omnimodal, no published video duration cap)
Structured segmentation output
JSON-native, schema-conditioned
Structured Outputs supported; no native temporal segmentation
Structured Outputs supported; no native temporal segmentation
Multimodal prompting (image+text)
Yes (within structured segment output)
Yes (OCR is marquee Gemini 3.x capability)
Yes (omnimodal OCR support)
Per-definition time windows
Yes
General-purpose frontier model
General-purpose omnimodal model



