For everything you want to do with video.
Build features like semantic search and content recommenders, or novel applications that redefine what’s possible. Our state-of-the-art AI video understanding unlocks your video’s full potential.
Human-level understanding. For superhuman feats.
Experience semantic search and video-to-text capabilities that surpass anything you’ve tried before – video-native AI makes all the difference.
Search
Find specific moments within your videos by describing the scene in natural language.

Analyze
Generate text from videos - summary, chapters, highlights and more.

Embed

Find any scene in natural language.
Fast, precise, context-aware results that truly understand what you’re looking for. Search across speech, text, audio and visuals to explore your video in every dimension.
What do you want to find?
Try 'search' with a Sample App
What do you want to describe?
Try 'analyze' with a Sample App
What do you want to build?
Try 'embed' with a Sample App
Play for free.
Pay as you go.
For testing and building
Indexing limit
<10 hours
Environment
Shared
Org account
SSO/ SAML
For launching and growing
Indexing limit
Unlimited
Environment
Shared
Org account
SSO/ SAML
For scaling and services
Indexing limit
Unlimited
Environment
Dedicated
Org account
Included
SSO/ SAML
Included
Our stable of models.
Marengo 3.0

Sets new benchmarks in zero-shot text-to-video, text-to-image, and text-to-audio retrieval tasks with a single embedding model.
Outperforms Google's VideoPrism-G model by +10% on the MSR-VTT dataset and +3% on the ActivityNet dataset
Surpasses the SOTA image foundation model in zero-shot text-to-image retrieval tasks, showcasing its ability to understand and process visual content.
Pegasus 1.5

Processes the video input to generate rich embeddings from both video frames and audio speech recognition (ASR) data.
Maps the video embeddings to corresponding language embeddings, creating a shared space where video and text representations are aligned.
The large language model decoder takes the aligned embeddings and user prompts to generate coherent and contextually relevant text output.









