TRUSTED BY
Human-level understanding. For superhuman feats.
Experience semantic search and video-to-text capabilities that surpass anything you’ve tried before – video-native AI makes all the difference.
Search
Find specific moments within your videos by describing the scene in natural language.

Analyze
Generate text from videos - summary, chapters, highlights and more.

Embed

Our stable of models.
Marengo 3.0

Sets new benchmarks in zero-shot text-to-video, text-to-image, and text-to-audio retrieval tasks with a single embedding model.
Outperforms Google’s VideoPrism-G model by +10% on the MSR-VTT dataset and +3% on the ActivityNet dataset
Surpasses the SOTA image foundation model in zero-shot text-to-image retrieval tasks, showcasing its ability to understand and process visual content.
Pegasus 1.5

Processes the video input to generate rich embeddings from both video frames and audio speech recognition (ASR) data.
Maps the video embeddings to corresponding language embeddings, creating a shared space where video and text representations are aligned.
The large language model decoder takes the aligned embeddings and user prompts to generate coherent and contextually relevant text output.




