Product
Marengo 3.5: Built for the Structure of Video
Royce, Jeremy, Kihyun, Chris, Kwanseok, Cooper, Kyle, Roy, Dan
A video-native embedding model built around multimodal signals, temporal context, natural boundaries, uncertainty, and efficient representations.
A video-native embedding model built around multimodal signals, temporal context, natural boundaries, uncertainty, and efficient representations.

In this article
Join our newsletter
Receive the latest advancements, tutorials, and industry insights in video understanding
Search, analyze, and explore your videos with AI.
Aug 31, 2026
10min
Copy link to article
Marengo 3.0 established a conviction: video intelligence should be designed for the realities of video, not adapted from image models or optimized for a narrow benchmark. Marengo 3.5 builds on that foundation, with greater attention to how different signals are represented, how context aligns with time, and how video is structured around meaningful events.
Video is not a sequence of images with an audio track attached. Meaning can emerge from what appears in a frame, what changes across time, what is spoken, what is written on screen, and what happens before or after a particular moment. Sometimes the most important signal is not visible at all: a voice outside the frame, a line of dialogue, or context associated with a specific point in time.
A multimodal embedding model represents these signals in a shared semantic space, where related content can remain close even when it comes from different modalities. A text query can retrieve a video moment, an image can retrieve another visual example, or multiple inputs can be combined to express something that neither could specify precisely on its own.
Marengo 3.5 is TwelveLabs’ latest video-native multimodal embedding model. It represents video, audio, images, text, and supported documents in that shared space, enabling retrieval within a modality, across modalities, and from composed inputs.
But representing the right signals is only part of understanding video. Their relationship to time matters too. Context should enrich the moment it describes. Natural transitions should separate distinct events rather than mix them into the same representation. And when an input or match is ambiguous, applications should have a signal that helps them decide what to do next.
Marengo 3.5 carries the conviction behind Marengo 3.0 forward: the structure of video should shape the representation itself. Multimodal signals, temporal context, natural boundaries, and efficiency are treated as parts of the same design problem rather than features layered onto a generic embedding.
Marengo 3.5 broadens what a single embedding can represent—adding flexible dimensions, uncertainty, and richer multimodal composition—without giving up the retrieval foundation established by Marengo 3.0. The comparison covers familiar video and multilingual-text workloads alongside newer audio, visual-document, and composed-input tasks.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
All retrieval tasks are evaluated using nDCG@10, composed retrieval tasks using R@5, and classification and question-answering tasks using Hit@1.
Model | Provider | Default embedding dimension |
|---|---|---|
Marengo 3.5 | TwelveLabs | 512 |
Marengo 3.0 | TwelveLabs | 512 |
Gemini Embedding 2 | 3,072 | |
Amazon Nova Multimodal Embeddings | Amazon Web Services | 3,072 |
The charts shorten Amazon Nova Multimodal Embeddings to Amazon Nova. Scores are task-group means unless noted; higher is better, and unreported values remain N/A rather than being estimated.
New capabilities without giving up core retrieval
Across five standard text-to-video retrieval datasets in MMEB-v2—MSR-VTT, MSVD, DiDeMo, VATEX, and YouCook2—Marengo 3.5 reaches 70.57, improving 6.29 points over Marengo 3.0 and recording the highest reported mean. Its MTEB Multilingual v2 score also rises from 57.03 to 63.44, showing that expanded multimodal capabilities arrive alongside stronger general video retrieval and multilingual text representation quality.
Marengo 3.5 also leads MMEB-v2 video classification, question answering, and moment retrieval, as well as all three MAEB audio groups. On visual documents it leads VisDoc-OOD at 67.58—9.79 points above the next-best reported model—while remaining competitive across the broader ViDoRe suites.
Composition now reaches across target modalities
Marengo 3.0 established strong composed visual retrieval: a reference image and a text modification retrieve the corresponding visual result. Marengo 3.5 remains close on the five public composed-retrieval tasks at 67.31 versus 68.42, while extending composition to queries whose target can be text.
In MMEB-v2 question answering, an image or video is paired with a natural-language question and used to rank candidate text answers. Marengo 3.5 reaches 53.32 on video question answering and 32.69 on image question answering—improvements of 15.26 and 15.81 points over Marengo 3.0. This is answer retrieval in a shared embedding space, not answer generation.
3. Evaluation on production retrieval tasks
Media and sports teams judge retrieval by how quickly an editor can move from an idea to the exact usable clip. Named TwelveLabs customers—including NFL Media, MLSE, Dyn Media, and SBS—describe workflows around finding game moments, searching across archives, and reusing scenes without relying entirely on manual logging.
The following suites are benchmark datasets, not customer footage. They focus on three recurring production patterns: searching sports video, combining a reference image with a text instruction, and using an image itself as the query.
Evaluation suite | Query | Target | # tasks |
|---|---|---|---|
Sports video retrieval | Sports concepts, actions, events, and entities | Video clips | 8 |
Composed video retrieval | Reference image + modifying text | Video clips | 5 |
Image-query retrieval | Reference image | Images or video frames | 3 |
The charts use each model’s default output: 512 dimensions for Marengo 3.5 and Marengo 3.0, and 3,072 for Gemini Embedding 2 and Amazon Nova Multimodal Embeddings. Scores are task means using R@5 or mAP@5.
Search for the moment the scoreboard cannot describe
Structured feeds capture scores and timestamps, but editorial search often begins with a detail that was never logged: a particular action, player, celebration, sponsor logo, or piece of on-screen text. Finding those moments requires understanding motion and identity together with OCR and broadcast graphics.
The evaluation spans diverse sports such as American football, baseball, basketball, ice hockey, and soccer, testing whether the representation transfers across different camera conventions, playing surfaces, and types of action.
Production retrieval summary
Task-weighted aggregate across retrieval suites and enriched metadata · higher is better unless noted
Combine who the user means with what they want to see
Composed retrieval is useful when neither an image nor text is precise enough alone. A reference image can specify identity or appearance, while text supplies the desired action or context—for example, a photo of LeBron James plus “finishing a dunk in traffic.”
Most tasks in the public composed-retrieval suite use image and text to retrieve another image. These five tasks instead evaluate the production-oriented pattern image + text → video, spanning sports, movies, TV person and name retrieval, and single- and multi-reference entity search.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
Start with what the user can show
Sometimes the clearest query is something the user can show: a cropped product, a headshot, or a logo. Image-query retrieval tests whether that example can locate the same visual concept inside a much larger scene or collection without requiring a complete text description.
The evaluation covers finding a generic object—including small objects—from a cropped reference; finding the same person in other frames from a separate profile photo; and finding a brand mark when the logo appears small or surrounded by visual clutter.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
4. Put external context on the timeline
A video may show a possession, a scene, or a news shot without naming the player, score, speaker, or story. Index-time metadata fusion—also called metatext fusion—attaches that external context to the moment it describes.
Timestamped text can be provided as video or audio is indexed. Marengo 3.5 aligns it with the corresponding media segment and adds it to the fused representation, while the visual and audio representations remain unmodified. This is useful for sports feeds, broadcast rundowns, scripts, scene descriptions, event logs, captions, transcripts, OCR, and model-generated descriptions. Because the context is fused at index time, later semantic retrieval does not require application-side score blending.
Basketball segment
[GAME] Cavaliers @ Lakers
[SCOREBOARD] Q1 · 02:02 · CLE 31–LAL 25
[PLAY_BY_PLAY] Davis · driving floating jump shot · missed
[CAPTION] Anthony Davis drives into the paint and attempts a short shot.
[TRANSCRIPT] …his 15th since the start of last year…
[ROSTER] Cavaliers: Mobley, Allen, Mitchell, …
Movie script or production record
[SCENE] Interior · apartment kitchen · night
[CHARS] Maya Chen; Daniel Ruiz
[ACTION] Maya places a sealed envelope on the table. Daniel hesitates before opening it.
[DIALOG] Maya: You deserve to know what happened. Daniel: Then tell me everything.
[PLOT] A private conversation reveals evidence that changes their investigation.
This context becomes especially useful when a query asks for facts the media does not reliably expose. In basketball, “Anthony Davis misses a short shot with 2:02 left in the first quarter” depends on player identity, game clock, score state, and play outcome—even when the scoreboard is obscured or the commentary omits them.
In a movie, a query may name a character directly, but the pixels do not label that person and the embedding model may not reliably connect a face to a fictional identity. Time-aligned script and production metadata bridge that gap by attaching character names, dialogue, actions, and plot context to the correct scene.
These tags are readable delimiters, not required fields; only the relevant context needs to be present. Accuracy and timing matter most. Keep the original metadata for exact filters, auditing, or display, and use the fused representation for semantic queries that combine visible, audible, and contextual clues.
Evaluation with and without timestamped metadata
Basketball search combines actions and player or game context with play-by-play, score state, rosters, captions, and transcripts.
Movie search uses scene descriptions, dialogue, and character context.
News search combines visual descriptions with spoken and topical context.
Soccer search adds match events and game-state information to the action on screen.
Together, the tasks test entity-, event-, and transcript-oriented queries whose answer is the correct video clip. The chart compares Marengo 3.5 and Gemini Embedding 2 with and without timestamped metadata at each model’s default output dimension.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
Strong media understanding, strengthened by metadata
Across basketball, movie, news, and soccer retrieval, Marengo 3.5 already outperforms Gemini Embedding 2 without metadata on every task group, averaging 55.84 versus 31.18. This stronger baseline shows that metadata fusion builds on the underlying video-and-audio representation rather than compensating for it.
When timestamped metadata is fused, Marengo 3.5 improves to 67.94, a gain of 12.10 points, and remains higher across all four task groups; Gemini Embedding 2 reaches 58.50. The resulting 9.44-point lead shows the value of combining strong media understanding with accurate, time-local context.
5. Choose the embedding size that fits the system
Matryoshka Representation Learning (MRL) makes embedding dimension a deployment choice. Smaller embeddings reduce storage, memory, bandwidth, and embedding-index cost; larger embeddings can preserve additional quality when the workload justifies the footprint.
The curves compare Marengo 3.5, Gemini Embedding 2, and Amazon Nova Multimodal Embeddings at the dimension settings reported in this evaluation. The x-axis uses a labeled log₂ scale so compact embeddings remain readable, while each panel uses its own labeled score range.
One view across eight evaluation groups
The aggregate gives equal weight to MMEB-v2 video retrieval, image retrieval, and VisDoc; MTEB Multilingual v2; three MAEB groups; and ViDoRe v1–v3. MMEB-v2 VisDoc and ViDoRe combine their component suites using task-count weighting before contributing to the eight-group mean.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
Strong quality at compact dimensions
Across the eight-group aggregate, Marengo 3.5 leads at every shared embedding size. At 256 dimensions, it scores 66.24, compared with 55.54 for Gemini Embedding 2 and 46.76 for Amazon Nova Multimodal Embeddings. At 512 dimensions, Marengo 3.5 reaches 67.38—higher than the largest reported 3,072-dimensional scores from Gemini Embedding 2 (62.35) and Amazon Nova Multimodal Embeddings (50.34).
Most of Marengo 3.5’s aggregate gain arrives at compact dimensions: it scores 63.75 at 128 dimensions and 66.24 at 256, then adds 1.14 points from 256 to 512. As supporting context, the 256- and 128-dimensional embeddings retain 98.3% and 94.6% of its 512-dimensional score. The detailed curves show where individual workloads flatten or continue to benefit from additional dimensions.
6. Know when the embedding is uncertain
An embedding normally places an input at a point in semantic space. Building on probabilistic embedding research such as ProLIP, Marengo 3.5 can also return an opt-in uncertainty signal that estimates how much plausible meaning surrounds that point. A broad query such as “a player scores” should be less certain than one that identifies the player, action, and game context.
Marengo 3.5 extends this signal across video, audio, images, text, and supported documents. The model learns to associate greater uncertainty with inputs that admit more plausible interpretations and with matches that are harder to separate. It can support clarification, abstention, reranking, or routing difficult cases to a more intensive workflow.
One signal, two decision points
Signal | What it measures | Example use |
|---|---|---|
Query uncertainty | Ambiguity in the input itself | Ask for a more specific query or choose a broader retrieval strategy |
Pairwise confidence | Stability of a particular query–result match | Filter, rerank, or escalate low-confidence candidates |
Uncertainty complements similarity; it does not replace it. It should not be treated as a universal probability of correctness. The right operating threshold depends on the corpus, task, and cost of an error, so decisions should be calibrated on representative labeled traffic.
Measuring uncertainty on hierarchical captions
HierarCaps pairs each of 1,000 manually reviewed test images with four valid captions ordered from a broad concept to a precise description. The evaluation measures whether embedding geometry follows that order, whether confidence ranks more reliable retrievals first, and how closely confidence matches observed retrieval accuracy.
All four models use the same images, captions, and retrieval pool. Hierarchy ordering uses each model’s caption embeddings and a shared empty-text anchor. For AURC and ECE, Marengo 3.5 uses its pairwise confidence signal; Marengo 3.0, Gemini Embedding 2, and Amazon Nova Multimodal Embeddings use raw cosine similarity as the confidence baseline. No evaluation-set normalization or post-hoc calibration is applied.
Measure | What it tests | Better |
|---|---|---|
Hierarchy ordering (τd) | Caption embeddings should progress from broad to specific relative to the shared anchor | Higher |
Area under the risk–coverage curve (AURC) | More reliable retrievals should remain as low-confidence cases are withheld | Lower |
Expected calibration error (ECE) | Reported confidence should match observed accuracy | Lower |
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
7. Let the video define its structure
A fixed window assumes every event lasts the same length. Video does not: a hard cut can change the subject in one frame, a dissolve can unfold gradually, and a sports broadcast can switch camera angles several times within seconds.
The indexing flow preserves this structure at each stage. Before embedding, the Marengo temporal segmentation stage identifies transition boundaries and divides the timeline into variable-length segments. Marengo 3.5 then represents each segment independently, keeping unrelated shots separate and giving change-dense video the temporal resolution it needs. When retrieval requires one representation for the complete video, a lightweight learned aggregation step we call the “Composer” combines the segment embeddings into one 512-dimensional video-level embedding.
Segment at natural boundaries
The boundary evaluation covers hard cuts and gradual transitions across three public datasets and a TwelveLabs evaluation set:
Dataset | Evaluation videos | What it tests |
|---|---|---|
ClipShots | 342 | Diverse short web video with camera shake, motion, occlusion, hard cuts, and gradual transitions |
BBC Planet Earth | 11 | Professionally edited wildlife documentary footage with broadcast-style transitions |
SportsShot | 240 | Basketball, football, and volleyball footage with rapid camera changes, zooms, and changes of view |
General video | 107 | TwelveLabs evaluation set spanning movies and media, sports, news, and general video with varied editing styles |
A predicted boundary counts as correct when it falls within 0.1 or 0.5 seconds of the annotated transition. The chart reports best-threshold boundary F1 using the same threshold grid for Marengo temporal segmentation, TransNetV2, and AutoShot.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
Accurate boundaries across varied video
Marengo temporal segmentation leads all four evaluation sets at both tolerances. At ±0.1 seconds, it reaches 80.62 on ClipShots, 97.04 on BBC Planet Earth, 98.46 on SportsShot, and 95.35 on General video—respectively 1.40, 0.07, 3.87, and 0.25 points above the next-best result.
The widest separation appears on SportsShot, where rapid camera changes make accurate localization especially important. Across professionally edited wildlife footage and the TwelveLabs general-video evaluation, Marengo temporal segmentation remains on par with strong baselines while taking the lead. Accurate boundaries protect the embeddings that follow: a missed cut can mix unrelated events, while a false cut can fragment one coherent moment.
Compose segments into one video representation
A simple baseline averages the segment embeddings. The comparison below tests mean aggregation against Composer across both Marengo temporal segmentation and TransNetV2 using five MMEB-v2 video-retrieval datasets.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
Composer strengthens whole-video retrieval
The lightweight Composer raises the five-dataset mean from 69.13 to 70.57 with Marengo temporal segmentation, a gain of 1.44 points, and from 69.13 to 70.42 with TransNetV2, a similar 1.29-point gain. The consistency shows that Composer efficiently strengthens whole-video retrieval without depending on a particular segmentation method.
8. Build around the structure of video
Marengo 3.5 advances the direction we set with Marengo 3.0: build representations around the structure of video rather than treating video as another input format.
The result is a model that is stronger across video, audio, images, text, and visual documents while also handling the parts of video retrieval that matter beyond a benchmark—time-aligned context, natural temporal boundaries, composed queries, and efficient representations.
Taken together, these improvements point toward a broader goal: representations that preserve enough of the structure and context of video for increasingly complex search and retrieval systems, without losing the simplicity of a shared embedding space.
TwelveLabs Team
Marengo 3.5 was a joint effort across multiple areas:
Research & Technical Direction: Dan Kim
Model Research & Training: Royce Han, Jeremy Kim, Kihyun You, Cooper Han, Kwanseok Kim, Kyle Park
Model Systems & Serving: Chris Jeon, Roy Kim
Product Management: Eric Kim, Travis Couture
Marengo 3.0 established a conviction: video intelligence should be designed for the realities of video, not adapted from image models or optimized for a narrow benchmark. Marengo 3.5 builds on that foundation, with greater attention to how different signals are represented, how context aligns with time, and how video is structured around meaningful events.
Video is not a sequence of images with an audio track attached. Meaning can emerge from what appears in a frame, what changes across time, what is spoken, what is written on screen, and what happens before or after a particular moment. Sometimes the most important signal is not visible at all: a voice outside the frame, a line of dialogue, or context associated with a specific point in time.
A multimodal embedding model represents these signals in a shared semantic space, where related content can remain close even when it comes from different modalities. A text query can retrieve a video moment, an image can retrieve another visual example, or multiple inputs can be combined to express something that neither could specify precisely on its own.
Marengo 3.5 is TwelveLabs’ latest video-native multimodal embedding model. It represents video, audio, images, text, and supported documents in that shared space, enabling retrieval within a modality, across modalities, and from composed inputs.
But representing the right signals is only part of understanding video. Their relationship to time matters too. Context should enrich the moment it describes. Natural transitions should separate distinct events rather than mix them into the same representation. And when an input or match is ambiguous, applications should have a signal that helps them decide what to do next.
Marengo 3.5 carries the conviction behind Marengo 3.0 forward: the structure of video should shape the representation itself. Multimodal signals, temporal context, natural boundaries, and efficiency are treated as parts of the same design problem rather than features layered onto a generic embedding.
Marengo 3.5 broadens what a single embedding can represent—adding flexible dimensions, uncertainty, and richer multimodal composition—without giving up the retrieval foundation established by Marengo 3.0. The comparison covers familiar video and multilingual-text workloads alongside newer audio, visual-document, and composed-input tasks.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
All retrieval tasks are evaluated using nDCG@10, composed retrieval tasks using R@5, and classification and question-answering tasks using Hit@1.
Model | Provider | Default embedding dimension |
|---|---|---|
Marengo 3.5 | TwelveLabs | 512 |
Marengo 3.0 | TwelveLabs | 512 |
Gemini Embedding 2 | 3,072 | |
Amazon Nova Multimodal Embeddings | Amazon Web Services | 3,072 |
The charts shorten Amazon Nova Multimodal Embeddings to Amazon Nova. Scores are task-group means unless noted; higher is better, and unreported values remain N/A rather than being estimated.
New capabilities without giving up core retrieval
Across five standard text-to-video retrieval datasets in MMEB-v2—MSR-VTT, MSVD, DiDeMo, VATEX, and YouCook2—Marengo 3.5 reaches 70.57, improving 6.29 points over Marengo 3.0 and recording the highest reported mean. Its MTEB Multilingual v2 score also rises from 57.03 to 63.44, showing that expanded multimodal capabilities arrive alongside stronger general video retrieval and multilingual text representation quality.
Marengo 3.5 also leads MMEB-v2 video classification, question answering, and moment retrieval, as well as all three MAEB audio groups. On visual documents it leads VisDoc-OOD at 67.58—9.79 points above the next-best reported model—while remaining competitive across the broader ViDoRe suites.
Composition now reaches across target modalities
Marengo 3.0 established strong composed visual retrieval: a reference image and a text modification retrieve the corresponding visual result. Marengo 3.5 remains close on the five public composed-retrieval tasks at 67.31 versus 68.42, while extending composition to queries whose target can be text.
In MMEB-v2 question answering, an image or video is paired with a natural-language question and used to rank candidate text answers. Marengo 3.5 reaches 53.32 on video question answering and 32.69 on image question answering—improvements of 15.26 and 15.81 points over Marengo 3.0. This is answer retrieval in a shared embedding space, not answer generation.
3. Evaluation on production retrieval tasks
Media and sports teams judge retrieval by how quickly an editor can move from an idea to the exact usable clip. Named TwelveLabs customers—including NFL Media, MLSE, Dyn Media, and SBS—describe workflows around finding game moments, searching across archives, and reusing scenes without relying entirely on manual logging.
The following suites are benchmark datasets, not customer footage. They focus on three recurring production patterns: searching sports video, combining a reference image with a text instruction, and using an image itself as the query.
Evaluation suite | Query | Target | # tasks |
|---|---|---|---|
Sports video retrieval | Sports concepts, actions, events, and entities | Video clips | 8 |
Composed video retrieval | Reference image + modifying text | Video clips | 5 |
Image-query retrieval | Reference image | Images or video frames | 3 |
The charts use each model’s default output: 512 dimensions for Marengo 3.5 and Marengo 3.0, and 3,072 for Gemini Embedding 2 and Amazon Nova Multimodal Embeddings. Scores are task means using R@5 or mAP@5.
Search for the moment the scoreboard cannot describe
Structured feeds capture scores and timestamps, but editorial search often begins with a detail that was never logged: a particular action, player, celebration, sponsor logo, or piece of on-screen text. Finding those moments requires understanding motion and identity together with OCR and broadcast graphics.
The evaluation spans diverse sports such as American football, baseball, basketball, ice hockey, and soccer, testing whether the representation transfers across different camera conventions, playing surfaces, and types of action.
Production retrieval summary
Task-weighted aggregate across retrieval suites and enriched metadata · higher is better unless noted
Combine who the user means with what they want to see
Composed retrieval is useful when neither an image nor text is precise enough alone. A reference image can specify identity or appearance, while text supplies the desired action or context—for example, a photo of LeBron James plus “finishing a dunk in traffic.”
Most tasks in the public composed-retrieval suite use image and text to retrieve another image. These five tasks instead evaluate the production-oriented pattern image + text → video, spanning sports, movies, TV person and name retrieval, and single- and multi-reference entity search.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
Start with what the user can show
Sometimes the clearest query is something the user can show: a cropped product, a headshot, or a logo. Image-query retrieval tests whether that example can locate the same visual concept inside a much larger scene or collection without requiring a complete text description.
The evaluation covers finding a generic object—including small objects—from a cropped reference; finding the same person in other frames from a separate profile photo; and finding a brand mark when the logo appears small or surrounded by visual clutter.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
4. Put external context on the timeline
A video may show a possession, a scene, or a news shot without naming the player, score, speaker, or story. Index-time metadata fusion—also called metatext fusion—attaches that external context to the moment it describes.
Timestamped text can be provided as video or audio is indexed. Marengo 3.5 aligns it with the corresponding media segment and adds it to the fused representation, while the visual and audio representations remain unmodified. This is useful for sports feeds, broadcast rundowns, scripts, scene descriptions, event logs, captions, transcripts, OCR, and model-generated descriptions. Because the context is fused at index time, later semantic retrieval does not require application-side score blending.
Basketball segment
[GAME] Cavaliers @ Lakers
[SCOREBOARD] Q1 · 02:02 · CLE 31–LAL 25
[PLAY_BY_PLAY] Davis · driving floating jump shot · missed
[CAPTION] Anthony Davis drives into the paint and attempts a short shot.
[TRANSCRIPT] …his 15th since the start of last year…
[ROSTER] Cavaliers: Mobley, Allen, Mitchell, …
Movie script or production record
[SCENE] Interior · apartment kitchen · night
[CHARS] Maya Chen; Daniel Ruiz
[ACTION] Maya places a sealed envelope on the table. Daniel hesitates before opening it.
[DIALOG] Maya: You deserve to know what happened. Daniel: Then tell me everything.
[PLOT] A private conversation reveals evidence that changes their investigation.
This context becomes especially useful when a query asks for facts the media does not reliably expose. In basketball, “Anthony Davis misses a short shot with 2:02 left in the first quarter” depends on player identity, game clock, score state, and play outcome—even when the scoreboard is obscured or the commentary omits them.
In a movie, a query may name a character directly, but the pixels do not label that person and the embedding model may not reliably connect a face to a fictional identity. Time-aligned script and production metadata bridge that gap by attaching character names, dialogue, actions, and plot context to the correct scene.
These tags are readable delimiters, not required fields; only the relevant context needs to be present. Accuracy and timing matter most. Keep the original metadata for exact filters, auditing, or display, and use the fused representation for semantic queries that combine visible, audible, and contextual clues.
Evaluation with and without timestamped metadata
Basketball search combines actions and player or game context with play-by-play, score state, rosters, captions, and transcripts.
Movie search uses scene descriptions, dialogue, and character context.
News search combines visual descriptions with spoken and topical context.
Soccer search adds match events and game-state information to the action on screen.
Together, the tasks test entity-, event-, and transcript-oriented queries whose answer is the correct video clip. The chart compares Marengo 3.5 and Gemini Embedding 2 with and without timestamped metadata at each model’s default output dimension.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
Strong media understanding, strengthened by metadata
Across basketball, movie, news, and soccer retrieval, Marengo 3.5 already outperforms Gemini Embedding 2 without metadata on every task group, averaging 55.84 versus 31.18. This stronger baseline shows that metadata fusion builds on the underlying video-and-audio representation rather than compensating for it.
When timestamped metadata is fused, Marengo 3.5 improves to 67.94, a gain of 12.10 points, and remains higher across all four task groups; Gemini Embedding 2 reaches 58.50. The resulting 9.44-point lead shows the value of combining strong media understanding with accurate, time-local context.
5. Choose the embedding size that fits the system
Matryoshka Representation Learning (MRL) makes embedding dimension a deployment choice. Smaller embeddings reduce storage, memory, bandwidth, and embedding-index cost; larger embeddings can preserve additional quality when the workload justifies the footprint.
The curves compare Marengo 3.5, Gemini Embedding 2, and Amazon Nova Multimodal Embeddings at the dimension settings reported in this evaluation. The x-axis uses a labeled log₂ scale so compact embeddings remain readable, while each panel uses its own labeled score range.
One view across eight evaluation groups
The aggregate gives equal weight to MMEB-v2 video retrieval, image retrieval, and VisDoc; MTEB Multilingual v2; three MAEB groups; and ViDoRe v1–v3. MMEB-v2 VisDoc and ViDoRe combine their component suites using task-count weighting before contributing to the eight-group mean.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
Strong quality at compact dimensions
Across the eight-group aggregate, Marengo 3.5 leads at every shared embedding size. At 256 dimensions, it scores 66.24, compared with 55.54 for Gemini Embedding 2 and 46.76 for Amazon Nova Multimodal Embeddings. At 512 dimensions, Marengo 3.5 reaches 67.38—higher than the largest reported 3,072-dimensional scores from Gemini Embedding 2 (62.35) and Amazon Nova Multimodal Embeddings (50.34).
Most of Marengo 3.5’s aggregate gain arrives at compact dimensions: it scores 63.75 at 128 dimensions and 66.24 at 256, then adds 1.14 points from 256 to 512. As supporting context, the 256- and 128-dimensional embeddings retain 98.3% and 94.6% of its 512-dimensional score. The detailed curves show where individual workloads flatten or continue to benefit from additional dimensions.
6. Know when the embedding is uncertain
An embedding normally places an input at a point in semantic space. Building on probabilistic embedding research such as ProLIP, Marengo 3.5 can also return an opt-in uncertainty signal that estimates how much plausible meaning surrounds that point. A broad query such as “a player scores” should be less certain than one that identifies the player, action, and game context.
Marengo 3.5 extends this signal across video, audio, images, text, and supported documents. The model learns to associate greater uncertainty with inputs that admit more plausible interpretations and with matches that are harder to separate. It can support clarification, abstention, reranking, or routing difficult cases to a more intensive workflow.
One signal, two decision points
Signal | What it measures | Example use |
|---|---|---|
Query uncertainty | Ambiguity in the input itself | Ask for a more specific query or choose a broader retrieval strategy |
Pairwise confidence | Stability of a particular query–result match | Filter, rerank, or escalate low-confidence candidates |
Uncertainty complements similarity; it does not replace it. It should not be treated as a universal probability of correctness. The right operating threshold depends on the corpus, task, and cost of an error, so decisions should be calibrated on representative labeled traffic.
Measuring uncertainty on hierarchical captions
HierarCaps pairs each of 1,000 manually reviewed test images with four valid captions ordered from a broad concept to a precise description. The evaluation measures whether embedding geometry follows that order, whether confidence ranks more reliable retrievals first, and how closely confidence matches observed retrieval accuracy.
All four models use the same images, captions, and retrieval pool. Hierarchy ordering uses each model’s caption embeddings and a shared empty-text anchor. For AURC and ECE, Marengo 3.5 uses its pairwise confidence signal; Marengo 3.0, Gemini Embedding 2, and Amazon Nova Multimodal Embeddings use raw cosine similarity as the confidence baseline. No evaluation-set normalization or post-hoc calibration is applied.
Measure | What it tests | Better |
|---|---|---|
Hierarchy ordering (τd) | Caption embeddings should progress from broad to specific relative to the shared anchor | Higher |
Area under the risk–coverage curve (AURC) | More reliable retrievals should remain as low-confidence cases are withheld | Lower |
Expected calibration error (ECE) | Reported confidence should match observed accuracy | Lower |
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
7. Let the video define its structure
A fixed window assumes every event lasts the same length. Video does not: a hard cut can change the subject in one frame, a dissolve can unfold gradually, and a sports broadcast can switch camera angles several times within seconds.
The indexing flow preserves this structure at each stage. Before embedding, the Marengo temporal segmentation stage identifies transition boundaries and divides the timeline into variable-length segments. Marengo 3.5 then represents each segment independently, keeping unrelated shots separate and giving change-dense video the temporal resolution it needs. When retrieval requires one representation for the complete video, a lightweight learned aggregation step we call the “Composer” combines the segment embeddings into one 512-dimensional video-level embedding.
Segment at natural boundaries
The boundary evaluation covers hard cuts and gradual transitions across three public datasets and a TwelveLabs evaluation set:
Dataset | Evaluation videos | What it tests |
|---|---|---|
ClipShots | 342 | Diverse short web video with camera shake, motion, occlusion, hard cuts, and gradual transitions |
BBC Planet Earth | 11 | Professionally edited wildlife documentary footage with broadcast-style transitions |
SportsShot | 240 | Basketball, football, and volleyball footage with rapid camera changes, zooms, and changes of view |
General video | 107 | TwelveLabs evaluation set spanning movies and media, sports, news, and general video with varied editing styles |
A predicted boundary counts as correct when it falls within 0.1 or 0.5 seconds of the annotated transition. The chart reports best-threshold boundary F1 using the same threshold grid for Marengo temporal segmentation, TransNetV2, and AutoShot.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
Accurate boundaries across varied video
Marengo temporal segmentation leads all four evaluation sets at both tolerances. At ±0.1 seconds, it reaches 80.62 on ClipShots, 97.04 on BBC Planet Earth, 98.46 on SportsShot, and 95.35 on General video—respectively 1.40, 0.07, 3.87, and 0.25 points above the next-best result.
The widest separation appears on SportsShot, where rapid camera changes make accurate localization especially important. Across professionally edited wildlife footage and the TwelveLabs general-video evaluation, Marengo temporal segmentation remains on par with strong baselines while taking the lead. Accurate boundaries protect the embeddings that follow: a missed cut can mix unrelated events, while a false cut can fragment one coherent moment.
Compose segments into one video representation
A simple baseline averages the segment embeddings. The comparison below tests mean aggregation against Composer across both Marengo temporal segmentation and TransNetV2 using five MMEB-v2 video-retrieval datasets.
MMEB-v2
Video, image, and visual-document evaluation · higher is better unless noted
Composer strengthens whole-video retrieval
The lightweight Composer raises the five-dataset mean from 69.13 to 70.57 with Marengo temporal segmentation, a gain of 1.44 points, and from 69.13 to 70.42 with TransNetV2, a similar 1.29-point gain. The consistency shows that Composer efficiently strengthens whole-video retrieval without depending on a particular segmentation method.
8. Build around the structure of video
Marengo 3.5 advances the direction we set with Marengo 3.0: build representations around the structure of video rather than treating video as another input format.
The result is a model that is stronger across video, audio, images, text, and visual documents while also handling the parts of video retrieval that matter beyond a benchmark—time-aligned context, natural temporal boundaries, composed queries, and efficient representations.
Taken together, these improvements point toward a broader goal: representations that preserve enough of the structure and context of video for increasingly complex search and retrieval systems, without losing the simplicity of a shared embedding space.
TwelveLabs Team
Marengo 3.5 was a joint effort across multiple areas:
Research & Technical Direction: Dan Kim
Model Research & Training: Royce Han, Jeremy Kim, Kihyun You, Cooper Han, Kwanseok Kim, Kyle Park
Model Systems & Serving: Chris Jeon, Roy Kim
Product Management: Eric Kim, Travis Couture
Platform
Enterprise
©
2026
TwelveLabs, Inc. All Rights Reserved
Platform
Enterprise
©
2026
TwelveLabs, Inc. All Rights Reserved
Platform
Enterprise
©
2026
TwelveLabs, Inc. All Rights Reserved





