Partnerships
How Mimir and TwelveLabs Make Video Archives Searchable
Richard Bentley, Simon Lecointe
Mimir and TwelveLabs make video archives searchable without human-created tags, with semantic search, scene-level compliance flags, and automated metadata.
Mimir and TwelveLabs make video archives searchable without human-created tags, with semantic search, scene-level compliance flags, and automated metadata.

In this article
Join our newsletter
Receive the latest advancements, tutorials, and industry insights in video understanding
Search, analyze, and explore your videos with AI.
Jul 23, 2026
8 minutes
Copy link to article
Somewhere in a broadcaster's warehouse sit rooms full of tape: DigiBeta, DVCam, VHS, film reels, LTO tapes, three-quarter inch. The metadata lives in handwritten notebooks on the shelves, and the only person who can say what's actually on any given tape is the person who wrote those notes by hand. That archive is real, and it's the exact problem this integration is built to solve: hundreds of thousands of hours of content, sitting unused, because nobody can find anything in it.
Getting that content off legacy physical media and into a system is its own project. But even once it's digitized, the underlying problem just changes shape. The bottleneck was never the format, but the metadata: whether the system actually knows what's happening inside the content, not just where the file lives.
Mimir and TwelveLabs solve two different halves of that problem. Mimir is a serverless media production and collaboration platform — the system of record where content lives, gets organized, and moves through production. TwelveLabs is the intelligence layer that understands what's inside that content: every scene, every word spoken, every object and action, indexed and made searchable.
Mimir gives organizations a place to manage media at scale. TwelveLabs gives that media the kind of understanding that used to require a person watching every frame. Together, they turn a media archive from a place where content is simply stored into one where it can be found, understood, and repurposed.
Everything described in this blog was recently demoed live in a webinar with Mimir and TwelveLabs. If you’d rather watch the demo than read about the workflows, the on-demand recording is available now.
Semantic Search in Mimir: Finding Footage That was Never Tagged
Search for “Macron” in Mimir, and it will pull results from clip titles, transcripts, and face recognition. Add on qualifiers like “wearing a tie” and “says the word monopoly,” stack those into an overlap search, and a full archive narrows down to one specific, exact moment.
While it’s a powerful tool, it depends on previously created metadata. Semantic search powered by TwelveLabs’ Marengo, a multimodal media embedding model, works differently. Type “stairs,” “people running,” “voting,” “people eating,” or “helicopters” into the same search bar, and now results surface without ever relying on a single human-created tag. Marengo reads intent as much as keywords, routing a query toward visual, transcript, or audio signals depending on what was asked.

The two approaches can be combined, too. A query like “goals scored left to right, player in a white shirt, against a team in yellow, in the rain” is fully within reach of TwelveLabs’ semantic search on its own. Teams working with especially large archives can use Mimir's filters upfront to narrow the field before running semantic search, shortening time to results even further.
It’s the difference between hoping the right tag exists and knowing the content has been indexed at all — a way to shine a light into archives that have sat dark for decades.
Automated Compliance Logging: Flagging Sensitive Footage at the Scene Level
Manually screening footage for violence, nudity, or graphic material doesn’t scale against a growing archive, but regulators, such as Ofcom in the UK, enforce strict penalties if content that could cause harm or offense to viewers is broadcast. Finding the moments that need a second look shouldn't require watching every asset start to finish.
Inside Mimir, TwelveLabs’ Pegasus, a video language model built for granular segmentation, scans content and breaks it into scene-level segments, flagging the ones that match a defined set of compliance categories: violence, gore, criminal activity, drug and alcohol references. Flags come with a timestamp, severity level, and stated reason for the flag, jumping reviewers straight to the relevant frame.

The same segmentation approach isn’t limited to compliance, either. Pegasus can be pointed at entirely different questions, like finding the recap at the top of a show or the end credits at the close of one, and those segmentation passes can be layered on top of each other.
Metadata Enrichment at Archive Scale
Run against a single clip, Pegasus generates both an asset-level overview summary and a detailed, scene-by-scene breakdown: people, objects, and taxonomy-aligned metadata, pulled out in a single pass, with support for up to two hours of video in one context window. Once that metadata lands in Mimir, it’s immediately searchable through standard lexical search, no semantic toggle required.
When done well, logging historical footage also means understanding the context of the moment it was shot. A clip from 50 years ago means something different to someone who lived through it than to someone encountering it cold. Automated metadata generation doesn't replace that judgment — it merely removes the bottleneck of needing a human to sit through hours of tape before any of it becomes searchable. That metadata belongs to the customer and lives in Mimir's own database; TwelveLabs doesn't persist what it generates.
On average, it takes three to four minutes to index and generate metadata from an hour of content once its proxy is online in Mimir. For live news workflows specifically, editors don’t have to wait for a recording to finish at all: Mimir generates proxy segments as the content is captured, so teams can start reviewing footage and building clips within seconds of the recording’s start time. Actual processing time for a finished, rendered clip depends on the source format, as formats that support a simple rewrap complete faster than ones that require a full transcode.
From Natural Language Query to Finished Cut
A request for “soccer goals” returns ranked results pulled from across the archive. Narrowing further to a specific player works the same way TwelveLabs recognizes people generally: uploading an image of the player lets TwelveLabs match them as a recognized entity across the archive, rather than relying on a name or tag someone entered in advance. Asking for goal celebrations featuring that player then returns matches ranked by relevance, with the qualifying segment highlighted directly on each clip's timeline so there's no guessing which part of a longer clip actually matched the search.

From there, a handful of selected clips can be summarized in natural language to generate a usable clip description. Selected clips can also be sent directly into Mimir Cutter, where they land as an assembled timeline rather than a folder of separate files.
In a newsroom, that timing matters more than in almost any other context. Sports typically run at the end of a broadcast, partly so editors have as much time to cut a highlight reel from a game that sometimes hasn’t even finished when the broadcast goes to air. Being able to ask for goal celebrations by name or by image and get an assembled, trimmable timeline back in moments reduces the editorial work to deciding what goes in the voiceover.
How the Mimir and TwelveLabs Integration Works
The integration runs on an orchestration layer sitting between the two platforms. When new content lands in Mimir, an event fires, an orchestration layer (built on tools like EventBridge, Step Functions, or Lambda) triggers TwelveLabs to index it, and results flow back into Mimir automatically. Mimir Cutter workflows operate the same way, calling TwelveLabs for clip curation and metadata, then assembling results through Cutter's own API.
Indexing runs on proxy media, not full-resolution files, and is storage-agnostic. Mimir renders media in place, whether it lives in Amazon S3, GCP, on-prem, or a local NAS. TwelveLabs pulls a lightweight proxy, generates embeddings, and discards the proxy once it's been processed.
Indexing can also run fully on demand rather than automatically on ingestion, the same way a broadcaster might choose not to transcribe every asset that comes through the door. And on identifying people, the two platforms work together rather than duplicating effort: TwelveLabs maintains an entity library covering people, objects, brands, and logos, with active work underway to connect that library directly to the customer’s own person-recognition data, enabled by Mimir’s face-recognition, based on AWS.
The Bigger Picture
Video has always been one of the richest sources of information an organization can have, and one of the hardest to actually use. A camera captures every face, gesture, object, and word, but almost none of that richness has historically been searchable unless someone sat down and wrote it out by hand. That's the gap this integration closes.
Metadata enrichment that used to take an archivist hours now runs automatically. Compliance review that used to require watching full assets end to end now surfaces flagged moments directly. Clip curation that used to mean a request ticket to the archive team now happens in a search bar.
Nobody has to predict which tags will be necessary in five years or which frame someone might search for a decade from now. The archive gets watched, understood, and made findable through a specific gesture or phrase whether or not a human ever thought to look for it. For anyone sitting on decades of legacy physical media, that's the shift: when content stops being something you have to already know about to use, and starts being something you can simply ask for.
If you’d like to go deeper on how to use TwelveLabs inside your MAM, check out our whitepaper: Bringing Semantic Video Intelligence to Enterprise Workflows.
If you’re ready to get started today, our sales team would love to discuss what an integration could look like for your specific workflow, storage setup, and content volume.
Somewhere in a broadcaster's warehouse sit rooms full of tape: DigiBeta, DVCam, VHS, film reels, LTO tapes, three-quarter inch. The metadata lives in handwritten notebooks on the shelves, and the only person who can say what's actually on any given tape is the person who wrote those notes by hand. That archive is real, and it's the exact problem this integration is built to solve: hundreds of thousands of hours of content, sitting unused, because nobody can find anything in it.
Getting that content off legacy physical media and into a system is its own project. But even once it's digitized, the underlying problem just changes shape. The bottleneck was never the format, but the metadata: whether the system actually knows what's happening inside the content, not just where the file lives.
Mimir and TwelveLabs solve two different halves of that problem. Mimir is a serverless media production and collaboration platform — the system of record where content lives, gets organized, and moves through production. TwelveLabs is the intelligence layer that understands what's inside that content: every scene, every word spoken, every object and action, indexed and made searchable.
Mimir gives organizations a place to manage media at scale. TwelveLabs gives that media the kind of understanding that used to require a person watching every frame. Together, they turn a media archive from a place where content is simply stored into one where it can be found, understood, and repurposed.
Everything described in this blog was recently demoed live in a webinar with Mimir and TwelveLabs. If you’d rather watch the demo than read about the workflows, the on-demand recording is available now.
Semantic Search in Mimir: Finding Footage That was Never Tagged
Search for “Macron” in Mimir, and it will pull results from clip titles, transcripts, and face recognition. Add on qualifiers like “wearing a tie” and “says the word monopoly,” stack those into an overlap search, and a full archive narrows down to one specific, exact moment.
While it’s a powerful tool, it depends on previously created metadata. Semantic search powered by TwelveLabs’ Marengo, a multimodal media embedding model, works differently. Type “stairs,” “people running,” “voting,” “people eating,” or “helicopters” into the same search bar, and now results surface without ever relying on a single human-created tag. Marengo reads intent as much as keywords, routing a query toward visual, transcript, or audio signals depending on what was asked.

The two approaches can be combined, too. A query like “goals scored left to right, player in a white shirt, against a team in yellow, in the rain” is fully within reach of TwelveLabs’ semantic search on its own. Teams working with especially large archives can use Mimir's filters upfront to narrow the field before running semantic search, shortening time to results even further.
It’s the difference between hoping the right tag exists and knowing the content has been indexed at all — a way to shine a light into archives that have sat dark for decades.
Automated Compliance Logging: Flagging Sensitive Footage at the Scene Level
Manually screening footage for violence, nudity, or graphic material doesn’t scale against a growing archive, but regulators, such as Ofcom in the UK, enforce strict penalties if content that could cause harm or offense to viewers is broadcast. Finding the moments that need a second look shouldn't require watching every asset start to finish.
Inside Mimir, TwelveLabs’ Pegasus, a video language model built for granular segmentation, scans content and breaks it into scene-level segments, flagging the ones that match a defined set of compliance categories: violence, gore, criminal activity, drug and alcohol references. Flags come with a timestamp, severity level, and stated reason for the flag, jumping reviewers straight to the relevant frame.

The same segmentation approach isn’t limited to compliance, either. Pegasus can be pointed at entirely different questions, like finding the recap at the top of a show or the end credits at the close of one, and those segmentation passes can be layered on top of each other.
Metadata Enrichment at Archive Scale
Run against a single clip, Pegasus generates both an asset-level overview summary and a detailed, scene-by-scene breakdown: people, objects, and taxonomy-aligned metadata, pulled out in a single pass, with support for up to two hours of video in one context window. Once that metadata lands in Mimir, it’s immediately searchable through standard lexical search, no semantic toggle required.
When done well, logging historical footage also means understanding the context of the moment it was shot. A clip from 50 years ago means something different to someone who lived through it than to someone encountering it cold. Automated metadata generation doesn't replace that judgment — it merely removes the bottleneck of needing a human to sit through hours of tape before any of it becomes searchable. That metadata belongs to the customer and lives in Mimir's own database; TwelveLabs doesn't persist what it generates.
On average, it takes three to four minutes to index and generate metadata from an hour of content once its proxy is online in Mimir. For live news workflows specifically, editors don’t have to wait for a recording to finish at all: Mimir generates proxy segments as the content is captured, so teams can start reviewing footage and building clips within seconds of the recording’s start time. Actual processing time for a finished, rendered clip depends on the source format, as formats that support a simple rewrap complete faster than ones that require a full transcode.
From Natural Language Query to Finished Cut
A request for “soccer goals” returns ranked results pulled from across the archive. Narrowing further to a specific player works the same way TwelveLabs recognizes people generally: uploading an image of the player lets TwelveLabs match them as a recognized entity across the archive, rather than relying on a name or tag someone entered in advance. Asking for goal celebrations featuring that player then returns matches ranked by relevance, with the qualifying segment highlighted directly on each clip's timeline so there's no guessing which part of a longer clip actually matched the search.

From there, a handful of selected clips can be summarized in natural language to generate a usable clip description. Selected clips can also be sent directly into Mimir Cutter, where they land as an assembled timeline rather than a folder of separate files.
In a newsroom, that timing matters more than in almost any other context. Sports typically run at the end of a broadcast, partly so editors have as much time to cut a highlight reel from a game that sometimes hasn’t even finished when the broadcast goes to air. Being able to ask for goal celebrations by name or by image and get an assembled, trimmable timeline back in moments reduces the editorial work to deciding what goes in the voiceover.
How the Mimir and TwelveLabs Integration Works
The integration runs on an orchestration layer sitting between the two platforms. When new content lands in Mimir, an event fires, an orchestration layer (built on tools like EventBridge, Step Functions, or Lambda) triggers TwelveLabs to index it, and results flow back into Mimir automatically. Mimir Cutter workflows operate the same way, calling TwelveLabs for clip curation and metadata, then assembling results through Cutter's own API.
Indexing runs on proxy media, not full-resolution files, and is storage-agnostic. Mimir renders media in place, whether it lives in Amazon S3, GCP, on-prem, or a local NAS. TwelveLabs pulls a lightweight proxy, generates embeddings, and discards the proxy once it's been processed.
Indexing can also run fully on demand rather than automatically on ingestion, the same way a broadcaster might choose not to transcribe every asset that comes through the door. And on identifying people, the two platforms work together rather than duplicating effort: TwelveLabs maintains an entity library covering people, objects, brands, and logos, with active work underway to connect that library directly to the customer’s own person-recognition data, enabled by Mimir’s face-recognition, based on AWS.
The Bigger Picture
Video has always been one of the richest sources of information an organization can have, and one of the hardest to actually use. A camera captures every face, gesture, object, and word, but almost none of that richness has historically been searchable unless someone sat down and wrote it out by hand. That's the gap this integration closes.
Metadata enrichment that used to take an archivist hours now runs automatically. Compliance review that used to require watching full assets end to end now surfaces flagged moments directly. Clip curation that used to mean a request ticket to the archive team now happens in a search bar.
Nobody has to predict which tags will be necessary in five years or which frame someone might search for a decade from now. The archive gets watched, understood, and made findable through a specific gesture or phrase whether or not a human ever thought to look for it. For anyone sitting on decades of legacy physical media, that's the shift: when content stops being something you have to already know about to use, and starts being something you can simply ask for.
If you’d like to go deeper on how to use TwelveLabs inside your MAM, check out our whitepaper: Bringing Semantic Video Intelligence to Enterprise Workflows.
If you’re ready to get started today, our sales team would love to discuss what an integration could look like for your specific workflow, storage setup, and content volume.
Platform
Enterprise
©
2026
TwelveLabs, Inc. All Rights Reserved
Platform
Enterprise
©
2026
TwelveLabs, Inc. All Rights Reserved
Platform
Enterprise
©
2026
TwelveLabs, Inc. All Rights Reserved






