Building with AI

Condé Nast cuts video search time by 99% with multimodal AI

Condé Nast reduced video discovery time from 250 minutes to under two minutes using Amazon Bedrock and TwelveLabs Marengo for semantic search across 140,000 clips.

Condé Nast partnered with the AWS Generative AI Innovation Center to overhaul its video discovery process, reducing search time from an average of 250 minutes to under two minutes per task. The media giant deployed a multimodal semantic search system across its library of more than 140,000 videos, leveraging Amazon Bedrock and Amazon OpenSearch Service to enable intent-based retrieval.

What happened

Editorial teams at brands including Vogue, GQ, Vanity Fair, and Wired previously relied on manual scrubbing and static metadata like titles and descriptions to locate video assets. This approach created significant operational drag, as staff spent hours searching for specific visual or audio moments that keyword searches could not identify. The reliance on institutional knowledge also created single points of failure, leaving vast amounts of archive content undiscoverable when key personnel were unavailable.

To address this structural inefficiency, Condé Nast worked with AWS to build a solution that understands video content beyond text labels. The new system allows editors to search using natural language queries such as "beginner yoga content with calming backgrounds" or "behind-the-scenes fashion week moments." By analyzing visual elements, audio tracks, and transcripts simultaneously, the platform surfaces relevant clips with precise timestamps, eliminating the need to watch entire files.

The implementation resulted in a 99.2 percent reduction in content discovery time during a May 2026 benchmarking workshop. The company estimates annual operational savings of approximately $800,000 due to increased productivity and faster response times to advertiser requests. Billy Keenly, Global Senior Director of Creative Optimization at Condé Nast, noted that the collaboration helped clarify the business case for adopting these advanced workflow goals.

How it works

The architecture separates the compute-heavy ingestion pipeline from the low-latency query serving tier, allowing each to scale independently. When a video is uploaded to Amazon S3, an event triggers an ingestion service running on Amazon ECS with AWS Fargate. This service validates the file, extracts metadata, and splits the video into segments. Each segment is processed asynchronously through the TwelveLabs Marengo embedding model via Amazon Bedrock, which generates multimodal vector embeddings that capture semantic meaning across visual, audio, and transcript data.

These embeddings are stored in Amazon S3 for durability and indexed in Amazon OpenSearch Service for fast k-nearest neighbor (k-NN) similarity search. On the query side, user requests pass through an Application Load Balancer to a frontend service, which converts natural language into vector embeddings. The search service then performs a similarity search against the OpenSearch index and enriches results with metadata from Amazon DocumentDB before returning precise clip timestamps to the user.

Key details

  • Library size: The system indexes and searches over 140,000 videos.
  • Performance gain: Discovery time dropped from 250 minutes to under 2 minutes per task.
  • Core model: TwelveLabs Marengo embedding model handles joint encoding of visual, audio, and transcript signals.
  • Infrastructure: Built on Amazon Bedrock for model access and Amazon OpenSearch Service for vector indexing.
  • Resilience: Multi-AZ deployment with synchronous replication for OpenSearch and primary-standby replicas for DocumentDB.
  • Cost impact: Estimated $800,000 in annual operational savings based on May 2026 benchmarks.

Why it matters

For engineering teams building search systems, this case study highlights the limitations of keyword-based retrieval for rich media. Traditional metadata fails to capture the nuanced intent of users who describe content by mood, action, or context rather than filename. Multimodal embeddings bridge this gap by creating a unified semantic space where text, image, and audio signals align, enabling more intuitive and accurate discovery.

The architectural decision to decouple ingestion from serving is critical for scalability. Processing embeddings for hundreds of thousands of videos is resource-intensive and can block real-time queries if handled monolithically. By using asynchronous invocation and event-driven orchestration, Condé Nast ensured that backfill operations and system updates did not impact the availability or latency of the live search experience for editorial staff.

What you can do

  • Evaluate multimodal embedding models like TwelveLabs Marengo for projects involving video or audio search.
  • Decouple your ingestion and serving pipelines to prevent heavy batch processing from affecting user-facing latency.
  • Use asynchronous invocation for embedding generation to handle large backlogs without blocking main application threads.
  • Implement hybrid search strategies that combine vector similarity with metadata filtering for more precise results.
  • Conduct user research to understand natural language query patterns before designing your semantic search abstraction layer.
  • Design for high availability from the start by distributing compute and data resources across multiple Availability Zones.

Tools from the Bytechap store

Keep reading

All stories