Building with AI

Deploying image and video generation with vLLM-Omni on SageMaker AI

A technical guide details how to deploy FLUX.2 and Wan2.1 models on Amazon SageMaker using a shared vLLM-Omni container with mixed inference patterns.

Amazon Web Services released a technical tutorial on September 28, 2026, demonstrating how to generate images and animate them into video using Amazon SageMaker AI. The guide details the deployment of two specific generative models, FLUX.2-klein-4B and Wan2.1-VACE-1.3B, utilizing the AWS vLLM-Omni Deep Learning Container to handle both real-time and asynchronous workloads within a unified infrastructure.

What happened

The publication outlines a workflow that transforms a text prompt into a still image and subsequently animates that image into a short video clip. This process relies on two distinct endpoints deployed from the same pinned version of the AWS vLLM-Omni Deep Learning Container (DLC). The first endpoint handles image generation using the FLUX.2-klein-4B model, while the second manages video generation using the Wan2.1-VACE-1.3B model. By using a single container image for both tasks, the architecture reduces variation in the serving stack while allowing each model to operate on instance types optimized for its specific computational demands.

The tutorial distinguishes between the inference patterns required for each task. Image generation is handled via a SageMaker AI real-time endpoint, which returns the generated PNG directly to the client. In contrast, video generation uses a SageMaker Asynchronous Inference endpoint. This approach is necessary because video creation is a longer-running operation that involves larger payloads. The asynchronous endpoint writes the final MP4 file to Amazon Simple Storage Service (Amazon S3), allowing the client to poll for the result rather than keeping a connection open.

This release serves as the second part of a series exploring specialized AWS DLCs. While the previous installment focused on real-time speech processing using bidirectional streaming, this guide isolates image and video generation to clarify the differences in payloads and response mechanisms. The provided solution includes both a command-line interface workflow and a Streamlit application, enabling developers to test the end-to-end pipeline from text prompt to animated video.

How it works

The core technology enabling this workflow is the AWS vLLM-Omni DLC, which extends the vLLM framework beyond text generation to support audio, images, and video through OpenAI-compatible APIs. The container includes routing middleware that reads the CustomAttributes header in incoming requests and forwards traffic to the appropriate vLLM-Omni route. For the image model, requests are routed to /v1/images/generations, while video requests go to /v1/videos/sync. This abstraction allows developers to interact with complex multimodal models using familiar API structures.

The data flow begins when the application sends a text prompt to the FLUX.2-klein endpoint. The model returns a base64-encoded PNG, which the application then resizes to match the target video dimensions and converts into a compact JPEG data URL. This reference image is embedded into a multipart request for the Wan VACE video model. Because the payload exceeds the 128,000-byte limit for inline asynchronous bodies, the application uploads the request to Amazon S3 and provides the location to the SageMaker endpoint. The endpoint processes the job, writes the resulting MP4 to a designated S3 output location, and the client retrieves the file once generation is complete.

Key details

  • The solution uses the omni-sagemaker-cuda-v1.6 pinned DLC image for both endpoints.
  • Image generation runs on an ml.g6.xlarge instance type via a real-time endpoint.
  • Video generation runs on an ml.g6e.xlarge instance type via an asynchronous endpoint.
  • The Wan VACE model uses variational autoencoder (VAE) tiling to reduce peak memory during decoding.
  • Default video settings produce 17 frames using 30 diffusion steps for quality preservation.
  • Successful responses and invocation failures are stored under separate Amazon S3 prefixes.

Why it matters

For engineering teams building generative media applications, this pattern offers a clear blueprint for managing mixed-latency workloads. Using real-time inference for fast tasks like image generation ensures immediate feedback, while asynchronous inference prevents timeouts and resource contention for slower video rendering. This separation allows systems to scale more efficiently, as each endpoint can be tuned independently based on its specific latency and throughput requirements.

Additionally, the use of a shared container image simplifies operational overhead. Maintaining a single DLC version across multiple models reduces the complexity of dependency management and security patching. Developers can swap or update models by changing the SM_VLLM_MODEL environment variable without rebuilding the underlying serving infrastructure. This modularity supports rapid experimentation with different model families while keeping the production environment stable.

What you can do

  • Clone the sagemaker-genai-hosting-examples repository from GitHub to access the sample code.
  • Ensure your AWS account has quota for ml.g6.xlarge and ml.g6e.xlarge instances in your target region.
  • Configure a SageMaker execution role with permissions to access Amazon S3 and create endpoints.
  • Run the provided deploy.py script to provision the real-time and asynchronous endpoints.
  • Use the generate.py command-line tool to test the pipeline with custom text and motion prompts.
  • Launch the included Streamlit application to visualize the image-to-video workflow in a browser interface.

Tools from the Bytechap store

$89

DocBento

Self-hosted document management that reads every scan and answers with page citations.

Live demo

Keep reading

All stories