Building with AI

Normalizing trajectory models enable exact likelihood in few-step diffusion

Apple researchers introduce Normalizing Trajectory Models to achieve high-quality image generation in four steps while retaining exact likelihood training.

Illustration of compressed diffusion steps represented by glowing circuits in a glass cube next to a stopwatch
Illustration generated for this article

Researchers from Apple, the University of Pennsylvania, and UIUC have introduced Normalizing Trajectory Models (NTM), a new architecture for generative AI that accelerates image synthesis. Published on October 8, 2026, this work addresses the efficiency bottleneck in diffusion models by enabling high-fidelity output in just four sampling steps. Unlike previous fast-generation methods, NTM maintains an exact likelihood framework throughout the process.

What happened

Standard diffusion models generate images by reversing a noise-adding process through many small Gaussian denoising steps. This approach is computationally expensive because it requires hundreds of sequential passes to produce a coherent result. When engineers attempt to speed up inference by compressing these steps into fewer, coarser transitions, the underlying mathematical assumptions often break down, leading to degraded image quality or artifacts.

To solve this, the research team developed NTM, which treats each reverse step as an expressive conditional normalizing flow. This allows the model to be trained with exact likelihood, preserving the probabilistic rigor that is typically lost in accelerated methods. The architecture combines shallow invertible blocks within individual steps with a deep parallel predictor that operates across the entire trajectory. This design creates an end-to-end network that can be trained from scratch or initialized from existing pretrained flow-matching models.

The resulting system achieves performance comparable to strong baselines on text-to-image benchmarks but requires only four sampling steps. A key feature of this approach is self-distillation, where a lightweight denoiser is trained on the score function induced by the model itself. This mechanism allows NTM to produce high-quality samples rapidly without sacrificing the theoretical benefits of likelihood-based training.

How it works

Normalizing flows are a class of generative models that transform a simple probability distribution into a complex one through a series of invertible functions. In NTM, each step of the generation trajectory is modeled as such a flow. Because these transformations are invertible, the model can compute the exact likelihood of the data at every stage. This stands in contrast to standard diffusion distillation or adversarial training, which often approximate the distribution and lose the ability to calculate precise probabilities.

The architecture uses a dual structure to manage complexity and speed. Shallow invertible blocks handle the local transformations within each step, ensuring that the mathematical properties of the flow are maintained. Simultaneously, a deep parallel predictor analyzes the entire trajectory, allowing the model to understand the global context of the generation process. This parallel processing capability is crucial for reducing the number of required steps from hundreds to just four.

Self-distillation further enhances efficiency. Instead of relying on external teachers or complex adversarial losses, NTM uses its own induced score function to train a lighter denoiser. This internal feedback loop refines the generation process, ensuring that the few-step output remains consistent with the high-quality distributions learned during training. The result is a system that balances speed, quality, and theoretical soundness.

Key details

  • NTM reduces image generation to four sampling steps while matching or outperforming strong baselines.
  • The model retains exact likelihood over the generative trajectory, unlike most few-step distillation methods.
  • Architecture combines shallow invertible blocks per step with a deep parallel predictor across the trajectory.
  • Supports training from scratch or initialization from pretrained flow-matching models.
  • Uses self-distillation via a lightweight denoiser trained on the model’s own induced score function.
  • Research conducted by teams from Apple, University of Pennsylvania, and UIUC, published October 8, 2026.

Why it matters

For engineers building generative applications, inference cost is a primary constraint. Reducing the number of sampling steps from hundreds to four drastically lowers latency and computational expense, making real-time generation more feasible. However, previous attempts to accelerate diffusion often compromised quality or removed the ability to measure likelihood, which is useful for tasks like anomaly detection and model evaluation. NTM offers a path to speed without losing these analytical capabilities.

The retention of exact likelihood also opens new possibilities for model optimization and debugging. Developers can precisely quantify how well the model fits the data, which is difficult with adversarial or approximate methods. This transparency can lead to more robust systems, particularly in domains where reliability and interpretability are critical. By combining the speed of distilled models with the rigor of normalizing flows, NTM represents a significant step toward more efficient and trustworthy generative AI.

What you can do

  • Review the NTM architecture to understand how invertible blocks can be integrated into existing diffusion pipelines.
  • Experiment with self-distillation techniques using your own model’s score function to reduce inference steps.
  • Evaluate whether exact likelihood calculation is beneficial for your specific use case, such as outlier detection.
  • Consider initializing new models from pretrained flow-matching checkpoints to leverage existing knowledge.
  • Benchmark four-step generation against your current baseline to assess potential latency improvements.
  • Explore parallel predictors in your trajectory modeling to improve global context awareness during generation.

Tools from the Bytechap store

$89

DocBento

Self-hosted document management that reads every scan and answers with page citations.

Live demo

Keep reading

All stories