Deploying generative recommenders with NVIDIA Dynamo-Triton and HSTU
NVIDIA demonstrates an end-to-end workflow for serving Hierarchical Sequential Transduction Unit models using PyTorch AOTI, FlexKV caching, and Dynamo-Triton to reduce inference latency.