AI agents

Implementing persistent agent memory with NVIDIA NeMo and AWS S3 Vectors

A technical guide details how to build scalable, consistent long-term memory for AI agents using NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors on EKS.

Engineers building multi-agent systems now have a concrete path to implement persistent, shared memory using the NVIDIA NeMo Agent Toolkit (NAT) backed by Amazon S3 Vectors. Published in October 2026, this technical guide demonstrates how to deploy these components on Amazon Elastic Kubernetes Service (EKS) to solve consistency and scaling challenges in production environments.

What happened

The article provides a step-by-step implementation strategy for creating a custom memory provider within NAT. While NAT supports existing backends like Redis and Zep, the guide argues that Amazon S3 Vectors is better suited for production multi-agent systems requiring elastic scale and strong write consistency. The authors walk through creating the necessary infrastructure, writing a custom plugin, and configuring the agent workflow.

The core of the implementation involves defining a MemoryEditor interface, which serves as the plugin contract for custom backends. Developers must implement three specific methods: add_items(), search(), and remove_items(). This plugin allows agents to store conversation history, user preferences, and long-term knowledge as MemoryItem objects, which include metadata fields for filtering and scoping queries.

Deployment relies on Amazon EKS for operational control over the agent lifecycle. The architecture uses the HorizontalPodAutoscaler to scale agent replicas based on CPU utilization. Each replica connects to the same S3 Vectors index using IAM Roles for Service Accounts (IRSA). This setup ensures that any pod handling a request can access the most recent memories without requiring cache invalidation strategies.

How it works

The system operates by wrapping standard agents with an auto_memory_agent component. This wrapper automatically handles the capture and retrieval of context, removing the need for the large language model to explicitly invoke memory tools during every turn. When an agent generates a response, the system embeds the content using a model like Amazon Titan Text Embeddings V2 and stores it in the S3 Vectors index.

S3 Vectors acts as the durable backend, supporting up to two billion vectors per index. It provides semantic retrieval using configurable distance metrics such as cosine or Euclidean similarity. Crucially, it offers strong write consistency, meaning that when one agent pod writes a memory, all other pods see it immediately. This eliminates the race conditions common in eventually consistent databases when multiple agents coordinate on a single task.

Metadata filtering allows agents to scope their searches effectively. Each vector can carry string, number, Boolean, or list metadata, enabling queries that target specific users, sessions, or topics. This structure supports different types of memory, including episodic, semantic, and procedural, allowing the system to retrieve only the most relevant context for a given interaction.

Key details

  • Plugin Interface: Custom memory backends must implement the MemoryEditor abstract interface with add_items, search, and remove_items methods.
  • Consistency Model: Amazon S3 Vectors provides strong write consistency, ensuring immediate visibility of new memories across all agent pods without cache management.
  • Scaling Mechanism: Agent replicas on Amazon EKS scale via HorizontalPodAutoscaler based on CPU usage, with all instances sharing a single S3 Vectors index.
  • Evaluation Metrics: The NAT evaluation harness (nat eval) measures improvements in groundedness, token usage, latency, and duplicate work reduction.
  • Infrastructure Limits: S3 Vectors supports up to two billion vectors per index, with no capacity planning required and pay-per-use pricing for storage and queries.
  • Access Control: Security is managed through AWS IAM policies per bucket and index, allowing for hard isolation between tenants or teams.

Why it matters

For software teams building complex agent workflows, memory management is often a bottleneck. Without a shared, consistent memory layer, agents repeat work, lose context between sessions, and struggle to coordinate. By offloading memory to S3 Vectors, developers gain a system that scales elastically without managing idle compute resources. This reduces operational overhead while improving the reliability of multi-agent collaborations.

The ability to quantify memory impact is also significant. The guide highlights that enabling memory typically improves groundedness by providing verifiable source context and reduces token usage by skipping re-derivation of known facts. While latency increases slightly due to vector queries, the trade-off often results in higher quality outputs and lower overall costs for repetitive tasks. This data-driven approach helps engineers tune parameters like top_k and similarity thresholds to balance cost and performance.

What you can do

  • Install NVIDIA NeMo Agent Toolkit version 1.6 or later and ensure your environment runs Python 3.11 or 3.12.
  • Create an Amazon S3 Vectors bucket and index with a metadata schema matching your agent’s memory requirements, using 1024 dimensions for Titan embeddings.
  • Implement the MemoryEditor interface in Python to connect NAT to your S3 Vectors index, handling embedding generation and metadata tagging.
  • Configure the auto_memory_agent wrapper in your NAT YAML configuration to enable automatic memory capture and retrieval for your agents.
  • Run comparative evaluations using nat eval with and without memory enabled to measure changes in accuracy, groundedness, and token consumption.
  • Clean up resources by deleting the EKS namespace, S3 Vectors index, and IAM roles when testing is complete to avoid ongoing charges.

Tools from the Bytechap store

Keep reading

All stories