Data & documents

Infino unifies agent data retrieval on Parquet to cut stack complexity

OpenSearch veterans launch Infino, an open source retrieval layer that lets AI agents query a single Parquet copy of data, replacing fragmented search and vector stacks.

A unified glass cube representing single-source data retrieval replacing fragmented database boxes.
Illustration generated for this article

A team of former OpenSearch engineers has launched Infino, an open source retrieval platform designed specifically for AI agents. Released in October 2026, the system allows agents to query structured and unstructured data from a single source of truth stored in Apache Parquet. This approach aims to replace the complex, multi-system data stacks currently required to support agentic workflows.

What happened

Infino emerged from stealth with a platform that addresses the fragmentation typical in modern AI data infrastructure. CEO Ekechi Nwokah argues that while agents have become the largest new consumer of data since web browsers, they are still served by infrastructure built for human-driven queries. Traditional setups require separate systems for SQL analytics, keyword search, and semantic vector search, each with its own ingestion pipeline and synchronization logic.

The company proposes collapsing this stack into a single retrieval layer. Instead of maintaining multiple copies of data across different databases, Infino keeps one copy in Apache Parquet format on object storage. Agents can then search, rank, filter, join, aggregate, and reason over this data directly. The core engine is released under the Apache-2.0 license, ensuring that the data remains accessible via standard Parquet readers even without the Infino software.

Nwokah emphasizes that this architecture removes the need for extensive glue code and multiple ETL jobs. By embedding retrieval functions directly into SQL, developers can express complex questions involving both keyword and semantic search in a single query. This reduces the number of tool calls an agent must make, lowering latency and token costs associated with context consumption.

How it works

Infino stores data as valid Parquet files with search indexes embedded beside the Parquet footer. This design allows any tool that reads Parquet, such as Spark or DuckDB, to access the raw data, while Infino provides the accelerated retrieval capabilities. The system supports both structured and unstructured queries, enabling agents to perform exact counts, joins, and filters alongside semantic searches without switching contexts.

To address the issue of retrieval loops, where agents repeatedly query and validate results, Infino integrates targeted inference models with the retrieval engine. These smaller, specialized models handle simpler tasks more cheaply and quickly than large frontier models. The platform allows developers to point at existing Parquet or JSON data, ingest it, and immediately expose it to agents via a unified interface.

Governance is centralized within this single data copy. Security engineers can define hard policies regarding who can read specific rows or columns, and what actions get logged. This contrasts with managing permissions across multiple Model Context Protocol (MCP) gateways and SaaS connectors, offering a more deterministic security posture for enterprise deployments.

Key details

  • Infino uses Apache Parquet on object storage as the single source of truth for all data types.
  • Search indexes are embedded within Parquet file footers, maintaining compatibility with standard Parquet readers.
  • The core engine is open source under the Apache-2.0 license, available on GitHub.
  • Retrieval functions are embedded in SQL, allowing combined keyword, semantic, and analytical queries in one statement.
  • The company claims the architecture is approximately 10.5 times cheaper than Elasticsearch and 23 times cheaper than OpenSearch for tested workloads.
  • Founders include Ekechi Nwokah, Vinay Kakade, Asif Makhani, and Murali Krishna, with backgrounds at Amazon, Google, and LinkedIn.

Why it matters

For software teams building AI agents, the current requirement to orchestrate multiple data systems creates significant operational overhead. Each additional database or search engine introduces latency, cost, and maintenance burden. By unifying these capabilities, Infino reduces the complexity of the data stack, allowing engineers to focus on agent logic rather than data synchronization. This simplification is critical as organizations scale from prototype agents to production systems handling billions of documents.

The cost implications are also substantial. Agents often issue many small queries to answer a single large question, consuming context window space and increasing API costs. A unified retrieval layer that returns comprehensive results in fewer calls can significantly reduce these expenses. Furthermore, the ability to enforce security policies centrally helps mitigate risks associated with agents accessing sensitive data across disparate systems.

What you can do

  • Evaluate if your current agent workflows involve multiple round-trips to different data sources like vector DBs and SQL warehouses.
  • Consider migrating static or semi-static datasets to Apache Parquet format to prepare for unified retrieval architectures.
  • Review your existing governance policies to see how they might be consolidated into a single access control layer.
  • Test the open source Infino engine on GitHub with a subset of your data to measure latency improvements.
  • Analyze the cost of your current retrieval stack, including ETL maintenance and infrastructure fees, against Infino’s claims.
  • Explore how embedding retrieval functions in SQL could simplify the prompt engineering required for your agents.

Tools from the Bytechap store

Keep reading

All stories