Portia: An open-source agent harness for verifiable data engineering
Portia is a new open-source tool that profiles data, generates SQL, and validates results without the AI model ever reading raw records.
Bài viết này chỉ có sẵn bằng tiếng Anh.
Released on October 7, 2026, Portia is an open-source agent harness designed to assist with real-world data engineering tasks. Created by Jad1908 and hosted on GitHub, this tool connects to major data warehouses and local file systems to profile tables, answer questions, and generate SQL pipelines. Unlike many AI assistants that process raw data directly, Portia ensures the underlying language model never accesses the actual dataset, relying instead on deterministic code for all numerical claims.
What happened
The project introduces a copilot for data work that integrates with Snowflake, BigQuery, PostgreSQL, or local folders containing CSV and Parquet files. Once connected, Portia profiles every table it encounters, measuring null rates, distinct counts, key coverage, and fan-out before any join operations are executed. It allows users to ask natural language questions about their data structure and content, such as checking if a specific identifier is unique or requesting revenue breakdowns by month. The system then builds the requested tables one decision at a time, recording each step in a YAML specification.
A critical architectural choice in Portia is the separation of reasoning from data access. The language model used for interaction has no filesystem access and no shell capabilities. Every number or statistic presented to the user is computed by deterministic code rather than generated by the model’s probabilistic next-token prediction. This design ensures that the tool can handle tables too large to fit into memory while keeping every claim verifiable. The final output is compiled into dbt-shaped SQL files that can run independently of Portia, ensuring the generated pipeline is not locked into the tool itself.
The system also includes a knowledge graph powered by Neo4j, which stores column lineage and measured overlaps between datasets. This graph helps the copilot determine which tables are relevant to a user’s query. Users can browse this graph locally to understand how columns relate across the pipeline. Additionally, Portia supports visual charts within the conversation interface. If the underlying model supports image input, the copilot can visually inspect a chart rendered in the user’s browser to verify its accuracy before reporting findings, though the image itself is never saved or used for numerical decisions.
How it works
Portia operates by indexing data sources and building a catalog of metadata. When a user asks a question, the copilot consults the knowledge graph to identify relevant tables and columns. It then formulates queries that are executed against the data source using the user’s own credentials. For database connections, queries run under the user’s role, meaning nothing is pulled down to the application server, and any new tables are created directly in the warehouse. For local files, the data remains on disk and is not copied. The model only sees the schema, summary statistics, and the results of these deterministic queries.
The tool runs on the Claude Agent SDK and can drive Claude Code unmodified. It requires Python 3.11 or higher, uv, and Docker. Installation involves cloning the repository, syncing dependencies, and starting a Neo4j container. Users must provide their own Anthropic API key or configure a local model provider like Ollama or llama.cpp. Because the instructions for the agent are approximately 15,000 tokens long, local models must support a context window of at least 32K tokens. The system checks this requirement before sending prompts and refuses to proceed if the context is insufficient.
Key details
- Data Privacy: The language model never reads raw data, has no filesystem access, and possesses no shell capabilities.
- Supported Sources: Connects to Snowflake, BigQuery, PostgreSQL, and local folders with CSV or Parquet files.
- Verification: Every numerical claim is computed by deterministic code, allowing users to check results independently.
- Output Format: Generates dbt-shaped SQL files stored in a
models/directory, which can run without Portia. - Knowledge Graph: Uses Neo4j to store column lineage and measure overlaps between columns across the pipeline.
- Local Model Support: Compatible with Ollama and llama.cpp, requiring a minimum 32K context window for local inference.
Why it matters
For software engineers and data teams, the primary value of Portia lies in its verifiability. Traditional AI-assisted coding tools often hallucinate syntax or logic, requiring significant manual review. In data engineering, incorrect SQL can lead to silent data corruption or misleading analytics. By forcing the model to rely on deterministic code for all measurements and preventing it from seeing raw data, Portia reduces the risk of hallucinated statistics. The "gate on zeros" feature further protects against common errors by refusing to write results if a query returns empty sets, non-unique grains, or entirely null columns.
The ability to compile specs into standard SQL also mitigates vendor lock-in. Teams can use Portia to accelerate the initial development of data pipelines but deploy the resulting SQL files in their existing production environments. This means the tool serves as a development accelerator rather than a runtime dependency. The integration with CI pipelines via the build --check command ensures that if a SQL file drifts from its original specification, the build fails, maintaining consistency between intent and implementation.
What you can do
- Install Portia using uv and Docker, ensuring you have Python 3.11+ available.
- Connect the tool to a test database or a folder of sample CSV files to explore its profiling capabilities.
- Use the
devtools.demodatamodule to generate a five-table demo project with known issues for testing. - Configure a local model like Qwen 3 8B via Ollama if you prefer not to use cloud-based APIs.
- Review the generated YAML specs and compiled SQL files to understand how the agent structures its decisions.
- Browse the Neo4j knowledge graph at localhost:7474 to visualize column lineage and overlaps.



