Rashomon verifies Claude Code agent actions to detect hidden failures
The open-source tool Rashomon independently records every command and subagent action in Claude Code, flagging discrepancies when the agent claims success despite underlying errors.
A new open-source utility called Rashomon provides an independent verification layer for Claude Code, the AI coding assistant. Released in alpha version 1.1.0 on October 5, 2026, this tool runs on macOS and Linux to track exactly what an AI agent does during a coding session. It addresses the growing concern among developers that AI agents may confidently report success even when background tasks or subagents have failed.
What happened
Rashomon acts as a silent observer during Claude Code sessions, maintaining a separate log of every tool call, file edit, and command execution. Unlike the standard interaction where developers rely solely on the agent’s final summary, Rashomon compares its own recorded data against the agent’s closing statement. If the agent claims all tests pass but Rashomon detects a failed command or error code, it flags the discrepancy immediately.
The tool is designed to remain unobtrusive during normal operation. It only interrupts the workflow when it detects a potential mismatch between the agent’s narrative and the actual system events. This includes tracking actions performed by subagents, which are helper instances spawned by the main AI to handle specific subtasks. These subagent activities are often invisible in the main conversation transcript, creating blind spots for developers who assume the primary agent has full visibility and control.
Developers can install Rashomon either as a native Claude Code plugin or as a standalone command-line utility. The project emphasizes privacy and local processing, ensuring that no code, prompts, or file contents are stored or transmitted externally. It functions purely as a local auditor, providing a second opinion on the agent’s performance without altering the agent’s behavior or blocking its actions.
How it works
Rashomon operates by hooking into Claude Code’s execution pipeline. When installed, it registers listeners for tool calls and command executions. For every action the agent takes, such as running a test suite or editing a file, Rashomon records the outcome, including exit codes and any errors generated. It specifically monitors for a list of forty-three failure-related keywords, such as "error," "failed," or "unable," in the agent’s final summary.
If a command fails but the agent’s summary does not contain any of these failure words, Rashomon generates a alert. This mechanism helps catch hallucinations where the AI might misinterpret a failed test run as a success or simply overlook a subagent’s error. The tool also tracks subagent activities separately, listing their declarations and executions even if they do not appear in the main chat history.
The reporting system is deterministic and transparent. It does not guess intent or analyze the semantic meaning of the code. Instead, it presents a factual comparison: the agent’s quoted summary versus the recorded technical outcomes. This allows developers to quickly identify whether a discrepancy exists and investigate the specific commands or subagents involved without sifting through extensive logs manually.
Key details
- Platform support: Currently supports macOS and Linux only; Windows is not yet supported in this alpha release.
- Installation methods: Available as a Claude Code plugin via the marketplace or as a standalone binary installed via curl or Go.
- Privacy model: Stores no prompts, responses, or file contents; only records command names, argument counts, and hostnames.
- Subagent visibility: Explicitly tracks and reports actions taken by subagents, which are typically hidden from the main conversation transcript.
- Failure detection: Uses a fixed list of forty-three failure keywords to check if the agent’s summary acknowledges any recorded errors.
- Network limitations: Does not observe live network traffic in this alpha version, though it can integrate with external sandbox audits like nono.
Why it matters
For software engineers relying on AI coding assistants, trust is a critical but fragile component of the workflow. As agents become more autonomous, spawning subagents and executing complex chains of commands, the risk of silent failures increases. An agent might successfully refactor a function but fail to run the associated test suite correctly, leading to broken builds that are difficult to trace back to the AI’s actions. Rashomon provides a necessary safety net by ensuring that the agent’s confidence matches reality.
This tool also highlights the opacity of multi-agent systems. When an AI delegates tasks to subagents, the primary interface often hides the details of those delegated actions. Without independent verification, developers may miss critical errors occurring in these hidden layers. By exposing subagent activity and cross-referencing it with the main agent’s summary, Rashomon helps maintain accountability and transparency in AI-assisted development.
Furthermore, the emphasis on local processing and zero data retention addresses significant security concerns. Many organizations are hesitant to adopt AI tools due to fears of code leakage or proprietary data exposure. Rashomon’s design, which avoids storing sensitive content and operates entirely offline, offers a model for how verification tools can be built without compromising security or privacy.
What you can do
- Install Rashomon using the Claude Code plugin marketplace commands or the standalone curl installer for macOS or Linux.
- Run a test session by intentionally breaking a test in your project and asking Claude Code to fix it, then check the Rashomon report for discrepancies.
- Use the
/rashomon:reportcommand within Claude Code orrashomon reportin the terminal to view detailed logs of agent and subagent actions. - Review the list of forty-three failure keywords to understand how Rashomon determines if an agent’s summary is misleading.
- Ensure your installation path is stable and permanent to avoid breaking the hook integration, as moving the binary can cause Claude Code to fail on tool calls.
- Monitor the project’s GitHub repository for updates on Windows support and enhanced network observation features in future releases.



