Développer avec l’IA

JevShield uses typed decisions to block prompt injections locally

A new open-source tool uses Ollama and Nimble to detect prompt injections with structured data instead of generative text, keeping security checks local.

Cet article est disponible uniquement en anglais.

Developer Rajat Rao released JevShield on October 1, 2026, an open-source security layer designed to protect large language model agents from prompt injection attacks. The project runs entirely on local infrastructure, leveraging Ollama’s decision-model API and the Nimble model to evaluate incoming content before it reaches the primary application logic.

What happened

JevShield introduces a firewall specifically built for LLM agents that need to process untrusted external content. Traditional security approaches often ask a generative model to output a text-based judgment, such as answering yes or no to whether a prompt is malicious. This method requires the application to parse free-form text, which can be brittle and error-prone. JevShield replaces this pattern with a structured decision engine that returns typed data rather than natural language.

The system is designed to detect five specific categories of threats: direct prompt injection, indirect prompt injection, jailbreaks, data exfiltration attempts, and malicious tool manipulation. By intercepting user inputs and retrieved documents before they influence the agent, the tool aims to enforce the principle that untrusted text should be treated as data, not as executable instructions. The repository includes a benchmark suite, a Streamlit dashboard for visualization, and a FastAPI interface for integration into existing Python applications.

How it works

The core mechanism relies on Ollama’s Jev-style decision API, specifically the /v1/systemone endpoint, rather than standard chat or generation endpoints. When content arrives, JevShield sends it to the Nimble model, which evaluates multiple security dimensions simultaneously. Instead of generating a sentence, the model returns structured objects including choices, scores, and noul values. A score represents an expected index within a range, while a noul value provides a direct probability estimate between zero and one.

These structured outputs feed into a deterministic Python policy engine. The engine normalizes the risks and applies configurable thresholds to make the final decision. If the maximum normalized risk is below 0.50, the content is allowed. If it falls between 0.50 and 0.85, it is flagged for review. Any risk score at or above 0.85 triggers a block. This separation ensures that the probabilistic model provides evidence, while the deterministic code enforces the security policy. The system also tracks taint, ensuring that content identified as untrusted remains isolated from privileged operations.

Key details

  • Uses Ollama’s /v1/systemone endpoint with the Nimble model for local inference.
  • Detects direct injection, indirect injection, jailbreaks, exfiltration, and tool manipulation.
  • Returns typed decisions including choice, score, and noul instead of generative text.
  • Applies a deterministic Python policy with default thresholds of 0.50 for review and 0.85 for block.
  • Includes a fail-closed behavior where evaluator failures result in blocked requests by default.
  • Provides a Streamlit dashboard and FastAPI REST endpoints for integration and monitoring.

Why it matters

For engineers building AI agents that interact with external data sources like emails, websites, or databases, prompt injection remains a critical vulnerability. Indirect injections, where malicious instructions are hidden inside retrieved documents, are particularly difficult to mitigate because the agent may interpret the document content as part of its instruction set. JevShield addresses this by creating a security boundary that evaluates content provenance and risk before the main agent processes it. This allows developers to treat external inputs as potentially hostile data rather than trusted commands.

The shift from generative judgments to typed decisions also improves reliability. Parsing natural language responses from a security model introduces ambiguity and potential failure points. By using structured data types, the application can handle security evaluations with the same precision as other backend logic. Furthermore, running these checks locally via Ollama ensures that sensitive data does not leave the infrastructure, addressing privacy concerns associated with cloud-based security APIs. This architecture supports a defense-in-depth strategy, complementing other controls like least-privilege tool access and sandboxing.

What you can do

  • Install Ollama version 0.35.0 or higher and pull the Nimble model to prepare the local environment.
  • Clone the JevShield repository and set up the Python virtual environment using the provided requirements file.
  • Run the included benchmark suite to evaluate detection accuracy against synthetic attack examples before deployment.
  • Integrate the FastAPI /check-content endpoint into your agent’s ingestion pipeline to scan retrieved documents.
  • Configure the JEVSHIELD_BLOCK_THRESHOLD and JEVSHIELD_REVIEW_THRESHOLD environment variables to match your risk tolerance.
  • Use the Streamlit dashboard to simulate direct and indirect injection attacks and observe the typed decision outputs.

Outils de la Boutique Bytechap

$89

DocBento

Gestion documentaire auto-hébergée qui analyse chaque scan et répond avec des citations de pages.

Démo en ligne

Continuer la lecture

Tous les articles