Security & privacy

ZeroLeaks releases Shield Small for local prompt injection detection

ZeroLeaks has released Shield Small, a lightweight 118M-parameter classifier that detects prompt injections and jailbreaks locally on CPU.

ZeroLeaks has released Shield Small, a compact classifier designed to detect prompt injections and jailbreak attempts in AI applications. Published on October 2, 2026, this model runs entirely offline on a standard CPU, offering developers a low-latency way to screen untrusted text before it reaches an agent.

What happened

Shield Small is a 118 million parameter model based on the intfloat/multilingual-e5-small architecture. It uses int8 ONNX weights totaling approximately 118 MB, making it small enough to deploy alongside application code without requiring specialized GPU hardware. The model functions as a binary classifier, labeling input text as either benign or an injection attempt, and provides a confidence score for each prediction.

The release targets a critical gap in agentic workflows where agents process retrieved documents, tool results, or user messages that may contain hidden instructions. By running locally, the classifier allows applications to inspect these inputs in real time. The system returns both a binary verdict and an injection score, leaving the final decision to block or review content to the application developer.

This version, labeled S15e, represents a focused effort to provide a non-commercial baseline for security screening. While the example code is MIT licensed, the model weights themselves are released under CC BY-NC 4.0, meaning commercial use requires a separate license from ZeroLeaks. This distinction highlights the dual nature of the release: accessible for research and personal projects, but controlled for enterprise deployment.

How it works

The model operates as a 12-layer BERT encoder with a binary classification head. It processes text in windows of 256 tokens, using a stride of 192 tokens to slide across longer inputs. For any given input, it scans up to 32 windows and returns the highest injection score found. This sliding window approach ensures that localized attacks within larger documents are not missed by global averaging.

When using the provided Python example, the system tokenizes the complete input. If the text exceeds 6,206 tokens, the classifier checks the first 28 windows and the last four, skipping the middle section. In such cases, the result includes a coverage.truncated: true flag. Developers must treat an unflagged result on truncated text as inconclusive rather than safe, since the omitted middle portion has not been analyzed.

The inference pipeline is straightforward. After downloading the model via Hugging Face Hub, developers can initialize the ShieldSmall class and pass strings for classification. The output includes a boolean flagged status and a numeric score. The default threshold for flagging is 0.5, but the documentation advises tuning this value based on specific application traffic to balance sensitivity and false positives.

Key details

  • Model size: 118 million parameters with 118 MB of int8 ONNX weights.
  • Architecture: 12-layer BERT encoder derived from intfloat/multilingual-e5-small.
  • Performance: Reports 84.3% mean category-balanced accuracy and 71.8% attack recall when combined with Shield rules v2.
  • Input handling: Processes 256-token windows with a 192-token stride; truncates inputs over 6,206 tokens.
  • License: Weights are CC BY-NC 4.0 (noncommercial); example code is MIT licensed.
  • Deployment: Runs offline on CPU using Python 3.11 or 3.12.

Why it matters

For engineers building agentic systems, prompt injection remains a persistent threat that traditional firewalls cannot address. These attacks embed malicious instructions within seemingly harmless data, such as a retrieved webpage or a customer support ticket. Shield Small provides a dedicated layer of defense that sits between data ingestion and agent processing, allowing teams to filter out dangerous inputs before they influence model behavior.

The ability to run this classifier locally on a CPU is significant for cost and latency. Many existing security solutions require API calls to external services, which introduces dependency risks and potential data privacy concerns. By keeping the screening process internal and offline, organizations can maintain stricter control over their data flow while adding minimal computational overhead to their infrastructure.

However, the tool is not a silver bullet. The reported false-positive rate of 7.2% indicates that legitimate security discussions or complex queries may occasionally be flagged. Furthermore, the model’s performance varies across languages, as most training data is English. Developers must integrate this classifier as part of a broader security strategy that includes access controls and tool policies, rather than relying on it as the sole defense mechanism.

What you can do

  • Download the S15e weights from Hugging Face and test the classifier against your own dataset of user inputs.
  • Tune the decision threshold above or below 0.5 based on your tolerance for false positives versus missed attacks.
  • Implement logic to handle truncated inputs by rejecting or manually reviewing texts that exceed 6,206 tokens.
  • Extract plain text from HTML or structured documents before passing them to the classifier to ensure accurate scoring.
  • Evaluate the model’s multilingual performance if your application supports non-English languages, as accuracy may vary.
  • Contact ZeroLeaks for a commercial license if you intend to deploy Shield Small in a production business environment.

Tools from the Bytechap store

Keep reading

All stories