PROOF, NOT SAMPLING

Stalwart AI Assurance

Find every way past your AI guardrails, prove the ones that hold and get a checked fix for the ones that don't. One API for guardrails, agent tool choice and model verification.

Products

5 in one API

Every finding

Reproduced + fixed

Evidence

Auditor-checkable

Access

One API key

A REAL EXAMPLE

We scanned the public guardrails that a whole industry copies

We pointed the Guardrail Bug Hunter at the public rule files of a leading open-source AI guardrail framework: the framework's own library of ready-made safety rails, the sample assistants its vendor publishes, and partner and community projects. 185 files from 12 repositories, scanned offline with no AI model in the loop.

Two files held real vulnerabilities, and each was reproduced in the framework's own runtime. The Bug Hunter wrote a corrective rule for the first and proved that it closes the hole without blocking anything else.

These are professional reference files that many teams copy, so every defect travels with every copy. Sampling would not have found them. Exhaustive analysis did.

185

guardrail files from 12 public repositories

2

vulnerable files, each reproduced in the framework's own runtime

32

tool calls an attacker could trigger just by steering the conversation

8.5 s

to scan the framework's 115 toolkit files

Finding 1 · Library rail

Hallucinations reported as jailbreaks

A rail that detects hallucinations raises them as jailbreaks. A team that blocks jailbreaks blocks every hallucination too, and a team that handles hallucinations never sees this rail's.

Fix written and proved
Finding 2 · Sample assistant

Three defects in a food-ordering assistant

Clearing the cart never runs, swapping an item checks the wrong value, and replies about the customer's order are left to unchecked AI-generated text.

Reproduced in the runtime
Finding 3 · Across the files

Tool calls anyone can trigger

32 tool calls in 28 files run on the user's apparent intent alone, so anyone who can steer the conversation can trigger them. Two of them place food orders with no check of who asked or whether they confirmed.

Reported with the route

Figures from our research paper, for a scan run on 3 October 2026.

EXECUTIVE BRIEFING

Proving what an AI agent cannot do

SLIDE 1 OF 4
HOW THE RAIL IS WIRED

Intent can be talked into anything. Context cannot. Actions need a policeman.

Three things are in play. They are not the same kind of thing.

IntentWhat was askedManipulableContextWhat your systems knowDeterministicThe gateReads bothA written ruleActionBooks, or refusesNothing else

Every hole has the same shape. A reachable intent, in front of an ungated action.

WHY PROOF

Testing samples. Proof covers every case.

Red-teaming and evaluation suites only find a hole if someone happens to try it. Stalwart AI Assurance checks every case, and gives you evidence either way.

THE QUESTIONSAMPLING & RED-TEAMINGSTALWART AI ASSURANCE
What is checkedThe prompts someone thought to tryEvery possible case
When nothing is foundA pass only means nothing failedA certificate anyone can re-check
When something breaksA transcript, if you are luckyThe exact case, reproduced
The fixManual prompt tweakingA fix that has already been checked
For your auditorHard to repeatDeterministic and replayable
GET STARTED

See what your AI cannot do

Read the API reference and try a call, or register your interest and our team will be in touch.

TECHNICAL DEMO

A finance case study, step by step

Our technical demo runs the engine on a regulated-finance policy set: a trade-booking approval gate, a below-threshold exemption that opens a bypass, and a consumer-duty advice rule. Pick a case to see the policy, the verdict and the evidence.

INTERACTIVE FORMAL VERIFICATION CONSOLE

Algebraic Guardrail Verification in Action

Tool & MCP Dispatch
define user request_trade_booking
  "book this contract note"

define flow trade_booking_gate
  user request_trade_booking
  if $four_eyes_approved
    bot book_trade
  else
    bot refuse_unapproved_booking
$verify-guardrail policy.rules --pin None (Unconditional Obligation)
VERDICT

MATHEMATICALLY SAFE (UNSAT)

The algebraic solver derives 0 = 1 in 5 steps. Over all possible 2^N runtime contexts, the action book_trade is provably unreachable unless four_eyes_approved is positively asserted.

Steps / Derivations5 Steps (Replayable)
Proof SystemGF(2) Algebra (C++)
Attack Hypothesis:Action = 1 ∧ Precondition = 0
Context Space:Quantified over all 2^N combinations
Independent Verification:Self-contained DAG receipt
VERIFIED SUITES

Algebraic verification benchmarks

Regulatory policies, tool dispatch gates and state lookbacks, decided by deterministic finite-field elimination.

Case 01 · Regulatory AI

FCA Consumer Duty Advice Guardrail (410 Variables)

410-variable synthesized GF(2) circuit under FCA COBS/Duty obligations. Proved that advising permissions cannot be bypassed via targeted-support advice flows.

Verified UNSAT (5-Step Receipt)
Case 02 · MCP Tool Gating

Middle-Office Trade Booking Dispatch Gate

Discovered an OR-accumulation hole in a sub-threshold booking exemption. Generated a minimally disruptive policy patch certified to block the vulnerability.

SAT Bypass Found & Patched
Case 03 · Stateful Guardrails

Multi-Step Velocity Chain State Lookback ($prev.v)

Stateful lookback verification enforcing dynamic transaction frequency caps across conversational turn state variables.

Verified UNSAT (8-Step Receipt)