LOW COST · RELEVANT · SAFE

Stalwart AI Assurance

AI agents that do real work through your tools. They pick the right tool from what your users actually say, never write unless your rules allow it, and need no model call to choose.

Choosing a tool

No model call

Writes

Only when your rules allow

Guardrail bugs

Found, with the fix

Access

One API key

A REAL EXAMPLE

We scanned the public guardrails that a whole industry copies

We pointed the Guardrail Bug Hunter at the public rule files of a leading open-source AI guardrail framework: the framework's own library of ready-made safety rails, the sample assistants its vendor publishes, and partner and community projects. 185 files from 12 repositories, scanned offline with no AI model in the loop.

Two files held real vulnerabilities, and each was reproduced in the framework's own runtime. Both are fixable by the Bug Hunter: it writes a corrective rule for each, and checks that the rule closes the hole without blocking anything else.

These are professional reference files that many teams copy, so every bug travels with every copy. Red-teaming would have had to guess the right conversation. The Bug Hunter found each one, with the conversation that triggers it.

185

guardrail files from 12 public repositories

2

vulnerable files, each reproduced in the framework's own runtime

32

tool calls an attacker could trigger just by steering the conversation

8.5 s

to scan the framework's 115 toolkit files

Finding 1 · Library rail

Hallucinations reported as jailbreaks

A rail that detects hallucinations raises them as jailbreaks. A team that blocks jailbreaks blocks every hallucination too, and a team that handles hallucinations never sees this rail's.

Fix written and checked
Finding 2 · Sample assistant

Three defects in a food-ordering assistant

Clearing the cart never runs, swapping an item checks the wrong value, and replies about the customer's order are left to unchecked AI-generated text.

Reproduced, fixable by the Bug Hunter
Finding 3 · Across the files

Tool calls anyone can trigger

32 tool calls in 28 files run on the user's apparent intent alone, so anyone who can steer the conversation can trigger them. Two of them place food orders with no check of who asked or whether they confirmed.

Reported with the route

Figures from our research paper, for a scan run on 3 October 2026.

EXECUTIVE BRIEFING

How an agent is kept to its rules

SLIDE 1 OF 4
HOW THE RAIL IS WIRED

Intent can be talked into anything. Context cannot. Actions need a policeman.

Three things are in play. They are not the same kind of thing.

IntentWhat was askedManipulableContextWhat your systems knowDeterministicThe gateReads bothA written ruleActionBooks, or refusesNothing else

Every hole has the same shape. A reachable intent, in front of an ungated action.

WHY STALWART

Low cost, relevant and safe by design, not by luck

A model left to run an agent pays for every decision, guesses at your domain, and writes whatever it decides. Stalwart AI Assurance puts your rules and your users' words in charge, and calls a model only where you choose.

WHAT MATTERSA MODEL ON ITS OWNSTALWART AI ASSURANCE
Choosing a toolA model call every turn, and a different answer some daysNo model call; the same request picks the same tool
Writing through your toolsWhatever the model decidesOnly when your rules allow it; otherwise it asks or refuses
Your domain's wordsPrompt tweaking, againLearned from your users' feedback, versioned
Bugs in your guardrailsFound when someone happens to try themHunted down, each with its conversation and the fix
Paying for a modelOn every decisionOnly where your setting sends a turn to one
GET STARTED

Put an agent to work, safely

Read the API reference and try a call, or register your interest and our team will be in touch.

TECHNICAL DEMO

A finance case study, step by step

Our technical demo runs the engine on a regulated-finance policy set: a trade-booking approval gate, a below-threshold exemption that opens a bypass, and a consumer-duty advice rule. Pick a case to see the policy, the verdict and the evidence.

INTERACTIVE FORMAL VERIFICATION CONSOLE

Algebraic Guardrail Verification in Action

Tool & MCP Dispatch
define user request_trade_booking
  "book this contract note"

define flow trade_booking_gate
  user request_trade_booking
  if $four_eyes_approved
    bot book_trade
  else
    bot refuse_unapproved_booking
$verify-guardrail policy.rules --pin None (Unconditional Obligation)
VERDICT

MATHEMATICALLY SAFE (UNSAT)

The algebraic solver derives 0 = 1 in 5 steps. Over all possible 2^N runtime contexts, the action book_trade is provably unreachable unless four_eyes_approved is positively asserted.

Steps / Derivations5 Steps (Replayable)
Proof SystemGF(2) Algebra (C++)
Attack Hypothesis:Action = 1 ∧ Precondition = 0
Context Space:Quantified over all 2^N combinations
Independent Verification:Self-contained DAG receipt
VERIFIED SUITES

Algebraic verification benchmarks

Regulatory policies, tool dispatch gates and state lookbacks, decided by deterministic finite-field elimination.

Case 01 · Regulatory AI

FCA Consumer Duty Advice Guardrail (410 Variables)

410-variable synthesized GF(2) circuit under FCA COBS/Duty obligations. Proved that advising permissions cannot be bypassed via targeted-support advice flows.

Verified UNSAT (5-Step Receipt)
Case 02 · MCP Tool Gating

Middle-Office Trade Booking Dispatch Gate

Discovered an OR-accumulation hole in a sub-threshold booking exemption. Generated a minimally disruptive policy patch certified to block the vulnerability.

SAT Bypass Found & Patched
Case 03 · Stateful Guardrails

Multi-Step Velocity Chain State Lookback ($prev.v)

Stateful lookback verification enforcing dynamic transaction frequency caps across conversational turn state variables.

Verified UNSAT (8-Step Receipt)