Your AI Agent Guardrail Might Be Switched Off: Fail-Open Policies and How to Test Them
7 min read Fullmakt Team
- business-case
- governance
- agents
- agentic-ai-gone-wrong
- observability
- traceability
- mcp
You wrote the rule. “Block anything that looks like DROP TABLE.” Review
approved it, the policy was saved, and the dashboard shows a policy attached to
the agent. Months later an agent runs a destructive query and the guardrail
never fires.
Nothing was bypassed. The rule never worked. A misspelled key, a regex that didn’t compile or a policy document that didn’t parse can all leave an AI agent with a guardrail that exists on paper and does nothing in practice.
This post is about that failure mode, why it’s easy to hit with agentic AI, and what a governance setup should do about it.
The failure: a policy that exists but doesn’t enforce
Most agent policy systems share a design question: what happens when the policy itself is broken? There are two answers.
- Fail closed. If the policy can’t be evaluated, deny the call. Safe, but a typo can take your agents offline.
- Fail open. If the policy can’t be evaluated, allow the call. Agents keep working, but the guardrail is silently gone.
Neither is wrong in every case. The failure is not knowing which one you have. Here are the ways a policy goes inert in practice:
- A misspelled key.
"hardDenny"instead of"hardDeny"is just an unknown property, ignored like any other. - A regex that doesn’t compile. One unescaped bracket and the rule can never match.
- A regex that times out. Patterns with catastrophic backtracking are usually cut off by a match timeout, and a timed-out match counts as “no match”.
- Double-escaping mistakes. When a policy is stored as a JSON string
containing regexes, backslashes need escaping twice.
\sbecomes\\\\s, and getting it wrong changes what the rule matches. - Matching the wrong field. A rule aimed at
querythat actually readssqlevaluates against an empty string and never fires. - A malformed enum value.
"action": "Denyy"can invalidate the whole document at evaluation time, not just that rule.
None of these produce a loud error. The agent keeps getting 200 OK, the
dashboard keeps showing a green “policy attached” badge, and the first sign of
trouble is the incident. This is the shape of many
agentic AI failures: the
control looked present, so nobody checked it.
Why this matters more for agents than for people
A human who is blocked by a bad rule complains. An agent just retries or routes around it, and a rule that never fires gives nobody a complaint to read.
- Agents act at machine speed. An unprotected destructive action can repeat many times before a person looks, the same dynamic behind runaway agent spend and duplicate actions.
- Policies drift. The ratchet effect means rules get edited by different people over time, and every edit is a chance to break an earlier one.
- Inputs are adversarial. With indirect prompt injection, the guardrail is the last line of defense, and it will be tested by someone who is trying to get through it.
What a trustworthy policy setup looks like
Whatever tooling you use, check it against five properties.
1. A documented failure mode. You should be able to answer “if my policy is unreadable, does the call go through?” without reading source code. If the answer is “allow”, you need compensating controls, below.
2. A side-effect-free way to test. You should be able to submit a policy, a tool name and sample arguments and get back a decision (allow, deny, or require approval) without executing anything. This turns a policy into something you can unit test.
3. A test suite of known-bad inputs. For every rule, keep at least one input that must be denied and one that must be allowed. Run them whenever the policy changes, the same discipline you apply to code.
4. Layered defenses. A regex rule is one layer. Pair it with least-privilege scoping so the credential can’t do the destructive thing, with approval for risky actions, and with behavioral limits such as calls per minute. If a regex silently fails, a scoped credential still holds.
5. An audit trail of decisions. You want evidence for what was evaluated, not only what ran. See what incident forensics needs from your audit trail.
How Fullmakt approaches it
Fullmakt is a credential broker for AI agents. Agents call it over MCP or A2A, and it applies a per-collection policy before attaching any credential and forwarding the request. We’ll be specific about the failure modes, because specifics are what you need when deciding who to trust.
- Policy rules. A policy can hard-deny calls, require human approval, apply behavioral limits (calls per minute or hour, errors per hour, new tools and hosts) and name response fields to redact before output returns to the model. Rules are evaluated in a fixed order: hard deny first, then behavior, then approval patterns.
- Honest defaults. Rules that can’t be evaluated are skipped. An invalid regex or empty pattern makes a rule inert, and an unparseable policy document is treated as allow-everything. That is fail-open at the policy-document level, so testing before you rely on a rule matters. When policy evaluation runs in a separate decision service, an unreachable service fails closed instead and the call is denied.
- A way to test. The policy decision service exposes a decision endpoint that returns the outcome for a given policy, tool and arguments without executing anything, so a policy can be tested like a function.
- Defense in depth. Policy is only one gate. The agent never holds the upstream secret, since credentials are resolved server-side from references, so a missed deny rule can’t hand the agent a broader key. The broker also supports approval flows for sensitive actions and redacts configured fields from responses.
- Evidence. Each call is written to an audit log tied to the agent identity, with credential references recorded and never their values.
The business case is not that policies are unbreakable. It’s that the failure is contained: a rule that silently misfires meets a narrowly scoped credential, an approval gate and a trace you can read afterward.
A pre-flight checklist for agent policies
- Do you know whether your policy engine fails open or closed on a bad policy?
- Is there a way to evaluate a policy against sample calls without side effects?
- Does every deny and approval rule have a must-match and a must-not-match test?
- Do you re-run those tests on every policy edit, not only the first?
- Would a missed rule still be stopped by a scoped credential or approval gate?
- Can you query which policy version made each past decision?
If you answered “no” to two or more, assume at least one of your guardrails is decorative today.
FAQ
What does “fail open” mean for an AI agent policy? If the policy can’t be parsed or a rule can’t be evaluated, the call is allowed instead of denied. Agents keep working but the guardrail is effectively off.
How do I test an AI agent policy before deploying it? Use a side-effect-free decision check: supply the policy, a tool name and sample arguments, and confirm the outcome is allow, deny or require approval. Keep one must-block and one must-allow case for every rule.
Why would a regex rule silently stop working? Common causes are a typo in the rule key, an invalid or double-escaped pattern, a pattern that times out, or a rule that targets a field the tool doesn’t send.
Should agent guardrails fail open or fail closed? Prefer fail-closed for destructive or regulated actions and accept the operational cost. Where you can’t, layer scoped credentials and approval gates so one broken rule isn’t the only protection.
Does Fullmakt fail open or closed? Rules that can’t be evaluated are skipped, and an unparseable policy allows calls. If a separate decision service is configured and unreachable, calls are denied. Test policies with the decision endpoint before relying on them.
What does Fullmakt cost? Usage-based pricing per operation, with no plans or minimums, plus a no-card sandbox for trying policies against your own agents.