Updated Sep 2026. The first version of this post described a proxy that failed open in two places. Those are fixed and the post now describes the code as it is today.
AI agents get real tools: file access, shell execution, API keys. Most safety layers live inside the agent's own process, where the agent (or a bug in it) can route around them. I wanted to see how far a layer outside the agent could go, so I built AgentBrake, a proxy that sits between an MCP client and an MCP server and checks every tool call before it runs.
This post walks through how it works, and which parts don't work yet.
1. Interception over stdio
MCP servers commonly talk JSON-RPC over stdin and stdout. AgentBrake exploits that: instead of being a library you import, it wraps the server process.
- Normal flow: IDE -> MCP server
- With AgentBrake: IDE -> AgentBrake -> MCP server
In src/proxy/interceptor.ts, AgentBrake spawns the server as a child process (without a shell, and without its own configuration variables in the child's environment) and takes over its standard input and output.
// Simplified concept
this.input.on("data", (data) => {
// 1. Buffer bytes and cut complete lines (a message may arrive in pieces, or several at once)
// 2. Parse each line; anything unparseable, or not a valid JSON-RPC object, is refused
// 3. If it is a tool call and passes every policy -> forward the parsed message
// 4. Otherwise -> answer with a JSON-RPC error instead of forwarding
});
Two properties follow from this design:
- Language agnostic. It doesn't matter whether the server is written in Python, Node.js or Go. The proxy only speaks JSON-RPC.
- The agent can't skip it. The proxy owns the process, so the traffic has to go through it.
The first version failed open
That "simplified concept" hides the part I got wrong. The original code split every chunk of input on newlines, parsed each piece, and in the catch branch forwarded the piece unchanged. A message split across two reads reached the server as two unchecked fragments, and any message that JavaScript's parser rejects but the server's parser accepts (a bare NaN, for example) skipped every policy. Both are the same mistake: when the proxy couldn't tell what a message was, it let it through.
The fix is in src/proxy/framing.ts and the interceptor. Lines are cut from a byte buffer (so a multi-byte character or a message split at any point is reassembled), with a size cap. Unparseable input, batches, malformed tool calls, an unknown policy action and a policy that throws are all blocked with a JSON-RPC error. What gets forwarded is the re-serialised parsed message, so the server sees exactly what the policies saw. A regression test splits a tool call at every byte boundary and checks the policy still applies.
The same lesson applies to configuration: a policy file that failed validation used to fall back to no policies. Now the proxy refuses to start, unless you opt in to the old behaviour with an environment variable.
2. Budgets without a tokenizer
The obvious way to limit an agent's cost is to count tokens. That needs a tokenizer, which is heavy and specific to each model. I went with a cruder idea, which I call the "arcade token" model: every tool call costs a fixed amount, and the budget is a total.
In src/policy/policies/BudgetPolicy.ts the check looks like this:
// BudgetPolicy.ts
const cost = this.toolCosts.get(toolName) || this.defaultCost;
const projectedSpend = this.currentSpend + cost;
if (projectedSpend > this.maxBudget) {
return { action: "block", reason: "Budget exceeded." };
}
The trade-off:
- Pros: no tokenization latency, and a limit like "$5 total" is easier to configure than "500k tokens".
- Cons: it's imprecise. A tool call with 100 lines of text costs the same as one with a single line.
What isn't there yet: the policy class supports per-tool costs, but the YAML config doesn't pass them through, so today every call costs the same flat default. It counts calls, not real spend or tokens.
3. Policies that read the arguments
Enforcement is a sequential loop in interceptor.ts: each message passes through a chain of policy classes, and the first violation decides the outcome.
The most useful policy is GranularAccessPolicy. Instead of only checking which tool is called, it checks the arguments, with regular expressions. It's data-loss prevention for agent tool calls.
- Scenario: an agent tries to write a file.
- The rule:
allow_if: { arguments: { path: "^/tmp/.*" } }
// GranularAccessPolicy.ts
if (!new RegExp(pattern).test(String(argValue))) {
return { action: "block", reason: "Arguments do not match ALLOW pattern." };
}
That lets you say "you may use write_file, but only inside /tmp". Regex allow and deny lists are simple and fast to evaluate, but only as good as the patterns you write. Patterns are compiled once at startup, rejected if they are longer than 512 characters or contain nested quantifiers such as (a+)+ (a cheap guard against catastrophic backtracking, not a proof), and an argument value over 8,192 characters is treated as a violation. Attackers can still write a value that means the same thing but doesn't match, so this is a filter, not a sandbox.
4. What works, and what doesn't
I'd rather say this plainly than let the README imply otherwise.
The circuit breaker works now. In the first version nothing reported a failed tool call to it, so it could never trip. The proxy now parses the server's responses and feeds tool errors back to each policy, so after a few consecutive failures a tool is cut off for a while.
Human approval still isn't built. A tool call that requires approval gets a "pending" error and nothing can approve it: no CLI command, no webhook, no Slack integration is connected. I looked at a file-based approve command and dropped it, because an agent with file access could approve its own request. The README now lists approval as not implemented, and a "sandbox" action is enforced as a plain block.
The budget is a call counter, as described above.
It is not a sandbox. It covers stdio MCP servers only, and a regular-expression rule can be bypassed by a value that means the same thing but reads differently.
What works end to end: the allow, deny and block policies, regex argument filtering, rate limiting, the call cap, the circuit breaker, and a config that refuses to load when it is invalid. The test count went from 22 to 85, mostly around framing and fail-closed behaviour.
Summary
AgentBrake is a control point outside the agent rather than a monitoring tool. The proxy design is the part I'd keep: it needs no changes to the agent, and it sees everything. The bigger lesson was in the failure mode. For a component whose only job is to say no, "I couldn't parse it, so I passed it through" is the worst possible default, and I only found it by asking what happens to input the code doesn't understand.