aiexpert
Home / Podcast / Ep. 36
36
Episode 36 · Sep 23, 2026 · 5 min · Wire

O harness ao redor do modelo decidiu esta semana se um agente sobrevivia à produção, não o modelo em si.

O harness ao redor do modelo decidiu esta semana se um agente sobrevivia à produção, não o modelo em si.

Hosted by AlanHosting
00:00 -05:16

Episode transcript

The script as aired, in full
Alan

Ninety-three point six percent. That's the hijack rate for MCP agents using nothing but poisoned tool metadata, without touching a single line of the model. [ref: a2m-trace-optimized-agent-hijacking-in-the-mcp-ecosystem]

Ada

And that's exactly the layer the whole week put on trial: the harness, not the model.

Alan

This is ai|expert Wire. This week, the harness around the model decided whether an agent survived production — not the model itself.

Alan

Vinoth Govindarajan, from OpenAI's technical staff, presented an agent reliability framework that doesn't talk about model quality. It talks about state ownership, serialization of concurrent mutations, execution scope, work limits, and validation at the user-visible edge. [ref: the-agent-harness-control-planes-invariants-and-approval-boundaries-for-producti]

The case he opens with is telling: the agent confirms it will remember a refund, and the system simply forgets. No error, no red screen — silent success is worse than a crash. [ref: the-agent-harness-control-planes-invariants-and-approval-boundaries-for-producti]

Ada

And this matters because UC Berkeley's research with Arena Intelligence showed that harness choice can cost five times more on the same model, with nearly identical success rates. Claude Code costs about twice as much as Pi running the same Claude Fable 5 — ninety-seven point eight percent success versus ninety-six point seven. [ref: harnesstax-how-much-does-the-harness-matter-for-coding-agents]

Alan

And the gap gets worse once you leave the local test. SWE-Serve, from NVIDIA with Berkeley, found that a third of patches pass every local test and fail end-to-end validation in production — twenty-three percentage points of difference once serving tests enter the count. [ref: swe-serve-benchmarking-agentic-engineering-for-production-inference-serving]

Ada

That's the opposite of what any platform squad wants to hear: the benchmark that approved the agent didn't measure the one thing that matters, which is it working with real traffic.

Alan

If the control layer decides whether the agent survives, the next question is how much it costs to armor that layer — and that's the second half of the week.

Ada

A heap overflow in libheif, chained with an SSO misconfiguration, opened up OpenAI employee accounts and internal repositories. From initial finding to proof of concept, seventy-two hours. OpenAI paid a six thousand five hundred dollar bounty. [ref: heap-overflow-and-sso-misconfiguration-compromised-openai-internal-repos]

Alan

The detail that matters to anyone auditing the model is the capability jump in the middle of the chain: Claude Opus 4.8 couldn't produce a working exploit with ASLR enabled across multiple sessions; Claude Opus 5.5 did it in three hours, on the same task. [ref: heap-overflow-and-sso-misconfiguration-compromised-openai-internal-repos]

Ada

And that's exactly the kind of failure Cloudflare Worker Previews tries to contain before production: every Git branch gets an isolated environment, with its own namespace for Durable Objects and Containers, so an agent's change doesn't leak into real traffic. [ref: cloudflare-worker-previews-isolated-environments-for-agent-changes]

Alan

But the announcement itself admits the gap: service bindings triggered from inside a Preview still call the production version of the bound Worker. Partial isolation is still attack surface. [ref: cloudflare-worker-previews-isolated-environments-for-agent-changes]

Ada

And the price isn't just security. Linear had to rebuild its entire CI pipeline because AI agents accelerated code volume — they cut PR wait time from over six minutes to just over five, with the test suite nearly quadrupling since early 2026. [ref: ai-coding-made-ci-a-bottleneck-linear-reworked-theirs-to-keep-up]

Alan

So the same pressure that exposed the harness also forced a rebuild of the validation around it. Security and speed competing for the same engineering budget.

Alan

The harness has become the company's new security perimeter, and protecting that perimeter costs almost as much as building it. In Friday's Edition, the full rundown of the week — including what this harness bottleneck costs in latency and compute budget. Good week.