O harness ao redor do modelo decidiu esta semana se um agente sobrevivia à produção, não o modelo em si.
Ninety-three point six percent. That's the hijacking rate for MCP agents using nothing but poisoned tool metadata, without touching a single line of the model. [ref: a2m-trace-optimized-agent-hijacking-in-the-mcp-ecosystem]
And that's exactly the layer the whole week put in question: the harness, not the model.
This is ai|expert Wire. This week, the harness around the model decided whether an agent survived production — not the model itself.
Vinoth Govindarajan, from OpenAI's technical staff, presented an agent reliability framework that doesn't talk about model quality. It talks about state ownership, serialization of concurrent mutations, execution scope, work boundaries, and validation at the user-visible edge. [ref: the-agent-harness-control-planes-invariants-and-approval-boundaries-for-producti]
The case he opens with is revealing: the agent confirms it will remember a refund, and the system simply forgets. No error, no red screen — silent success is worse than a crash. [ref: the-agent-harness-control-planes-invariants-and-approval-boundaries-for-producti]
And this matters because UC Berkeley's research with Arena Intelligence showed that the harness choice can cost five times more on the same model, with nearly identical success rates. Claude Code costs about double what Pi costs running the same Claude Fable 5 — ninety-seven point eight percent success versus ninety-six point seven. [ref: harnesstax-how-much-does-the-harness-matter-for-coding-agents]
And the gap gets worse once you leave the local test. SWE-Serve, from NVIDIA with Berkeley, found that a third of patches pass every local test and fail end-to-end validation in production — twenty-three percentage points of difference once serving tests enter the count. [ref: swe-serve-benchmarking-agentic-engineering-for-production-inference-serving]
That's the opposite of what any platform squad wants to hear: the benchmark that approved the agent didn't measure the one thing that matters, which is it working under real traffic.
If the control layer decides whether the agent survives, the next question is how much it costs to shield that layer — and that's the second half of the week.
A heap overflow in libheif, chained with a misconfigured SSO, opened up OpenAI employee accounts and internal repositories. From initial discovery to proof of concept, seventy-two hours. OpenAI paid a six thousand five hundred dollar bounty. [ref: heap-overflow-and-sso-misconfiguration-compromised-openai-internal-repos]
The detail that matters for anyone auditing a model is the capability jump mid-chain: Claude Opus 4.8 couldn't produce a working exploit with ASLR enabled across multiple sessions; Claude Opus 5.5 did it in three hours, on the same task. [ref: heap-overflow-and-sso-misconfiguration-compromised-openai-internal-repos]
And that's exactly the kind of failure Cloudflare Worker Previews tries to contain before production: every Git branch gets its own isolated environment, with its own namespace for Durable Objects and Containers, so an agent's change doesn't leak into real traffic. [ref: cloudflare-worker-previews-isolated-environments-for-agent-changes]
But the announcement itself admits the hole: service bindings triggered from inside a Preview still call the production version of the linked Worker. Partial isolation is still an attack surface. [ref: cloudflare-worker-previews-isolated-environments-for-agent-changes]
And the cost isn't just security. Linear had to rebuild its entire CI pipeline because AI agents accelerated code volume — they cut PR wait time from over six minutes to just over five, with the test suite nearly quadrupling since early 2026. [ref: ai-coding-made-ci-a-bottleneck-linear-reworked-theirs-to-keep-up]
So the same pressure that exposed the harness also forced the rebuilding of the validation around it. Security and speed competing for the same engineering budget.
The harness has become the company's new security perimeter, and protecting that perimeter costs nearly as much as building it did. In Friday's Edition, the full rundown of the week — including what this harness bottleneck costs in latency and compute budget. Have a good week.