O harness ao redor do modelo decidiu esta semana se um agente sobrevivia à produção, não o modelo em si.
Ninety-three point six percent. That's the hijack rate for MCP agents using nothing but poisoned tool metadata, without touching a single line of the model. [ref: a2m-trace-optimized-agent-hijacking-in-the-mcp-ecosystem]
And that's exactly the layer the whole week put on trial: the harness, not the model.
This is ai|expert Wire. This week, the harness around the model decided whether an agent survived production — not the model itself.
Vinoth Govindarajan, from OpenAI's technical staff, presented an agent reliability framework that doesn't talk about model quality. It talks about state ownership, serialization of concurrent mutations, execution scope, work limits, and validation at the user-visible edge. [ref: the-agent-harness-control-planes-invariants-and-approval-boundaries-for-producti]
The case he opens with is telling: the agent confirms it will remember a refund, and the system simply forgets. No error, no red screen — silent success is worse than a crash. [ref: the-agent-harness-control-planes-invariants-and-approval-boundaries-for-producti]
And this matters because UC Berkeley's research with Arena Intelligence showed that harness choice can cost five times more on the same model, with nearly identical success rates. Claude Code costs about twice as much as Pi running the same Claude Fable 5 — ninety-seven point eight percent success versus ninety-six point seven. [ref: harnesstax-how-much-does-the-harness-matter-for-coding-agents]
And the gap gets worse once you leave the local test. SWE-Serve, from NVIDIA with Berkeley, found that a third of patches pass every local test and fail end-to-end validation in production — twenty-three percentage points of difference once serving tests enter the count. [ref: swe-serve-benchmarking-agentic-engineering-for-production-inference-serving]
That's the opposite of what any platform squad wants to hear: the benchmark that approved the agent didn't measure the one thing that matters, which is it working with real traffic.
If the control layer decides whether the agent survives, the next question is how much it costs to armor that layer — and that's the second half of the week.
A heap overflow in libheif, chained with an SSO misconfiguration, opened up OpenAI employee accounts and internal repositories. From initial finding to proof of concept, seventy-two hours. OpenAI paid a six thousand five hundred dollar bounty. [ref: heap-overflow-and-sso-misconfiguration-compromised-openai-internal-repos]
The detail that matters to anyone auditing the model is the capability jump in the middle of the chain: Claude Opus 4.8 couldn't produce a working exploit with ASLR enabled across multiple sessions; Claude Opus 5.5 did it in three hours, on the same task. [ref: heap-overflow-and-sso-misconfiguration-compromised-openai-internal-repos]
And that's exactly the kind of failure Cloudflare Worker Previews tries to contain before production: every Git branch gets an isolated environment, with its own namespace for Durable Objects and Containers, so an agent's change doesn't leak into real traffic. [ref: cloudflare-worker-previews-isolated-environments-for-agent-changes]
But the announcement itself admits the gap: service bindings triggered from inside a Preview still call the production version of the bound Worker. Partial isolation is still attack surface. [ref: cloudflare-worker-previews-isolated-environments-for-agent-changes]
And the price isn't just security. Linear had to rebuild its entire CI pipeline because AI agents accelerated code volume — they cut PR wait time from over six minutes to just over five, with the test suite nearly quadrupling since early 2026. [ref: ai-coding-made-ci-a-bottleneck-linear-reworked-theirs-to-keep-up]
So the same pressure that exposed the harness also forced a rebuild of the validation around it. Security and speed competing for the same engineering budget.
The harness has become the company's new security perimeter, and protecting that perimeter costs almost as much as building it. In Friday's Edition, the full rundown of the week — including what this harness bottleneck costs in latency and compute budget. Good week.