Kernel profiling, power conversion, e segurança de imagem expõem a infraestrutura de IA como o novo gargalo da produção.
100 terabytes.
That's what Cloudflare pulled out of RAM by cutting unnecessary hashes. And it's what the entire industry is leaving on the table while a single GPU burns 5 kilowatts. [ref: cloudflare-saved-100tb-of-ram-through-algorithmic-optimization-in-rust]
This is the ai|expert Wire. The week infrastructure became the constraint — kernel profiling, power conversion, and a single image file exposed how much of AI's real cost lives below the model.
Google extended XProf this week with cycle-level kernel profiling for custom Pallas kernels on TPU v7. [ref: google-adds-cycle-level-kernel-profiling-to-xprof] Hardware counters at microsecond resolution. A demo matmul kernel dropped 30 percent just by overlapping HBM loads with compute.
That's the systems work nobody talks about. The model is done. Now you're fighting memory stalls and synchronization overhead, and you need to see it happening in real time.
NVIDIA opened native GPU kernel programming in Rust the same day. Two tracks — cuda-oxide for SIMT control, cutile-rs for higher-level Tile programming. [ref: nvidia-announces-native-gpu-programming-in-rust-with-two-kernel-writing-tracks] Both enforce memory safety at compile time instead of crashing in production.
But here's the thing: silicon cannot keep up anymore. GPUs are pulling 5 kilowatts within two years. Data centers are shifting to 800-volt distribution and deploying GaN semiconductors across every conversion stage because silicon's on-resistance hit its theoretical limit. [ref: ai-power-demands-push-gan-into-data-center-design]
GaN switches at about 3 megahertz. Silicon switches at roughly 1 megahertz — GaN is 3 times faster. The roadmap targets 10 megahertz.
The cost of that power conversion is now part of your inference bill. GaN can improve overall efficiency by 5 percent from AC to the GPU core. At gigawatt scale, that's real money.
Second story. Cloudflare reclaimed 100 terabytes by analyzing their consistent hashing algorithm mathematically. [ref: cloudflare-saved-100tb-of-ram-through-algorithmic-optimization-in-rust] The last 90,000 hashes per server were reducing error by 0.7 percent. Waste.
They cut hash count by 90 percent and compacted the data structures in Rust. Struct packing bought another 25 percent. That's not a model improvement. That's an architect who measured what was actually working and deleted the rest.
Cloudflare's Pingora Backend Router uses consistent hashing to route cacheable requests to servers by URL. The same pattern shows up in caching, task assignment, and distributed routing across the industry.
A heap overflow in libheif and an SSO misconfiguration exposed OpenAI's internal repositories. [ref: heap-overflow-and-sso-misconfiguration-compromised-openai-internal-repos] The image-processing vulnerability was unpatched. The forum accepted HEIC uploads. ImageMagick passed them to libheif.
The Hacktron team used Claude Opus 5 to develop a working exploit. When Anthropic released Claude Opus 5 that evening, a new session produced an ARM64 variant within three hours.
The escalation vector was the SSO. Sign in with OpenAI on the forum, take over an employee's Codex account, and Codex was connected to the GitHub organization. The researchers demonstrated access by having the compromised Codex open a pull request in OpenAI's internal monorepo, then stopped testing.
The entire timeline from initial discovery to repository access took less than 72 hours. Three researchers. Under three thousand dollars in tokens. The same libheif flaw affects Slack, Meta, GitHub Enterprise, and major frameworks.
The operational reality is that memory-corruption vulnerabilities are now compressible into days of work with AI assistance.
Linear halved CI runner time despite the test suite nearly quadrupling. [ref: linear-halved-ci-runner-time-despite-test-suite-quadrupling] Agents accelerated code shipping. Validation became the bottleneck.
They moved to third-party runners with faster CPUs. Switched to a native TypeScript compiler. Rewrote lint rules to use static analysis instead of building the full type graph. Reduced per-shard setup time from 110 to 140 seconds down to 72 to 72 seconds — roughly 44 percent.
The largest single optimization came from test execution. They split large test files, moved from four to eight shards, and introduced an opt-in vitest project with shared module registry. That one change was worth 17 percent in monthly savings.
But the real lesson is that when AI agents change the shape of your build workload, the bottleneck moves from code generation to validation. And the fix requires rethinking infrastructure, job dependencies, and setup overhead as a system.
Infrastructure is where the real cost lives now — not in the model, but in the kernel, the power supply, and the validation pipeline that has to keep up with agents that code faster than humans can review. The Wire, Monday. Good week.