Stampli compressed 243 hours of launch work into 77 hours—a 68% reduction—by shipping Deep Finance using OpenAI's Codex and ChatGPT Work. OpenAI published the figures on August 20, 2026. It's one of the most granular productivity benchmarks available for AI-assisted software and content delivery.
Deep Finance is a spend-intelligence layer for CFOs and VPs on Stampli's procure-to-pay platform. The launch required seven deliverable types in parallel: blog series, launch emails, webinar and deck, social and paid creative, PR Newswire release, web page, and sales enablement. Design resources and outside contractors were unavailable—the marketing team executed with what it had.
The team built a shared knowledge system in Codex rather than tackling tasks sequentially. Codex ingested product context, meeting notes, Jira tickets, GitHub history, and messaging guidelines, then produced review-ready drafts across all deliverable types. Every customer-facing asset got human review. Codex handled 90% of the hero animation before a contractor finished the opening scene. The 243-hour estimate is what those tasks would have cost without Codex; 77 hours is what they actually took.
The same infrastructure now runs day-to-day work. Before, keeping materials current meant interviewing PMs, reading Jira, reviewing GitHub, and manually building help center articles, one-pagers, and presentations. Stampli now uses agents connected to its source of truth to pull from product systems and meeting notes. Director of Product Marketing Melad Zahedi said the system multiplied output by 10×, from a couple pieces per week to hundreds.
Live querying revealed higher reliability demands. During an executive meeting, an employee used Codex to pull and analyze HubSpot metrics in seconds. Zahedi said the alternative was half a day of FP&A work. The actual time: 20 seconds of keystrokes. That's not a controlled benchmark—it's an in-meeting event, which means the latency and reliability bar was higher than a typical async workflow.
| Task | Before (without AI agents) | After (with Codex / ChatGPT) | Improvement |
|---|---|---|---|
| Deep Finance launch work | 243 h (estimated) | 77 h (actual) | 68% reduction |
| Marketing content output | A couple of pieces per week | Hundreds of pieces per week | ~10× multiplier |
| HubSpot metrics pull & analysis | ~Half a day of FP&A work | ~20 seconds of keystrokes | ~1,440× faster |
The case study omits prompt engineering overhead, error rates before human review, and how the team handles version drift when Jira or GitHub updates faster than Codex's context window. Teams replicating this will face those costs. The 68% figure is net output, not setup or maintenance effort.
For teams benchmarking AI-assisted velocity, Stampli's numbers work as a ceiling, not a median. The conditions were favorable: a small, technically willing marketing team, fixed scope, and strong product-context access. If your team is drowning in context-reconstruction before execution, a GPT-powered ingestion layer is worth modeling. Stampli gives you a specific number to use.