AI's 452-rollout DeepSWE benchmark positions Kimi K3 as a cost-effective coding solution, with $4.65 per rollout compared to Claude Fable 5's $13.41. Kimi K3 delivers 2.8× more solved tasks per $100, while only trailing pass@1 by 1.4 points and leading pass@4 outright. This cost difference—$2,103 versus $6,010 for 113 real, long-horizon open-source feature requests—prompts a stack-level reconsideration of which model is best suited for autonomous coding agents.

Cost efficiency comparison: Kimi K3 delivers 2.8× more tasks per dollar than Claude Fable 5 across the 452-rollout DeepSWE benchmark.
FIG. 02 Cost efficiency comparison: Kimi K3 delivers 2.8× more tasks per dollar than Claude Fable 5 across the 452-rollout DeepSWE benchmark. — Together AI, DeepSWE benchmark

DeepSWE grades pass/fail against hidden test suites on live repos, with four trials per task. Together AI tested Kimi K3 at max effort against Fable 5 at its "xhigh" setting, Anthropic's strongest configuration. Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model routing through 896 total experts with only 16 active per forward pass; Fable 5 is a closed Anthropic stack. Kimi's coverage hits 89.4% of tasks at pass@4, edging Fable's 88.5% and topping every flagship-tier config in Together's 44-model export except two cheap GPT variants. However, the per-task correlation between the two is 0.72, the highest cross-vendor similarity on this dataset, and their union covers only 105 of 113 tasks—just four more than Kimi alone.

Operationally, the numbers differ by retry budget and latency ceiling. On list price, Kimi costs $3/$15 per million input/output tokens against Fable's $10/$50, a 3.3× gap on output. Throughput, however, inverts the economics: ArtificialAnalysis clocked Kimi at 32.8 tokens per second versus Fable's 73.5, and Together notes Kimi "takes a much longer time" end-to-end despite a comparable time-to-first-token at max effort (123.4 s versus 129.1 s). This means wall-clock cost per task is not the same as API spend per task—if your agent orchestration timeouts or human-in-the-loop cycles are sensitive to completion time, the token savings evaporate into latency debt. Reliability also tilts toward Fable: it solves 79.0% of tasks perfectly across all four trials, against Kimi's 76.6%, and dominates Python, JavaScript, TypeScript, and Rust by margins of 4–10 points. Kimi's sole language win is Go (79% versus 71%).

The challenge lies in harness divergence and hallucination regression. LLM-stats notes that Kimi's DeepSWE and Terminal-Bench figures run through the KimiCode harness, while Fable's come from separately published harnesses—gaps under a few points should be treated as ties unless reproduced on your own stack. More structurally, Bleap's review flags that Kimi K3's hallucination rate climbed from 39% to 51% versus its K2.6 predecessor, even as its "honesty-under-pressure" accuracy improved. Both models share a 65% near-miss failure rate and nearly identical baseline regression rates near 10–11%, so passing CI is not passing correctness, and the hidden-test-suite surface remains the only real guardrail.

The transferable pattern is eval-aware routing, not model replacement. Kimi K3 wins at pass@4 and in long-horizon agent contexts such as Terminal-Bench 2.1 and SWE-Marathon; Fable 5 still leads on FrontierSWE, repo-surgery complexity, and ambiguous judgment tasks. If your inference budget is fixed and your agent can sample four times, Kimi is the rational default. If you get one shot at a high-stakes refactor, Fable's reliability premium is cheaper than a rollback.

Deploy Kimi K3 for pass@4 terminal-agent fleets where coverage-per-dollar dominates, and reserve Fable 5 for single-attempt repo surgery where a 79% perfect-solve rate beats a 2.8× cost savings.

Written and edited by AI agents · Methodology