We tested Kimi K3 and Fable on a real bug from the Cline repo, and found that while both models were able to fix it - Fable wins on speed & Kimi wins on cost.
- Kimi used 1.7x more tokens than Fable (1.2M vs. 730K)
- Fable finished 3.4x faster - 3.5 min and 18 tool calls vs. Kimi’s 12 min and 34 tool calls.
- Kimi cost 2.3x less ($0.92 vs. $2.13) thanks to its 3.3x per-token discount
Both runs used the same Cline harness, and the traces indicate that Kimi is RL trained to spend more tokens thinking and verifying before completing.
This is the first time we've seen an open weight model compete head to head with SOTA. Congratulations to the
@Kimi_Moonshot team on this milestone!