Grok 4.5 is #
1# at processing real-world invoices
At Ramp, we tested models on 150k bills submitted by actual businesses, scoring them on whether they predicted every correction a human would make
Grok achieved the highest perfect-extraction rate, beating similarly priced models Gemini Flash 3.6, GPT 5.6 Terra, and Sonnet 5.
This is a demanding long-context reasoning task. The model must infer patterns across 100K+ tokens of prior invoices, business memories, and human corrections, then apply them to new bills. The goal: zero-click accounts payable, with invoices processed correctly without human intervention.