Rogo's BigFinanceBench

Make additional thinking count.

Final-answer accuracy across output budgets and training checkpoints 550B reference · 32% 128k output tokens
Step 0 Step 10 Step 20 Step 30

Move across the plot to compare all four runs.

30B model, public-facing evaluation set. Output budgets are categorical. The 550B reference is a separate evaluation. Values transcribed from the reported figure ↗.