Reasoning Dial benchmark shows high effort spikes costs 3.4x for zero gain
A benchmark of 1,800 runs across four models evaluates API reasoning_effort settings, showing steep token and cost increases with measurable accuracy gains limited strictly to hard logic tasks. Leave…
A benchmark of 1,800 runs across four models evaluates API reasoning_effort settings, showing steep token and cost increases with measurable accuracy gains limited strictly to hard logic tasks.
Leave reasoning_effort set to none or low for basic data extraction, formatting, and counting tasks. Cranking the reasoning dial to high increases API costs x1.5 to x3.4 without delivering accuracy gains across most task types. Maximize reasoning effort strictly for complex, multi-step constraint satisfaction logic, where models like gpt-5.4-mini improve from 15% to 97.5% accuracy. Engineering teams should audit reasoning parameters per API endpoint rather than applying high-reasoning defaults globally across services.
Methodology
This review evaluates Abeera Lodhi's Reasoning Dial benchmark published in October 2026. The test suite systematically varied the reasoning_effort parameter across four settings (none, low, medium, and high) over 1,800 graded API calls across four models. The evaluation utilized 60 code-generated questions scored by a deterministic grader.
The review scope covers published dataset artifacts, pre-registered hypothesis documents, and open-source execution scripts from Lodhi's GitHub repository at https://github.com/Abeera81/reasoning-dial and associated Kaggle benchmark pages. Measured data includes output token counts, API cost structures, response latencies, and accuracy across task tiers. This v0 review draws on Lodhi's published benchmark results; independent benchmarks by Founderr Pulse are pending. This review does not cover production latency under heavy concurrent load, long-context retrieval, or unstructured multi-turn conversations.
What It Does
Systematic parameter testing across four tiers
The benchmark sweeps reasoning_effort through none, low, medium, and high settings while holding prompt templates, system instructions, and deterministic temperature parameters constant. Lodhi organized the benchmarks into public Kaggle tasks (dial-none, dial-low, dial-medium, and dial-high) to isolate output token variance across identical prompts.
Exposing API-level implementation quirks
Testing revealed erratic provider behaviors when handling reasoning controls. One model generated x14 more tokens at high relative to none. Another model rejected the none parameter entirely, throwing an HTTP 400 error. A third model generated hidden reasoning tokens at none despite explicit parameter instructions to bypass reasoning loops.
Measuring token inflation on basic tasks
On direct factual and counting prompts, high reasoning effort triggered severe hallucinations. When asked to count tools in the input string "You have a chisel and a drill", gpt-5.4-mini set to high returned an answer of 10,016 while consuming thousands of reasoning tokens.
What's Interesting / What's Not
Statistically isolated accuracy wins
Across 12 model-by-task test cells, high produced a statistically significant accuracy gain in exactly one scenario: gpt-5.4-mini handling complex logic puzzles, jumping from 15% to 97.5% accuracy. On a seven-person, seven-day logic puzzle ("Who gives the talk on Friday?"), none returned an immediate incorrect answer ("Cleo") using 18 output tokens at $0.00024. At high, the model used 1,333 tokens costing $0.006 to correctly identify "Fay".
Cost multipliers without precision improvements
In 7 out of 12 cells, setting reasoning to high did not improve accuracy at all. Instead, cost per correct answer increased by x1.5 to x3.4. Developers enabling extended thinking settings on routine queries pay substantially higher API fees for zero performance benefit.
Where the benchmark falls short
The benchmark relies on 60 synthetic, code-generated questions. While deterministic prompts eliminate ambiguity during grading, they do not replicate full enterprise workloads such as multi-document summarization, repository-wide code edits, or conversational state tracking. Furthermore, because model API backends change rapidly, these token multipliers represent a point-in-time snapshot of provider behavior.
Pricing
The complete main benchmark run across 1,800 API calls cost $4.20. Individual API call costs depended heavily on reasoning settings. For gpt-5.4-mini on logic puzzles, a prompt at none consumed 18 tokens ($0.00024), whereas high consumed 1,333 tokens ($0.006). Because API vendors bill directly for generated reasoning tokens, high-effort settings elevate costs by x1.5 to x3.4 per correct response.
Verdict
Engineering teams should treat reasoning_effort as a specialized parameter reserved for verified reasoning bottlenecks. Enabling high is justified only on complex multi-step constraint logic where base accuracy is demonstrably low. For factual extraction, formatting, and standard text processing, keeping the parameter at none or low prevents token waste and eliminates unnecessary x1.5 to x3.4 cost multipliers.
What We'd Test Next
In future iterations, we will benchmark reasoning effort parameters across agentic multi-turn loops, large code refactoring pipelines, and latency distribution under high concurrency. We will also test whether structured prompt engineering can achieve equivalent accuracy on logic puzzles without incurring provider-side thinking token overhead.
The investor read
Abeera Lodhi's benchmark highlights a structural inefficiency in developer LLM spending: naive application of high reasoning settings inflates token billing by x1.5 to x3.4 without accuracy gains across standard task categories. For AI infrastructure startups, this signals strong demand for dynamic parameter optimization middleware that routes prompts based on task complexity rather than static API configurations. Tooling platforms offering intelligent prompt routing and reasoning parameter management can deliver immediate cost reductions to enterprise buyers.
Pull quote: “In 7 out of 12 cells, setting reasoning to high did not improve accuracy at all.”
Every claim ties to a primary source. See our methodology.