Home›Read›Tools desk›Reasoning Dial benchmark shows high effort spikes costs 3.4x for zero gain
Tools·Oct 11, 2026

Reasoning Dial benchmark shows high effort spikes costs 3.4x for zero gain

A benchmark of 1,800 runs across four models evaluates API reasoning_effort settings, showing steep token and cost increases with measurable accuracy gains limited strictly to hard logic tasks. Leave…

A benchmark of 1,800 runs across four models evaluates API reasoning_effort settings, showing steep token and cost increases with measurable accuracy gains limited strictly to hard logic tasks.

Leave reasoning_effort set to none or low for basic data extraction, formatting, and counting tasks. Cranking the reasoning dial to high increases API costs x1.5 to x3.4 without delivering accuracy gains across most task types. Maximize reasoning effort strictly for complex, multi-step constraint satisfaction logic, where models like gpt-5.4-mini improve from 15% to 97.5% accuracy. Engineering teams should audit reasoning parameters per API endpoint rather than applying high-reasoning defaults globally across services.

Methodology

This review evaluates Abeera Lodhi's Reasoning Dial benchmark published in October 2026. The test suite systematically varied the reasoning_effort parameter across four settings (none, low, medium, and high) over 1,800 graded API calls across four models. The evaluation utilized 60 code-generated questions scored by a deterministic grader.

The review scope covers published dataset artifacts, pre-registered hypothesis documents, and open-source execution scripts from Lodhi's GitHub repository at https://github.com/Abeera81/reasoning-dial and associated Kaggle benchmark pages. Measured data includes output token counts, API cost structures, response latencies, and accuracy across task tiers. This v0 review draws on Lodhi's published benchmark results; independent benchmarks by Founderr Pulse are pending. This review does not cover production latency under heavy concurrent load, long-context retrieval, or unstructured multi-turn conversations.

What It Does

Systematic parameter testing across four tiers

The benchmark sweeps reasoning_effort through none, low, medium, and high settings while holding prompt templates, system instructions, and deterministic temperature parameters constant. Lodhi organized the benchmarks into public Kaggle tasks (dial-none, dial-low, dial-medium, and dial-high) to isolate output token variance across identical prompts.

Exposing API-level implementation quirks

Testing revealed erratic provider behaviors when handling reasoning controls. One model generated x14 more tokens at high relative to none. Another model rejected the none parameter entirely, throwing an HTTP 400 error. A third model generated hidden reasoning tokens at none despite explicit parameter instructions to bypass reasoning loops.

Measuring token inflation on basic tasks

On direct factual and counting prompts, high reasoning effort triggered severe hallucinations. When asked to count tools in the input string "You have a chisel and a drill", gpt-5.4-mini set to high returned an answer of 10,016 while consuming thousands of reasoning tokens.

What's Interesting / What's Not

Statistically isolated accuracy wins

Across 12 model-by-task test cells, high produced a statistically significant accuracy gain in exactly one scenario: gpt-5.4-mini handling complex logic puzzles, jumping from 15% to 97.5% accuracy. On a seven-person, seven-day logic puzzle ("Who gives the talk on Friday?"), none returned an immediate incorrect answer ("Cleo") using 18 output tokens at $0.00024. At high, the model used 1,333 tokens costing $0.006 to correctly identify "Fay".

Cost multipliers without precision improvements

In 7 out of 12 cells, setting reasoning to high did not improve accuracy at all. Instead, cost per correct answer increased by x1.5 to x3.4. Developers enabling extended thinking settings on routine queries pay substantially higher API fees for zero performance benefit.

Where the benchmark falls short

The benchmark relies on 60 synthetic, code-generated questions. While deterministic prompts eliminate ambiguity during grading, they do not replicate full enterprise workloads such as multi-document summarization, repository-wide code edits, or conversational state tracking. Furthermore, because model API backends change rapidly, these token multipliers represent a point-in-time snapshot of provider behavior.

Pricing

The complete main benchmark run across 1,800 API calls cost $4.20. Individual API call costs depended heavily on reasoning settings. For gpt-5.4-mini on logic puzzles, a prompt at none consumed 18 tokens ($0.00024), whereas high consumed 1,333 tokens ($0.006). Because API vendors bill directly for generated reasoning tokens, high-effort settings elevate costs by x1.5 to x3.4 per correct response.

Verdict

Engineering teams should treat reasoning_effort as a specialized parameter reserved for verified reasoning bottlenecks. Enabling high is justified only on complex multi-step constraint logic where base accuracy is demonstrably low. For factual extraction, formatting, and standard text processing, keeping the parameter at none or low prevents token waste and eliminates unnecessary x1.5 to x3.4 cost multipliers.

What We'd Test Next

In future iterations, we will benchmark reasoning effort parameters across agentic multi-turn loops, large code refactoring pipelines, and latency distribution under high concurrency. We will also test whether structured prompt engineering can achieve equivalent accuracy on logic puzzles without incurring provider-side thinking token overhead.

The investor read

Abeera Lodhi's benchmark highlights a structural inefficiency in developer LLM spending: naive application of high reasoning settings inflates token billing by x1.5 to x3.4 without accuracy gains across standard task categories. For AI infrastructure startups, this signals strong demand for dynamic parameter optimization middleware that routes prompts based on task complexity rather than static API configurations. Tooling platforms offering intelligent prompt routing and reasoning parameter management can deliver immediate cost reductions to enterprise buyers.

Pull quote: “In 7 out of 12 cells, setting reasoning to high did not improve accuracy at all.”

Sources · how we verified
  1. I Turned the Reasoning Dial to 'High' on 4 Models. It Fixed One Thing and Billed Me for Everything. ↗

Every claim ties to a primary source. See our methodology.

Reported by the Riley desk on Founderr Pulse’s Tools beat. Every factual claim is tied to a primary source and linked; anything that can’t be stood up doesn’t run. Founderr (RIKHATH LLC) is the accountable publisher and corrects in place. How we work · About · File a correction.
R
Riley

The Riley desk covers tools — what founders are building with, switching to, and abandoning. Every claim is sourced and linked. Operated by Founderr (RIKHATH LLC) See the desk →

Founderr Pulse — free & independent. The desk for people who build & back.
Reasoning Dial benchmark shows high effort… · Founderr Pulse