- Google produces reasoning ‘dial’ for Gemini 2.5 Flash.
- Allows developers to control how much AI processing occurs, potentially solving overthinking problem.
- Prioritisation of efficient reasoning over larger models.
Google’s latest AI enhancement tackles a costly problem: artificial intelligence that wastes resources pondering simple questions. The tech giant has added a “reasoning dial” to its Gemini 2.5 Flash model, allowing developers to control exactly how much computational effort the system spends ”thinking’ before generating answers. The practical solution addresses what industry experts describe as a growing inefficiency in advanced AI systems. When asked basic questions like “How many provinces does Canada have?” some reasoning models still burn through expensive computing resources as if solving complex engineering problems.
“The model overthinks,” admits Tulsee Doshi, Director of Product Management at Gemini. “For simple prompts, the model does think more than it needs to.”
Released on April 17, the new feature arrives as AI companies increasingly focus on “reasoning” capabilities – training models to work through problems logically rather than simply pattern-matching from training data. While this approach improves performance on complex tasks, it has created an unexpected side effect of computational waste.
The economics of AI overthinking
The reasoning dial confronts a practical business problem rather than just a technical one. According to Google’s documentation, when reasoning is fully activated, generating outputs becomes approximately six times more expensive.
For developers building commercial applications, this cost multiplier can quickly become prohibitive. Nathan Habib, an engineer at Hugging Face who studies reasoning models, sees the problem as widespread in the industry. “In the rush to show off smarter AI, companies are reaching for reasoning models like hammers even where there’s no nail in sight,” he explained to MIT Technology Review. Habib shared an example with MIT Technology Review of a leading reasoning model that, when asked to work through an organic chemistry problem, began sputtering “Wait, but…” hundreds of times, resembling a computational meltdown while driving up processing costs.
Kate Olszewska, who evaluates Gemini models at DeepMind, confirmed Google’s systems can experience similar issues, sometimes getting “stuck in loops” that consume resources without improving answers.
Sliding scale for thought

The sliding scale of ‘thought’ allows customisation based on task complexity:
- Simple translations or factual queries can run with minimal reasoning budget
- Probability calculations or scheduling problems benefit from moderate reasoning
- Complex engineering or programming challenges may warrant maximum reasoning resources
Jack Rae, principal research scientist at DeepMind, acknowledges the boundaries remain fuzzy: “It’s really hard to draw a boundary on, like, what’s the perfect task right now for thinking.” The ambiguity makes the adjustable dial approach particularly valuable for real-world applications.
Shifting the AI development paradigm
The feature potentially marks a pivot in how AI companies approach model improvements. Instead of the traditional focus on increasing model size – which has dominated development since 2019 – Google is optimising how models use existing capabilities.
“Scaling laws are being replaced,” Habib notes, suggesting that future advances may come from more efficient reasoning rather than ever-larger models. The implications extend beyond technical performance to environmental impact. As reasoning models gain popularity, their energy consumption grows accordingly. Google’s approach could help mitigate what researchers have identified as a concerning trend: inferencing (generating responses) now contributes more to AI’s carbon footprint than the initial training process.
Competitive landscape
Google isn’t alone in pursuing reasoning capabilities. DeepSeek R1, an “open weight” model released earlier this year, triggered market volatility by demonstrating powerful reasoning at potentially lower costs. Unlike Google’s proprietary system, DeepSeek makes its internal settings publicly available for developers to run locally.
Despite this competition, Google DeepMind’s chief technical officer Koray Kavukcuoglu maintains that proprietary models will retain advantages in specialised domains: “Coding, math, and finance are cases where there’s a high expectation from the model to be very accurate, to be very precise, and to be able to understand complex situations.”
Future implications
The reasoning dial exemplifies a maturing AI industry grappling with practical constraints rather than just pushing technical boundaries. Google’s solution acknowledges that efficiency matters as much as raw performance. “Reasoning is the key capability that builds up intelligence,” states Kavukcuoglu. “The moment the model starts thinking, the agency of the model has started.”
For businesses deploying AI solutions, the ability to fine-tune reasoning budgets could democratise access to advanced capabilities and maintain control over cost. Google claims Gemini 2.5 Flash delivers “comparable metrics to other leading models for a fraction of the cost and size” – a proposition that becomes more compelling when developers can optimise reasoning resources for specific applications.
Author
View all postsDashveenjit is an experienced tech and business journalist with a determination to find and produce stories for online and print daily. She is also an experienced parliament reporter with occasional pursuits in the lifestyle and art industries.