Why Do AI Costs Fall the More U.S. Businesses Use Them in 2026?
Most U.S. founders assume that scaling up AI usage means a scarier bill at the end of the month. DBS Bank CEO Piyush Gupta recently flipped that logic on its head, describing what he calls the 'token paradox': the more an organization uses AI, the cheaper each unit of that AI becomes. For American startups and mid-sized companies watching their OpenAI, Anthropic, or Google Cloud invoices climb, this is the counterintuitive insight that changes how you should budget for AI in 2026.
What is the Concept
The token paradox describes a simple but underappreciated dynamic in generative AI economics: as demand for AI models grows, providers compete harder, infrastructure gets optimized, and the cost per token (the basic unit of text an AI model processes) keeps falling. Gupta's point at DBS was that banks shouldn't fear AI adoption costs because heavier use accelerates the very efficiencies that bring those costs down, echoing the economic idea known as Jevons paradox, where increased efficiency actually drives higher overall consumption at lower unit cost.
For a U.S. business, this means the AI budget line isn't a fixed, ever-rising cost center the way a software license often is. It behaves more like cloud computing did in the 2010s: unit prices decline steadily even as total capability and usage go up, provided you're structuring contracts and workloads to capture that decline instead of locking into stale pricing.
Why It Matters in United States (2025–2026 Context)
American SMEs and enterprises have poured billions into generative AI pilots since 2023, and CFOs are now asking a hard question: does this scale profitably? The token paradox gives a data-backed answer. Major U.S. model providers, including OpenAI, Anthropic, and Google, have cut their per-token API pricing multiple times since GPT-4's 2023 launch, even as model capability improved. That trend directly affects how a Chicago logistics company or an Austin SaaS startup should plan its AI line item for 2026.
This matters most for founders and CTOs who paused AI rollouts over cost anxiety. If unit economics genuinely improve with scale, the businesses that lean in now, in cities like San Francisco, New York, and Denver, capture a widening cost advantage over competitors who wait, because vendor pricing and internal efficiency both compound in their favor.
How AI Is Changing This
Three forces are compressing AI costs simultaneously. First, model providers are shipping smaller, distilled models (like GPT-4o-mini class models) that match older flagship performance at a fraction of the compute cost. Second, inference hardware from Nvidia and custom silicon from hyperscalers like Google and Amazon is making each unit of computation cheaper to run. Third, intense competition among OpenAI, Anthropic, Google, Meta's open-weight models, and newer entrants keeps pricing pressure constant across the U.S. market.
The result is a compounding curve: better models, cheaper compute, and fiercer competition all push token prices down together. Businesses that route workloads intelligently, using cheaper models for simple tasks and premium models only where accuracy truly requires it, benefit fastest from this trend.
Real-World Examples
U.S. financial institutions are already living this pattern. JPMorgan Chase has scaled an internal generative AI assistant to hundreds of thousands of employees, betting that broader usage justifies the infrastructure investment as per-query costs decline. Bank of America's Erica virtual assistant has processed billions of client interactions, a volume that would have been cost-prohibitive under 2022-era AI pricing but is now routine. Klarna, though European, offers a widely cited case that American SaaS leaders reference directly: its AI customer service system handles the workload of hundreds of agents at a fraction of the per-interaction cost of two years earlier.
These examples mirror exactly what Gupta described at DBS: organizations that pushed usage up first are the ones now negotiating the best volume pricing and seeing the steepest cost declines per transaction.
Practical Insights / Actions
Apply what we call the Token Deflation Curve framework: instead of budgeting AI spend as a fixed monthly cost, track cost-per-completed-outcome (per resolved support ticket, per generated report, per qualified lead) and expect that number to fall quarter over quarter as you scale usage and providers compete for your volume. If it isn't falling, that's a signal you're locked into stale pricing or an inefficient model choice, not that AI is inherently expensive.
Three concrete moves for U.S. founders in 2026: negotiate usage-based API contracts with quarterly price reviews rather than annual locks; route routine tasks to smaller, cheaper models and reserve premium models for high-stakes work; and avoid single-vendor lock-in so you can shift workloads toward whichever provider is currently winning the token price war. RP SoftTech works with U.S. SMEs to audit exactly this kind of AI workflow spend and rebuild it around falling-cost architecture rather than fixed licensing.
Future Outlook
Expect the token paradox to keep playing out through 2027 as open-weight models from Meta and others put a hard ceiling on what closed providers can charge, and as inference-specific chips further cut compute costs. The competitive advantage for U.S. businesses will shift away from simply having access to AI, since that becomes commoditized, and toward how well a company integrates AI into proprietary workflows and data. Cost will stop being the barrier; execution will be.
Founders who treat 2026 as the year to scale usage, not retreat from it, are positioned to ride the falling-cost curve DBS's CEO described, turning what looks like a spending risk into a compounding operational advantage.
Conclusion
The token paradox isn't a banking-specific insight, it's a preview of how AI economics work everywhere, including for U.S. startups and SMEs weighing their 2026 budgets. The businesses that scale usage deliberately, track cost-per-outcome, and stay vendor-flexible will see their AI costs fall exactly as Gupta predicted. If you're unsure whether your current AI spend is following this curve or working against it, RP SoftTech offers a free AI cost audit to map your usage against where token pricing is heading next.
Frequently Asked Questions
What is the 'token paradox' in AI costs?
It's the observation, highlighted by DBS CEO Piyush Gupta, that as organizations use AI models more heavily, the cost per token (unit of AI processing) tends to fall due to provider competition, model efficiency gains, and cheaper compute, rather than rising as usage scales.
Why are AI costs falling for U.S. businesses in 2026?
Providers like OpenAI, Anthropic, and Google have repeatedly cut API pricing since 2023 while improving model quality, driven by competition, smaller efficient models, and cheaper inference hardware, making per-token costs lower even as total AI capability increases.
How can small businesses in the U.S. reduce AI token costs?
Route simple tasks to smaller, cheaper models and reserve premium models for complex work, negotiate usage-based API contracts with regular price reviews, and avoid locking into a single vendor so you can shift to whoever offers the best current pricing.
Will AI costs keep dropping in the future?
Most signals point to continued declines through 2026 and 2027 as open-weight models pressure closed-provider pricing and inference-specific chips reduce compute costs, though businesses still need efficient workflow design to fully capture those savings.