This article first ran in Cascade Insights® CEO Sean Campbell’s LinkedIn newsletter, The Human Side of AI. Subscribe there for weekly takes on AI and the people who use it.
Anyone who spent part of their childhood in a school computer lab playing The Oregon Trail learned a few hard lessons early. The frontier was dangerous, the river crossings were a gamble, and at some point someone in your party was going to die of dysentery. The game also taught an additional lesson, which is that the frontier was expensive. You spent your savings on oxen, food, and spare wagon axles before you left Independence, and a bad river crossing could wipe out half of what you brought with you in one turn.
That’s a useful frame for AI right now. If you’ve been watching your AI bill lately, you’d be forgiven for thinking AI is getting more expensive. Teams are burning through token limits, long agent runs cost real money, and I’ve had the moment myself where a complex task hits a limit and the question becomes whether to spend another fifty dollars to let it finish.
To me, the real story about AI spend at a team or small company level depends on what that money buys. Familiar capabilities are getting cheaper, while pushing into new territory can still cost plenty. And a cheap run can waste an expensive hour.
Cost per finished, usable piece of work is the measure that matters, including the time spent directing, reviewing, and correcting it. Otherwise, falling token prices can make wasted human attention harder to see.
The frontier was always expensive
The first settlers on the Oregon Trail paid a high price in money, time, and lives. The people who came later benefited from routes that had been scouted, with known water, established stops, and hard-won knowledge about which crossings to avoid. Somebody else had already paid for the learning.
Technology works the same way, and my first laptop is a good example. In 1998 I was teaching 3 and 5 day networking classes on Windows NT 4.0, and my company handed me a Toshiba that cost around $4,000. The exact model number has been lost to memory, but it was most likely a Tecra 8000. It had a Pentium II processor running at a few hundred megahertz, a hard drive measured in single-digit gigabytes, and it weighed about 7 pounds once you added the power brick and external floppy drive. That was serious hardware at the time, and the company paid serious money for it because the work required it, believe it or not, given those specs. Today a few hundred dollars buys far more capability at a fraction of the weight. Nobody who bought that Toshiba was wrong to buy it. They were paying the frontier premium at the time.
The real question is whether being early is worth it for the work you’re doing.
Yesterday’s frontier gets cheap, fast
The decline in AI prices is dramatic, but the unit matters. Stanford’s AI Index tracked the price of models matching GPT-3.5’s score on MMLU, a knowledge benchmark: from $20 per million tokens in November 2022 to seven cents by October 2024. That’s comparable performance on one test, rather than proof of equal quality on every task.
Epoch AI’s September 2026 analysis gets closer to the question a budget owner cares about. It estimates that the cost of achieving a fixed level of performance across five benchmarks fell about 13x per year since 2023. It counts the tokens used to achieve the result, including reasoning, rather than just the sticker price per token.
Earlier Epoch estimates of 50x or even 200x annual declines tracked per-token prices at benchmark thresholds. A separate study found roughly 5x to 10x declines in benchmark cost at fixed performance. These estimates use different data and methods. They aren’t separate price curves for cheap models and premium models. In this research, the cost-performance frontier means the cheapest way to reach a given score, which may involve a small model.
The practical lesson holds without projecting today’s task price two years into the future. What once required premium capability can become cheap enough for everyday use. That’s the hand-me-down frontier. You may get it from a new small model, an older model, or a lower reasoning setting. You’re inheriting capability, not necessarily using old software.
Cheaper help means more help
In 1865, the economist William Stanley Jevons noticed something counterintuitive about coal. As steam engines got more efficient, Britain didn’t use less coal. It used far more, because efficiency made coal worth using for things that hadn’t made sense before. That pattern now carries his name, Jevons paradox.
AI gives us plenty of reasons to expect a similar effect. When a capable model costs pennies, people start using it for work nobody would have paid a person to do, or even thought to automate: checking every contract clause, drafting a hundred variations, monitoring a competitor’s site every day. That’s one reason bills can go up while prices fall. Some of the interesting uses over the next few years will come from people doing smart things with inexpensive capability.
The frontier itself can still be pricey
Reasoning models can consume many more tokens per task, so a cheaper token doesn’t automatically mean a cheaper answer. The same price-performance researchers found that the cost of evaluating the best-performing available models rose roughly 3x to 18x per year across the benchmarks they studied. That measures the cost of chasing the best score as the best score keeps improving.
For work that genuinely needs the strongest model or a long agent run, the bill may hold steady or climb. Sometimes that extra fifty dollars buys a useful result you couldn’t otherwise get. Sometimes it buys another hour of circling the same problem.
What you get keeps growing, with a caveat
METR measures the difficulty of tasks an AI agent can complete, using how long those tasks take human experts. Its widely cited time horizon uses a 50 percent success rate, mainly on well-specified software, machine-learning, and cybersecurity tasks. The original trend doubled about every seven months from 2019 into 2025; its January 2026 update estimated about 4.3 months for the post-2023 period. That’s real progress, with considerable uncertainty. It doesn’t mean an agent reliably handles the equivalent stretch of your team’s messy working day.
Four patterns in what you pay and what you get
For a team trying to make sense of its AI spending, I see four recurring patterns. Judge them against your current workflow for the same kind of work, counting human time alongside the machine bill.
The Workhorse handles routine work like summarizing, classifying, and first drafts at a sensible cost. You don’t need everything the strongest model can do. You need enough capability to meet the task’s finish line reliably.
The Frontier Premium buys a result that cheaper options can’t deliver well enough. It’s the $4,000 Toshiba of the moment. When it’s pointed at the right problem, being early is worth it.
The Token Treadmill is the almost-good-enough loop: long runs circling the same problem, revision after revision, agents launched without enough context, and spend climbing while finished work doesn’t. It’s the Oregon Trail river crossing where you chose to ford, lost two oxen, and still ended up on the wrong bank. Tokenmaxxing lives here too, when usage itself becomes the goal.
The Hand-Me-Down Frontier delivers work that used to require premium spending at today’s lower cost. A workflow that was hard to justify becomes practical, or the same budget buys more usable work. The saving only counts if extra correction and review don’t swallow it.
The cost the token bill doesn’t show
A cheap model still needs a clear goal and a defined finish line. It can still produce work that’s almost good enough, revision after revision. The Token Treadmill exists at every price point. If anything, cheaper tokens make it easier to get on, because one more run feels free. A team running a workhorse model in circles can waste as much time as a team running the frontier in circles. They’ll just have a smaller machine bill, at least at the start, which makes the waste harder to spot.
The amount of checking depends on the work and the consequences of getting it wrong. Tests can catch some errors automatically. Better models and better workflows can reduce review time per result. But if the volume grows faster than the time saved, the total demand on people’s attention still rises.
How people think about the help matters too. BCG reported a study of more than 1,200 managers in which those given the “AI employee” framing identified 18 percent fewer errors. That warns us about misplaced trust; it doesn’t establish a fixed cost of supervision. Cheap output that nobody reviews well can move the cost somewhere harder to see.
Two organizations with the same AI bill can end up with very different results. One has made the work easier. The other has given everyone more to check.
Run the experiment
Here’s a simple way to find out where you actually stand. Pick a routine workflow, like summarizing calls, drafting follow-ups, analyzing customer support interactions, processing a set of financial transactions, or building content. Take a representative task, including the awkward edge cases, and define what an acceptable result looks like before comparing models.
Run that batch through your current model choices alongside cheaper (older) alternatives. Try a smaller model or less reasoning as well as an older model. Give each the context it needs and use the same acceptance criteria. Count failures, retries, AI cost, and the human minutes spent reviewing and correcting. Then divide the total cost by the number of finished results you’d actually use. An inexpensive run that needs three attempts and twenty minutes of repair and handholding may be the expensive option.
A second model can help check the work, but it can share the first model’s blind spots. Use source material, tests, or a qualified human to settle whether the result is right. If a mistake would be costly, make the checks match that risk.
Don’t make this a one-time exercise. Capability and pricing shift fast enough that January’s answer may be wrong by June, or February. Rerun the comparison when the work changes, or when a new model or price makes switching worth considering.
Relying on a single model for everything makes this harder, and arguably is somewhat foolhardy. It’s convenient, but without a comparison you don’t know where you’re overpaying or missing a better fit. You don’t need to keep a dozen models in rotation, but you shouldn’t rely on a single AI partnership either. You need to regularly test credible alternatives and have ways to deploy those to your team.
Do you know which pattern you’re actually in?
This is the question I think every leaders should be thinking through.
Managers need to know where their team’s work falls, task by task, and whether they’re rewarding finished work or visible busyness. Individual contributors need to ask it of themselves, too: is this run producing something I’ll use, or is it motion that feels like progress? The organizations that get this right will move faster in more meaningful ways than those who don’t.
Where this is headed
As more capability becomes affordable, the useful question is what it costs to get the work finished, what is the ROI for that work output.
The decision about when to pay the frontier premium deserves the same thought the settlers gave to when to leave Independence. Sometimes going first is worth it. Often, waiting for the route to be mapped gets you most of the value at a fraction of the cost.
What caught my eye in the world of AI this week
The Week of the Always-On Agent
- OpenAI introduces dots Always-on personal agents powered by GPT-6 Astra, each with its own cloud computer, rolling out first to ChatGPT Pro and Business Premium users.
- ChatGPT Space A shared workspace for people and their dots, announced at the same DevDay keynote.
- OpenAI is retiring custom GPTs Creation of new custom GPTs ended September 25, and existing ones are scheduled to stop running December 11, with plugins as the replacement. Every company that standardized on custom GPTs just inherited a migration project no one planned for…
- Meta launches Muse for Small Business New skills and connectors tie Meta’s Muse agent to Facebook and Instagram business accounts, ad accounts, and tools like Shopify, QuickBooks, Stripe, and Canva.
- Muse announces a batch of new features
- Introducing Manus 2.0 A rebuilt agent framework, a desktop app upgraded into Manus Studio with video editing and game development environments, and a new standalone personal agent app called Cue.
Rogue Agents, and Who Polices the Labs
- OpenAI agents posted 53 user images to the internet without the lab’s knowledge Research agents uploaded images ChatGPT users had provided to public image-hosting sites as unlisted links, and OpenAI says it can’t trace them back to the users who supplied them.
- OpenAI’s misalignment reports The running log of incidents published under the disclosure framework OpenAI launched September 16, alongside its statement that the industry hasn’t solved alignment well enough to keep scaling at maximum speed much longer.
- OpenAI parts ways with three researchers over information shared with an outside safety group OpenAI says the three violated its policies on handling sensitive information, and the WSJ reports they were safety researchers who shared confidential material with a third-party AI safety organization. Disclosure on the company’s terms.
- There are plenty of laws on the books to check the AI giants A New York Times opinion piece arguing that product liability, tort, and consumer protection law already apply to AI companies, and that state attorneys general should enforce them.
Accountants, Robots & Middle Managers
- Frontier models now outscore junior accountants, per Mercor Mercor’s human-baseline study reports AI models outperforming the accountants it tested on its benchmark tasks.
- Eighteen months ago the best models trailed the average accountant, per Andrew Curran Curran quotes the Mercor post: the best models once fell short of the average accountant’s roughly 37% score, they now ace the same tasks, and the authors say they considered not publishing.
- A new research finding (Ethan Mollick)
- What Work Can Robots Do?
Tools & Toys
- Give Opus 5.5 a dashboard before a long task A tip to have Opus 5.5 vibe-code an HTML dashboard before it runs a long task on its own, then hand it a dedicated dashboard-builder subagent set to medium effort.
- Numerous.ai An add-on for Google Sheets and Excel that runs AI prompts inside cells through spreadsheet functions.
- Destroy Any Website A free browser game by Hugo Duprez that turns any URL into a destructible pixel-art level.
Markets, Fraud & the Writing Debate
- State of Markets II The second edition of a16z’s State of Markets deck on AI and the private and public technology markets.
- An AI messaging scam cost Intesa’s private bank €95 million Fraudsters impersonated Intesa Sanpaolo’s CEO on WhatsApp and used a cloned lawyer’s voice to get Fideuram’s chairman to wire €95 million abroad, about €36 million of which is still missing. Business email compromise with a cloned voice attached.
- I’m a college professor, and writing isn’t as important as we think A professor’s case, in the Times opinion section, that higher education overrates writing.