The free lunch was useful. It was never the business model.

For the past few years, artificial intelligence has arrived inside organizations under unusually forgiving economics.

Capabilities were bundled into existing software. Premium models were subsidized. "Unlimited" often meant that somebody else was absorbing the marginal cost. An employee could summon increasingly powerful computation without anyone asking what each request cost.

That arrangement helped AI spread.

It also taught organizations the wrong lesson.

The emerging pricing models from major enterprise software vendors point toward a different future: premium models, agents, and high-value inference increasingly measured and charged according to use. Oracle is already moving in this direction. SAP is following. Workday is preparing its own transition.

The free lunch was useful.

It was never the business model.

And as the subsidy disappears, something deeper than an IT budgeting problem comes into view.

Organizations are beginning to exchange capacity they own for capacity they rent.

That distinction matters.

Signal

Labor has traditionally been expensive but relatively legible.

Hire another person and the organization understands, within reasonable bounds, the annual cost of that additional capacity. Salary, benefits, equipment, management overhead, and perhaps some expected productivity.

AI rearranges the equation.

An agent may cost almost nothing while lightly used and become expensive when embedded across thousands of transactions. Its consumption changes according to workload, model selection, context length, retries, tool calls, reasoning depth, and the number of other agents involved.

The marginal cost is tiny.

Until it is multiplied by everything.

Companies are already discovering this dynamic as experimental AI programs become operational systems. Budgets expected to last a year can disappear in months because adoption itself changes the economics.

A useful agent gets used more.

A better model gets asked to do more.

Automation makes previously uneconomic tasks cheap enough to perform.

Then new tasks appear because the capability exists.

This is the inference paradox: falling unit costs do not necessarily reduce spending. They can expand the universe of economically sensible computation so quickly that total consumption rises instead.

The cheaper intelligence becomes, the more intelligence we consume.

Pattern

This is not primarily a software-pricing story.

It is a capacity-planning story.

Consider what happens when an organization replaces part of a human workflow with an AI agent.

At first, the economics look obvious.

A task that required an hour now requires five minutes. A report that consumed a day appears in seconds. A developer finishes work faster. A consultant processes ten documents instead of two.

Productivity rises.

But something else has happened quietly.

A fixed portion of organizational capacity has been converted into a variable one.

The organization may no longer need the same labor hours. But it now depends upon infrastructure whose price, availability, model behavior, rate limits, and usage patterns are controlled elsewhere.

The organization has traded one dependency for another.

That is not necessarily a bad trade.

It is simply a trade that should be visible.

Yet most organizations still place AI primarily inside the technology budget. Tokens sit beside SaaS licenses. Copilots appear beside Microsoft or Salesforce subscriptions. Model access becomes another line on the software spreadsheet.

That accounting treatment obscures what AI is actually replacing.

If AI is performing work, its economics belong beside the economics of work.

Tokens are becoming a form of labor capacity.

Not identical to labor. Not interchangeable with it.

But increasingly part of the same equation.

Implication

This changes how organizations should think about both AI and workforce design.

Stacia Garr's recent HBR framework offers a useful starting point: understand actual AI economics, prepare for pricing volatility, and bring AI spending into workforce planning rather than leaving it isolated inside IT.

The larger implication is that organizations need to understand their AI elasticity.

Not simply:

How much are we spending on AI?

But:

How much AI consumption can this process economically support before the underlying decision changes?

Imagine an agent that currently costs $0.18 to perform a task that previously required $8 of labor.

That sounds extraordinary.

But the important number is not $0.18.

The important number is the threshold.

Would the process still make sense at $0.50?

At $1?

At $3?

What if the organization begins running it ten times more frequently because the cost initially appears negligible?

What if one agent begins calling four others?

What if a frontier model becomes unnecessary for 80 percent of those requests?

What if the organization eliminates the human capacity required to perform the work manually?

Those questions reveal the actual system.

The objective is not perfect forecasting. That is impossible.

The objective is establishing boundaries.

Once organizations understand those boundaries, AI spending becomes manageable in the same way other variable inputs are manageable.

There can be thresholds.

There can be routing.

There can be substitution.

There can be contingency.

There can be a decision rule.

The Load-Bearing Agent

Another issue appears as AI moves deeper into operations.

Some agents become load-bearing.

A company experiments with an AI workflow for six months. It works. Employees adapt around it. Processes change. Documentation disappears because the agent "knows" what to do. Eventually nobody remembers exactly how the original workflow operated.

Then something changes.

The model price increases.

A rate limit appears.

A vendor changes behavior.

A model is retired.

An API fails.

The organization discovers that what began as an efficiency tool has quietly become infrastructure.

This is where resilience matters.

Critical AI workflows should have the same questions asked of them that we ask of other infrastructure:

What happens if it disappears tomorrow?

Can the process degrade gracefully?

Can another model perform the work?

Can a human temporarily step back in?

Is the knowledge required to do that still inside the organization?

The goal is not to maintain duplicate staffing for every automated process.

It is to avoid building brittle systems.

A resilient organization should be able to throttle, route, substitute, or humanize important workflows without operational free-fall.

The New Capacity Stack

The most interesting organizations may ultimately stop thinking in binary terms about humans versus AI.

Instead, they will manage a portfolio of capacity.

Some work will remain human-led because judgment, trust, accountability, ambiguity, or relationship matter.

Some work will become AI-led with human supervision.

Some work will become high-volume automation where inexpensive models are perfectly sufficient.

And some tasks will occasionally justify frontier intelligence because the incremental judgment is worth the incremental cost.

This begins to resemble an infrastructure stack:

  • Human judgment
  • Frontier models
  • Efficient commercial models
  • Open-weight or locally hosted models
  • Deterministic automation

The mistake is sending everything to the top.

Organizations that route every task through the most capable model available may discover that they have recreated the classic cloud-computing problem: extraordinary infrastructure used without economic discipline.

The winners will route intelligence according to the value of the decision.

Use expensive reasoning where reasoning matters.

Use cheap inference where scale matters.

Use software where intelligence is unnecessary.

Use people where responsibility matters.

That is organizational design.

Action

The next step is not another AI strategy workshop.

It is making the economics visible.

Start with a handful of important workflows.

Measure what AI is actually consuming and what work it replaces or accelerates. Calculate the economics at current prices, then identify the maximum cost the workflow could tolerate before the decision changes.

Make those thresholds explicit.

Give teams visibility into consumption.

Introduce soft limits before hard ones are necessary.

Route routine work toward cheaper capable models.

Reserve frontier models for places where their additional judgment changes the outcome.

Identify AI systems that have already become load-bearing and create a credible fallback for each.

And during the next workforce-planning cycle, put AI capacity on the same page as employees, contractors, outsourcing, and automation.

Not because tokens are people.

Because they increasingly compete for the same work.

Undersong

The coming adjustment in AI pricing is sometimes described as a cost shock.

That may be true.

But it is also a useful correction.

Subsidized AI encouraged organizations to experiment without asking too many questions about the meter. That was probably necessary. Entire categories of work had to be explored before anyone could know where the technology belonged.

Now the meter is becoming visible.

That is not the end of the AI opportunity.

It is the beginning of understanding what the opportunity actually costs.

The organizations best positioned for the next phase will not necessarily be those using the most AI.

They will be the ones that understand where intelligence creates value, how much that intelligence is worth, and what happens when its price changes.

The free lunch was always temporary.

The more important question is whether the organization has learned how to read the check.