The AI Cost Paradox

The unit price of Artificial Intelligence  (AI) continues to fall. Models are becoming more efficient, competition is increasing and the cost of processing a thousand tokens can be considerably lower than it was only a few years ago. The big AI companies have all announced recently that  they have lowered their token costs. Great news, right?

It is tempting to conclude that an organization’s total AI bill should fall. In practice, many businesses are discovering the opposite.

Lower token prices make more AI use cases economically viable. More employees gain access to Copilot, more applications incorporate AI, and more business processes are delegated to agents. Each individual interaction may cost less, but the number and complexity of those interactions can grow much faster than the unit price declines.

That is the AI cost paradox: AI becomes cheaper to use, so organizations use much more of it!

For Dynamics 365 customers, the important question is no longer, “What does a token cost?” It is, “How much AI work are we generating, and how much of that work creates business value?”

Why AI Consumption Continues to Rise

Early business AI use was relatively simple. A limited group of users might ask a Copilot to summarize a record, draft an email, or answer a straightforward question. These interactions typically involve one prompt and one response.

The current environment is very different. Organizations are consuming more AI through:

  • More frequent Copilot interactions across a larger user population
  • Agents that operate without waiting for a user prompt
  • Retrieval-Augmented Generation (RAG) searches across CRM records and documents
  • Multi-step reasoning and planning
  • Tool calls into Dynamics 365, Power Platform and external systems
  • Automated workflows triggered by AI decisions
  • Larger context windows containing more customer and business information

One business request can therefore generate multiple model calls, searches and actions. An agent asked to qualify a lead might retrieve CRM data, review documents, compare the lead with existing accounts, calculate a score, update a record, assign an owner and start a workflow. What looks like one interaction to the user may involve a chain of consumption behind the scenes

A Simple Scale Comparison

Imagine an organization’s AI use in 2024:

  • 100 users
  • 20 prompts per user each day
  • Mostly short prompts and simple completions
  • Limited retrieval or workflow activity

Now compare that with the same organization in 2026:

  • 1,500 users with access to AI
  • Multiple agents working across sales, service and operations
  • RAG searches across Dynamics 365 and connected knowledge sources
  • Tool calling and workflow actions
  • Longer prompts containing customer history, documents and instructions

Even if the effective cost per thousand tokens has fallen substantially, the organization is processing far more tokens, and paying for retrieval, storage, agent activity and connected services. Total spending can rise dramatically while the unit economics improve.

This is not necessarily a failure. Higher AI expenditure can be entirely justified when it generates greater productivity, faster service, or additional revenue. The problem arises when consumption grows without visibility, governance, or a clear connection to business value.


Understanding the Full AI Cost Picture

Token pricing remains important, but it is only one part of the cost model. Consider the following:

Prompt and Completion Tokens

Prompt tokens include the user’s question, system instructions, conversation history and any business information supplied to the model. Completion tokens are generated in the answer. Long instructions, repeated history and excessive retrieved content increase the amount processed during every interaction.

Context-window Consumption

A larger context window allows the model to consider more information at once. It does not make every piece of that information useful. Duplicated records, irrelevant activities and obsolete documents can expand the context without improving the answer.

Embedding, Indexing and Retrieval

RAG systems convert content into embeddings, store it in an index and search it when questions are submitted. Duplicated or low-value information can create recurring costs during preparation, storage, retrieval and prompt construction.

Agent and Workflow Execution

An agent may perform several reasoning steps and call multiple tools before completing one task. Actions in connected systems can also trigger Power Automate flows, APIs, or other consumption-based services. Agent cost should therefore be assessed by completed business outcomes, not just by visible user interaction.

Copilot Credits and Consumption-based Pricing

Microsoft increasingly relies on Copilot Credits as the primary currency for agent capabilities. Depending on licensing, organizations draw from either prepaid capacity or pay-as-you-go billing, with consumption driven by agent architecture, query volume, and feature complexity. While Microsoft offers native reporting and environment-level governance in Copilot Studio, specific meters and pricing models evolve regularly.

The strategic takeaway, however, remains constant: as AI shifts to consumption-based pricing, data inefficiency transforms directly into an operating expense.

The biggest problem is token waste

Organizations shouldn't aim to minimize AI interactions across the board. The goal is to eliminate consumption that yields no business value.

Common sources of waste include:

  • Repeated retrieval of duplicate customer records
  • Old or irrelevant activities included in prompts
  • Multiple versions of the same document
  • Poor metadata that causes broad, unfocused searches
  • Agents repeating unsuccessful steps
  • Overly large prompts and unnecessary conversation history
  • Users verifying answers because the underlying data is unreliable

The last item is particularly important. A technically inexpensive answer becomes costly when an employee must search Dynamics 365, compare records and correct the result before taking action.


How Dynamics 365 customers can respond

Effective AI cost governance begins with measurement. Organizations should monitor consumption by environment, agent, process and business outcome, not merely as one monthly total.

They should also identify which Dynamics 365 records and fields are repeatedly supplied to AI. Duplicate accounts and contacts, incomplete records, stale activities and inconsistent classifications can increase processing while weakening results.

Before scaling Copilot or autonomous agents, establish baseline metrics for consumption, response quality, and human-in-the-loop intervention. From there, prioritize cleaning the records and tightening the retrieval logic that drive high-frequency tasks. Spending caps and threshold alerts protect budgets from runaway surprises, but clean data and precise retrieval eliminate the root cause of compute waste.


Final Takeaway

Falling token prices do not guarantee a lower AI bill. Wider adoption, larger contexts, RAG, agents and workflow automation can increase total consumption far faster than unit costs decline.

The goal should not be to use less AI. It should be to ensure that AI processes the right information, takes purposeful actions and produces outcomes users can trust.

The biggest AI cost problem is no longer simply token pricing. It is token waste!