Each week we find a new topic for our readers to learn about in our AI Education column.
“You photograph something and then the photograph is split up into millions of tiny pieces and they go whizzing through the air down to your TV set where they’re all put together again in the right order.” – Mike Teavee, Willy Wonka & The Chocolate Factory.
For years, the Gene Wilder-starring version of Roald Dahl’s children’s classic, Charlie and the Chocolate Factory, was where I received my understanding of how broadcast television works. It wasn’t exactly right, of course, but it isn’t far off the mark—information recorded by broadcasters is subdivided into tiny bits which are then projected through the air itself as waves, picked up, and reconstituted on screens in people’s houses.
There is a similar process at work behind a lot of the software we use, even artificial intelligence, that is almost perfectly explained by little Mike Teavee, that 9-year-old character in a 55-year-old film: tokenization. This week’s AI Education is going to focus on AI tokens and why they’ve become so crucial to the internal workings and external economics of the technology.
Tokens sit at the intersection of how artificial intelligence works and how artificial intelligence is bought. They are pieces of information that models process, but they are also becoming units of consumption, measures of performance and components of enterprise technology bills. That combination is turning tokens into something resembling the kilowatt-hours of the AI economy: a way of measuring how much machine intelligence an organization consumes. One recent explanation of the emerging token economy makes precisely that analogy, noting that businesses are increasingly purchasing AI as a metered service.
For financial institutions, tokens are more than technical trivia. Understanding tokens helps explain why one AI application can cost dramatically more than another, why agentic AI can produce unexpectedly large bills, why model speed is often expressed in tokens per second, why enormous context windows are useful but potentially expensive and why financial organizations may eventually have to manage AI consumption with the same discipline they apply to cloud computing.
What Is a Token?
In computing generally, “token” has several meanings. A token can be a security credential, an authentication object or a unit produced when software parses a programming language. The common idea is that a token represents something else in a form that a computer system can efficiently recognize and manipulate. Artificial intelligence borrows that basic concept. An AI token is a small unit of information that a model processes.
For a large language model, tokens are usually pieces of text. A token can be an entire word, part of a word, punctuation, a space or another character sequence. The model does not receive a sentence in quite the same form that a human reads it. A tokenizer first breaks the material into tokens and converts those tokens into numerical identifiers that can be processed mathematically.
This process is tokenization. IBM describes it simply as converting words in a prompt into tokens and notes that there is no fixed word-to-token relationship: a word can become one token or several, spaces and punctuation can matter, and different models and languages can tokenize identical-looking material differently. A commonly used rule of thumb for English is that one token represents approximately four characters, or roughly three-quarters of a word. OpenAI similarly estimates that 100 tokens amount to about 75 English words, although the actual number depends on the content, language and tokenizer.
Tokens are not limited to ordinary text, either. Multimodal AI systems can tokenize images, audio, video and other information. NVIDIA describes tokens more broadly as units of data processed by AI systems during training and inference. Different tokenization methods can transform text, visual information, audio or other data into representations that models can manipulate. An AI token should not be confused with a cryptocurrency token. A crypto token is a blockchain-based digital asset. An AI token, in the sense discussed here, is a computational unit. The two share a name but not a technology or economic function.
How AI Uses Tokens
Tokens are fundamental to both training and inference. During training, enormous quantities of information are tokenized and fed into a model. A language model learns statistical relationships among those tokens. Simplifying considerably, it repeatedly attempts to predict tokens from the tokens that precede them and adjusts its internal parameters as it learns from errors.
Inference is what happens after the model has been trained and somebody actually uses it. A user’s prompt becomes input tokens. The model processes those tokens and predicts an appropriate continuation. Its response becomes output tokens. More sophisticated reasoning models complicate that picture. They can also consume reasoning tokens as they work through a problem before producing the answer visible to the user. Those tokens may never appear on the screen, meaning a short answer can represent considerably more computation than its visible length suggests. OpenAI’s documentation, for example, distinguishes input, output, cached and reasoning tokens and notes that reasoning tokens can count toward billed output usage.
This makes token counts an imperfect but useful approximation of AI work.
Imagine asking an AI assistant to summarize a 100-page investment report. The report, the instructions supplied to the model, relevant conversation history and possibly retrieved research all become input. The model then consumes additional computational resources to reason about that material and produces output tokens containing the summary. Every step has an economic consequence.
Tokens, Speed and Efficiency
Tokens are also central to measuring AI performance. One common metric is tokens per second: how quickly a model produces tokens during inference. Another is time to first token, measuring how long the user waits between submitting a request and receiving the beginning of a response. There is also inter-token latency—the delay between successive pieces of generated output.
Those measures affect the experience of using AI. A customer-service assistant that pauses conspicuously after every question may feel broken even if its answers are accurate. An investment-research system analyzing thousands of documents may be less sensitive to conversational latency but much more sensitive to total throughput.
NVIDIA consequently describes the economics of AI infrastructure partly in terms of producing more tokens at lower cost. Faster hardware, memory, networking and optimized software can increase token throughput and reduce the cost of generating each token. Efficiency, however, does not simply mean maximizing tokens per second. It can also mean using fewer tokens to accomplish the same useful work. Sending an entire 300-page document to a model when only three paragraphs are relevant consumes unnecessary input tokens. Asking an expensive frontier reasoning model to classify a simple transaction can waste computational capacity. Allowing an agent to repeatedly reread an enormous conversation history can similarly inflate usage.
Google’s 2026 guidance on AI tokenomics warns that bloated context increases cost and latency and can even reduce model effectiveness. It recommends techniques including careful context selection, task decomposition, reusable instructions and strict controls on autonomous loops. Thus, token efficiency is analogous to fuel efficiency. More consumption does not necessarily mean more useful work.
Token Limits and Context Windows
Every model also has limits on how many tokens it can handle. The context window is the amount of information the model can consider within a particular interaction. Depending on the architecture and implementation, that context can contain system instructions, the user’s prompt, previous conversation, documents supplied to the model, retrieved information, tool descriptions and generated output. A larger context window therefore gives an AI system more working material.
This can be enormously useful in finance. An AI with sufficient context might analyze a lengthy prospectus, several research reports, a client’s financial plan and portfolio information together. A smaller context window might require the application to divide, summarize or selectively retrieve that information. But context capacity should not be confused with an instruction to fill the window.
A one-million-token context window means the system can process an enormous quantity of information; it does not mean every query should contain one million tokens. More context can mean more cost and latency, while irrelevant context can make it harder for the model to focus on the important material. The related phrase token limit can refer to several different restrictions. It can mean the context-window limit, an output-token ceiling, a usage allocation or a provider’s rate limit, such as tokens permitted per minute. The important point is to determine which limit is being discussed rather than assuming all token limits mean the same thing.
How AI Companies Charge for Tokens
Tokens have become especially important because AI providers have turned them into billing units. Traditional enterprise software is often purchased through licenses, subscriptions or per-seat fees. Cloud computing introduced more granular consumption pricing for storage, processing and network resources. Generative AI takes usage pricing another step by metering the information a model processes.
A common API pricing structure separately charges for input tokens and output tokens, often quoting prices per million tokens. Output generally costs more because generation is computationally intensive and sequential. Cached input can be cheaper because information already processed by the system can sometimes be reused rather than recomputed from scratch. This produces a deceptively simple equation:
AI cost = input-token consumption + output-token consumption + other model or infrastructure charges.
In practice, things get more complicated. Different models carry different rates. Reasoning can add hidden computational consumption. Long conversations repeatedly incorporate earlier context. Retrieval-augmented generation can insert thousands of tokens of source material. Tool definitions consume context. Agentic systems can make multiple model calls while pursuing one user request. Some platforms hide these mechanics behind subscriptions or bundled enterprise contracts. Others expose them directly. IBM, for example, offers foundation-model services with pay-as-you-go pricing measured by token consumption while also offering hourly hosting arrangements for some deployments. That is why the sticker price per million tokens tells only part of the story.
Token Spend and Tokenomics
AI token spend is the money an organization spends consuming tokens across its AI systems. As enterprises deploy more copilots, agents, research tools, customer-service systems and automated workflows, that spend can spread across vendors, departments and applications.
The challenge increasingly resembles FinOps—the discipline developed to understand and govern variable cloud-computing expenditures. Deloitte argues that conventional technology total-cost-of-ownership models need adjustment because AI introduces highly variable consumption economics. Organizations may combine SaaS products, external APIs and self-hosted models, each with different token economics. Model choice, infrastructure, context size, reasoning intensity and usage patterns all influence the resulting cost.
That leads to tokenomics. The word already has another meaning in cryptocurrency, where tokenomics describes the economic structure of a digital token. In enterprise AI, however, the term increasingly describes the economics of producing and consuming AI tokens: what tokens cost, where they are being spent, what business value they generate and how an organization can optimize the relationship between consumption and outcomes. Tokenomics therefore goes beyond reducing bills. The real question is value per token.
An expensive workflow that saves thousands of employee hours may represent excellent token economics. A cheap chatbot nobody uses may represent terrible economics. Likewise, simply minimizing token usage can become counterproductive if the result is lower-quality analysis, inadequate context or repeated failures. The emerging enterprise tooling reflects this shift. On Oct. 6, 2026, Stacklet launched Token Custodian, designed to attribute token consumption to teams, projects, agents, applications and cost centers and apply policies to that spending. The product illustrates how token management is evolving from a developer concern into an enterprise-governance function.
Tokenmaxxing—and Token Baiting
One of the stranger consequences of this new economics is tokenmaxxing. Tokenmaxxing is informal technology-industry jargon for encouraging workers to consume more AI tokens on the theory that heavy AI usage indicates experimentation, adoption or productivity. IBM reports that the term surged in popularity during 2026 as some organizations began treating token consumption almost as an AI-adoption KPI. The problem is obvious: consumption is not productivity.
An employee who consumes ten times as many tokens may be doing ten times as much AI-enabled work—or may simply be using an inefficient workflow. Tokenmaxxing can create a perverse incentive similar to judging an employee’s productivity by electricity consumption. Research from the Stanford Digital Economy Lab reinforces that warning. Researchers examining agentic coding tasks found that agentic workloads could consume vastly more tokens than ordinary code-chat tasks, that repeated runs could differ by as much as 30-fold in consumption and, crucially, that higher token usage did not consistently translate into greater accuracy.
Token baiting, meanwhile, should be used much more cautiously. Unlike tokenmaxxing, it does not yet have a stable, broadly recognized technical definition across the AI industry. The phrase is sometimes used informally to describe behavior, product design or prompts that encourage unnecessarily large AI interactions—and therefore additional token consumption—but it should presently be treated as emerging jargon rather than an established AI concept.
If the term becomes more common, the important economic idea will be familiar to financial professionals: incentives matter. A provider paid according to consumption has an economic interest in consumption. Buyers therefore need to understand whether AI systems are engineered to maximize successful outcomes or merely activity.
Why Tokens Should Matter to Us
That question makes token economics particularly relevant to financial services. Banks, insurers, asset managers, broker-dealers and wealth managers are moving from experimental chatbots toward AI systems embedded in workflows. Those applications may summarize documents, analyze portfolios, prepare client communications, conduct surveillance, assist compliance personnel, research securities, process claims, monitor transactions or operate as autonomous agents. Each additional AI workflow creates another stream of token consumption.
For financial institutions, therefore, token governance should eventually become part of AI governance. Firms will need to know which models are being used, which departments are consuming them, how much those workflows cost, what information enters the context window and what measurable value comes out.
There is also a risk-management dimension. Financial firms cannot optimize purely for the cheapest token. A less expensive model that makes more errors in a compliance workflow may be economically disastrous. Conversely, using the most expensive frontier model for every mundane classification task wastes resources without necessarily improving results.
The objective is not minimum token spend. It is appropriate spending for the required accuracy, latency, security, explainability and business value. Financial institutions should therefore think in terms of token budgets, model routing, caching, context management and cost attribution. A routine document classification might go to a smaller model. A complicated financial-planning analysis might justify a more capable reasoning model. Retrieval systems can supply only relevant sections of documents rather than entire libraries. Repetitive instructions can sometimes be cached. Outputs can be constrained when verbosity adds no value. That is essentially capital allocation applied to artificial intelligence.
What About Wealth Managers?
For wealth management, token economics may be especially consequential because advisory work is context intensive. Good financial advice requires information: household circumstances, assets, liabilities, tax considerations, goals, risk tolerance, portfolio holdings, account history, planning assumptions, market information and sometimes years of client communications.
AI can potentially synthesize all of that. But the more context an advisory system receives, the more tokens it may consume. Imagine an AI assistant preparing for a client meeting. It might retrieve the client’s financial plan, CRM notes, portfolio holdings, previous meeting summaries, emails, research on major positions and recent market developments. An agent could then perform multiple reasoning steps, invoke analytical tools and draft talking points.
From the advisor’s perspective, the result might be a two-page briefing. From the AI system’s perspective, producing those two pages could require processing a very large number of tokens. That creates a new dimension for evaluating wealthtech. Advisory firms should ask not merely whether an AI application works, but how its economics scale across hundreds or thousands of clients. What is the token cost per household served? Per client meeting prepared? Per financial plan updated? Per compliance review? Per service request resolved?
Those measurements can ultimately be compared with outcomes: advisor hours saved, response times reduced, clients served per professional, new revenue generated, compliance errors avoided or client satisfaction improved. That is the more mature version of tokenomics.
Tokens, in other words, may become for AI what basis points are for asset management: a small unit that becomes enormously important when multiplied across an enterprise.The irony is that the goal should not be to maximize them. Nor should financial firms reflexively minimize them. The objective is to turn the right number of tokens, produced by the right model and supplied with the right context, into something more valuable than the resources consumed. For financial services, that is where token economics ultimately becomes ordinary economics.






