What a Token Is on Your AI Coding Assistant Bill

A Token Is Roughly Four Characters and Every One Is Billed
A token is a chunk of text about four characters long, or about three quarters of an English word. Anthropic states the rough estimate plainly as 1 token to 4 characters, or 0.75 words in English.
If you pay per token, the arithmetic is direct. A 50,000-token input on Claude Opus 5 costs $0.25 at $5 per million tokens, and 15,000 output tokens cost $0.375 at $25 per million.
Subscription tools hide that arithmetic behind credits. GitHub Copilot Pro is $10 per month for 1,500 monthly AI credits, and Cursor Pro is $20 per month with usage-based billing past the included amount.
So the number that moves your bill is not the question you type. It is the volume of code the assistant reads before it answers.
Two Meters Are Running and They Are Not the Same Meter
- ● Subscriptions meter credits per tier
- ● APIs meter input and output tokens
- ● Credit value is set by the vendor
Every AI coding tool sits on a model that charges per token. What differs is whether the vendor passes that unit through to you or wraps it in something else.
Subscription tools wrap it. GitHub documents Copilot Pro at $10 per month, with 1,000 base AI credits plus a 500 credit flex allowance. Copilot Pro+ is $39 per month for 7,000 total credits, and Copilot Max is $100 per month for 20,000.
Team tiers follow the same shape. Copilot Business seats run $19 per user per month with 1,900 credits, and Enterprise seats run $39.
API access passes the unit through instead. You see input tokens, output tokens, and cache tokens as separate line items, and any charge traces back to a specific request.
The difference shows up on a bad day. A credit plan degrades to blocked or slower requests when the allowance runs dry, while a token plan simply keeps billing.
| Product | What the meter counts | Entry paid price | What happens past the limit |
|---|---|---|---|
| GitHub Copilot Pro | AI credits per month | $10 per month for 1,500 credits | Allowance resets monthly |
| GitHub Copilot Pro+ | AI credits per month | $39 per month for 7,000 credits | Allowance resets monthly |
| GitHub Copilot Business | AI credits per seat | $19 per seat per month | Pooled per user at 1,900 credits |
| Cursor Pro | Agent request limits | $20 per month | Usage-based billing continues |
| Claude Pro | Session usage caps | $20 per month, or $17 on annual | Limits reset on a rolling window |
| Claude API | Input, output, and cache tokens | Pay as you go, no seat fee | Billing continues per token |
Figures come from the GitHub Copilot plans documentation, the Cursor pricing page, and claude.com as of Aug 2026. Confirm current pricing on the official site before you budget.
Your Codebase Is the Expensive Part Not Your Question
- ● A 10 kB page is about 2,500 tokens
- ● Tool definitions add 286 to 675 tokens
- ● The bash tool adds another 325 tokens
Developers reach for shorter prompts when a bill looks high. That instinct targets the smallest line on the invoice.
Anthropic publishes useful reference points for input volume. An average web page of 10 kB is roughly 2,500 tokens, a 100 kB documentation page is roughly 25,000 tokens, and a 500 kB research paper is roughly 125,000.
Source files land in the same range. A 40 kB module pulled into context is a five figure token count before the model writes a single character of reply.
Tool definitions add a quieter surcharge. Declaring tools adds between 286 and 675 input tokens depending on the model, the bash tool adds about 325 tokens, and the text editor tool adds about 700.
Server side tools bill separately again. Web search on the Claude API runs $10 per 1,000 searches, on top of the tokens those results consume once they land in your context.
Caching Changes the Arithmetic More Than Model Choice
Prompt caching is the largest lever most developers never touch. It stores a processed chunk of your prompt so the next request reads it instead of reprocessing it.
The published multipliers make the tradeoff easy to check. A 5-minute cache write costs 1.25x the base input price, a 1-hour write costs 2x, and a cache read costs 0.1x.
Run the break-even yourself. A 5-minute cache pays for itself after a single read, and a 1-hour cache pays for itself after two.
Anthropic gives a worked example on Opus 5. A one-hour session with 50,000 input and 15,000 output tokens costs $0.705, and the same session with 40,000 of those input tokens served from cache costs $0.525.
That is a 25 percent cut with no change to the model or the work. Batch processing stacks on top, with a 50 percent discount on both input and output for jobs that can wait.
A Tokenizer Change Can Move Your Bill Without You Changing Anything
Tokens are not a fixed physical unit. They are the output of a tokenizer, and vendors change tokenizers between model versions.
Anthropic flags this directly. Claude 4.7 and later models use a newer tokenizer that produces about 30 percent more tokens for the same text.
A per-million rate is therefore only half of a price comparison. The same file can meter differently on two models from the same vendor.
Rate cuts move the other way. Claude Sonnet 5 sits at $2 per million input and $10 per million output, against $3 and $15 for Sonnet 4.6.
The budgeting lesson is to measure your own workload after any model switch. Treat a published rate as one input to the estimate rather than the estimate itself.
Three Numbers That Turn Tokens Into Dollars
You can forecast a monthly bill with three figures and no spreadsheet gymnastics.
The first is context size per request, meaning how much code you send. The second is requests per working day. The third is your cache hit rate, since cached input bills at a tenth of the base rate.
Multiply context size by requests, split the result into cached and uncached shares, then apply the per-million rates. Output tokens matter less than most people expect, because replies are short next to the context that produced them.
Try it against a realistic day. Twenty requests at 60,000 input tokens each is 1.2 million tokens, which costs $6.00 on Opus 5 at $5 per million with no caching. The same day costs about $1.02 if 80 percent of that input arrives as cache reads.
Which Billing Model Fits the Way You Actually Code
- ● Steady daily use - flat subscription
- ● Bursty agent runs - metered tokens
- ● Shared team context - caching first
Steady daily editing in one language: A flat subscription wins. Copilot Pro at $10 per month or Cursor Pro at $20 per month costs less than metered access for predictable inline work.
Occasional deep agent runs: Metered tokens win. Long multi-file sessions drain a credit allowance quickly, and a pay as you go account bills only the sessions you actually run.
A team sharing one large repository: Fix caching before you shop for a plan. Shared system prompts and repository context are exactly the material a 1-hour cache is built for, at 2x to write and 0.1x to read.
A solo developer watching spend closely: Start metered for a month to learn your real context size. Then switch to a subscription once you know whether your usage sits above or below the seat price.
A regulated or budget-capped environment: Choose the plan with per-request visibility. Token line items can be attributed to a project, while pooled credits usually cannot.
What to Check Before You Pick a Plan
Read the allowance definition rather than the headline price. A plan advertising a credit count is telling you about its own unit, not about tokens.
Check whether overage blocks or bills. Copilot resets a monthly allowance while Cursor continues on usage-based billing, and those two failure modes suit different teams.
Check which models the tier reaches. Cheaper tiers often route to smaller models, and a $1 per million model against a $5 per million model changes the economics more than the seat fee does.
Our guides on Copilot pricing and Claude Code pricing walk through those tier details, and the per-seat versus usage-based comparison covers the same decision at team scale.
The Unit Behind Every Line on the Invoice
A token is small, mechanical, and uninteresting on its own. It becomes interesting once you notice that your assistant reads thousands of them for every one you write.
That asymmetry is the whole story of an AI coding bill. Context is the cost, output is the rounding error, and caching is the discount hiding in plain sight.
If your bill surprised you this month, measure what goes into the prompt before you shorten what comes out. Our breakdown of what prompt caching means for coding costs is the natural next read, and what assistants see in your repository explains where all that context comes from.
FAQ
How many characters are in one token?
A token is roughly four characters of text, or about three quarters of an English word. Anthropic gives 1 token as approximately 4 characters or 0.75 words in English, though the exact count varies by language and content type.
What uses more tokens, my question or my codebase?
Almost always the input side. Your question might be 30 tokens while the files and conversation history sent with it run into tens of thousands. That is why prompt caching matters more than shortening your prompts.
Are AI credits the same thing as tokens?
Credits are the vendor unit for a subscription tier. GitHub Copilot Pro includes 1,500 monthly AI credits at $10 per month, and what one credit consumes depends on the model and the request. Tokens are the raw unit the model itself meters.
Does prompt caching actually lower the bill?
Yes, when the same context is reused. On the Claude API a cache read costs 0.1x the base input price, so a 5-minute cache pays for itself after one read and a 1-hour cache after two reads.
Can the same code cost more tokens on a different model?
It can. Claude 4.7 and later models use a newer tokenizer that produces about 30 percent more tokens for the same text, so identical work can meter differently across model versions even at the same per-million rate.
Sources
- Anthropic docs: prompt caching — checked 2026-09-07
- GitHub Copilot plans — checked 2026-09-07
- Anthropic pricing — checked 2026-09-07
Some links may be affiliate links. We may earn a commission at no extra cost to you.
This article was written with AI assistance. It is researched and fact-checked, not based on personal hands-on testing unless explicitly stated.
Comments
Post a Comment