Home › Glossary

What is an AI token?

Plain-English definition · Updated 2026-10-06. Numbers dated; verify with the vendor.

AITop is an independent guide. Prices and features change fast — always check the vendor's page.

The 30-second answer

A token is the unit of text that AI models actually read and write — roughly three-quarters of an English word — and it is the unit every metered AI price is counted in. API bills, context limits, rate limits, "credits": underneath all of them is a token counter. In 2026, mid-tier model APIs run roughly $0.15–$3 per million input tokens, and the trap is on the output side: reasoning models can burn 3–10× more tokens thinking than a plain chat ever did.

What it actually means

Models don't read letters or words; they read tokens — chunks of text produced by a tokenizer, common enough to get their own entry. "Unbelievable" might be three tokens; "the" is one. A practical rule of thumb for English: 1,000 words ≈ 1,300 tokens. Everything a model does — your prompt in, its answer out, any documents you attach — is counted in these units, both for pricing and for capacity limits.

The word shows up in three different price tags, and conflating them is how bills surprise people. Input tokens (what you send) are usually 3–5× cheaper than output tokens (what the model writes). And increasingly there is a third, invisible line: reasoning tokens — the scratch-work a "thinking" model generates before answering, which you pay for even though you never see it.

Why it matters when you're picking a tool

Because "free" and "unlimited" are both token statements in disguise. A flat subscription buys you tokens with softer edges: rate limits, usage caps and throttling that all start counting when the vendor decides you've had your share. A metered API bills exactly what you spend — predictable per unit, unpredictable per month. Neither is a scam; they are different bets on your own usage, and you should pick the one your workload actually fits.

Three questions settle it: How heavy is your use (a daily summary is tokens-light; an agent that reads ten documents per answer is tokens-heavy)? Do you need cost predictability (finance) or cost efficiency (developers)? And what happens when you hit the cap — throttle, overage, or a polite lockout?

The 2026 reality check

Two shifts define token economics right now. First, reasoning models: since late 2024 the serious models "think" before answering, and that thinking is billed — real-world agent tasks routinely consume several times more output tokens than the visible answer contains. Second, pricing stratified: budget models now handle bulk work at a few cents per million tokens while frontier reasoning commands dollars, and vendors like Perplexity expose their own metered rates on top (its Sonar API starts around $1 per million tokens — see our AI search comparison). The per-unit price of intelligence keeps falling; the total bill keeps rising, because agents use so much more of it.

Quick checklist

Where you'll hit it

Tokens are the hidden axis of almost every comparison on this site: Perplexity vs Google AI Mode vs ChatGPT Search (where API access is priced per token), the best AI chatbots ranking (where flat plans hide different caps), and Claude Code vs Cursor (where agent coding burns more tokens than any chat ever will). Related term: the full glossary.

The bottom line

A token is a quarter-word with a price tag. Learn to see it and every AI invoice becomes readable: input cheap, output dear, thinking billed invisibly. Choose flat plans for human-paced chatting, meters for machines that work while you sleep — and never evaluate an AI price without asking what it's counting.