ToolNest

Month 1 ยท LLM Applications

Day 2 โ€” Tokens, Context Windows and the Cost of Everything

Published September 14, 2026

Every surprise on an LLM bill and every mysteriously truncated prompt traces back to the same three concepts: tokens, the context window, and pricing per token class. Day 2 builds the mental model and then verifies it with a script, because intuition about tokens is famously wrong โ€” especially across languages.

By the end of the day you should be able to look at any prompt, estimate its token count and cost within an order of magnitude, and explain exactly what happens when a conversation outgrows the window.

The short answer

A token is a word fragment produced by the model's tokenizer โ€” English averages roughly one token per short word, while Chinese often splits into several tokens per character-cluster. Cost equals tokens times price, input tokens are cheaper than output, and an overflowing context window means truncation or errors โ€” solved by summarizing or retrieving, never by hoping.

What a token actually is

Models do not read words or characters; they read tokens โ€” chunks of text produced by a byte-pair-encoding-style tokenizer that learned common fragments from training data. Common English words tend to be one token; rarer words split into pieces; whitespace often attaches to the following word.

The practical consequence is language-dependent pricing. Chinese text frequently decomposes into more tokens than the equivalent English, because tokenizers over-index on English corpora โ€” the same meaning can cost two or three times as much in Chinese. Testing this with a tokenizer library on both languages is the fastest way to make it stick.

Context windows and what happens when they overflow

The context window is everything the model can attend to in one call: system prompt, conversation history, retrieved documents, and the generation in progress. Exceeding it produces an error or silent truncation, depending on the provider โ€” neither is a design.

The engineering answers, in rising order of sophistication: truncate old turns (a sliding window), summarize compressed history and re-inject it, or move knowledge out of the window entirely and retrieve it on demand โ€” which is exactly the problem RAG, arriving later in this series, exists to solve.

  • Input tokens: everything sent to the model โ€” cheaper per token.
  • Output tokens: generated autoregressively, more compute โ€” priced higher.
  • Prompt caching: providers reuse computation for identical prefixes, which is why stable, long system prompts are cheap to repeat.

Today's hands-on task

Two small scripts anchor the concepts. First, a tokenizer script: feed it a sentence in English and one in Chinese carrying the same meaning, print both token counts, and compute the ratio. Second, a pricing table: for four or five mainstream models, note the input price, output price, and context window, then calculate what 1,000 conversations of a fixed shape would cost on each.

The deliverable is not the scripts โ€” it is the habit of estimating cost before calling. An agent that makes ten tool-loop calls per request multiplies every prompt overhead tenfold; you only see that if you count tokens by default.

How an interviewer asks about today

Token fluency signals production experience more than almost any buzzword.

Day 2 interview questions
QuestionWhat a strong answer covers
Why does Chinese cost more tokens than English?Tokenizer corpora skew English; Chinese characters split into more fragments โ€” same meaning, more tokens, higher cost and latency.
The context window is full โ€” what do you do?Sliding-window truncation, summarization of old turns, externalize knowledge to retrieval; name the trade-off each makes.
Why is output pricier than input?Generation is sequential and compute-heavy per token; input processing is parallel. Price reflects the compute shape.

Common mistakes on day two

  • Assuming one word equals one token โ€” estimates built on that are off by 30% in either direction, which compounds across agent loops.
  • Pasting the whole conversation into every call forever, then wondering why latency and cost creep. History is a design decision, not a default.
  • Comparing model prices without checking the context window โ€” a cheap model with a small window can force expensive workarounds.

Frequently asked questions

How do I estimate an LLM API cost before running anything?
Count tokens with a tokenizer for your real prompt shape, multiply by the model's input price, add the expected output tokens at the output price, then multiply by call volume โ€” remembering agent loops can multiply calls per request several times over.
Is a bigger context window always better?
No. Beyond cost, very long contexts dilute attention โ€” relevant details get lost in the middle. Retrieval that feeds only what matters usually beats a bigger window, and costs far less.
What is prompt caching and when does it help?
Providers can cache the computation for identical prompt prefixes. Long, stable system prompts benefit most; conversations that mutate their prefix every turn benefit least.

Part of a public learning journal โ€” general educational content, not professional advice. See our disclaimer.