30 Days of AIDay 10 of 30 · Week 2: Applied AI & APIs
Course outline▾

Week 3 · Infrastructure & Hosting

  1. Day 15Introduction to AI Infrastructure: Hardware, Runtimes, and ComputeComing soon · 2026-10-15 · 7 PM IST
  2. Day 16Local Model Execution: Running Open-Weight LLMs Securely (Ollama, vLLM)Coming soon · 2026-10-16 · 7 PM IST
  3. Day 17Containerizing AI Workloads: Writing Production Dockerfiles for Python APIsComing soon · 2026-10-17 · 7 PM IST
  4. Day 18Docker Compose for Multi-Container AI Stacks (Web UI, Vector DB, LLM Engine)Coming soon · 2026-10-18 · 7 PM IST
  5. Day 19Kubernetes for AI 101: Pods, Deployments, and Services for Model ServingComing soon · 2026-10-19 · 7 PM IST
  6. Day 20Persistent Storage in Kubernetes: Managing State, Weights, and Vector IndicesComing soon · 2026-10-20 · 7 PM IST
  7. Day 21High-Performance Networking: Configuring Ingress and Egress for AI ClustersComing soon · 2026-10-21 · 7 PM IST

Week 4 · Enterprise Workflows

  1. Day 22Autonomous Agents: From Passive LLMs to Goal-Driven ExecutionComing soon · 2026-10-22 · 7 PM IST
  2. Day 23Tool Use & Function Calling: Connecting LLMs to APIs, Databases, and ShellsComing soon · 2026-10-23 · 7 PM IST
  3. Day 24Multi-Agent Orchestration: Supervisor, Worker, and Evaluator PatternsComing soon · 2026-10-24 · 7 PM IST
  4. Day 25Automated CI/CD: Automating the Software Lifecycle for AI ModelsComing soon · 2026-10-25 · 7 PM IST
  5. Day 26Zero-Trust Security for AI: Sandboxing Ephemeral Execution & Model EgressComing soon · 2026-10-26 · 7 PM IST
  6. Day 27Observability & Tracing for Agentic Systems (Telemetry, Logs, and Metrics)Coming soon · 2026-10-27 · 7 PM IST
  7. Day 28Human-in-the-Loop Architecture: Machine Second, Human First in PracticeComing soon · 2026-10-28 · 7 PM IST
  8. Day 29Managing Technical Debt, Drift, and Model Governance in Enterprise ITComing soon · 2026-10-29 · 7 PM IST
  9. Day 30The 10-Year Horizon: Architecting IT Strategy for the AI-Native EnterpriseComing soon · 2026-10-30 · 7 PM IST
Open the course page →

Day 10: Tokenization, Context Windows and Cost Optimization

2026-10-10 · 8 min read

Watch the video lesson, or subscribe on YouTube for a new lesson every day.

LLMs do not read letters or whole words. They read tokens: subword chunks of text that are converted into numbers. Understanding tokens matters because they decide two things at once: how much the model can remember, and how much your AI app costs to run.

In plain terms: Think of a taxi meter that counts in small chunks of text, not in kilometres. Every chunk you send, and every chunk the model writes back, adds to the fare.

LLMs read tokens, not letters "Network latency is rising"Network12950 ·latency40841 ·is374 ·rising18717 4 tokens"microservices"micro21962 services8173 2 tokens"The server restarted at 03:15."The791 ·server3622 ·restarted61454 ·at520 ·220 032839 :25 15868 .13 9 tokensReal counts, from a standard tokenizer. · marks a leading space.Rare words split into pieces: micro + services.
A tokenizer splits text into subword chunks and maps each to a number. Common words are one token, rare or compound words split into pieces, and numbers and punctuation often take several.

As a rule of thumb, one token is about three quarters of an English word, so 100 tokens is roughly 75 words. Common words are a single token. Rare or compound words are split into pieces, and numbers, symbols and code often take several tokens each. Different models use different tokenizers, so the exact counts vary.

The context window

The context window is the model's short-term memory limit: the total number of tokens it can consider at once. A model with a 128,000-token window can hold roughly the equivalent of a few hundred pages of text in a single call. Everything counts towards that limit: the system instructions, the chat history, any documents you paste in, and the reply itself. Go past it and the oldest content is dropped, so the model appears to forget earlier instructions.

The context window is the model's short-term memory 0128,000 tokens rulesold chat historydocumentsnowWhen the window is full, the oldest text is dropped. The model forgets it.128,000 tokens is roughly a few hundred pages of text.A rule of thumb: 1 token is about three quarters of an English word.Everything counts: the rules, the history, the documents and the answer.
The context window is how much the model can hold at once. Instructions, chat history, documents and the reply all count. When it overflows, the oldest text is dropped and effectively forgotten.

Token economics

Providers bill you for input tokens (what you send) and output tokens (what the model generates). Output tokens usually cost several times more, often three to four times, because producing text takes more GPU work than reading it. Prices differ by provider and change often, so always check the current price list.

You pay for input and output, and output costs more One call: 2,000 tokens in, 500 tokens out Tokensinput 2,000output 500 Cost (if output is 4× the price) inputoutputOutput is 20% of the tokens but about half of the bill.Writing costs more than reading. Prices vary, so check your provider.
Providers bill separately for input and output tokens, and output tokens usually cost several times more. In this illustrative call, output is a fifth of the tokens but about half of the cost.

Optimization

Small habits add up quickly at scale:

  • Trim the history. Do not send the whole chat every time. Keep the recent turns, and summarize or drop old ones.
  • Remove redundant whitespace and noise from injected text and code.
  • Use prompt caching for static system instructions, so repeated long prompts cost less.
  • Cap the answer length with a maximum-tokens setting.
  • Route easy tasks to smaller models and keep the flagship model for hard problems.
Five ways to cut the bill Trim old chat history−18%Remove extra whitespace−6%Cache the static system prompt−20%Cap the answer length−12%Route easy tasks to a small model−14%monthly bill Illustrative savings. Measure your own traffic before you tune.
Trimming chat history, removing whitespace, caching the static prompt, capping answer length and routing easy tasks to smaller models each take a bite out of the bill. The percentages here are illustrative.

Here is the same idea in Python: counting tokens with a real tokenizer, and estimating a call's cost. The prices are made-up units, only to show the shape of the calculation.

import tiktoken

enc = tiktoken.get_encoding("cl100k_base")          # a common tokenizer; others split text differently
for text in ["Network latency is rising", "microservices"]:
    tokens = enc.encode(text)
    print(len(tokens), "tokens:", [enc.decode([t]) for t in tokens])

# Cost of one call. The prices below are made-up units: output costs 4x as much as input.
PRICE_IN, PRICE_OUT = 1.0, 4.0                       # per 1,000 tokens (illustrative)
def call_cost(tokens_in: int, tokens_out: int) -> float:
    return tokens_in / 1000 * PRICE_IN + tokens_out / 1000 * PRICE_OUT

print("2,000 in + 500 out:", call_cost(2000, 500))   # 2.0 for input + 2.0 for output
print("trim the history to 1,200 in:", call_cost(1200, 500))

Coming Up Next

Day 11: Embeddings and vector representations: how machines map meaning.

#Tokenization #GenerativeAI #CostOptimization #SoftwareArchitecture #CloudEconomics