Course outline▾
Week 1 · The Foundations
- Day 1Demystifying AI — From Buzzword to Business Logic
- Day 2How Machines Actually Learn — Supervised, Unsupervised and Reinforcement Learning
- Day 3Inside Neural Networks — The Engine of Modern Deep Learning
- Day 4The AI Project Lifecycle — From Raw Data to Production Deployment
- Day 5The Math Behind the Magic — Why Linear Algebra and Probability Matter
- Day 6Data Preprocessing — Cleaning the Messy Reality of Enterprise Data
- Day 7Measuring Success — Understanding Accuracy, Precision, Recall and F1 Scores
Week 2 · Applied AI & APIs
- Day 8Introduction to LLMs & The Modern AI API Landscape
- Day 9Advanced Prompt Engineering — Few-Shot, Chain-of-Thought and Structured JSON
- Day 10Tokenization, Context Windows and Cost Optimization
- Day 11Embeddings & Vector Representations: How Machines Map MeaningComing soon · 2026-10-11 · 7 PM IST
- Day 12Vector Databases: Storing and Searching Enterprise KnowledgeComing soon · 2026-10-12 · 7 PM IST
- Day 13Retrieval-Augmented Generation (RAG): Chatting with Proprietary DocumentsComing soon · 2026-10-13 · 7 PM IST
- Day 14RAG Evaluation & Hallucination GuardrailsComing soon · 2026-10-14 · 7 PM IST
Week 3 · Infrastructure & Hosting
- Day 15Introduction to AI Infrastructure: Hardware, Runtimes, and ComputeComing soon · 2026-10-15 · 7 PM IST
- Day 16Local Model Execution: Running Open-Weight LLMs Securely (Ollama, vLLM)Coming soon · 2026-10-16 · 7 PM IST
- Day 17Containerizing AI Workloads: Writing Production Dockerfiles for Python APIsComing soon · 2026-10-17 · 7 PM IST
- Day 18Docker Compose for Multi-Container AI Stacks (Web UI, Vector DB, LLM Engine)Coming soon · 2026-10-18 · 7 PM IST
- Day 19Kubernetes for AI 101: Pods, Deployments, and Services for Model ServingComing soon · 2026-10-19 · 7 PM IST
- Day 20Persistent Storage in Kubernetes: Managing State, Weights, and Vector IndicesComing soon · 2026-10-20 · 7 PM IST
- Day 21High-Performance Networking: Configuring Ingress and Egress for AI ClustersComing soon · 2026-10-21 · 7 PM IST
Week 4 · Enterprise Workflows
- Day 22Autonomous Agents: From Passive LLMs to Goal-Driven ExecutionComing soon · 2026-10-22 · 7 PM IST
- Day 23Tool Use & Function Calling: Connecting LLMs to APIs, Databases, and ShellsComing soon · 2026-10-23 · 7 PM IST
- Day 24Multi-Agent Orchestration: Supervisor, Worker, and Evaluator PatternsComing soon · 2026-10-24 · 7 PM IST
- Day 25Automated CI/CD: Automating the Software Lifecycle for AI ModelsComing soon · 2026-10-25 · 7 PM IST
- Day 26Zero-Trust Security for AI: Sandboxing Ephemeral Execution & Model EgressComing soon · 2026-10-26 · 7 PM IST
- Day 27Observability & Tracing for Agentic Systems (Telemetry, Logs, and Metrics)Coming soon · 2026-10-27 · 7 PM IST
- Day 28Human-in-the-Loop Architecture: Machine Second, Human First in PracticeComing soon · 2026-10-28 · 7 PM IST
- Day 29Managing Technical Debt, Drift, and Model Governance in Enterprise ITComing soon · 2026-10-29 · 7 PM IST
- Day 30The 10-Year Horizon: Architecting IT Strategy for the AI-Native EnterpriseComing soon · 2026-10-30 · 7 PM IST
Day 10: Tokenization, Context Windows and Cost Optimization
2026-10-10 · 8 min read
Watch the video lesson, or subscribe on YouTube for a new lesson every day.
LLMs do not read letters or whole words. They read tokens: subword chunks of text that are converted into numbers. Understanding tokens matters because they decide two things at once: how much the model can remember, and how much your AI app costs to run.
In plain terms: Think of a taxi meter that counts in small chunks of text, not in kilometres. Every chunk you send, and every chunk the model writes back, adds to the fare.
As a rule of thumb, one token is about three quarters of an English word, so 100 tokens is roughly 75 words. Common words are a single token. Rare or compound words are split into pieces, and numbers, symbols and code often take several tokens each. Different models use different tokenizers, so the exact counts vary.
The context window
The context window is the model's short-term memory limit: the total number of tokens it can consider at once. A model with a 128,000-token window can hold roughly the equivalent of a few hundred pages of text in a single call. Everything counts towards that limit: the system instructions, the chat history, any documents you paste in, and the reply itself. Go past it and the oldest content is dropped, so the model appears to forget earlier instructions.
Token economics
Providers bill you for input tokens (what you send) and output tokens (what the model generates). Output tokens usually cost several times more, often three to four times, because producing text takes more GPU work than reading it. Prices differ by provider and change often, so always check the current price list.
Optimization
Small habits add up quickly at scale:
- Trim the history. Do not send the whole chat every time. Keep the recent turns, and summarize or drop old ones.
- Remove redundant whitespace and noise from injected text and code.
- Use prompt caching for static system instructions, so repeated long prompts cost less.
- Cap the answer length with a maximum-tokens setting.
- Route easy tasks to smaller models and keep the flagship model for hard problems.
Here is the same idea in Python: counting tokens with a real tokenizer, and estimating a call's cost. The prices are made-up units, only to show the shape of the calculation.
import tiktoken
enc = tiktoken.get_encoding("cl100k_base") # a common tokenizer; others split text differently
for text in ["Network latency is rising", "microservices"]:
tokens = enc.encode(text)
print(len(tokens), "tokens:", [enc.decode([t]) for t in tokens])
# Cost of one call. The prices below are made-up units: output costs 4x as much as input.
PRICE_IN, PRICE_OUT = 1.0, 4.0 # per 1,000 tokens (illustrative)
def call_cost(tokens_in: int, tokens_out: int) -> float:
return tokens_in / 1000 * PRICE_IN + tokens_out / 1000 * PRICE_OUT
print("2,000 in + 500 out:", call_cost(2000, 500)) # 2.0 for input + 2.0 for output
print("trim the history to 1,200 in:", call_cost(1200, 500))
Coming Up Next
Day 11: Embeddings and vector representations: how machines map meaning.
#Tokenization #GenerativeAI #CostOptimization #SoftwareArchitecture #CloudEconomics