Course outline▾
Week 1 · The Foundations
- Day 1Demystifying AI — From Buzzword to Business Logic
- Day 2How Machines Actually Learn — Supervised, Unsupervised and Reinforcement Learning
- Day 3Inside Neural Networks — The Engine of Modern Deep Learning
- Day 4The AI Project Lifecycle — From Raw Data to Production Deployment
- Day 5The Math Behind the Magic — Why Linear Algebra and Probability Matter
- Day 6Data Preprocessing — Cleaning the Messy Reality of Enterprise Data
- Day 7Measuring Success — Understanding Accuracy, Precision, Recall and F1 Scores
Week 2 · Applied AI & APIs
- Day 8Introduction to LLMs & The Modern AI API LandscapeComing soon · 2026-10-08 · 7 PM IST
- Day 9Advanced Prompt Engineering: Few-Shot, Chain-of-Thought, and Structured JSONComing soon · 2026-10-09 · 7 PM IST
- Day 10Tokenization, Context Windows, and Cost OptimizationComing soon · 2026-10-10 · 7 PM IST
- Day 11Embeddings & Vector Representations: How Machines Map MeaningComing soon · 2026-10-11 · 7 PM IST
- Day 12Vector Databases: Storing and Searching Enterprise KnowledgeComing soon · 2026-10-12 · 7 PM IST
- Day 13Retrieval-Augmented Generation (RAG): Chatting with Proprietary DocumentsComing soon · 2026-10-13 · 7 PM IST
- Day 14RAG Evaluation & Hallucination GuardrailsComing soon · 2026-10-14 · 7 PM IST
Week 3 · Infrastructure & Hosting
- Day 15Introduction to AI Infrastructure: Hardware, Runtimes, and ComputeComing soon · 2026-10-15 · 7 PM IST
- Day 16Local Model Execution: Running Open-Weight LLMs Securely (Ollama, vLLM)Coming soon · 2026-10-16 · 7 PM IST
- Day 17Containerizing AI Workloads: Writing Production Dockerfiles for Python APIsComing soon · 2026-10-17 · 7 PM IST
- Day 18Docker Compose for Multi-Container AI Stacks (Web UI, Vector DB, LLM Engine)Coming soon · 2026-10-18 · 7 PM IST
- Day 19Kubernetes for AI 101: Pods, Deployments, and Services for Model ServingComing soon · 2026-10-19 · 7 PM IST
- Day 20Persistent Storage in Kubernetes: Managing State, Weights, and Vector IndicesComing soon · 2026-10-20 · 7 PM IST
- Day 21High-Performance Networking: Configuring Ingress and Egress for AI ClustersComing soon · 2026-10-21 · 7 PM IST
Week 4 · Enterprise Workflows
- Day 22Autonomous Agents: From Passive LLMs to Goal-Driven ExecutionComing soon · 2026-10-22 · 7 PM IST
- Day 23Tool Use & Function Calling: Connecting LLMs to APIs, Databases, and ShellsComing soon · 2026-10-23 · 7 PM IST
- Day 24Multi-Agent Orchestration: Supervisor, Worker, and Evaluator PatternsComing soon · 2026-10-24 · 7 PM IST
- Day 25Automated CI/CD: Automating the Software Lifecycle for AI ModelsComing soon · 2026-10-25 · 7 PM IST
- Day 26Zero-Trust Security for AI: Sandboxing Ephemeral Execution & Model EgressComing soon · 2026-10-26 · 7 PM IST
- Day 27Observability & Tracing for Agentic Systems (Telemetry, Logs, and Metrics)Coming soon · 2026-10-27 · 7 PM IST
- Day 28Human-in-the-Loop Architecture: Machine Second, Human First in PracticeComing soon · 2026-10-28 · 7 PM IST
- Day 29Managing Technical Debt, Drift, and Model Governance in Enterprise ITComing soon · 2026-10-29 · 7 PM IST
- Day 30The 10-Year Horizon: Architecting IT Strategy for the AI-Native EnterpriseComing soon · 2026-10-30 · 7 PM IST
Day 7: Measuring Success — Understanding Accuracy, Precision, Recall and F1 Scores
2026-10-07 · 13 min read
Watch the video lesson, or subscribe on YouTube for a new lesson every day.
Saying that an AI model is "99% accurate" sounds impressive, and in real companies it is often a dangerous illusion. Imagine a busy network switch that drops packets on only 0.1% of its traffic. A lazy model that always predicts "no dropped packets" would be 99.9% accurate, yet it would never find a single fault. To judge an AI fairly, the measurement has to match what a mistake really costs.
In plain terms: A smoke alarm that never rings is right almost every day of the year, because most days there is no fire. It is also useless. Accuracy alone cannot tell you that, which is why we need better measures.
The Confusion Matrix
Every yes-or-no prediction a model makes falls into one of four boxes, laid out as a small two-by-two grid called the confusion matrix:
- True positive (TP): something really happened, and the model caught it. For example, an outage occurred and was flagged.
- True negative (TN): nothing happened, and the model correctly stayed quiet.
- False positive (FP), a false alarm: the model raised an alert when nothing was wrong. Statisticians call this a Type I error.
- False negative (FN), a silent failure: something real happened and the model missed it. This is a Type II error.
Here is a model that flags suspicious packets, tested on 100,000 packets of which 100 were really dropped:
Precision and Recall
Two questions come out of that grid, and they pull in different directions.
- Precision asks: when the model raises an alarm, how often is it right? It equals true positives divided by everything the model flagged (TP ÷ (TP + FP)). Aim for high precision when false alarms are expensive, for example an email spam filter, where wrongly hiding an executive's email disrupts the business.
- Recall asks: of all the real events, how many did the model find? It equals true positives divided by all real events (TP ÷ (TP + FN)). Aim for high recall when missing an event is serious, for example medical diagnosis, fraud detection, a failing piece of critical infrastructure, or flagging invalid EEPROM memory data on a network transceiver before it causes an outage.
In our packet example the model flagged 120 packets, and 80 of them were real: precision is 80 ÷ 120 = 67%. There were 100 real drops and the model caught 80: recall is 80%.
The trade-off
You usually cannot have both at their best. Make the alarm stricter and false alarms fall (precision goes up), but more real faults slip through (recall goes down). Make it more sensitive and you catch more, but you also raise more false alarms. Where you set the dial is a business decision.
The F1 Score
When the classes are very unbalanced, as with rare faults, people often want one number that combines precision and recall. That is the F1 score, the harmonic mean of the two:
F1 = 2 × (precision × recall) ÷ (precision + recall)
Why not just take the ordinary average? Because the harmonic mean punishes lopsided results. Suppose a model has 98% precision but only 4% recall. A simple average gives 51%, which sounds acceptable, but the F1 score is only 7.7%, which honestly says that this model misses almost everything.
The same numbers in Python
Here is the packet example as a few lines of Python. It is plain Python, and I have run it:
# 100,000 network packets, 100 of them really dropped
tp, fn, fp, tn = 80, 20, 40, 99_860 # a model that flags suspicious packets
accuracy = (tp + tn) / (tp + tn + fp + fn)
precision = tp / (tp + fp) # of the alerts, how many were real?
recall = tp / (tp + fn) # of the real faults, how many did we catch?
f1 = 2 * precision * recall / (precision + recall)
print(f"accuracy {accuracy:.2%}") # 99.94%
print(f"precision {precision:.1%}") # 66.7%
print(f"recall {recall:.1%}") # 80.0%
print(f"F1 {f1:.1%}") # 72.7%
# The lazy model that always says "no fault" scores 99.9% accuracy but its recall is 0%.
print(f"lazy accuracy {(99_900) / 100_000:.1%}, lazy recall {0 / 100:.0%}")
Which one should you use?
- If false alarms are costly, watch precision.
- If misses are costly, watch recall.
- If events are rare and both mistakes matter, use the F1 score, and always look at the confusion matrix itself, not only one number.
- Treat plain accuracy with care. It is only trustworthy when the classes are roughly balanced and every mistake costs about the same.
Coming Up Next
Day 8: Introduction to LLMs and the modern AI API landscape.
#MachineLearning #DataScience #MLOps #ModelEvaluation #AIInfrastructure #TechEducation #DataAnalytics #EnterpriseAI