30 Days of AIDay 7 of 30 · Week 1: The Foundations
Course outline▾

Week 3 · Infrastructure & Hosting

  1. Day 15Introduction to AI Infrastructure: Hardware, Runtimes, and ComputeComing soon · 2026-10-15 · 7 PM IST
  2. Day 16Local Model Execution: Running Open-Weight LLMs Securely (Ollama, vLLM)Coming soon · 2026-10-16 · 7 PM IST
  3. Day 17Containerizing AI Workloads: Writing Production Dockerfiles for Python APIsComing soon · 2026-10-17 · 7 PM IST
  4. Day 18Docker Compose for Multi-Container AI Stacks (Web UI, Vector DB, LLM Engine)Coming soon · 2026-10-18 · 7 PM IST
  5. Day 19Kubernetes for AI 101: Pods, Deployments, and Services for Model ServingComing soon · 2026-10-19 · 7 PM IST
  6. Day 20Persistent Storage in Kubernetes: Managing State, Weights, and Vector IndicesComing soon · 2026-10-20 · 7 PM IST
  7. Day 21High-Performance Networking: Configuring Ingress and Egress for AI ClustersComing soon · 2026-10-21 · 7 PM IST

Week 4 · Enterprise Workflows

  1. Day 22Autonomous Agents: From Passive LLMs to Goal-Driven ExecutionComing soon · 2026-10-22 · 7 PM IST
  2. Day 23Tool Use & Function Calling: Connecting LLMs to APIs, Databases, and ShellsComing soon · 2026-10-23 · 7 PM IST
  3. Day 24Multi-Agent Orchestration: Supervisor, Worker, and Evaluator PatternsComing soon · 2026-10-24 · 7 PM IST
  4. Day 25Automated CI/CD: Automating the Software Lifecycle for AI ModelsComing soon · 2026-10-25 · 7 PM IST
  5. Day 26Zero-Trust Security for AI: Sandboxing Ephemeral Execution & Model EgressComing soon · 2026-10-26 · 7 PM IST
  6. Day 27Observability & Tracing for Agentic Systems (Telemetry, Logs, and Metrics)Coming soon · 2026-10-27 · 7 PM IST
  7. Day 28Human-in-the-Loop Architecture: Machine Second, Human First in PracticeComing soon · 2026-10-28 · 7 PM IST
  8. Day 29Managing Technical Debt, Drift, and Model Governance in Enterprise ITComing soon · 2026-10-29 · 7 PM IST
  9. Day 30The 10-Year Horizon: Architecting IT Strategy for the AI-Native EnterpriseComing soon · 2026-10-30 · 7 PM IST
Open the course page →

Day 7: Measuring Success — Understanding Accuracy, Precision, Recall and F1 Scores

2026-10-07 · 13 min read

Watch the video lesson, or subscribe on YouTube for a new lesson every day.

Saying that an AI model is "99% accurate" sounds impressive, and in real companies it is often a dangerous illusion. Imagine a busy network switch that drops packets on only 0.1% of its traffic. A lazy model that always predicts "no dropped packets" would be 99.9% accurate, yet it would never find a single fault. To judge an AI fairly, the measurement has to match what a mistake really costs.

100,000 network packets, 100 of them dropped Lazy model: always says "no fault"Accuracy99.9%Faults caught0 of 100 Useful model: flags suspicious packetsAccuracy99.94%Faults caught80 of 100 99.9% accurate and completely useless.When faults are rare, accuracy hides the problem.
A lazy model that always predicts 'no fault' scores 99.9% accuracy while catching nothing. A useful model has about the same accuracy but actually finds most of the faults.

In plain terms: A smoke alarm that never rings is right almost every day of the year, because most days there is no fire. It is also useless. Accuracy alone cannot tell you that, which is why we need better measures.

The Confusion Matrix

Every yes-or-no prediction a model makes falls into one of four boxes, laid out as a small two-by-two grid called the confusion matrix:

  • True positive (TP): something really happened, and the model caught it. For example, an outage occurred and was flagged.
  • True negative (TN): nothing happened, and the model correctly stayed quiet.
  • False positive (FP), a false alarm: the model raised an alert when nothing was wrong. Statisticians call this a Type I error.
  • False negative (FN), a silent failure: something real happened and the model missed it. This is a Type II error.

Here is a model that flags suspicious packets, tested on 100,000 packets of which 100 were really dropped:

Model says: faultModel says: normalReal faultReal normal True positivefault caught80A real fault, and the model caught it. False negativefault missed20A real fault the model missed: a silent failure. False positivefalse alarm40A false alarm: nothing was wrong. True negativecorrectly quiet99,860Normal traffic, correctly left alone. The confusion matrixFour outcomes sum to 100,000 packets.
Every prediction lands in one of four boxes: a fault caught, a fault missed, a false alarm, or normal traffic correctly left alone.

Precision and Recall

Two questions come out of that grid, and they pull in different directions.

  • Precision asks: when the model raises an alarm, how often is it right? It equals true positives divided by everything the model flagged (TP ÷ (TP + FP)). Aim for high precision when false alarms are expensive, for example an email spam filter, where wrongly hiding an executive's email disrupts the business.
  • Recall asks: of all the real events, how many did the model find? It equals true positives divided by all real events (TP ÷ (TP + FN)). Aim for high recall when missing an event is serious, for example medical diagnosis, fraud detection, a failing piece of critical infrastructure, or flagging invalid EEPROM memory data on a network transceiver before it causes an outage.
Real faults: 100Flagged: 12020missed80caught40false alarms Precision: of everything flagged, 80 of 120 were real = 67%Recall: of all real faults, 80 of 100 were caught = 80%Precision and recall, as two questionsSame overlap (80), two different denominators.
Precision asks how many of the model's alerts were real. Recall asks how many of the real events the model caught. Both use the same overlap but divide by different totals.

In our packet example the model flagged 120 packets, and 80 of them were real: precision is 80 ÷ 120 = 67%. There were 100 real drops and the model caught 80: recall is 80%.

The trade-off

You usually cannot have both at their best. Make the alarm stricter and false alarms fall (precision goes up), but more real faults slip through (recall goes down). Make it more sensitive and you catch more, but you also raise more false alarms. Where you set the dial is a business decision.

How strict the alarm is →scorelenientstrict RecallPrecision Lenient: catches nearly everything, but many false alarmsStrict: few false alarms, but real faults slip throughYou choose the balance to fit the cost of each mistakeThe precision and recall trade-off
Making an alarm stricter raises precision but lowers recall, and the other way round. The right balance depends on what each kind of mistake costs.
Which mistake costs more? A false alarmFavour precisionspam filtersauto-blocking A missed eventFavour recallfraudmedical, outages Both matterUse F1 scorerare eventsa balance Pick the metric that matches the real-world cost.
Choose the metric by asking which mistake hurts more: false alarms point to precision, missed events point to recall, and rare events with both risks point to the F1 score.

The F1 Score

When the classes are very unbalanced, as with rare faults, people often want one number that combines precision and recall. That is the F1 score, the harmonic mean of the two:

F1 = 2 × (precision × recall) ÷ (precision + recall)

Why not just take the ordinary average? Because the harmonic mean punishes lopsided results. Suppose a model has 98% precision but only 4% recall. A simple average gives 51%, which sounds acceptable, but the F1 score is only 7.7%, which honestly says that this model misses almost everything.

Precision 98%, recall 4%: how good is it? Precision98% Recall4% Simple average51% F1 score7.7% The simple average says 51%, which sounds fair.F1 says 7.7%, which is the truth.F1 = 2 × precision × recall ÷ (precision + recall)
The F1 score is a harmonic mean, so it punishes a model that is strong on one measure and weak on the other. A simple average would hide the problem.

The same numbers in Python

Here is the packet example as a few lines of Python. It is plain Python, and I have run it:

# 100,000 network packets, 100 of them really dropped
tp, fn, fp, tn = 80, 20, 40, 99_860     # a model that flags suspicious packets

accuracy  = (tp + tn) / (tp + tn + fp + fn)
precision = tp / (tp + fp)               # of the alerts, how many were real?
recall    = tp / (tp + fn)               # of the real faults, how many did we catch?
f1        = 2 * precision * recall / (precision + recall)

print(f"accuracy  {accuracy:.2%}")       # 99.94%
print(f"precision {precision:.1%}")      # 66.7%
print(f"recall    {recall:.1%}")         # 80.0%
print(f"F1        {f1:.1%}")             # 72.7%

# The lazy model that always says "no fault" scores 99.9% accuracy but its recall is 0%.
print(f"lazy accuracy {(99_900) / 100_000:.1%}, lazy recall {0 / 100:.0%}")
Accuracyall correct ÷ all predictionsfine when the classes are balanced Precisioncorrect alerts ÷ all alertshow far you can trust the alarm Recallcaught events ÷ all real eventshow much you do not miss F1 scorebalance of precision and recallrare events, imbalanced data Four numbers to rememberNone is best everywhere. Pick the one that matches your risk.
Accuracy, precision, recall and F1 each answer a different question. Knowing which question matters for your problem is the real skill.

Which one should you use?

  • If false alarms are costly, watch precision.
  • If misses are costly, watch recall.
  • If events are rare and both mistakes matter, use the F1 score, and always look at the confusion matrix itself, not only one number.
  • Treat plain accuracy with care. It is only trustworthy when the classes are roughly balanced and every mistake costs about the same.

Coming Up Next

Day 8: Introduction to LLMs and the modern AI API landscape.

#MachineLearning #DataScience #MLOps #ModelEvaluation #AIInfrastructure #TechEducation #DataAnalytics #EnterpriseAI