30 Days of AIDay 8 of 30 · Week 2: Applied AI & APIs
Course outline▾

Week 3 · Infrastructure & Hosting

  1. Day 15Introduction to AI Infrastructure: Hardware, Runtimes, and ComputeComing soon · 2026-10-15 · 7 PM IST
  2. Day 16Local Model Execution: Running Open-Weight LLMs Securely (Ollama, vLLM)Coming soon · 2026-10-16 · 7 PM IST
  3. Day 17Containerizing AI Workloads: Writing Production Dockerfiles for Python APIsComing soon · 2026-10-17 · 7 PM IST
  4. Day 18Docker Compose for Multi-Container AI Stacks (Web UI, Vector DB, LLM Engine)Coming soon · 2026-10-18 · 7 PM IST
  5. Day 19Kubernetes for AI 101: Pods, Deployments, and Services for Model ServingComing soon · 2026-10-19 · 7 PM IST
  6. Day 20Persistent Storage in Kubernetes: Managing State, Weights, and Vector IndicesComing soon · 2026-10-20 · 7 PM IST
  7. Day 21High-Performance Networking: Configuring Ingress and Egress for AI ClustersComing soon · 2026-10-21 · 7 PM IST

Week 4 · Enterprise Workflows

  1. Day 22Autonomous Agents: From Passive LLMs to Goal-Driven ExecutionComing soon · 2026-10-22 · 7 PM IST
  2. Day 23Tool Use & Function Calling: Connecting LLMs to APIs, Databases, and ShellsComing soon · 2026-10-23 · 7 PM IST
  3. Day 24Multi-Agent Orchestration: Supervisor, Worker, and Evaluator PatternsComing soon · 2026-10-24 · 7 PM IST
  4. Day 25Automated CI/CD: Automating the Software Lifecycle for AI ModelsComing soon · 2026-10-25 · 7 PM IST
  5. Day 26Zero-Trust Security for AI: Sandboxing Ephemeral Execution & Model EgressComing soon · 2026-10-26 · 7 PM IST
  6. Day 27Observability & Tracing for Agentic Systems (Telemetry, Logs, and Metrics)Coming soon · 2026-10-27 · 7 PM IST
  7. Day 28Human-in-the-Loop Architecture: Machine Second, Human First in PracticeComing soon · 2026-10-28 · 7 PM IST
  8. Day 29Managing Technical Debt, Drift, and Model Governance in Enterprise ITComing soon · 2026-10-29 · 7 PM IST
  9. Day 30The 10-Year Horizon: Architecting IT Strategy for the AI-Native EnterpriseComing soon · 2026-10-30 · 7 PM IST
Open the course page →

Day 8: Introduction to LLMs & The Modern AI API Landscape

2026-10-08 · 9 min read

Watch the video lesson, or subscribe on YouTube for a new lesson every day.

A large language model (LLM) is the combination of a huge Transformer neural network (the design you will meet in detail later) and racks of high-performance GPUs. For a business leader, the big question today is less about how an LLM works inside and more about how to use it, and that means understanding the modern API landscape.

In plain terms: An API is a waiter between your app and a kitchen. Your app writes an order (the prompt), the waiter carries it to the kitchen (the model), and the finished dish (the answer) comes back. You never have to own or run the kitchen.

Your appor agentAPIthe front doorLLMhosted modelAnswerstreamed backOne LLM request, from start to finish Theserverrestartedat03:15The model writes one token at a time, and streams them backYour app sends text. The API returns generated text.
An app or agent sends a prompt to the model's API. The model generates the answer one token at a time and streams it back.

What an LLM API call looks like

You send the model a message, and you get generated text back. The request is a small piece of structured data. This is the typical shape (field names vary a little between providers, so treat it as an illustration):

import json

# The shape of a typical LLM API request. Field names differ a little between providers.
request = {
    "model": "your-chosen-model",
    "messages": [
        {"role": "system", "content": "You are a careful IT assistant. Answer in two sentences."},
        {"role": "user", "content": "Summarize this incident: the VPN gateway restarted at 03:15."},
    ],
    "max_tokens": 200,     # a cap on the length of the answer, which also caps its cost
    "temperature": 0.2,    # lower means more predictable wording
    "stream": True,        # receive the answer token by token
}
print(json.dumps(request, indent=2))

The landscape in 2026

APIs used to be about one application talking to another. Increasingly they are being built so that AI models and autonomous agents can use them directly, and that changes what a company has to manage.

  • Closed commercial APIs (such as OpenAI, Anthropic and Google): you manage no infrastructure and you get immediate access to state-of-the-art reasoning. Providers increasingly offer very large context windows, with some now reaching a million tokens or more, and prompt caching, which makes repeated long instructions cheaper.
  • API gateways as AI control layers: a gateway sits in front of the models and manages who can call them, which APIs an AI system is allowed to discover and use, what policies apply, and what agents actually do, so that risks can be caught.
  • The routing approach: many production apps send simple tasks to cheaper, faster models and reserve the expensive flagship model for hard, multi-step reasoning. This controls cost and reduces dependence on one vendor.
AppsandagentsAI gatewaypoliciesloggingroutingcost controlSmall, fast, cheapsimple tasksFlagship modelhard reasoningSelf-hosted modelprivate data A gateway routes each request to the right modelCheap for easy work, flagship for hard work, private for sensitive data.
Many production apps put a gateway in front of their models. It applies policies, logs usage, and routes simple tasks to cheaper models while keeping the flagship model for hard reasoning and a self-hosted model for private data.
Prompt caching: pay less for repeated instructions Call 1full pricesystem prompt (2,000 tokens)new questionCall 2cachedsame prompt, reused from cachenew questionA big system prompt is billed in full once, then at a lower rate.Prices differ by provider. The bars are illustrative.
When many calls start with the same long system prompt, providers can cache it so repeated reads cost less. The bars are illustrative, since exact prices vary by provider.
AIgateway AuthenticationRate limitsPolicy checksLoggingSmart routingCost trackingWhat a gateway does for AI trafficIt also watches what autonomous agents try to do.
A modern API gateway acts as a control layer for AI: it checks who is calling, limits usage, applies policies, records what happens, routes requests and tracks cost, including actions started by agents.

Putting it together

Think of an AI app as three layers: your application or agent, a gateway that applies the rules, and one or more models. Open-weight models, which you can run yourself, belong in this picture too. They are useful when data must stay inside your own walls, and we will build exactly that in Week 3.

Questions worth asking about any LLM API:

  • How much does it cost per token, and is output priced higher than input?
  • How large is the context window, and is prompt caching available?
  • Where does my data go, and is it used for training?
  • What happens if the provider changes or retires the model?

Coming Up Next

Day 9: Advanced prompt engineering: few-shot examples, chain-of-thought and structured JSON.

#LargeLanguageModels #GenerativeAI #CloudComputing #TechStrategy #APIs