Subscribe to Newsletter

Module 1: Building and Deploying AI Applications

By the end of this module you will have made the seven core decisions that turn an LLM demo into an AI application people can rely on. This is the largest area on Ng's map and the one job postings mention most.

Start module Module 1 of 5 · 7 lessons

Source for this module: Ng’s Aug 21 letter, What You Need to Know to Build and Deploy AI Applications In Real Life.

1.1

LLM Foundations

Ng lists what you need to understand about the model itself: how it tokenizes input and generates output, when to reach for a multimodal model, what to put in the context window, and how to reason about cache hits, knowledge cutoff, reasoning effort, sampling parameters, and tool calling. That understanding is what lets you choose the right model or mix of models, and know when to reach for fine-tuning or self-hosting.

The failure mode: you pick the biggest model for everything, pay ten times what you need to, and still get caught out when it cannot see last month's product change because of its knowledge cutoff.

The practical version of this skill is a budget. Every call has a cost in tokens and a cost in seconds, and both scale with what you put in the context window. Prompt caching cuts the cost of a repeated prefix. Reasoning effort trades latency for quality. Sampling temperature trades consistency for variety.

You do not need to memorize the arithmetic. You need to know which lever to pull when a stakeholder says it is too slow, or too expensive.

HANDS-ON: TRIAGE
  1. Write MODEL.md. For each of Triage's four steps (classify, retrieve, draft, confidence check) pick a model tier and write one line on why. Classification probably wants a small fast model. Drafting probably wants a larger one.
  2. Estimate tokens per ticket and cost per 1,000 tickets. Put the number in the file, not in your head.
Watch: Let’s build the GPT Tokenizer, Andrej Karpathy.

You have this skill when: you can look at a latency or cost complaint and name the specific lever you would change first.

1.2

Grounding Models With Data

An LLM only knows what is in its context. Ng's point in the second letter is that RAG with vector search was an early answer to that problem and the menu has grown since. You now choose between putting data in the prompt and letting the model retrieve it on demand with tools, and you choose the representation that fits your data and your queries: a vector index, a knowledge graph, or a semantic layer over structured records. You also turn messy documents into inputs the model can use, and build pipelines that keep that data clean and fresh.

The failure mode: you embed every document in the company, retrieve the top five chunks, and the model answers confidently from a policy that was replaced eight months ago.

Think about it from the query side first. If the question is what does our refund policy say, a vector index over docs works. If the question is what plan is this customer on and what did they buy, that is a database query, so give the model a tool that runs it rather than a pile of embedded records. Most real applications need both.

HANDS-ON: TRIAGE
  1. List the ten most common support questions. Mark each one as doc (vector search), row (tool call), or relationship (graph).
  2. Build retrieval for the doc questions only. The other two are deliberately out of scope this week.
  3. Add a last_updated field to every chunk and exclude anything older than the current product version. Freshness is the part everyone skips.
Watch: Retrieval-augmented generation, clearly explained.

You have this skill when: you can look at a question and say which retrieval method answers it, and why the other two would not.

1.3

Building Agentic Systems

Ng draws the range clearly. At one end, a workflow: a fixed sequence of LLM calls you designed. At the other, an agent harness where the model decides its own next step in a loop. Your job is to choose the architecture, pick the tools the model can call, decide the memory setup, manage context over long sessions, and know when a task genuinely needs multiple agents. Then you harden it with guardrails and an answer for risks like data exfiltration.

The failure mode: you build a fully autonomous agent for a task that was three predictable steps, then spend a month debugging why it sometimes skips step two.

Anthropic's guidance on this is blunt and Ng agrees with it. Start with the simplest thing that works. If you can write the flow as code with LLM calls inside it, write it as code. Reach for an agent loop only when the path through the task cannot be known in advance.

HANDS-ON: TRIAGE
  1. Write Triage's four steps as a plain Python pipeline. They are a workflow, not an agent, and pretending otherwise costs you a month.
  2. Give the model a choice in exactly one place: the confidence check, where it can call approve_for_human or escalate.
  3. Draw the loop for that one step and nothing else.
Watch: Building more effective AI agents, Anthropic.

You have this skill when: you can defend, for any AI feature, why it is a workflow or why it needs an agent, in one sentence each.

1.4

Evaluation-Driven Development: Building the Eval Set

Ng is explicit that this is the skill that separates the best builders. He calls driving a disciplined evals and error analysis loop the most important trait of someone great at building AI systems, and says it is hard to master because the right approach changes by project and by stage.

The failure mode is the one nearly every team hits: you change a prompt, it feels better on the three examples you tried, you ship it, and the thing you did not test gets worse.

An eval is a set of inputs with a way to score the outputs. That is all it is. The skill is in choosing what to measure. Ng's method is to look at real traces and outputs, do exploratory analysis on them, and combine that with product judgment to decide what matters. For Triage, did it pick the right product area matters more than was the reply eloquent.

HANDS-ON: TRIAGE
  1. Pull 30 real or realistic tickets. For each, write the expected product area, the expected urgency, and whether it should escalate. Save as evals/tickets.jsonl.
  2. Run Triage on all 30 and record the score.
  3. Write the number at the top of SPEC.md. That number is now the thing you improve.
Watch: Evals, error analysis, and better prompts, with Hamel Husain.

You have this skill when: you refuse to change a prompt without a number to compare before and after.

1.5

Evaluation-Driven Development: Error Analysis and Choosing Your Judge

Once you have a score, the second half of the skill is knowing what kind of eval to run and how to read what fails. Ng lists three options: deterministic code-based checks, an LLM as judge, and a human in the loop. He also says you need to evaluate your evals, because a judge that scores the wrong thing steers the whole project the wrong way.

Use code when the answer is checkable. Did the classifier output one of the five allowed labels. Did the reply include the ticket number. Use an LLM judge when the answer is a quality call, and spot-check that judge against a human on a sample. Use humans for the cases where being wrong is expensive.

Error analysis is the part most teams skip. Take the failures, sort them into buckets by cause, and count. That is what makes progress systematic rather than random. If 60 percent of Triage's misses are billing tickets tagged as account tickets, you do not need a better model. You need three examples of billing tickets in the prompt.

HANDS-ON: TRIAGE
  1. Run your 30-ticket eval and open every failure. Assign each one a cause in a spreadsheet.
  2. Fix the largest bucket only. Re-run. Update the score in SPEC.md.
  3. Write a second eval, an LLM judge for reply quality, and check it against your own rating on ten replies. If it disagrees with you more than twice, fix the judge before you trust it.
Watch: LLM Evals: Common Mistakes, Hamel Husain.

You have this skill when: your first move after a bad score is to bucket the failures, not to swap the model.

The Code: Your daily unfair advantage in software engineering.

Join 350,000+ software engineers, tech leads, and CTOs who start their morning with The Code.

Subscribe to Newsletter
1.6

Operating in Production

Ng names three things that make operating AI software different from operating normal software: unpredictability, cost, and latency. So you need observability that shows how the system performs on real usage, drift detection, and a fast response to model failures and security incidents like prompt injection. Regression testing and CI need statistical evals rather than pass or fail tests, calibrated to how bad a mistake would be.

The failure mode: the model provider ships a silent update, your classification accuracy drops 15 points, and you find out from a customer three weeks later.

Observability for AI means logging the full trace of every call: the prompt, the retrieved context, the output, the tool calls, latency, and tokens. Without that you cannot do error analysis on production data, which means the eval loop from 1.4 stops working the day you launch.

HANDS-ON: TRIAGE
  1. Log every ticket's full trace to a file or a tracing tool.
  2. Define three alerts in OPS.md: escalation rate outside its normal band, p95 latency above your budget from 1.1, daily token spend above your budget from 1.1.
  3. Sample 20 production traces a week and add the interesting ones to evals/tickets.jsonl. This is how the eval set stays alive.
Watch: A 10-minute walkthrough of Langfuse, open source LLM observability.

You have this skill when: you can answer how is it doing in production with a number from the last 24 hours, not a feeling.

1.7

Machine Learning Foundations

Ng says every engineer he knows who is good at building with LLMs also understands machine learning and deep learning at some depth, and that many applications still need a classic trained model, whether you train it or someone else did. The concepts he calls out as still essential are bias and variance, error analysis, and engineering your data.

The failure mode: you use an LLM at two cents a call to do something a 5MB classifier does in a millisecond, and you cannot explain why its accuracy moves around from day to day.

You do not need to derive backpropagation. You need three ideas. Bias means your model is too simple to capture the pattern, so more data will not help and a better model will. Variance means it memorized the training data, so more data or a simpler model will help. And the data you feed in sets the ceiling for everything downstream. These three are exactly why the error analysis in 1.5 works, because you are diagnosing which problem you have before you pick a fix.

HANDS-ON: TRIAGE
  1. Train a small text classifier on your 30-plus labeled tickets. Logistic regression on TF-IDF is fine.
  2. Compare its accuracy and cost to the LLM classifier from 1.1. Write the comparison in MODEL.md.
  3. If the small model is within five points, use it, and save the LLM for drafting.
Watch: Bias and variance, Andrew Ng (DeepLearning.AI).

You have this skill when: you can look at an AI failure and say whether it is a data problem, a model capacity problem, or an overfitting problem.

END OF MODULE 1

By this point you should have:

  • A model choice and a cost budget per step (MODEL.md).
  • Retrieval built for the questions that need it, with freshness enforced.
  • Triage written as a workflow with exactly one small agentic step.
  • A 30-ticket eval set with a score you are tracking in SPEC.md.
  • Failures bucketed by cause, and an LLM judge you have checked against yourself.
  • Full traces logged and three alerts defined (OPS.md).
  • A small classical model benchmarked against the LLM.