Module 1: Building and Deploying AI Applications
By the end of this module you will have made the seven core decisions that turn an LLM demo into an AI application people can rely on. This is the largest area on Ng's map and the one job postings mention most.
Source for this module: Ng’s Aug 21 letter, What You Need to Know to Build and Deploy AI Applications In Real Life.
LLM Foundations
Ng lists what you need to understand about the model itself: how it tokenizes input and generates output, when to reach for a multimodal model, what to put in the context window, and how to reason about cache hits, knowledge cutoff, reasoning effort, sampling parameters, and tool calling. That understanding is what lets you choose the right model or mix of models, and know when to reach for fine-tuning or self-hosting.
The failure mode: you pick the biggest model for everything, pay ten times what you need to, and still get caught out when it cannot see last month's product change because of its knowledge cutoff.
The practical version of this skill is a budget. Every call has a cost in tokens and a cost in seconds, and both scale with what you put in the context window. Prompt caching cuts the cost of a repeated prefix. Reasoning effort trades latency for quality. Sampling temperature trades consistency for variety.
You do not need to memorize the arithmetic. You need to know which lever to pull when a stakeholder says it is too slow, or too expensive.
- Write
MODEL.md. For each of Triage's four steps (classify, retrieve, draft, confidence check) pick a model tier and write one line on why. Classification probably wants a small fast model. Drafting probably wants a larger one. - Estimate tokens per ticket and cost per 1,000 tickets. Put the number in the file, not in your head.
You have this skill when: you can look at a latency or cost complaint and name the specific lever you would change first.
Grounding Models With Data
An LLM only knows what is in its context. Ng's point in the second letter is that RAG with vector search was an early answer to that problem and the menu has grown since. You now choose between putting data in the prompt and letting the model retrieve it on demand with tools, and you choose the representation that fits your data and your queries: a vector index, a knowledge graph, or a semantic layer over structured records. You also turn messy documents into inputs the model can use, and build pipelines that keep that data clean and fresh.
The failure mode: you embed every document in the company, retrieve the top five chunks, and the model answers confidently from a policy that was replaced eight months ago.
Think about it from the query side first. If the question is what does our refund policy say, a vector index over docs works. If the question is what plan is this customer on and what did they buy, that is a database query, so give the model a tool that runs it rather than a pile of embedded records. Most real applications need both.
- List the ten most common support questions. Mark each one as doc (vector search), row (tool call), or relationship (graph).
- Build retrieval for the doc questions only. The other two are deliberately out of scope this week.
- Add a
last_updatedfield to every chunk and exclude anything older than the current product version. Freshness is the part everyone skips.
You have this skill when: you can look at a question and say which retrieval method answers it, and why the other two would not.
Building Agentic Systems
Ng draws the range clearly. At one end, a workflow: a fixed sequence of LLM calls you designed. At the other, an agent harness where the model decides its own next step in a loop. Your job is to choose the architecture, pick the tools the model can call, decide the memory setup, manage context over long sessions, and know when a task genuinely needs multiple agents. Then you harden it with guardrails and an answer for risks like data exfiltration.
The failure mode: you build a fully autonomous agent for a task that was three predictable steps, then spend a month debugging why it sometimes skips step two.
Anthropic's guidance on this is blunt and Ng agrees with it. Start with the simplest thing that works. If you can write the flow as code with LLM calls inside it, write it as code. Reach for an agent loop only when the path through the task cannot be known in advance.
- Write Triage's four steps as a plain Python pipeline. They are a workflow, not an agent, and pretending otherwise costs you a month.
- Give the model a choice in exactly one place: the confidence check, where it can call
approve_for_humanorescalate. - Draw the loop for that one step and nothing else.
You have this skill when: you can defend, for any AI feature, why it is a workflow or why it needs an agent, in one sentence each.
Evaluation-Driven Development: Building the Eval Set
Ng is explicit that this is the skill that separates the best builders. He calls driving a disciplined evals and error analysis loop the most important trait of someone great at building AI systems, and says it is hard to master because the right approach changes by project and by stage.
The failure mode is the one nearly every team hits: you change a prompt, it feels better on the three examples you tried, you ship it, and the thing you did not test gets worse.
An eval is a set of inputs with a way to score the outputs. That is all it is. The skill is in choosing what to measure. Ng's method is to look at real traces and outputs, do exploratory analysis on them, and combine that with product judgment to decide what matters. For Triage, did it pick the right product area matters more than was the reply eloquent.
- Pull 30 real or realistic tickets. For each, write the expected product area, the expected urgency, and whether it should escalate. Save as
evals/tickets.jsonl. - Run Triage on all 30 and record the score.
- Write the number at the top of
SPEC.md. That number is now the thing you improve.
You have this skill when: you refuse to change a prompt without a number to compare before and after.
Evaluation-Driven Development: Error Analysis and Choosing Your Judge
Once you have a score, the second half of the skill is knowing what kind of eval to run and how to read what fails. Ng lists three options: deterministic code-based checks, an LLM as judge, and a human in the loop. He also says you need to evaluate your evals, because a judge that scores the wrong thing steers the whole project the wrong way.
Use code when the answer is checkable. Did the classifier output one of the five allowed labels. Did the reply include the ticket number. Use an LLM judge when the answer is a quality call, and spot-check that judge against a human on a sample. Use humans for the cases where being wrong is expensive.
Error analysis is the part most teams skip. Take the failures, sort them into buckets by cause, and count. That is what makes progress systematic rather than random. If 60 percent of Triage's misses are billing tickets tagged as account tickets, you do not need a better model. You need three examples of billing tickets in the prompt.
- Run your 30-ticket eval and open every failure. Assign each one a cause in a spreadsheet.
- Fix the largest bucket only. Re-run. Update the score in
SPEC.md. - Write a second eval, an LLM judge for reply quality, and check it against your own rating on ten replies. If it disagrees with you more than twice, fix the judge before you trust it.
You have this skill when: your first move after a bad score is to bucket the failures, not to swap the model.
The Code: Your daily unfair advantage in software engineering.
Join 350,000+ software engineers, tech leads, and CTOs who start their morning with The Code.
Operating in Production
Ng names three things that make operating AI software different from operating normal software: unpredictability, cost, and latency. So you need observability that shows how the system performs on real usage, drift detection, and a fast response to model failures and security incidents like prompt injection. Regression testing and CI need statistical evals rather than pass or fail tests, calibrated to how bad a mistake would be.
The failure mode: the model provider ships a silent update, your classification accuracy drops 15 points, and you find out from a customer three weeks later.
Observability for AI means logging the full trace of every call: the prompt, the retrieved context, the output, the tool calls, latency, and tokens. Without that you cannot do error analysis on production data, which means the eval loop from 1.4 stops working the day you launch.
- Log every ticket's full trace to a file or a tracing tool.
- Define three alerts in
OPS.md: escalation rate outside its normal band, p95 latency above your budget from 1.1, daily token spend above your budget from 1.1. - Sample 20 production traces a week and add the interesting ones to
evals/tickets.jsonl. This is how the eval set stays alive.
You have this skill when: you can answer how is it doing in production with a number from the last 24 hours, not a feeling.
Machine Learning Foundations
Ng says every engineer he knows who is good at building with LLMs also understands machine learning and deep learning at some depth, and that many applications still need a classic trained model, whether you train it or someone else did. The concepts he calls out as still essential are bias and variance, error analysis, and engineering your data.
The failure mode: you use an LLM at two cents a call to do something a 5MB classifier does in a millisecond, and you cannot explain why its accuracy moves around from day to day.
You do not need to derive backpropagation. You need three ideas. Bias means your model is too simple to capture the pattern, so more data will not help and a better model will. Variance means it memorized the training data, so more data or a simpler model will help. And the data you feed in sets the ceiling for everything downstream. These three are exactly why the error analysis in 1.5 works, because you are diagnosing which problem you have before you pick a fix.
- Train a small text classifier on your 30-plus labeled tickets. Logistic regression on TF-IDF is fine.
- Compare its accuracy and cost to the LLM classifier from 1.1. Write the comparison in
MODEL.md. - If the small model is within five points, use it, and save the LLM for drafting.
You have this skill when: you can look at an AI failure and say whether it is a data problem, a model capacity problem, or an overfitting problem.
By this point you should have:
- A model choice and a cost budget per step (
MODEL.md). - Retrieval built for the questions that need it, with freshness enforced.
- Triage written as a workflow with exactly one small agentic step.
- A 30-ticket eval set with a score you are tracking in
SPEC.md. - Failures bucketed by cause, and an LLM judge you have checked against yourself.
- Full traces logged and three alerts defined (
OPS.md). - A small classical model benchmarked against the LLM.
Module 2: Software Engineering Fundamentals
Five decisions that decide whether Triage survives contact with real users. Ng's argument fits in one sentence: a developer who vibe codes without fundamentals did not steer the agent to make the right decisions.