Subscribe to Newsletter

Module 3: Using Coding Agents

By the end of this module you will have used a coding agent to build the parts of Triage you have not built yet, and you will know how much to delegate, how to check the work, and how to set the agent up so it gets better over time.

Start module Module 3 of 5 · 5 lessons

Source for this module: Ng’s Sep 4 letter, How to Use Coding Agents Effectively from Planning to Execution and Monitoring. Each lesson links to the deeper hands-on version in our companion course, How Engineering Teams at Fortune 1000 Companies Code with AI.

3.1

The Workflow: Planning, Execution, Deployment and Monitoring

After interviewing dozens of top AI engineers, Ng's team found one consistent workflow. Planning: brainstorm, research, understand the existing codebase, write a spec covering requirements and technical design, then generate an execution plan and review it for bad assumptions, security gaps, and overengineering. Execution: the agent builds, you verify with automated or human checks, and you calibrate how much autonomy to give it. Deployment: ship behind CI or human gates, then use agents to watch logs and propose fixes.

Ng points out that this is the same workflow we used before agents. What changed is where the human effort goes. Far less on code, far more on deciding what to build, designing the architecture, writing the spec, and checking the output. And the spec scales with risk. A greenfield prototype's spec can be a quick prompt. A brownfield system with real users needs a spec written and verified with care.

The failure mode: you give an agent a one-line prompt for a feature that touches production data, and spend two days undoing what it decided on your behalf.

HANDS-ON: TRIAGE
  1. Merge SPEC.md, MODEL.md, DATA.md, and ARCH.md into one spec for the feature you have not built yet: the approval UI's edit-and-send flow.
  2. Give the spec to your agent in plan mode and have it produce an execution plan.
  3. Read the plan and find one assumption it made that you did not state. Then fix the spec, not the plan.
Watch: Stop vibe coding everything: the case for spec-driven development.

Deeper version: lesson 5.3 of our agentic coding guide, on spec, review, and ship.

You have this skill when: the length of your spec matches the blast radius of the feature.

3.2

Directing the Workflow and Enabling Agent Autonomy

Two of Ng's sub-skills belong together. Directing the workflow means deciding how much human and how much agent effort each step gets, and when to go back a step. He frames it as a tradeoff between speed, cost, technical risk, and human effort. Enabling autonomy is the dial: watch the agent turn by turn, delegate a chunk, or set a goal and let it loop until it succeeds. It also covers managing the agent's context as the build moves through phases, running agents in parallel on a decomposed task, and setting permissions so a fast-moving agent cannot leak data or delete things.

The failure mode Ng calls out directly is the hype one. He says the value of very long autonomous runs has been amplified beyond reality, and that the practical version is a highly iterative process where skilled intervention gets much better results than walking away.

So the rule is simple. Autonomy goes up as verification goes up. If the agent can run a test that proves it is done, let it loop. If the only check is your eyes, stay in the loop.

HANDS-ON: TRIAGE
  1. Split the approval UI work into three pieces and place each in a quadrant: verifiable or not, high risk or low.
  2. Run the delegate-and-loop piece with the eval as its stopping condition. Run the interactive piece turn by turn. Write down which produced better code per minute of your attention.
  3. Set the agent's permissions so it cannot touch .env or run migrations without asking.
Watch: Claude Code best practices, Anthropic.

Deeper version: lesson 4.6 on goal loops and lesson 1.6 on permissions.

You have this skill when: you can say, for any task, which quadrant it is in and why.

3.3

Reviewing the Work

Ng's framing: the output of a coding agent is uncertain. You do not know in advance what good ideas it will have or what bugs it will write, so review is where you find out. The skill is designing verification matched to the task. Behavioral and functional tests. User flow tests, possibly with the agent supplying screenshots as evidence. Eval sets with an LLM judge for qualitative output. Then deciding how much of that to automate so the agent can check its own work, and evaluating the tests themselves to make sure they measure what you meant.

The failure mode: the agent reports that all tests pass, and it is true, because it edited the tests.

Ng's distinction between reviewing behavior and reviewing code is the one to hold on to. You will rarely read every line an agent writes. You will always check what the software does. That is why the eval set and the three tests from 2.4 matter more than a diff.

HANDS-ON: TRIAGE
  1. Have a fresh-context subagent review the approval UI code your first agent wrote, looking for exactly three things: edited tests, secrets in the client, and blocking model calls.
  2. Test the user flow yourself and have the agent capture a screenshot of the approved reply as evidence.
  3. Add any bug you find to the eval or the test suite so it cannot come back.
Watch: LGTM, ship it: the AI code review problem.

Deeper version: lesson 3.4 on second-opinion review with a fresh-context subagent.

You have this skill when: tests pass makes you ask which tests, and whether they changed.

The Code: Your daily unfair advantage in software engineering.

Join 350,000+ software engineers, tech leads, and CTOs who start their morning with The Code.

Subscribe to Newsletter
3.4

Customizing the Agent and Its Environment

This is the sub-skill that compounds. Ng's list: integrate skills, plugins, and MCP servers, and prune them when a new model makes one unnecessary. Use hooks to automate the repeatable parts. Maintain the agent's standing context with the codebase, architecture assumptions, code style, and data access patterns. Preserve state across sessions and across parallel agents. Accumulate what the agent learns, through post-run retrospectives. Keep the codebase navigable, clear out agent-generated debt now and then, and in a team, coordinate context across everyone's agents.

The failure mode: every session starts from zero, the agent re-learns your schema each time, and three engineers' agents each invent a different naming convention.

Every markdown file you wrote in this course has a second job now. MODEL.md, DATA.md, ARCH.md, and SECURITY.md are the agent's context. Put them where the agent reads them.

HANDS-ON: TRIAGE
  1. Write CLAUDE.md for the repo. Keep it short and point to the spec files rather than repeating them.
  2. Add a hook that runs the eval after any change to the prompts/ directory, and connect one MCP server (the ticket database is the obvious one).
  3. Run a five-minute retrospective with the agent after your next session. What did it get wrong, and what should CLAUDE.md say to prevent it. Commit the change.
Watch: Building agents with Model Context Protocol, full workshop with Anthropic’s Mahesh Murag.

Deeper versions: 2.3 on CLAUDE.md, 4.1 on hooks, 4.2 on skills, and 4.3 on MCP servers.

You have this skill when: your agent's second session on a task is faster than its first because of something you wrote down.

3.5

Coding Agent Foundations

Ng puts this last because it is what makes every other decision in this module good. You should understand how coding agents work: how they search a codebase, how they manage their context window, what adding tools or MCP servers does to that context, how agents and subagents interact, and the basic fact that an agent is a harness wrapped around an LLM. That understanding makes the agent less of a black box and helps you spot the failure modes: overengineering a simple task, losing rigor when there is no explicit verification step, stopping short of the goal, or taking an action that could destroy files or production data.

The failure mode: the agent starts making strange edits 40 minutes in, and you cannot tell whether it is the model, the context filling up, or a tool returning garbage.

Two ideas cover most of it. Context is finite, and every tool result, file read, and server definition spends some of it, so an agent that has been running a long time is working with a worse memory than when it started. And the agent only knows what is in its context, so when it goes off track the fix is almost always to change what it can see, not to argue with it.

HANDS-ON: TRIAGE
  1. Run one long session on a Triage task and watch for the moment quality drops. Note the turn count.
  2. Do the same task with a subagent handling the file exploration and reporting back a summary. Compare.
  3. Write what you noticed at the bottom of CLAUDE.md under how this agent behaves.
Watch: Effective context engineering for AI agents, from Anthropic’s guide.

Deeper versions: 1.2 on the agent loop and 1.3 on context engineering.

You have this skill when: you can look at a misbehaving agent and name which of Ng's four failure modes you are seeing.

END OF MODULE 3

By this point you should have:

  • One merged spec, and an agent-generated plan you corrected at the spec level.
  • Each piece of work placed in the autonomy quadrant, with permissions set.
  • A second-opinion review, a user flow test with screenshot evidence, and bugs turned into tests.
  • CLAUDE.md, a hook, an MCP server, and one retrospective committed.
  • A first-hand observation of context degradation, and how a subagent helps.