A practical, engineering-first reference for taking an idea to a verified, production-ready MVP with role-based coding agents.
Write a clear problem statement, measurable acceptance criteria, and why the change matters. For a greenfield MVP, also define one core user flow and an explicit out-of-scope list.
Create an editable plan artifact (for example tasks/todo.md) before writing code. Assign models by role: planner/architect, implementer, reviewer, verifier. Keep context lean. Use sub-agents or parallel worktrees for exploration, clear context between major steps, and prefer small, atomic, reversible units of work.
Package each step with an objective, precise scope (files or directories), hard constraints, exact verification commands the agent must run, evidence the agent must produce, and a definition of done. If a step cannot be reviewed or reverted cleanly in minutes, split it.
Send the task packet to the implementing agent. Require code, files changed, a minimal diff summary, commands actually run with full output, unresolved risks, and status against every acceptance criterion. Prefer the agent running verification commands inside the loop.
A review role (same or different model) checks correctness, regressions, edge cases, project rules (AGENTS.md, CLAUDE.md, or .cursorrules), and whether the evidence is sufficient. Red-team high-risk changes with a second model when security, auth, data, or payments are touched.
Do not accept a change until lint, type-check, tests, and project-specific commands have been run by the agent or CI and the results appear in the output. Prefer the smallest focused test set that exercises the change.
Run targeted runtime and scenario checks that automation cannot cover. The human still owns final behavioral judgment for non-trivial UX or business logic.
Keep commits tied to completed task packets. Record verification commands and results in the change note. Prefer small, reviewable diffs.
Note: This is still one tool (or set of tools) critiquing another. Explicit roles plus evidence requirements make the process auditable and repeatable. The human remains accountable for the outer loop: architecture, constraints, and final ownership.
Put durable rules in AGENTS.md (or CLAUDE.md / .cursorrules), not in every prompt. That file is the lasting place for conventions, required commands, model roles, review expectations, and definition of done.
Keep each prompt focused only on the current task, scope, and acceptance criteria.
Objective: Add a new endpoint for exporting user settings as JSON.
Scope: backend route only. Files: app/server/users.js and tests/server/users.test.js.
Constraints: preserve current auth flow; do not change UI or adjacent routes.
Commands (agent must run and paste full output):
npm test -- users.test.js
npm run lint -- app/server/users.js
npm run typecheck
Acceptance criteria:
- GET /users/:id/settings/export returns 200 and valid JSON for an authenticated user
- unauthorized requests return 401
- unit tests cover happy path, 401, and error paths
Definition of done:
- all commands above pass
- no files outside Scope are modified
Required evidence:
- files changed + minimal diff summary
- full command output for each verification command
- status against every acceptance criterion
- unresolved risks (or "none")
AGENTS.md.