Download status
The handbook manuscript and landing-page content structure are ready for review. The download button will not claim a file exists until the approved, accessible PDF has been supplied and tested.
Ten chapters, in production order
- Choose the job: user, outcome, boundary and non-goals.
- Design the agent: instructions, tools, handoffs and state.
- Build reliable context: retrieval, metadata, freshness and access.
- Write safe tools: contracts, identity, validation and retries.
- Connect with MCP: architecture, authorization and trust boundaries.
- Handle memory: usefulness, retention, privacy and deletion.
- Apply guardrails: approvals, refusal, escalation and safe failure.
- Evaluate: datasets, rubrics, assertions and regression gates.
- Deploy: versioning, staged release, rollback and cost controls.
- Operate: traces, incidents, feedback and continuous evaluation.
Preview: the production-readiness test
An agent is not ready because the happy path works. It is ready when its owners can explain what it may do, prove how it behaves and recover when a dependency or model decision fails.
- The user and business outcome are measurable.
- Tool permissions are narrower than the agent's conversational scope.
- Every high-impact action has an approval or equivalent control.
- The evaluation set includes normal, edge, refusal and failure cases.
- Traces reveal decisions without retaining unnecessary sensitive data.
- An operational owner, rollback trigger and refresh date are recorded.
Companion artifacts
Architecture canvas
Map users, data, models, tools, approvals and trust boundaries.
Evaluation worksheet
Record cases, expected behavior, measures, results and errors.
Threat-model canvas
Identify assets, entry points, abuse paths and controls.
Release checklist
Verify ownership, monitoring, rollback and change management.
Research foundation
The downloadable edition must pass editorial, technical, accessibility and file-integrity review before release. Last reviewed 12 August 2026.