Theme: The Responsibility
Length: ~700 words
Hook: Practical — give people a playbook for the worst case
Your agent will eventually do something wrong in production. Not "might." Will.
The question isn't whether it will happen. The question is: when it happens, will you know what to do?
Most teams have incident response plans for infrastructure outages. Few have incident response plans for agent failures. The two are different — agent failures involve decisions, not just downtime.
Here's the playbook.
Step 1: Detect (Automated)
You need automated detection before you can respond. This means:
- Policy violation alerts — the governance layer flags an action that violated policy
- Cost ceiling alerts — the agent exceeded its budget
- Evaluation drift alerts — the agent's quality scores are dropping
- User feedback spikes — complaints or error reports increase
- Anomalous behavior detection — the agent is calling tools in an unusual pattern
- Pause the agent. Not "shut down" — pause. You want to investigate, not destroy evidence.
- Block the affected tool. If the agent sent wrong emails via NotificationAPI, block NotificationAPI for this agent. Don't block the entire agent if only one tool is the problem.
- Preserve the audit trail. Don't clear logs. Don't restart services. The audit trail is your evidence.
- Notify the sponsor. The human who approved this agent's deployment needs to know immediately.
- Read the audit trail. What did the agent perceive? What did it reason? What did it do? What policy checks passed or failed?
- Identify the root cause. Was it a prompt issue? A tool failure? A policy gap? An adversarial input? A model regression?
- Determine the blast radius. How many users were affected? How many actions were taken? What data was accessed or modified?
- Classify the severity. SEV1 (data loss, financial impact, regulatory violation), SEV2 (significant user impact), SEV3 (minor impact, isolated).
- Undo the agent's actions. If it sent wrong emails, send corrections. If it processed wrong refunds, reverse them. If it modified wrong data, restore from backup.
- Notify affected users. Be honest. "Our AI agent made an error. Here's what happened. Here's what we're doing about it."
- Patch the agent. Fix the prompt, the tool, or the policy that caused the failure. Don't redeploy yet.
- Document the incident. Timeline, root cause, blast radius, remediation steps. This is for the post-mortem.
- Add evaluation cases. The specific scenario that caused the failure becomes a permanent test case. If the agent ever does this again, the evaluation suite catches it before production.
- Add or modify policies. If the agent did something it shouldn't have, the policy was too permissive. Tighten it.
- Add monitoring. If detection was slow, add alerts for this specific pattern.
- Review similar agents. If you have other agents with similar capabilities, check if they have the same vulnerability.
- Internal post-mortem. Within 1 week, conduct a blameless post-mortem. Focus on what went wrong in the system, not who went wrong.
- External communication. For SEV1 incidents, publish a public incident report. This builds trust. Hiding incidents destroys trust.
- Update the incident response plan. Every incident teaches you something. Update the plan with what you learned.
If you don't have automated detection, you'll find out about the incident from a customer on LinkedIn. That's the worst way to learn.
Step 2: Contain (Immediate)
Once detected, contain the damage immediately:
Time to contain: under 5 minutes. If it takes longer, your detection-to-action pipeline is too slow.
Step 3: Assess (Within 1 hour)
Now figure out what happened:
Step 4: Remediate (Within 4 hours for SEV1)
Fix the immediate damage:
Step 5: Prevent (Within 1 week)
Make sure this can't happen again:
Step 6: Communicate (Ongoing)
The template:
Every agent in production should have this incident response plan documented — specific to that agent. Who to notify. Which tools to block. How to undo actions. What the escalation path is.
If you don't have this plan, the first incident will be chaotic. Chaos extends the damage. Chaos destroys trust.
Does your team have an agent incident response plan — or will you figure it out when it happens?