Phase 7: Responsible AI, Governance & Risk · 45 min · Incident response runbooks · NIST AI RMF · Python
AI Incident Response — the PM's Role
Engineering fixes the model. Legal manages exposure. Comms manages the story. The PM is the only person whose job is all three at once.
Hiring signal: Frontier labs and enterprise AI teams explicitly test 'how do you ship responsibly under uncertainty' in AI PM interview loops, and a public model failure is the sharpest version of that question. PMs who can show a structured first-hour/first-day/first-week incident plan, not a vague 'we'd loop in the right people,' signal they've actually thought about the failure mode before it happens to them.
What you will learn
- Lead the first hour of a public AI incident: containment, fact-finding, and who to notify in what order
- Draft a first-day public statement that is honest, specific, and doesn't overpromise a fix timeline the team can't hit
- Distinguish what a PM owns during an incident from what belongs to engineering, legal, and trust & safety
- Run a first-week root-cause and follow-up process that closes the loop publicly, not just internally
The Problem
It's 9:14am. A screenshot is circulating on social media of your company's customer-support chatbot confidently telling a user that a prescription drug is safe to combine with alcohol — it isn't, and the interaction can be dangerous. The screenshot has 4,000 shares in ninety minutes and a health journalist has already DM'd your support account asking for comment. Your Slack has six people typing at once: an engineer wants to know if they should pull the model right now, legal wants to know if anyone said anything to the user yet, your VP wants a one-line summary for the CEO, and support is asking whether they should stop routing conversations to the AI at all.
Nobody in that thread is wrong to be asking their question. The problem is that if everyone acts on their own question independently, you get exactly the outcome that turns a bad hour into a bad month: an engineer silently patches a prompt without telling anyone, legal drafts a statement that contradicts what support already told the user, and the CEO finds out from Twitter instead of from you. The PM's job in an AI incident isn't to write the fix or the legal language — it's to be the one person holding the containment decision, the comms decision, and the fix timeline in the same head at the same time, because everyone else is only holding one piece.
The First Hour: Containment and Fact-Finding
The first hour has exactly one goal: stop the bleeding and get true facts, in that order. Not "figure out why," not "draft the public statement" — those come later and get worse if rushed.
- Contain first, explain later. If the failure is reproducible and ongoing (not a one-off), the PM's first real decision is whether to disable the feature, roll back to a previous model/prompt version, or add an emergency guardrail (a keyword filter, a stricter system prompt, human review before send) — a decision made with engineering, not for them. Containment doesn't require knowing root cause; it requires knowing the failure is real and ongoing.
- Get the facts nailed down before anyone talks publicly. How many users saw this. Is it reproducible or a one-off. Is it still happening right now or already stopped. A PM who tells support "we're not sure yet, don't confirm or deny details" for 30 minutes while facts get nailed down looks far better in six months than one who let someone guess in public and got the guess wrong.
- Notify in a specific order, not a group chat blast: engineering lead (can we contain it), legal/trust & safety (does this trigger disclosure obligations — see Lesson 3's regulatory triage, incident reporting can be a legal requirement not just a courtesy), then executive stakeholders (a factual one-paragraph summary, not a guess at the fix timeline), then a holding statement for anyone customer-facing (support, social) so they have something accurate to say rather than improvising.
"We're looking into it" is a real statement, not a stall tactic
The instinct under pressure is to say nothing until there's a complete answer. That's the wrong tradeoff during an active incident — silence for hours while a screenshot spreads reads as either not caring or not knowing, and both are worse than a short, honest holding statement: "We're aware of this report and are investigating. [Feature] has been temporarily [disabled/restricted] while we do." That sentence takes five minutes to get right and buys the hours you actually need to get the real answer right.
Unlock the full lesson
You've read the first 2 sections. The rest of this lesson covers The First Day: Comms and Stakeholder Management, The First Week: Root Cause, Fix, and Closing the Loop Publicly, Build It, What to Practice — plus a hands-on lab, quiz, and project artifact.
Create a free account to unlock Phase 0 and Phase 1 of every course — no credit card.
Browse all courses · View pricing · DeVenture Academy