What This Playbook Covers
Security incidents move quickly. Therefore, teams need a clear response process before an alert becomes a larger business problem. This playbook explains how to identify, contain, investigate, fix, and recover from Microsoft 365 security incidents using a practical, repeatable process.
The goal is not just to close an alert. Instead, the goal is to confirm what happened, stop further damage, protect evidence, restore control, and reduce the chance of the same issue happening again. As a result, leaders get a clearer view of risk, response teams act with less confusion, and the organization can show what actions were taken.
How Identity Risk Fits Into Incident Response
Many Microsoft 365 incidents begin with identity. For example, an attacker may steal a password, trigger MFA fatigue, use a risky sign-in location, or reuse a session token. Because of this, identity risk signals often provide the first sign that a user account needs review.
Microsoft Entra User Risk and Sign-In Risk help teams understand whether an account or a specific sign-in attempt may be unsafe. User risk focuses on whether the identity itself may be compromised. Sign-in risk focuses on whether a specific access attempt should be trusted. In other words, these signals help response teams decide whether to challenge access, block access, reset credentials, revoke sessions, or begin a deeper investigation.
When This Playbook Is Used
Use this playbook when Microsoft 365 activity suggests that an account, device, mailbox, file, or admin action may be unsafe. Some events will turn out to be false positives. However, every event should follow a clear triage path so teams can separate noise from real risk.
Suspicious Account Activity
Users may trigger unexpected sign-ins, MFA fatigue events, impossible travel alerts, or other unusual activity. As a result, the response team should review sign-in logs, user risk, device details, and recent account changes before deciding whether the account still appears safe.
Phishing Or Account Compromise
A user may report phishing, open a malicious link, approve an unsafe prompt, or show signs that an attacker has accessed the account. In this case, the team should act quickly to stop active sessions, secure the account, and check for mailbox rules, forwarding, or other signs of abuse.
Potential Data Exposure
Attackers may access, share, change, or remove sensitive data without approval. Therefore, the team should review file access, sharing activity, audit logs, and any related alerts before deciding whether notification, legal review, or compliance action is needed.
Execution Steps
The response process should move in order. First, confirm whether the event is real. Next, stop further damage. After that, preserve evidence and investigate root cause. Finally, restore control and improve the safeguards that failed or drifted.
Identify The Incident
Validate the alert before treating the event as confirmed compromise. Review sign-in logs, user risk, Defender alerts, mailbox activity, device status, and recent admin actions. As a result, the team can define the first scope of affected users, data, apps, and systems.
Contain The Threat
Stop the attacker from doing more damage. Disable or restrict compromised accounts, revoke active sessions, reset passwords, block malicious senders, and isolate affected devices when needed. In addition, document each containment action so the team can explain what changed and why.
Preserve Evidence
Save the records needed for review before making broad changes. This may include sign-in logs, audit logs, email records, Defender alerts, file activity, screenshots, and user reports. Because evidence can age out or change, teams should preserve it early in the response process.
Investigate Root Cause
Determine how the incident started and whether the attacker still has access. For example, review whether the issue came from phishing, password spray, token abuse, risky sign-in activity, weak Conditional Access coverage, or a device issue. Then confirm whether the same risk appears in other accounts.
Fix And Recover
Remove unsafe access, correct weak settings, restore secure access, and confirm that affected systems are stable. This may include resetting credentials, removing mailbox rules, closing risky sessions, updating Conditional Access, and checking device health. After that, validate the fix before returning the account or system to normal use.
Review And Improve
Close the incident only after the team captures lessons learned. Identify what worked, what failed, and what should change. As a result, the organization can improve policies, alert paths, access controls, user training, and response steps before the next incident occurs.
Microsoft 365 Evidence Checklist
Evidence helps the team understand what happened and supports later decisions. In addition, clear evidence allows leaders, legal teams, and auditors to review the response without relying on memory or assumptions.
Identity Evidence
Collect Entra sign-in logs, user risk details, sign-in risk details, MFA events, Conditional Access results, password reset history, session revocation actions, and recent changes to authentication methods.
Email And Data Evidence
Review message trace, phishing reports, mailbox rules, forwarding settings, email quarantine details, SharePoint activity, OneDrive sharing, and audit records tied to sensitive files or folders.
Device And Admin Evidence
Review Defender alerts, device timeline details, Intune compliance state, device isolation actions, admin audit logs, role changes, app consent events, and any privileged activity tied to the incident.
Operational Requirements
A playbook only works when the right people can act at the right time. Therefore, each organization should define ownership, access, records, and communication paths before an incident occurs.
Defined Response Roles
Assign clear owners for triage, containment, evidence handling, investigation, recovery, communication, and leadership updates. This prevents delays when several teams need to act at once.
Logging And Visibility
Confirm that Microsoft 365 logs, security alerts, and audit records are available before the team needs them. Without enough log detail, teams may struggle to prove what happened or confirm whether the threat ended.
Controlled Admin Access
Limit privileged access during response and use approved admin accounts for sensitive actions. In addition, track each change so the team can separate attacker activity from response activity.
Common Response Mistakes
Incident response often fails because teams skip basic steps under pressure. However, the most common mistakes are avoidable when the response process is clear and practiced.
Only Resetting The Password
A password reset may not end the threat if active sessions, tokens, mailbox rules, or malicious authentication methods remain in place. As a result, teams should also revoke sessions and review recent account changes.
Changing Systems Before Saving Evidence
Teams may remove the data they need if they fix systems before saving logs and screenshots. Therefore, responders should preserve key records before broad cleanup begins.
Closing Without Root Cause
An incident should not close just because the alert stopped. Instead, the team should confirm how the event started, what controls failed, and what changes will reduce repeat risk.
Expected Outcome
After the team completes the playbook, the organization should have more than a closed ticket. It should have a clear record of what happened, what changed, and what must improve.
- The team confirms the incident, defines the scope, and stops further damage.
- The team preserves key evidence for review, leadership decisions, and possible audit needs.
- The team restores secure access and validates that affected accounts, devices, and systems are under control.
- The team identifies control gaps and updates policies, alerts, or response steps to reduce repeat risk.
- Leadership receives a clear summary of the event, response actions, remaining risk, and follow-up work.
