TL;DR: Treat burnout as a measurable systems defect, not an attitude problem: pull the paging data, delete or fix whatever pages without needing a human, resize the rotation, then report the before-and-after numbers so leadership sees engineering work rather than mood management.
How to approach it
Frame the story as diagnosis, intervention, verification. Take employee reports seriously and use paging data to understand the working conditions behind them. People do not need a dashboard to justify asking for support. If you have never owned a rotation, say so and answer hypothetically with the same structure.
A strong answer
Start with detection. Look for workload patterns alongside what people report: pages per engineer per week, night pages between midnight and 6am, acknowledgement times drifting upward, shift-swap requests clustering, and the same two names on every weekend. Add the human signals: someone answering Slack on approved leave, juniors afraid to acknowledge without a senior awake, retros where paging tops the grievance list every sprint.
Say the quiet part out loud: in a five-person rotation, more than a couple of actionable night pages per person per week means the system is defective, because every page spends the next day's engineering quality.
Then the interventions, ordered by effort:
- Fund alert hygiene first. Pull the top ten firing alerts by volume. Delete the ones nobody acts on, deduplicate flapping ones, set thresholds that survive brief spikes, attach runbooks so a page contains its own first step. This takes engineering time; measure the volume reduction without assuming a fixed saving.
- Severity discipline. Page only for "needs a human now". Everything else routes to a ticket or a business-hours queue. An SMS and a Jira ticket are different products; many organisations accidentally ship everything as the former.
- Rotation structure. Size the rota so primary plus secondary exists, spread the hardest service across trained people instead of one hero, and use the offshore-onshore split properly: handover at evening in Bangalore hands the working day to onshore, converting 3am pages into working-hours pages.
- Close causes, not symptoms. Any page that recurs twice earns a postmortem action item with an owner. Track pages removed per month as a team metric alongside pages created.
Then verification: re-pull the same numbers after six weeks. Night pages per engineer and the share of pages that were actionable are the two that matter, and quoting your before-and-after is what turns this answer from opinion into evidence.
What interviewers probe next
"What if paging metrics look healthy but someone says the rotation is harming them?" Listen privately and arrange support or workload relief with the appropriate manager. Paging counts miss sleep disruption, service complexity and personal circumstances; they cannot invalidate a person's report.
"Leadership refuses headcount for a bigger rotation. Now what?" Prioritize alert review and severity routing within an explicit engineering allocation, then measure the effect. Cross-train adjacent teams into the secondary slot, automate the top recurring page, and only then revisit headcount with the residual number.
"Is some suffering not just part of operations?" Pages are the job; unactionable 3am pages are waste. Calling waste toughness is how teams lose exactly the experienced responders they need most mid-incident.
Common mistakes
Fixing morale with food and pep talks while the noisy alerts stay live. Candidates describe the pizzas; interviewers wait for the deleted alerts that never come.
Claiming the goal is zero pages. Healthy systems still page when production breaks; the win is rare, brief, actionable pages handled by a rested responder.
Using only paging counts to claim the problem is solved. Pair before-and-after workload measures with confidential feedback and check whether people can actually recover between shifts.
Stopping at "we encouraged people to take breaks", which changes nothing about the pager and therefore nothing about the problem.