All posts
Career

Three Pages a Night for Two Weeks and You're Starting to Hate a Job You Used to Love

On-call advice is written for the person who sets the rotation. If you can't change it, the useful question is whether yours is actually broken — and how you'd prove it.

12 views
Photo: Vladyslav Huivyk / Pexels
A conversation with a TrueTalk advisor about: Three Pages a Night for Two Weeks and You're Starting to Hate a Job You Used to Love

3:12am, and it's the disk alert again

You know which one it is before your eyes focus. Same host, same threshold, same clear-out-some-logs non-fix you've now performed eleven times. You ack it. You watch the graph come back down. You lie there at 3:40 with your pulse still up, and somewhere around four you open a job board on your phone with the brightness turned all the way down so you don't wake your partner.

Third night this week. You're eight days into a fourteen-day shift.

One-in-three rotation. So this is a third of your year.

Every post you've found was addressed to someone else

Search anything about on-call and you get the same well-meaning genre: alert hygiene, runbooks, error budgets, follow-the-sun coverage, blameless post-mortems. Most of it is correct. All of it is written for the person who owns the rotation.

You don't own the rotation. You're one of three people covering a system that would need five to cover comfortably, and the reason there are three is a headcount decision made in a meeting you weren't in. A checklist of things your employer isn't doing is a tidier description of the thing keeping you awake.

The question you're actually asking at 4am is narrower, and almost nobody writes about it. Is this rotation genuinely abnormal, or am I just tired and dramatic? And if it is abnormal, can it change from where I sit — or is the only lever I have the resignation?

Take them in that order. The second depends entirely on the first.

Two weeks of counting beats six months of resenting

Right now you have a feeling. Feelings don't survive contact with a manager who has his own quarter to worry about. Numbers do, and you're an engineer, so this part is the easy part.

For your next two shifts, log five things per page: when it fired, whether it woke you, what you did, whether the thing it warned about was actually happening, and whether a human was genuinely required or a script could have done it.

Four numbers fall out of that.

Pages per shift, split into business hours and out-of-hours. Different animals. Averaging them hides the entire problem.

Wake-ups per week. Wake-ups specifically, not pages. A page at 10:40pm while you're still up and a page at 3:12am are different events, and only one of them is taking your health.

Actionability. Of the pages that woke you, how many needed a decision only a person could make? If it's under half, you don't have an on-call problem. You have a monitoring problem wearing an on-call costume.

Concentration. Sort by alert name and by service. If the night pages turn out to be concentrated — on one alert, or on a service whose owner is somebody else — that's the single most useful thing your log can tell you.

That last one changes a conversation. "On-call is bad" gets sympathy. "Twenty-two out-of-hours pages across two shifts, fourteen of them from one alert on one service, three of the twenty-two needed a human — here's the list" gets a ticket with a name on it.

Google's Site Reliability Engineering book, free to read on their site, puts a ceiling of no more than two events per 8-to-12-hour on-call shift, on the reasoning that handling one properly — including the postmortem — averages about six hours. Useful as a reference point, mostly because a company famous for scale wrote down a ceiling at all. Worth saying plainly, though: that book describes teams with resources yours does not have.

Four shapes of broken, and they need different arguments

Once you have the counts, it's worth asking which of these you're looking at.

Noisy. Lots of pages, few of them real. Mostly fixable from where you sit, and the most maddening, because every one of those alerts was written by someone who thought it mattered and deleting other people's alerts feels like vandalism. Do it anyway — with the log in hand, and in the open. Propose thresholds. Propose demoting a page to a ticket. Propose deletion, and let people object in writing.

Genuinely on fire. Few alerts, all real, all at 3am. The system is actually breaking. No amount of tuning helps and you should stop trying, because this is a reliability-work-versus-feature-work argument and you will not win it alone. What you can do is make the cost visible in hours, in the same place the feature work is tracked, every sprint.

Understaffed. The pages are reasonable. There just aren't enough humans, so the shift never really ends. The hardest of the four, because nothing inside your control touches it, and every proposal comes back with sympathy attached.

No handover. Page volume is survivable, the alerts are honest, and nothing gets written down between shifts. You come on cold on Monday morning into somebody's half-finished Saturday incident and spend two hours reconstructing what was already known by someone who's now asleep. Cheap to fix, and almost nobody fixes it, because it never shows up in an alert count. Your log won't catch this one. Your memory will.

The diagnosis matters because it tells you what to ask for. Asking for headcount when your real problem is one misconfigured alert makes you look like you're complaining. Asking for alert tuning when the real problem is that three people cannot cover a 24/7 system makes you look like you've solved it, and buys your employer another year.

Ask for one thing

The temptation is to bring the whole grievance. Don't. Managers act on single, cheap, specific asks.

Say the numbers came out concentrated — one alert, fourteen of your twenty-two night pages, none of them requiring a person. Then your ask is: this alert stops paging out-of-hours as of Monday, it files a ticket instead, and the owning team is on the thread. That's a decision your manager can make in a two-line reply. If it turns out the alert has to keep paging, you've learned something real about how the business values that service, and you can ask the follow-up question: then who else goes into the rotation?

One more thing worth checking, quietly, before any of this: what your contract actually says about on-call. Whether standby hours are paid, what the compensation is, what the rest-period rules are. This varies enormously by country and by contract, and an article cannot tell you which regime you're in. Read yours. If it looks like it isn't being followed, that's a question for someone qualified in employment law where you live, not for a forum.

The reassurance I'd push back on

You will be told that any culture can be fixed from below with enough patience and good data.

Sometimes it can't. If the noisiest service belongs to a team that doesn't report to your manager, and their manager has no incentive to care, you have no lever. If headcount is frozen, alert tuning buys you a better bad rotation. And if you write a clear proposal, your manager agrees warmly, and then nothing happens — twice — that is your answer. Nobody's being hostile. That's just how this org says no.

So set the date before you start. One quarter. A written proposal with a named owner and a date on it, and one defined thing that has to have moved by the end.

If it hasn't moved, you already made the decision, and you made it in daylight.

Separately: if you're still waking at three on the weeks you aren't carrying the pager, book the GP. That stopped being a rotation problem a while ago, and chronic sleep disruption gets treated medically rather than organisationally.

Turning a log into three sentences

Using the numbers above, it goes something like this:

"Over my last two shifts I took twenty-two out-of-hours pages. Fourteen of them came from one alert on one service, and three of the twenty-two needed a person to decide anything. I'd like that alert to file a ticket instead of paging overnight from Monday, with the owning team on the thread."

No adjectives. No history. Nothing about how you feel about any of it. It fits in a Slack message and your manager can answer it in a line.

Getting from two weeks of messy notes to those three sentences is the awkward bit — particularly saying it without sounding accused or accusing. DevOps Dan is a persona on TrueTalk: a senior DevOps engineer specialised in CI/CD, infrastructure automation and monitoring. An AI you can argue with about alert thresholds at 4am, not a colleague. First conversation's free, then it's paid. Paste in the alert names and the counts and work out which of the four shapes you're in.

Quit on a Saturday, not at 4:40am

Leaving a job over on-call is a completely legitimate thing to do. People do it constantly and most of them don't regret it.

Just don't decide it on the third night with a job board open and the brightness down. Decide it rested, with two weeks of numbers in front of you — and preferably with somewhere to go where you've already asked, in the interview, exactly how many times a week the person doing your job gets woken up.

on-callburnoutdevopssrecareer