Five Years In, Watching a Model Do Your Job in Eleven Seconds
An eleven-second draft can still leave costly verification, hidden context, and accountability with you. Audit last month's work before making a career bet.

Eleven seconds, then the mortgage
You pasted in a ticket budgeted at a day and a half mostly to see what would happen. Eleven seconds later, a model had produced about eighty percent of what you planned to write. Its variable names were better than yours. It missed the timezone detail and invented a helper that does not exist. You found both errors in nine minutes, then sat there doing arithmetic about the mortgage.
At five years in, you are not junior enough to be cheap or senior enough to be structural. Vendors say it is over; colleagues on the internet say nothing has changed. You have good reasons to distrust both camps. The inevitability pitch sells software, while the colleague dismissing the change may be defending a career shape that took a decade to build.
Doing nothing carries a cost because you have a couple of years of slack in which to reposition deliberately instead of in a panic. If the sceptics are right, a close look at your work costs little. If they are wrong, you looked while you still had room to move. That does not justify an irreversible decision. Do not sell the house, retrain as an electrician, or accept a management job you would hate on the strength of a forecast nobody can make reliably. Whether AI replaces software engineers remains unknown, and confidence is not evidence.
Audit the fourteen things you did
The answerable question concerns your own work. Open last month's board and read the tickets you closed rather than keeping a two-week diary. Ten minutes gives you a more honest list than you are likely to write while watching yourself work. Use the fourteen things you actually did last month, whatever your exact total happens to be.
The list contains a few hours of genuinely writing new code and a great mass of everything else. That includes reviewing somebody's PR, sitting in a call to work out what the product manager means by just show the total, chasing a flaky test, reproducing a bug that appears for one customer, deciding whether a schema change is worth the migration, and explaining to support why a request takes three weeks rather than three hours.
Score every item three ways. Avoid asking whether a model could do it because a partial yes applies almost everywhere and gives you little guidance.
Is the output cheap to verify? When a wrong answer announces itself through a failed test, an unrendered page, or a type error, the task is more exposed. Generation is useful where checking is cheap. When verifying a wrong answer requires two days and risks a customer incident, you still have to perform the expensive work.
Does it need context outside the repo? An endpoint may retry because of a 2023 conversation with a partner who bills you for failed calls, even though the code never records that reason. Political history, a contract, and what the VP said in April can decide an implementation. Work resting on that knowledge is less exposed and often invisible on your CV.
Who carries it when it is wrong? Somebody signs off, takes the 2am call, or sits with the customer. That accountability does not transfer to a model, and organisations have reason to be cautious about pretending it does.
Across fourteen items, you will probably find a short group that scores badly on all three tests and a shorter group that scores well on all three. The latter gives you a specific direction grounded in your company and codebase. Treat the scores as an audit of today's work, not a prophecy.
The measurements are still unsettled
METR's 2025 randomised trial found experienced open-source developers took longer with AI assistance than without while believing they were faster. Its published study covered 16 developers, 246 tasks, and mature projects where participants averaged five years of prior experience (METR paper).
METR later marked that result as out of date. Its February 2026 follow-up involved 57 developers, more than 800 tasks, and 143 repositories, and pointed the other way. METR calls it only very weak evidence, partly because 30% to 50% of developers said they chose not to submit tasks they did not want to do without AI (METR update). Selection makes the follow-up difficult to interpret, while your own feeling of speed is unreliable in either direction. That is why the audit examines the shape of your work instead of claiming a speed forecast.
The Stack Overflow developer survey has asked working developers about these tools for several years. Read the survey itself rather than a summary of a summary. Twenty minutes with the source will usually be less alarming and less dismissive than the commentary around it.
Four moves, each with a cost
You could specialise deeply in embedded, real-time, compiler, safety-certified, or regulated work where general tools struggle. Tie the choice to physical constraints or regulation rather than a framework. A narrow market has fewer employers, and a niche that fades can leave you worse off than a generalist.
You could move into management, which scores well on the accountability test. It is a different job rather than a promotion: you stop making things, your hands-on skills fade, and returning can be difficult. If fear rather than interest sends you there, you may become bad at work you dislike and feel stuck in it.
You could leave the industry. That is legitimate and more expensive than it sounds when you are five years in with a mortgage. Expect a pay cut and junior status in another field, with no evidence that the destination is safer rather than merely discussed less.
You could stay and deliberately seek work that scores well on verification cost, outside context, and accountability. Becoming excellent at reviewing and integrating is often a sensible answer, and you can start on Monday without announcing a career change. It still leaves you exposed to changes in tooling and company decisions.
None of the four paths guarantees safety. Specialisation can strand you, management can trap you in the wrong job, leaving resets your earnings and seniority, and staying can leave the wider employment risk untouched. Make a reversible adjustment first when the forecast is uncertain.
The scores also move. Verification in particular could become cheaper, changing the whole analysis, so repeat the audit in a year. A strong task portfolio cannot protect you if your company cuts the team and your name appears on the list. Keep the other half practical: build savings, maintain a current CV, and stay in contact with people who would take your call.
The dread may not disappear after the audit. It can become smaller and more specific, leaving you worried about four particular tickets rather than the whole trade. You may also score your own work badly because you are too pessimistic about tasks that bore you and too generous about tasks you enjoy.
AI Researcher is a TrueTalk persona whose areas are artificial intelligence, machine learning, and technology trends. Bring the fourteen items and challenge the scores line by line. The free tier includes ten conversations and one hundred messages per day. Use that argument to test your assumptions, not to outsource the career decision.
The eighty-percent capability is not going backwards, but the missing twenty percent still required you to know where to look. Much of the paid work at this stage is not typing. Neither fact tells you whether software engineers will be replaced or when the remaining work will change, so do not turn it into a forecast.
