Somebody on your team is going to say "we should just rewrite this" in the next planning cycle. Maybe it's already happened. The backend has logic buried in controllers, nobody trusts the test suite because there isn't one, and the roadmap is waiting on an architecture that doesn't exist yet. That's the exact scenario a delivery lead at Brocoders described a client bringing to them: a working product, real users, and a backend where "all the logic sits in the controllers and no test coverage anywhere" — with a hiring plan riding on the answer (Brocoders).
Most of the advice on this decision collapses into "rewrites are risky, refactor when you can." True, and not that useful. The more interesting question is what actually separates the systems where a rewrite pays off from the ones where it quietly becomes a career event.
The real variable isn't code quality, it's whether the behavior is recoverable
Brocoders frames this as a "Specification Recovery Test" — can you write down what the system does without reading the code, from docs, tests, or people who still remember? If yes, refactor. If the behavior only exists inside code nobody wrote comments for and the original authors are gone, you're looking at a rebuild, because there's nothing left to preserve incrementally (Brocoders).
Razoyo draws a sharper, narrower line: refactor by default, and reserve a full rewrite for two situations — the data model literally can't express what the business now sells, or the architecture has a ceiling that gets worse as volume grows. Everything else, including code you find genuinely unreadable, is scheduled as a refactoring problem, not an existential one (Razoyo). That's a useful corrective to the instinct that ugly code justifies a rewrite. Ugly code is a staffing and velocity problem. A data model that can't represent multi-tenant customers is a different category of problem entirely.
Teamseven's version of the same logic names the three ways a rewrite actually fails in practice: the old system encodes years of edge cases nobody documented, the business doesn't pause while you rebuild, and nothing ships until everything ships — so the "clean start" date keeps moving (Teamseven). If you've run a rewrite before, that third one is the one that actually kills teams. A refactor failing mid-flight leaves you with a working system that's still ugly. A rewrite failing mid-flight leaves you with two systems, one of which doesn't exist yet.
What AI changed, and what it didn't
The practical argument for refactoring over rewriting used to rest entirely on economics: understanding and testing a legacy system consumed the same scarce engineering time that visible features needed, so teams kept deferring it (Sonar). That math is shifting. On The Next Commit podcast, Tim Ottinger describes pointing an AI agent at an unfamiliar codebase to map its architecture and derive an initial test safety net in minutes rather than days — reaching "70 or 80%" coverage, which he calls "probably enough to do some refactoring" (Sonar). Version control turns a large structural refactor into a disposable experiment: if it doesn't improve things, you abandon it and try again.
Brocoders' framing is sharper on what this does and doesn't solve: AI made writing software dramatically cheaper, but understanding software still costs what it always did. The decision now turns on how recoverable the specification is, not on how ugly the code looks (Brocoders). An AI agent can generate a test harness fast. It cannot invent the business rule that one customer gets billed differently because of a deal cut three years ago that lives only in a conditional nobody's touched since.
This is also why the incoming-engineer version of this decision has a mandated delay built in. A documented first-month plan for someone inheriting an unfamiliar codebase explicitly bars refactoring or rewriting in week one — the first move is just getting the app running, deployable, and rollback-capable, then pinning current behavior with characterization tests in week two. Only in week four, with that evidence in hand, do you write the keep-fix-replace memo (Axonbuild). The decision gets worse, not better, when it's made in week one on gut feeling.
The decision you can actually schedule
Run the recoverability test module by module, not system-wide. Most real systems land in the middle: one or two subsystems hit an actual architectural ceiling and justify a bounded rewrite, while the rest just need tests, ownership, and time. That's the strangler pattern Teamseven recommends — replace piece by piece, keep shipping value the whole way, and if priorities shift halfway through you're left with a working system instead of an abandoned bet (Teamseven). The module-by-module framing matters more than the final verdict. Teams that litigate "rewrite or refactor" as one binary choice for the whole codebase are usually arguing about the wrong unit of analysis.
