Reviving a legacy system, one careful step at a time

Reviving a legacy system, one careful step at a time

Somebody hands you a system. The developer who built it has moved on, the agency that looked after it has stopped answering, or it arrived inside a business you bought. It still runs, mostly. Nobody is quite sure how. And it does something important enough that you cannot simply turn it off.

We take these on all the time. We have inherited everything from VB6 applications older than some of our team to Laravel systems abandoned five major versions ago, and the process below is the one we follow every time. None of it is glamorous. All of it is the difference between a revival and a second failure.

One thing this post is not: a case for rebuilding. Whether to fix, modernise or replace is a decision with its own logic, and we have written a framework for making it. This post is about what happens once the answer is "keep it alive and make it better", which, in our experience, it usually is.

The state these systems actually arrive in

It helps to be clear about the starting point, because it is rarely just "old code". A neglected system usually arrives with some combination of the following. The code exists only on the server it runs on, or in a repository the previous developer still controls. Deployments happen by copying files by hand, if anyone still dares deploy at all. There are no tests. Passwords and API keys are written into the code itself. The server has not been patched in years, and nobody wants to reboot it because nobody is confident it will come back up.

If several of those describe your system, take some comfort in this: none of it is unusual, and none of it is fatal. Every one of those problems has a known fix. What matters is the order you fix them in, because the instinct to start improving the code straight away is exactly the instinct that breaks these projects.

Step one: get control of what exists

The first change we make is never a feature. Before anything else, we secure the raw materials of the system, because until they are in hand every other plan is provisional.

That means the code goes into source control, recovered from the live server if the repository is missing or being withheld. It means an inventory of every credential the system depends on: the domain, the DNS, the hosting account, the email service, the payment provider, the third-party APIs. In an inherited system these are usually scattered across the inboxes of people who left years ago, and chasing them down is dull work that pays for itself the first time something expires.

And it means verifying the backups by actually restoring one. Not checking that a backup job runs. Restoring one, and confirming the result is a working system. An unrestored backup is a rumour, and we have met more than one business whose nightly backup had been quietly failing for months.

Step two: make it run somewhere that is not production

A surprising number of legacy systems can only be worked on in production, because production is the only place they run. Sometimes the code depends on the exact versions of things installed on that one server. Sometimes it famously "only builds on the old laptop", and the old laptop is a genuine single point of failure for the business.

So we build a development environment that any of our engineers can spin up, and a staging copy of the live system loaded with realistic data. From that point on, no change meets a customer before it has run somewhere safe. This sounds basic because it is. It is also the single step that most reduces the fear of touching the system, and fear is what has usually kept it frozen for years.

Step three: build the safety net before touching anything

Three things go in before we change what the system does.

Error monitoring first, because a neglected system is almost always failing quietly somewhere, and has been for months. The team has learned to work around the failures without reporting them. Monitoring replaces folklore with a list, and the list is usually a surprise to everyone, including us.

Then tests around the behaviour we are about to touch. With no specification and no original developer, we test what the system does, not what anyone remembers it being supposed to do. Years of pricing rules, exceptions and hard-won fixes live in that behaviour, and the tests pin it down so that later changes cannot silently undo it.

Finally, a repeatable deployment. One command, same result every time, easy to roll back. The riskiest part of a neglected system is often not the code at all. It is the deploy, performed from memory, by hand, by whoever is bravest that day.

Step four: improve it in the order that reduces risk fastest

Only now does the improvement work start, and the order matters more than the ambition.

Security goes first. If the system depends on an operating system, framework or library that no longer receives patches, that is the one problem that gets worse while you think about it, so upgrades and patching lead the queue. Framework upgrades happen one major version at a time, with behaviour held deliberately identical. An upgrade the users notice has gone wrong.

Next, the single points of failure the audit uncovered: the one server, the one cron job, the one person. Then a quick win or two that the team can feel, like the report that takes forty minutes, or the crash everyone has a workaround for. Fixing something the staff complain about daily buys more goodwill for the project than any amount of invisible plumbing, and goodwill is a resource these projects run on.

The deeper work happens gradually: the worst parts of the system get rebuilt one at a time behind interfaces the rest of the code can rely on, while the whole thing keeps running. The industry calls this the strangler pattern. We think of it as replacing the worst room in the house without demolishing the street.

What we deliberately do not do

We do not start with a rewrite, for all the reasons set out in fix, modernise or replace: a rebuild throws away years of embedded business logic and asks you to rediscover it in production, at full price.

We do not refactor for taste. Code that offends the eye but works is paying its way. Code nobody can safely change is not, and the difference between those two is where the budget should go.

And we do not mix upgrades with behaviour changes. When a change breaks something, you want one suspect, not two.

What this looks like from the business side

From the outside, a well-run revival is mostly uneventful, which is the point. The system keeps running throughout. Work arrives in small slices, each one deployed and visible, so progress is something you can see rather than something you are asked to take on trust. There is no switchover weekend and no month of silence followed by a big reveal.

Commercially it usually runs as an audit up front, then a support arrangement that keeps the system monitored and patched, with the modernisation work planned in slices alongside. Priced that way, the work is stoppable at any point, and every slice leaves the system better than the last. That property matters more than any individual improvement, because it means the business controls how far the revival goes.

Common questions

Can you take over an application built by another developer or agency?

Yes, and it is a large part of what we do. We need access to the code, or failing that the server it runs on, plus whoever controls the domain and hosting accounts. Cooperation from the previous developer helps but is not required. We have recovered systems where they were no longer contactable at all.

The developer who built our system has left and there is no documentation. Is that recoverable?

Almost always. The documentation you are missing was probably never going to tell the whole truth anyway. The code is the one description of the system that cannot be out of date, and years of pricing rules and edge cases live in it. We read the code, pin the important behaviour down with tests, and write down what we find as we go.

Do we have to stop using the system while it is modernised?

No. The whole point of modernising incrementally is that the system keeps running throughout. Work lands in small slices, each one tested and deployed on its own, so the business never has to pause and there is no risky switchover day.

How long does it take to modernise a legacy system?

The first phase, getting the system under control and a safety net around it, typically takes a few weeks. What follows depends on what the audit finds and how far the business wants to go, and it runs in months of steady slices rather than one long project. Every slice leaves the system better than the last, so you decide how far to take it.

Codebased Limited. Registered in England and Wales, company number 14451447. Registered office: Rookery Offices, Rookery Farm, Wheaton Aston, Stafford, ST19 9QF.