
Put Customer Output First
Rick ElmoreCEO · Simply NotedWe build our own handwriting robots at Simply Noted, six patents pending, and the control software running the first generation of machines was starting to show its age a couple years in. Rising maintenance costs, sure, but also just fear, nobody wanted to touch the code because one person understood most of it.
The rule we landed on was simple: if a system touches customer facing output (the actual note quality, timing, tracking) we rebuild it properly, even if that's slower and more expensive up front. If it's internal tooling that nobody outside the company ever sees, we patch it and move on until it's genuinely blocking something. That single filter kept us from burning months rebuilding an internal reporting tool nobody complained about, while forcing us to prioritize the robotics control layer that clients actually feel the effects of.
It also meant we sequenced changes without ever shutting down a production line. We'd stand up the rebuilt piece next to the old one, run both, and only cut over once the new one had proven itself on real orders for a few weeks. Nothing broke customer facing, and that's the only metric that mattered to us.
Rick Elmore, Founder/CEO, Simply Noted (simplynoted.com)
Target Costly Isolated Components
Val NarodetskyCEO · OdesaThe first thing I do with an inherited stack is map maintenance spend against revenue dependency, not against how ugly the code is. Ugly code that quietly processes orders every hour is not the problem. The problem is usually a system nobody can name an owner for that still costs real money every month to keep breathing.
My sequencing signal is simple. I look for the component with the highest cost of change and the lowest number of upstream dependencies, and I touch that one first. Low dependency means I can rebuild or retire it without pulling the whole business down with it, and high cost of change means the savings show up fast enough to fund the next piece.
I retire anything with no owner and no traffic. I replace anything where a mature off-the-shelf option already does the job, because rebuilding commodity functionality is how budgets disappear. I rebuild only where the system carries logic that is specific to how that business makes money.
On cost, the framing I bring to a board is total ownership over a few years, not the quoted build price, because maintenance, staffing, and integration work usually dwarf the initial figure. Where we help clients, the development economics run roughly 60 to 70 percent below U.S. equivalents, which changes what is worth rebuilding versus tolerating for another year.
Address Greatest Failure Exposure
Assaf SternbergFounder & CEO · TiroflxI prioritize the system where failure creates the greatest operational exposure, not necessarily the one everyone complains about most. In manufacturing execution, a dated tool may still be tolerable until it affects supplier records, quality documentation, compliance visibility, or shipment coordination. My decision rule is: replace first where the combination of failure risk and business dependency is highest. Sequence matters because replacing everything at once creates its own risk. Modernization should reduce fragility without making daily operations the test environment.
Let Maintenance Burden Set Priorities
Oscar MoncadaCo-founder and CEO · Stratus10The signal I watch is where engineering time goes. When most of a team's effort on a legacy system is maintenance instead of improvement, it's time to decide that system's future, and the decision depends on whether it's core to the business.
If it's proprietary and something the company can't live without, rebuild it. Every month spent patching is money that could go toward a better version of the thing you actually depend on. If there's a better, more secure option on the market, like a newer version or a managed service, replace it. If the application logic is fine but it's running on old hardware or an outdated operating system, such as a self-hosted database that needs manual replication, backup, and recovery, move it to a managed service and let the provider carry that burden. Anything nobody depends on anymore is a candidate to retire.
Sequencing is where disruption gets avoided. We don't treat a legacy environment as one unit. We assess each component individually for feasibility, cost, time, and security implications, because not every piece needs full modernization, and the ones that don't can wait. The components eating the most maintenance time move to the front of the line, and the dates come out of that assessment, not the other way around.
The rule I'd reuse is simple: are you maintaining something because it's valuable or because you can't get rid of it? Answering that honestly for each component tells you what to rebuild, what to replace, and what can wait.
Protect Delayed Revenue Streams
Abhinav GuptaFounder · ProfitjetsWhen dealing with legacy systems with increasing maintenance costs, I prioritize work by estimating the impact of delayed cash or revenue streams. Typically, we work from the perspective of which component, if delayed, would impact cash or revenue the most. While there are flexible rules around this, such as situations that require immediate resolution, a person (referred to as a "decision owner") is tasked with enforcing the rule to help focus teams on the most important work.
Modernize Low-Risk Periphery First
Girish SongirkarDelivery Manager, Enterprise Software Engineering · ArionerpThe only decision-making guidance regarding legacy system upgrades is the concept of the Operational Friction Threshold, which dictates that one should start with modules where a rising cost of maintaining the system coincides with a lower risk of an operational failure and use that momentum to tackle the risky core. When acquiring legacy architecture with a high cost overhead one needs to plan the order in which to tackle the modules using more than technical debt as a criterion. Hence, I apply a risk-based approach that involves plotting every legacy module on two axes: the degree of operational impact on its core and the cost of modernization.
Using my method, I identify four quadrants of legacy footprints and decide which ones to retire, replace, and reconstruct. High-impact, low-cost modules need to be reconstructed right away so that a quick payback in the form of operational return gain can be achieved. Low-impact, expensive modules should be retired immediately or their functions have to be frozen in order to stop draining the resources. The real sequence applies to all other modules. My rule of thumb is to start with modifications of peripheral modules with a low degree of risk.
When we renovate peripheral modules first, we help our engineers stabilize the infrastructure and launch data integration pipelines without disrupting the manufacturing or delivery operations. For example, integrating information processes of a secondary reporting module or a non-core vendor procurement system is more practical than trying to reconstruct the core engine all at once.
Measure the True Bottleneck
Shane LarrabeePresident/Founder · FatLab Web SupportStart with whatever is actually costing you, and measure it before you pick a fix. When we reviewed load across our servers this summer, one site was the only real hot spot in the fleet: a nonprofit's headless setup, where a separate front end pulled every page out of WordPress through an API. Each site build fired about 3,711 queries at the back end.
The obvious move was a cache. We added one and timed it: 0.93 seconds with it on, 0.87 with it off. The cost wasn't the database at all. The API layer was rebuilding its whole query schema on every request, and nothing underneath it could fix that. That one measurement moved the plan from tuning to rebuilding.
My sequencing rule is to hold the data still and change only the layer on top of it. We rebuilt the front end as a native WordPress theme that reads the existing content exactly as it's stored. No renamed fields, no migrated data, no restructured settings. Every change had to pass one test: if we re-imported the production database tomorrow, would this still render with no manual fix-up? If the answer was no, we found another way to do it.
That rule is what kept the switch uneventful. The rebuild went live on September 16, and we retired the old stack the next day. Because the content never moved, there was nothing to reconcile at launch, and the load problem left with the pattern that caused it.
Use Human Intervention as a Signal
Aigars PilmanisFounder · VolRadarSequence by blast radius, not by how much the code annoys you. I run VolRadar, a bootstrapped options and volatility analytics platform, and everything I maintain falls on one side of a hard line: either it sits on the path that has to publish end-of-day data every trading session, or it doesn't. The rule I work to is that nothing on that path gets rebuilt in place -- it runs alongside the old version until both produce the same output, then cuts over. Everything off that path can be retired far more aggressively. And the signal I trust isn't maintenance cost, it's how often something needs a human to babysit it. Cost tells you what's expensive; babysitting tells you what's fragile.
Confront Critical Strategic Mismatches
Srinivasan NarayananSolution Delivery Lead Oracle Cloud Manufacturing & Supply Chain · Milwaukee ToolWhen I inherit legacy systems with increasing maintenance costs, I avoid making decisions based on cost alone. A system may be expensive to maintain, but if it supports a critical operation with no viable alternative, retiring it too quickly can create more risk than value.
I start by assessing scalability: can the system support expected business growth, higher transaction volumes, new locations, integrations, or changing customer needs? I then review the full cost of ownership, including maintenance effort, licensing, infrastructure, vendor support, and the hidden cost of recurring workarounds. Another important factor is whether the organization has the right skills, knowledge, and resources to run and enhance the system reliably. A platform that depends on one or two people with specialized knowledge creates a significant continuity risk.
Most importantly, I validate whether the system supports the long-term strategic objectives of the business. If it cannot enable the future operating model, digital capabilities, data visibility, or growth plans, it becomes a strong candidate for replacement or modernization.
The decision rule that has helped me most is to prioritize systems with high business dependency but declining strategic fit. These systems need attention first because the risk of doing nothing continues to grow. However, I sequence the change through phased transition plans, parallel testing, stakeholder readiness, and clear fallback options. This protects critical operations while allowing the business to modernize in a controlled way.
Retire Tools Far From Revenue
Victor SmushkevichFounder · Tested MediaMy rule is to sequence by distance from the moment a lead turns into a booked job. The closer a system sits to that moment, the later I touch it, and the more carefully I run old and new side by side.
What I keep running into in agency work is a pile of old tools nobody owns. A form plugin from years ago, a CRM full of dead workflows, tracking scripts that overlap. Rising maintenance cost hides in that pile, and most of it sits far from revenue.
So I ask one question about each piece. If this went dark for a week, who would notice first? If nobody would, retire it now. That's the cheapest win, and it cuts the maintenance bill right away. If only staff would notice, replace it next, because a mistake stays internal. If a lead or a paying job would notice, it goes last. I don't swap it in place. I build the new version next to the old one, route a small slice of traffic through it, and cut over only when the two match.
Rebuild is the most expensive answer, so it has to earn its spot. I reach for it only when replacing would just move a broken process onto new software.
Fix Broad Downstream Failure Points
Siim KostabiCEO · PagelootAround 2021, Pageloot's early codebase had grown into something where adding a simple feature could take two weeks because everything touched everything else. We had to decide what to actually fix first without breaking 20,000+ active users.
The signal we used: find the thing where a single bug causes the most downstream failures. Not the oldest code, not the ugliest code. The one where a problem in module A breaks C, F, and G simultaneously. That's what you rebuild first, because it's also the thing where a fix in module A improves C, F, and G simultaneously.
The second filter was maintenance hours per quarter. We tracked roughly how long each area took to patch and babysit. Anything consuming more than 20% of an engineer's week for something that wasn't growing the product got flagged. Retire if users won't notice, replace if the function is still needed but the implementation is brittle, rebuild only if the logic itself is worth keeping and the cost of rebuilding is less than two quarters of maintenance drag.
What we didn't do was let perceived business criticality drive the order. The "critical" systems usually had the most eyes on them already. The dangerous ones were the mid-tier modules everyone assumed someone else understood. Sequence by blast radius and maintenance cost, not by how important the feature feels in a roadmap conversation.
Follow the Blocker Horizon
Evgeny LeonovChief Technology Officer · Ronas IT | Software Development CompanyWe sequence legacy work by blocker horizon: what stops a critical business flow now, what has a dated path to failure, and what is merely old. We fix the first group, schedule the second, and leave the third alone until its risk becomes concrete. Age and maintenance cost can trigger a review, but they don't decide the order.
Once the order is set, retire a capability that no current operation uses. Rebuild only the smallest area where coupling makes an isolated change impossible. A full rewrite needs evidence that staged replacement can't contain the risk.
We applied the rule on Lainappi when Google announced the shutdown of Firebase Dynamic Links. The provider's deadline turned a background dependency into scheduled work with a known failure date. We moved deep linking to AppsFlyer as a self-contained module with one entry point, so the migration touched little outside that boundary. Google also deprecated Container Registry on the same product. The registry move stayed inside the CI configuration, confirming that a contained replacement was enough.
When a dated service sits behind a stable boundary, replace it there and keep the rest of the system untouched.
Validate Data Counts Before Cutover
KEITH YUNXI ZHUChief Executive · TKEG Expat INCTKEG Expat is a corporate-services firm, and when we moved our public website off the no-code platform to our own server-rendered site for SEO and performance, we did not retire the old system. It stays in place as the data source, a sync imports its data into MySQL and the new site renders from that mirror, so we rebuilt the part customers see while the no-code platform remains a shell for parts of our data and operations layer.
Where we did retire something, we checked the replacement against the plan before the switch. For example, in September 2026 we retired our blog into Product Guides by routing all 39 posts through 301 redirects, and before the switch the live redirect rows matched the plan, every post had exactly one row and the redirect map built 39 keys. After that, the deployment and the switch ran in a fixed order, with a rollback copy kept.
In August 2026 a rate limit on the platform's API made a routine sync write empty tables, and the run still logged success and published a blank blog. Because the legacy source was still live, we restored the data from it. The lesson we took from it is to look at row counts instead of trusting a job that says "success", and to keep the legacy source live.
Start With Reversible Changes
Rahul AgrawalFounder & CEO · QuickIntellSequence by blast radius: change first whatever is cheapest to undo, and touch the system of record last.
In healthcare operations, the EHR and practice-management system are the systems of record. Much of the rising maintenance cost sits around them: scripts, portal logins, spreadsheets and manual re-entry that move information between the EHR, payers and clearinghouses. Replacing the core system is the riskiest move with the longest payback. Replacing the brittle glue around it is lower risk and pays back sooner.
So we add a new layer beside the legacy system rather than inside it. At QuickIntell, our agents read from the existing EHR and billing systems, prepare the work, and hold high-impact updates for staff approval before anything is written back. Teams can retire a manual step or a fragile script without touching the core system, and the old path stays available until the new one has proven itself.
The signal I use is reversibility. If a component fails, can we spot it quickly and undo it cleanly? If yes, it can move early. If not, it waits until monitoring and a rollback path are in place.
Migrate Independently Verified Targets
When maintenance costs rise, I do not rank legacy systems by the size of the invoice. I rank them by business criticality, change isolation and reversibility.
A costly system with unclear dependencies stays in place until we can observe what it actually does. A smaller component with a clean interface becomes the first migration target because it lets the team prove the replacement method without putting payments, customer delivery or other critical operations at risk.
My decision rule is: replace the first component whose behavior can be independently verified and rolled back within one operating cycle. Run the new path in parallel, compare outputs, keep the old path recoverable, and only then cut over.
If the team cannot define equivalence and rollback before development begins, the system is not ready to retire—however expensive it feels.
Flag Reconciliation Labor Early
Todd HarmonFounder & Owner · BathGemsMaintenance expense becomes misleading when a legacy system is cheap to run but expensive to verify. In merchandising, the real warning sign is spreadsheet work reconciling price, availability, finishes, and delivery data across disconnected records. That labor masks risk because people make the process appear dependable until volume or staffing changes.
We treat recurring reconciliation as a rebuild signal when it alters decisions or financial reporting. First retire duplicate sources that create competing versions of truth. Then replace tools that cannot publish reliable data at speed. Rebuild the core only when its data model forces manual interpretation, because automation on top of ambiguity makes errors travel faster.
