Observability before the rewrite
Most “we need to rewrite” meetings are really missing telemetry. Instrument the path customers feel — then decide what deserves a new stack.

Friday afternoon. The app feels slow. Support is loud. Someone opens a slide deck titled “Platform 2.0.” New framework. New service cut. New datastore. It looks decisive.
It rarely starts with the boring question:
What do we actually know about how this fails in production?
That is not a buzzkill. That is the difference between a migration with a map and a migration with a mood.
The rewrite urge is a symptom
Rewrites are expensive in calendar time and in trust. While the new system is unfinished, the old one still serves customers — usually with the same blind spots that triggered the panic.
If you cannot answer these three questions with evidence, you are not ready to rewrite:
- Where does latency actually accumulate on the critical path?
- Which failure modes are customer-visible vs. merely noisy?
- Who owns the path when it breaks at 2 a.m.?
Missing answers are an observability problem first. Architecture second.
Instrument the path customers feel
Start with the journeys that make money or create tickets — checkout, invite, sync, report generation. Not every hop in the diagram. The hops customers wait on.
For each journey:
- Golden signals — latency, traffic, errors, saturation on the owning service
- Trace the request across the 2–4 hops that matter
- Log with correlation IDs you can paste into a ticket without a scavenger hunt
You do not need a perfect platform. You need enough light to stop arguing from anecdotes.
A team that can show “p95 checkout is 800ms, 600ms is inventory lookup” has a different conversation than a team that says “it feels slow.” Same app. Different adulthood.
Separate “bad code” from “bad load shape”
Teams often blame the runtime when the real issue is a query plan, a lock, a fan-out, or a retry storm. Traces make that argument short.
Once the bottleneck is named, the fix is usually incremental: an index, a batch, a cache with a clear invalidation story, a timeout that actually times out.
That work is less glamorous than a rewrite deck. It ships faster. It teaches the team how the system behaves under stress — knowledge a greenfield rewrite will have to relearn anyway.
When a rewrite is justified
Observability can also prove the opposite: the boundary is wrong, the model cannot evolve, or the operational burden is structural. In that case, rewrite with a thinner wedge:
- Carve the hottest path behind a stable interface
- Dual-run with compare metrics
- Cut over by traffic slice, not by “big bang weekend”
You are still rewriting — but with a map, not a mood.
Takeaway
Before you approve a rewrite, buy clarity. Instrument the customer path. Name the bottleneck. Then choose incremental repair or a deliberate migration.
Rooms that skip this step usually pay twice: once for the rewrite, and again for the incidents the new system inherits.
Found this insightful? Like or share with your team:
Spread good engineering craft & architecture lessons.
Comments
Email is not published. Keep it professional.