Architecture & DataEnglish

Observability before the rewrite

Most “we need to rewrite” meetings are really missing telemetry. Instrument the path customers feel — then decide what deserves a new stack.

1views

0

Cold crystalline distributed tracing beams analyzing glowing hot monolithic legacy core

Friday afternoon. The app feels slow. Support is loud. Someone opens a slide deck titled “Platform 2.0.” New framework. New service cut. New datastore. It looks decisive.

It rarely starts with the boring question:

What do we actually know about how this fails in production?

That is not a buzzkill. That is the difference between a migration with a map and a migration with a mood.

The rewrite urge is a symptom

Rewrites are expensive in calendar time and in trust. While the new system is unfinished, the old one still serves customers — usually with the same blind spots that triggered the panic.

If you cannot answer these three questions with evidence, you are not ready to rewrite:

  1. Where does latency actually accumulate on the critical path?
  2. Which failure modes are customer-visible vs. merely noisy?
  3. Who owns the path when it breaks at 2 a.m.?

Missing answers are an observability problem first. Architecture second.

Instrument the path customers feel

Start with the journeys that make money or create tickets — checkout, invite, sync, report generation. Not every hop in the diagram. The hops customers wait on.

For each journey:

  • Golden signals — latency, traffic, errors, saturation on the owning service
  • Trace the request across the 2–4 hops that matter
  • Log with correlation IDs you can paste into a ticket without a scavenger hunt

You do not need a perfect platform. You need enough light to stop arguing from anecdotes.

A team that can show “p95 checkout is 800ms, 600ms is inventory lookup” has a different conversation than a team that says “it feels slow.” Same app. Different adulthood.

Separate “bad code” from “bad load shape”

Teams often blame the runtime when the real issue is a query plan, a lock, a fan-out, or a retry storm. Traces make that argument short.

Once the bottleneck is named, the fix is usually incremental: an index, a batch, a cache with a clear invalidation story, a timeout that actually times out.

That work is less glamorous than a rewrite deck. It ships faster. It teaches the team how the system behaves under stress — knowledge a greenfield rewrite will have to relearn anyway.

When a rewrite is justified

Observability can also prove the opposite: the boundary is wrong, the model cannot evolve, or the operational burden is structural. In that case, rewrite with a thinner wedge:

  • Carve the hottest path behind a stable interface
  • Dual-run with compare metrics
  • Cut over by traffic slice, not by “big bang weekend”

You are still rewriting — but with a map, not a mood.

Takeaway

Before you approve a rewrite, buy clarity. Instrument the customer path. Name the bottleneck. Then choose incremental repair or a deliberate migration.

Rooms that skip this step usually pay twice: once for the rewrite, and again for the incidents the new system inherits.

Found this insightful? Like or share with your team:

Spread good engineering craft & architecture lessons.

1views

0

Comments

Email is not published. Keep it professional.

  1. No comments yet.