Event-driven SaaS: when queues become the product
Idempotency, poison messages, and the ops work you cannot ignore once async is in the critical path.

Queues feel like free scalability in a design doc. In production they feel like a second product — one that fails quietly until a customer asks why their export has been “processing” for six hours.
The moment customer-visible state depends on a message, the queue is not infrastructure garnish. It is part of the UX.
Make delivery semantics explicit
At-least-once is the common default. That means idempotent handlers or you will double-charge, double-email, or double-provision.
“Exactly once” on a slide is not exactly once in your code. Design for duplicates. Test for duplicates. Log when you dedupe so support can explain what happened without guessing.
Poison messages are a product bug
A message that fails forever is not an ops curiosity — it is a stuck customer.
You need:
- a dead-letter path with ownership (a human name, not
#alerts) - a way to inspect payload safely (PII redacted, still useful)
- a replay that does not surprise side effects
If replay is scary, your handlers were never idempotent. Fix that before the next campaign doubles traffic.
Observe the lag
Queue depth without consumer lag and age-of-oldest-message is vanity metrics. They look green while customers wait.
Alert on customer-facing delay: time from “user clicked export” to “file ready,” not only on “messages in flight.” The queue cares about throughput. Your customer cares about done.
Takeaway
If the queue is on the path to “done” for a customer, treat it like a user-facing service — SLOs, runbooks, and all.
Async is not “fire and forget.” It is “fire, track, and explain when it smolders.”
Found this insightful? Like or share with your team:
Spread good engineering craft & architecture lessons.
Comments
Email is not published. Keep it professional.