LivePositively

When Event-Driven Architecture Is the Wrong Choice: A Decision Framework

ad

adam


7 minutes

When Event-Driven Architecture Is the Wrong Choice: A Decision Framework

The Proposal That Prompted This Article

A few years ago I reviewed an internal RFC titled, roughly, "Decoupling Our Services with an Event Backbone." It proposed routing nearly all inter-service communication — including checkout, including authentication callbacks — through Kafka. The stated benefits were the familiar litany: loose coupling, independent scaling, resilience, replayability. The RFC was well-written, the author was talented, and the architecture would have been a disaster.

Not because event-driven architecture is bad. Because the RFC, like most adoption arguments I've seen, evaluated events against no alternative — it listed what events provide without pricing what they take away, and without asking whether each specific interaction actually had the shape events reward. The word "decoupling" did an enormous amount of unexamined work.

Nearly a decade of operating both styles has left me with a specific, checkable set of signals for when asynchronous eventing is the wrong tool for a given interaction — plus the cases where it is unambiguously right. This article is that checklist, with the reasoning attached. The goal isn't to talk you out of Kafka; it's to make "we should use events here" a claim that has to survive contact with four questions.

What Events Actually Buy You

To evaluate honestly, state the genuine benefits precisely — vagueness is where bad adoptions hide:

1. Temporal decoupling. Producer and consumer don't need to be up at the same time. The consumer can be down for an hour and catch up. This is real and valuable when the interaction tolerates that hour.

2. Fan-out without producer knowledge. New consumers subscribe without the producer changing or even knowing. Genuinely powerful for cross-cutting concerns: audit, analytics, search indexing, cache invalidation.

3. Load leveling. A queue absorbs bursts and lets consumers drain at their own rate — the classic and correct use of buffering, when latency during the drain is acceptable.

4. Replayability. With a log (not just a queue), you can re-derive downstream state, backfill a new consumer, or recover from a consumer bug by reprocessing. This is the benefit teams underrate before they have it and can't live without after.

Every one of those has an italicized condition. The framework below is essentially those conditions, inverted into warning signals.

The Four Costs Nobody Puts on the Slide

Events don't remove coupling; they move it — from compile-time interfaces and synchronous availability into schemas, ordering, timing, and operational surface. The four payments:

Debuggability collapses from a stack trace to an investigation. A synchronous call chain fails with a stack trace and a status code, in one place, at one time. An eventful workflow fails as an absence: the order sits in `pending` forever because a consumer three hops away dropped a message last Tuesday. Tracing tooling helps; it does not restore the property that cause and effect share a timestamp and a log line.

Ordering and duplication become everyone's problem. At-least-once delivery plus consumer retries means every consumer must be idempotent, forever, with no exceptions — and partial ordering means "the update arrived before the create" is a case your code handles, not a bug you file. These aren't advanced topics; they're the entry fee, paid by every team that touches the bus, including the intern's first consumer.

Schema evolution becomes a distributed governance problem. A synchronous API change breaks loudly at the boundary, in CI, against a contract. An event schema change breaks quietly, later, in consumers the producer has never heard of — that fan-out-without-producer-knowledge benefit, reread as a liability. Schema registries and compatibility rules tame this, but "tame" means a standing governance function, not a solved problem.

Workflow state goes off the books. In a synchronous orchestration, "where is order 12345 stuck?" has an answer: in the orchestrator, in one state machine. Spread the same workflow across five consumers reacting to each other's events and the workflow state exists only as an inference over five databases and a topic's consumer offsets. Choreography enthusiasts call this emergent; on-call engineers at 3 a.m. call it something else.

None of these is fatal. All of them are permanent operating costs, and the framework's core question is whether a given interaction's shape earns them back.

Signal 1: The Caller Needs the Answer

The clearest signal, and the one most often bulldozed by enthusiasm: if the caller cannot proceed without the result, the interaction is synchronous no matter what infrastructure you route it through.

Put the checkout's payment authorization on a bus and the user is still standing there, waiting. You haven't removed the synchronous dependency — you've disguised it, and now the user's wait includes broker hops, consumer poll intervals, and a response-correlation mechanism (a reply topic? polling? websockets?) that you built to simulate the request/response semantics you actually needed. You've reimplemented RPC, badly, with more moving parts and worse tail latency.

The test is behavioral, not architectural: What does the caller do while waiting? If the answer is "block, spin, or poll," the interaction is a request. Give it a request's infrastructure — a synchronous call with timeouts, retries, and a circuit breaker — and let it fail fast and loudly when it fails. "But the payment service might be down" is not an argument for a queue; a queued payment authorization during an outage isn't resilience, it's a user staring at a spinner while their intent ages in a buffer.

Signal 2: The Workflow Has a Deadline and an Owner

Some multi-step processes are workflows in the strong sense: they must complete within a bounded time, someone is accountable when they don't, and partial completion is a business problem requiring compensation. Order fulfillment. KYC verification. Money movement.

Pure event choreography — each service reacting to the previous service's event, no central coordinator — fits these badly, for a reason that only surfaces operationally: choreography has no natural place to stand to answer "where is it stuck and who fixes it?" The workflow's existence is distributed across consumers; its deadline is enforced by nobody; its failure modes are compensated by whoever notices.

The boring, correct answer for deadline-and-owner workflows is explicit orchestration — a saga orchestrator or a durable-execution engine (Temporal and friends) that owns the state machine, enforces timeouts, and runs compensations. The orchestrator can still invoke steps via queues where individual steps benefit from buffering; the point is that workflow state lives somewhere with a name, an owner, and a dashboard. Choreography is for relationships between peers who genuinely don't need to know about each other; it is not a management structure for a process with a deadline.

mermaid

flowchart TD

Q1{Does the caller block on the result?} -->|yes| SYNC[Synchronous call: timeout, retry, breaker]

Q1 -->|no| Q2{Bounded completion time + accountable owner?}

Q2 -->|yes| ORCH[Orchestrated workflow - steps may use queues]

Q2 -->|no| Q3{Same team, deployed together?}

Q3 -->|yes| CALL[Direct call or in-process - revisit at org split]

Q3 -->|no| Q4{Multiple consumers, tolerant of staleness?}

Q4 -->|yes| EDA[Events: this is the sweet spot]

Q4 -->|no| Q1a[Reexamine - the interaction is probably a request in disguise]

Signal 3: Two Services, One Team, One Deploy Train

Event advocates promise organizational decoupling: teams evolve independently, integrate through the bus. The promise is real when there are independent teams to decouple. When one team owns both the producer and the only consumer, deploys them from one pipeline, and coordinates changes in one standup, the bus is providing organizational insulation between an organization and itself.

Meanwhile the costs from earlier are all still being paid — idempotency in every consumer, schema ceremony, the debugging tax — to solve a coordination problem that a function call, or at most a direct HTTP call, solves for free. I think of this as Conway arbitrage in reverse: architecture buying independence the org chart doesn't need yet, at prices the org chart pays immediately.

The honest version of this decision: default to the direct call, and write down the trigger that flips it — "when a second team wants this data" or "when these services stop deploying together." Decoupling is cheap to add at the moment it's needed (introduce the topic, dual-publish, migrate the consumer) and expensive to carry for years before it's needed. Architectural options, like financial ones, have carrying costs; buy them when the underlying risk is real.

Signal 4: You Need to Read Your Writes

Eventual consistency is the fine print of every event-driven diagram, and the signal here is specific: if a user's next action depends on seeing the effect of their previous one, an event-shaped pipeline between the two is a bug generator.

The canonical incident: user updates their shipping address, immediately hits "place order," and the order service — fed by an `address-updated` event it hasn't consumed yet — ships to the old address. No component malfunctioned. The architecture worked as designed; the design was wrong for the interaction. Read-your-writes violations are especially poisonous because they're intermittent (consumer lag dependent), user-visible (they look like data loss), and unfixable at the consumer (the consumer did nothing wrong).

Patches exist — session pinning to the source of truth, version tokens the client carries forward, UI-side optimism — and each is a piece of complexity spent to reintroduce the consistency the synchronous design had natively. Sometimes that spend is right (the fan-out benefits elsewhere justify it). But it must be priced in during the decision, not discovered via a support ticket. Any workflow where a human takes consecutive dependent actions deserves a synchronous read path to the source of truth, whatever the rest of the architecture does.


Read This Next