Theme: The Shift

Length: ~800 words

Hook: Contrarian — the microservices playbook doesn't map to agents


When microservices became the default architecture, we developed a playbook. Circuit breakers. Retries with exponential backoff. Saga patterns for distributed transactions. Graceful degradation. Health checks. Service meshes.

This playbook is excellent — for deterministic services that do the same thing every time you call them.

Now teams are applying this playbook to AI agents. It's not working. Here's why.

Microservices are deterministic. Agents are not.

When you call a microservice /api/orders, you get the same response for the same input. Always. If it fails, you retry. If it's down, you circuit-break. The failure modes are known and bounded.

When you call an agent, you might get a different response every time. Not because it's broken — because it's reasoning. The same query with different context produces different actions. Retrying might give you a completely different result, not a corrected one. Circuit breaking doesn't help because the agent isn't down — it's just wrong.

Microservices don't make decisions. Agents do.

A microservice receives a request and processes it. It doesn't decide whether to process it. It doesn't choose which tools to use. It doesn't reason about trade-offs.

An agent receives a goal and decides what to do. Which tools to call. Whether to ask for more information. Whether to escalate. Whether to act at all. This means the failure surface isn't just "the service crashed" — it's "the agent made the wrong decision."

You can't retry a wrong decision. You can't circuit-break a bad judgment call. You need a different safety model.

Microservices have bounded blast radius. Agents don't.

A microservice typically operates within its domain. The orders service creates orders. The payment service processes payments. If the orders service fails, orders aren't created. The blast radius is contained.

An agent with access to multiple tools has an unbounded blast radius. It can call any tool it has access to, in any order, for any reason. An agent that can read data, write data, send emails, and call external APIs can do all of those things in a single decision chain. A failure isn't "one service is down" — it's "the agent took a series of actions across multiple systems, some of which were wrong."

Microservices are stateless. Agents are stateful.

The best practice for microservices is statelessness. Each request is independent. You can scale horizontally by adding instances. If one crashes, another handles the next request.

Agents are inherently stateful. They have memory. They have context. They have conversation history. They have goals they're pursuing across multiple steps. You can't just "add another instance" — the new instance doesn't have the context.

This changes how you think about scaling, recovery, and failover. With microservices, you restart the pod. With agents, you need to reconstruct the decision context.

What we need instead.

The microservices playbook gave us operational patterns. For agents, we need a different set:

  1. Policy-as-code, not just retries. Instead of retrying failed calls, prevent wrong calls. The agent's tools are gated by policies that are checked before execution, not after failure.
    1. Audit trails, not just logs. Microservice logs tell you what happened. Agent audit trails need to tell you why it happened — what the agent perceived, what it reasoned, what alternatives it considered.
      1. Evaluation gates, not just health checks. A microservice is healthy if it responds. An agent is healthy if it produces correct, safe, appropriate outputs. That requires evaluation, not just uptime checks.
        1. Human-in-the-loop, not just graceful degradation. When a microservice fails, you degrade gracefully. When an agent is uncertain, you escalate to a human. The fallback isn't a cached response — it's a human decision.
          1. Cost ceilings, not just rate limits. Microservices don't spend money per request. Agents do — every LLM call costs. You need cost ceilings per agent, per session, per action — not just rate limits per endpoint.
          2. The bigger point.

            We're not saying microservices patterns are useless for agents. Observability, deployment automation, infrastructure-as-code — these still apply. But the safety patterns, the failure handling, the operational model — these need rethinking.

            Agents are a new category. They need a new playbook.

            What patterns from microservices have you tried applying to agents — and where did they break?