Systems Design for Advanced Beginners: Anatomy of Infrastructure from Simple Monolith to Production Scale

Systems Design for Advanced Beginners: Anatomy of Infrastructure from Simple Monolith to Production Scale

Table of Contents

You have shipped a few small websites. You can stand up a CRUD app in Django, Laravel, or Gin in an afternoon. But then you look at the architecture behind a real platform - Stripe, GitHub, or a mid-sized marketplace - and wonder: what is actually behind the curtain, and how does it survive tens of millions of users?

The gap between “I can build a web app” and “I can design a system” is not a leap in coding skill. It is a leap in mental model: understanding how independent components compose, how they fail, and what every architectural choice trades away.

This post is based on Robert Heaton’s essay “Systems Design for Advanced Beginners,” expanded through the lens of a backend engineer who works daily with Go, PostgreSQL, and distributed systems. I will use one running example - a marketplace called Steveslist - to walk from the edge layer through webhooks, databases, search engines, pub/sub, and finally a data warehouse.

The problem statement: when an app outgrows pet-project scale

Assume Steveslist made it. Five years in, it has two consumer-facing products - a web app and mobile apps - plus a public API that lets third-party programmers build power tools (say, creating hundreds of listings programmatically).

graph TD
    subgraph "Client Layer"
        Web["Web Browser (SPA)"]
        Mobile["Smartphone App (iOS/Android)"]
        SDK["Client Libraries / Scripts"]
    end

    subgraph "Edge Layer"
        LB["Load Balancer / API Gateway"]
    end

    subgraph "Application Layer"
        API["Steveslist API Servers (Stateless)"]
    end

    Web -->|HTTPS| LB
    Mobile -->|HTTPS| LB
    SDK -->|HTTPS + API Key| LB
    LB --> API

Behind that application layer sits a set of subsystems most users never see:

  • Webhooks - proactively push notifications when events happen (e.g. “order placed”).
  • Password authentication - secure login.
  • SQL database - the primary data store; must scale and stay reliable.
  • Free-text search - powers the search box that accepts fuzzy queries like “used TV” or “motorbike”.
  • Internal tools - admin and support consoles.
  • Cron jobs - scheduled work (invoicing, data sync).
  • Pub/Sub - asynchronous reaction to trigger events.
  • Big data analytics - enormous aggregation queries over the full dataset.

One definition before we go further: what is a server? For our purposes, a server is a computer that runs on a network, listens for connections from other computers, performs some action, and usually returns data. A web server listens for HTTP requests; a database server listens for queries and reads/writes data. This skips a career’s worth of detail, but it carries us to the end.

The edge layer: from SPA to REST/JSON API

Single-Page Apps and the trap in the word “single”

The Steveslist web app is a Single-Page App (SPA). The “single” means the browser almost never fully reloads the page as you click around.

The browser makes its first HTTP request; the server returns a nearly empty skeleton HTML page plus a large JavaScript bundle. That JavaScript runs in the browser, updates the view in response to user actions, and when it needs data it fires an asynchronous AJAX request to a URL and updates the view from the response.

sequenceDiagram
    autonumber
    participant B as "User's Browser"
    participant S as "Steveslist Servers"

    B->>S: GET / (initial HTTP request)
    S-->>B: Skeleton HTML + JS file references
    B->>S: GET /static/app.bundle.js
    S-->>B: JavaScript bundle (~2MB)
    Note over B: JS boots, router mounts
    B->>S: GET /api/v1/listings (AJAX, JSON)
    S-->>B: JSON payload
    Note over B: UI updates in place, no full reload

SPAs are a lot of work to build and maintain - client-side routing, state management, caching, and now server-side rendering. But the UX is smooth, which is why almost every large consumer product adopts this model.

REST/JSON APIs and endpoint reuse

The mobile apps (iOS/Android) are architecturally identical to the web app: they send HTTP requests, the server processes and returns an HTTP response, and the client updates its UI. Because the mobile apps perform the same operations (create listing, send message), they hit the exact same URLs as the web app. You only write the native frontend layer; the backend is shared.

Third-party programmers interact through API endpoints. To fetch a user’s listings:

GET https://api.steveslist.com/v1/listings

The server responds with JSON - a structured format any language can parse:

{
  "listings": [
    {
      "id": 2178123867,
      "name": "Used TV",
      "country": "US",
      "city": "San Francisco",
      "price_amount": 1000,
      "price_currency": "usd"
    }
  ]
}

API authentication: the API key as a programmer’s password

Users authenticate to the API using an API key - a long, random string that is essentially a password for code. It is shown on the Settings page and attached as an HTTP header on every request. The server checks whether the key maps to a user and, if so, acts on that user’s behalf.

  • Go
  • Python
package main

import (
	"bytes"
	"encoding/json"
	"fmt"
	"net/http"
)

type Listing struct {
	Name          string `json:"name"`
	Country       string `json:"country"`
	City          string `json:"city"`
	PriceAmount   int    `json:"price_amount"`
	PriceCurrency string `json:"price_currency"`
}

func createListing(apiKey string) error {
	body, _ := json.Marshal(Listing{
		Name:          "Used TV",
		Country:       "US",
		City:          "San Francisco",
		PriceAmount:   1000,
		PriceCurrency: "usd",
	})

	req, _ := http.NewRequest("POST",
		"https://api.steveslist.com/v1/listings",
		bytes.NewReader(body))
	// API key travels in a custom header, never in the URL query string.
	req.Header.Set("X-Steveslist-API-Key", apiKey)
	req.Header.Set("Content-Type", "application/json")

	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		return err
	}
	defer resp.Body.Close()

	fmt.Println("status:", resp.StatusCode)
	return nil
}
import requests

url = "https://api.steveslist.com/v1/listings"
payload = {
    "name": "Used TV",
    "country": "US",
    "city": "San Francisco",
    "price_amount": 1000,
    "price_currency": "usd",
}
api_key = "YOUR_API_KEY_GOES_HERE"

resp = requests.post(
    url,
    json=payload,
    # Custom header keeps the key out of logs, browser history and proxies.
    headers={"X-Steveslist-API-Key": api_key},
    timeout=10,
)
print(resp.status_code, resp.json())

Finally, client libraries wrap the API so programmers never touch HTTP details. They write steveslist.Listing.create(...); the library builds the correct HTTP request. This is by far the most common way people use an API, and the reason you should ship SDKs for every language you can think of.

Webhooks: pushing events instead of letting clients poll

Push vs pull

Suppose a seller wants full automation: on each sale, send a thank-you email and instruct the warehouse to ship. The naive approach is for the seller to poll the API: “Any new sales? Any new sales?” This is brutally inefficient - it generates a flood of empty requests and hammers the server.

The industry-standard answer is webhooks: an HTTP request we send to the user’s server whenever an event happens. The user registers a URL and deploys a web server to receive those notifications.

sequenceDiagram
    autonumber
    actor Buyer as "Buyer"
    participant SL as "Steveslist Server"
    participant WS as "Seller's Webhook Server"

    Buyer->>SL: POST /checkout (buy item)
    Note over SL: Transaction complete,
look up seller webhook URL SL->>WS: POST /hooks (JSON + HMAC signature) Note over WS: Verify signature,
process the order WS-->>SL: 200 OK

Tip

Webhooks do not only fire on purchases. Steveslist also sends webhooks when a user receives a message, when a listing is removed by an admin, or when a buyer complains. That lets sellers automate not just listing, but selling and shipping too.

Security: signing with HMAC-SHA256

A webhook endpoint is public on the internet. Anyone who knows the URL can fire fake webhooks - and if the seller is careless, an attacker can trick them into shipping free merchandise. A hard-to-guess URL is not security; obscurity is not security.

To let sellers verify that a webhook really came from us, we cryptographically sign the payload using HMAC. When a seller enables webhooks, we generate a random shared secret key and give it to them. On every delivery, we combine the secret with the payload and run it through HMAC-SHA256 to produce a deterministic signature.

graph LR
    Payload["Webhook payload
{\"id\": 123, \"action\": \"item_sold\"}"] --> HMAC["HMAC-SHA256"] Secret["Shared secret key
(known only to us and the seller)"] --> HMAC HMAC --> Sig["Signature
234gj98d49j8..."]

The receiver recomputes the signature the same way and compares. Match means accept; mismatch means reject. Since only we and the seller know the secret, a valid signature proves the webhook came from us.

Warning

All signature-verification code is written and maintained by the seller. We can provide examples and documentation, but we cannot force them to verify correctly - or at all. This is why vendors like Stripe and GitHub publish very detailed signature-verification docs.

  • Go
  • Python
package webhook

import (
	"crypto/hmac"
	"crypto/sha256"
	"encoding/hex"
	"net/http"
)

var sharedSecret = []byte("123mhu23jy8xdwgmd...")

// VerifySignature recomputes the HMAC over the raw body and compares
// it against the signature header using a constant-time comparison.
func VerifySignature(r *http.Request, body []byte) bool {
	got, err := hex.DecodeString(r.Header.Get("X-Steveslist-Signature"))
	if err != nil {
		return false
	}

	mac := hmac.New(sha256.New, sharedSecret)
	mac.Write(body)
	want := mac.Sum(nil)

	// Constant-time compare prevents timing side-channels.
	return hmac.Equal(got, want)
}
import hmac
import hashlib

SHARED_SECRET = b"123mhu23jy8xdwgmd..."

def verify_signature(raw_body: bytes, signature_hex: str) -> bool:
    expected = hmac.new(
        SHARED_SECRET,
        raw_body,
        hashlib.sha256,
    ).hexdigest()
    # compare_digest avoids timing side-channels.
    return hmac.compare_digest(expected, signature_hex)

Reliability: at-least-once, retries, and the DLQ

The harder question: what happens when a webhook fails? Cannot connect to the server? Server returns an error? Server hangs for twenty seconds and then disconnects without saying anything?

Steveslist chooses to guarantee at-least-once delivery: if a send fails, we retry (many but not infinitely many times) until we are sure an attempt succeeded. This can occasionally deliver the same webhook twice - and it is the seller’s responsibility to write idempotent handlers instead of shipping five TVs for one order.

Proper retries use exponential backoff with jitter, and after the attempt budget is exhausted, deliveries go to a Dead Letter Queue (DLQ) for manual investigation.

graph TD
    Event["Event occurs"] --> Send["Send webhook HTTP"]
    Send --> Check{Connection OK
and status 2xx?} Check -->|Yes| Done["Mark delivered"] Check -->|No| Count{Retries left?} Count -->|Yes| Backoff["Wait exponential backoff
(1s, 2s, 4s, 8s...+jitter)"] Backoff --> Send Count -->|No| DLQ["Push to Dead Letter Queue"] DLQ --> Alert["Alert & manual handling"]
// Exponential backoff with jitter for webhook delivery attempts.
func backoffDelay(attempt int) time.Duration {
	base := time.Second << uint(attempt) // 1s, 2s, 4s, 8s...
	if base > 5*time.Minute {
		base = 5 * time.Minute // cap the ceiling
	}
	// Full jitter spreads retries so we don't hammer a recovering server.
	jitter := time.Duration(rand.Int63n(int64(base / 2)))
	return base/2 + jitter
}

State and storage: the single source of truth

The RDBMS as source of truth

Steveslist’s primary store is an RDBMS (PostgreSQL/MySQL). Every new write lands here first, before going anywhere else. We call it the source of truth. Queries that demand accuracy and freshness - “list this user’s open listings,” “verify this user’s password” - read from here.

What makes an RDBMS strong is ACID transactions: a sequence of operations either fully succeeds or fully rolls back. A money transfer must debit account A and credit account B inside one transaction - there is no half-state.

Databases run on computers like any other program. As data grows, disk and memory fill up, and queries slow down because the engine must scan more and more records. Running on a single machine with a giant disk only postpones the problem. You will have to think about horizontal scaling.

Connection pooling: PgBouncer

Before sharding, there is a widely misunderstood bottleneck: connection limits. PostgreSQL runs each connection as a separate process (fork-on-connect), costing ~5-10MB RAM and expensive context switches. With 200 app servers each opening a pool of 20 connections, you have slammed 4000 connections into a database that will die.

The answer is a connection pooler like PgBouncer sitting between the app and the database. It keeps a small, tight pool to the database and multiplexes thousands of clients through it.

Tip

In transaction pooling mode, PgBouncer assigns a server connection for the duration of a single transaction, then reclaims it. The catch: you cannot use session-state prepared statements, LISTEN/NOTIFY, or advisory locks spanning transactions. Audit your app before enabling it.

Read replicas and replica lag

Reads dominate most workloads. We replicate data across machines to both tolerate failure and spread read load.

graph TD
    App["Application Servers"] -->|Writes| Primary[("`PostgreSQL
    Primary (RW)`")]
    App -->|Reads| R1[("`Replica 1
    (RO)`")]
    App -->|Reads| R2[("`Replica 2
    (RO)`")]
    Primary -->|Streaming replication| R1
    Primary -->|Streaming replication| R2

Replica lag is the sharp edge. Reading from a replica does not guarantee seeing data just written to the primary - data can lag from milliseconds to seconds (longer under primary load). The classic symptom is “I just changed my avatar but it still shows the old one” - a read-your-own-writes violation.

Strategies to handle it:

  • Read-your-writes routing: after a user writes, pin them to the primary (or read from the primary for a short window).
  • Sticky sessions / session pinning: keep a client on the same replica.
  • Synchronous replication: the primary only returns after replicas acknowledge - safer but slower, and risky if a replica falls behind (write stalls).

Sharding and the write bottleneck

When write volume exceeds one machine - or data no longer fits on one machine - replication does not help (every replica receives the same write workload). This is where you shard: split data into chunks, each on a different machine. For Steveslist we shard by user, so all of one user’s data lives on the same machine.

This requires a routing layer: either app servers know the shard mapping, or a central database router does. App-side routing has fewer hops (faster); a central router is easier to update in one place.

Warning

Sharding is one of the most expensive operational decisions you can make. You pay for it with painful cross-shard queries, effectively impossible distributed transactions, complex rebalancing, and a large operational burden. Many systems never need sharding. Exhaust vertical scaling, indexing, caching, and native PostgreSQL partitioning first.

Migrating data between shards (when one fills up) usually follows a safe double-write procedure: write to both old and new shards, backfill, verify, read from the new shard, then stop double-writing and delete the old one. Detailed but entirely logical and doable.

Full-text search: why LIKE kills your database

A SQL database answers precise, well-defined queries beautifully: list user #145122’s listings from the last 90 days, or count new listings per day in San Francisco. These use sharp operators like =, >, and GROUP BY.

But a Google-style search - “used TV” - is a disaster. The naive attempt is LIKE:

SELECT * FROM items
WHERE description LIKE '%used TV%';

This matches only that exact substring. It misses “TV that is used,” “second-hand TV,” and any typo. Worse:

Warning

LIKE '%keyword%' with a leading % cannot use a B-tree index. The engine is forced into a full table scan - reading every row, every text column. On a table with tens of millions of rows, one such query can scan gigabytes, thrash the buffer pool, and peg the database CPU at 100% until the connection queue overflows. This is a textbook way for an analyst to kill production.

You could try to hand-write a giant OR query covering every permutation of words - but it stays slow, fragile, and still misses edge cases. And even when it returns results, you have no idea how to rank by relevance. A search engine needs “best first,” not “alphabetical.”

The inverted index: the core of a search engine

What SQL lacks is an inverted index: a structure mapping each term (tokenized, stemmed, lowercased) to the list of documents containing it.

graph LR
    subgraph "Inverted Index"
        T1["tv"] --> D1["doc 1, doc 7"]
        T2["used"] --> D2["doc 1, doc 4"]
        T3["car"] --> D3["doc 4, doc 9"]
    end
    Query["Query: 'used tv'"] --> T1
    Query --> T2
    T1 --> Merge["Intersect posting lists
+ compute relevance score"] T2 --> Merge

With this structure, the query “used tv” only needs to look up two posting lists and intersect them - roughly constant complexity in corpus size, not linear like a full scan. This is the foundation of Elasticsearch, Meilisearch, and Lucene.

Critically: Elasticsearch does not replace PostgreSQL. It is less reliable (more prone to losing data at scale), slower to write, and offers no transactional semantics. That is exactly why we use both.

Keeping them in sync: dual-write anti-pattern vs transactional outbox

Here is the hardest part: how do you sync data from PostgreSQL into Elasticsearch?

The first instinct is dual-write: in the service, write to Postgres, then immediately write to Elasticsearch.

// ANTI-PATTERN: dual-write without atomicity.
func CreateListing(ctx context.Context, l Listing) error {
	if err := pg.Insert(ctx, l); err != nil {
		return err
	}
	// If this step fails or the process crashes between the two writes,
	// Postgres and ES drift apart forever. There is no cross-system rollback.
	return es.Index(ctx, l)
}

Dual-write is a classic anti-pattern: the two systems share no transaction, so there is no atomicity. If the process crashes between the steps, or ES is briefly down, the data drifts permanently with no reliable way to know. You can reverse the order and write ES first, but then you drift the other way.

The correct answer is the Transactional Outbox. Write the domain data and an event row in the same PostgreSQL transaction. A separate process then reads the outbox table and pushes events to ES. Because the outbox write shares the transaction with the domain data, the two are always consistent.

graph TD
    Client["Client request"] --> Tx["BEGIN TX"]
    Tx --> InsertOrder["INSERT INTO orders"]
    Tx --> InsertOutbox["INSERT INTO outbox_events"]
    InsertOrder --> Commit["COMMIT (atomic)"]
    InsertOutbox --> Commit
    Commit --> Relay["Outbox Relay Worker
(poll or CDC)"] Relay --> ES["Elasticsearch / Kafka"] Relay --> MarkSent["Mark event processed"]
BEGIN;
INSERT INTO orders (id, user_id, name) VALUES ($1, $2, $3);
-- Same transaction guarantees consistency between the two.
INSERT INTO outbox_events (aggregate_id, event_type, payload)
VALUES ($1, 'listing.created', $2);
COMMIT;

Change Data Capture (CDC) - using Debezium to read the PostgreSQL WAL - is a subtler variant: instead of the app writing an outbox table, a connector reads the write-ahead log at a low level and emits events. It barely affects app performance, but it is more operationally complex and couples you to the WAL format.

Tip

For small datasets with low freshness requirements, periodic reindexing (every 15 minutes) is a perfectly reasonable compromise. Craigslist famously told users “your listing will be visible in search within 15 minutes” - the telltale signature of a batch sync pipeline from SQL to a search engine.

Asynchronous processing: pub/sub and event-driven design

Why not just do it synchronously?

Actions have consequences - a new signup triggers a welcome email; a new listing notifies everyone with a matching search alert; a declined card triggers a reminder. Technically, all of these could run synchronously by the server handling the triggering action.

But that is usually a bad idea. A reaction (finding and notifying thousands of interested users) can be slow. Running it synchronously means the user waits for every side effect to finish before getting a response. Latency balloons, UX suffers, and a failing side effect can take down the main request with it.

Pub/sub at the architecture level

We decouple the trigger from the reaction with pub/sub. When an action occurs, the code publishes an event (NewListingCreated, SubscriptionCardDeclined). Anyone who wants to react writes a consumer subscribed to that event type.

graph LR
    Server["Steveslist Server"] -->|publish NewUserSignup| Broker["Message Broker
(Kafka / RabbitMQ / Redis Streams)"] Broker -->|push/pull| C1["SendWelcomeEmailConsumer"] Broker -->|push/pull| C2["NewUserSpamCheckConsumer"] Broker -->|push/pull| C3["AnalyticsIngestConsumer"]

The benefits:

  • Non-critical work runs asynchronously, keeping the user experience snappy.
  • Clean separation: the publisher does not care who subscribes or what they do.
  • If a consumer fails (say the email system hiccups), the broker records it and retries later.

Two common delivery mechanisms: push (the broker pushes events to consumers) and pull (consumers poll the broker). Redis Streams, Kafka, and RabbitMQ differ in semantics, throughput, ordering guarantees, and durability - choose by requirement, not by trend.

Idempotency under at-least-once processing

This is where 90% of systems ship bugs. Most message queues guarantee at-least-once: if a consumer crashes after processing but before acking, the broker redelivers the message. That means the consumer must be idempotent - handling the same message twice must not cause duplicate side effects.

The standard fix is to use the event ID as an idempotency key and record that it was processed in the same transaction as the side effect.

  • Go
  • Python
// Idempotent consumer: dedupe by event ID inside the same transaction
// as the side effect, so a redelivered message is a no-op.
func (h *Handler) Handle(ctx context.Context, evt Event) error {
	tx, _ := h.db.Begin(ctx)
	defer tx.Rollback(ctx)

	// INSERT ... ON CONFLICT DO NOTHING returns 0 rows if already processed.
	tag, err := tx.Exec(ctx,
		`INSERT INTO processed_events (event_id) VALUES ($1)
		 ON CONFLICT (event_id) DO NOTHING`, evt.ID)
	if err != nil {
		return err
	}
	if tag.RowsAffected() == 0 {
		return nil // already handled, ack and move on
	}

	if err := applySideEffect(ctx, tx, evt); err != nil {
		return err
	}
	return tx.Commit(ctx)
}
def handle(event, db):
    """Idempotent consumer: dedupe by event ID in the same transaction
    as the side effect, so a redelivered message is a no-op."""
    with db.transaction() as tx:
        inserted = tx.execute(
            """
            INSERT INTO processed_events (event_id)
            VALUES (%s)
            ON CONFLICT (event_id) DO NOTHING
            """,
            (event.id,),
        ).rowcount

        if inserted == 0:
            return  # already handled, ack and move on

        apply_side_effect(tx, event)
    # Transaction commits here; event id + effect are atomic.

Warning

An idempotency key is not only local dedupe. For outbound webhooks (calls to third-party APIs), send an Idempotency-Key header so the receiver can dedupe too. This is how Stripe avoids double-charging when a client retries.

Separating OLTP and OLAP: never run reports on the production database

To understand and optimize the business, we compute complex statistics over the whole dataset: “how many listings per day, segmented by country and city?” or “how many users who signed up in a given month created a listing within 90 days?”.

These are aggregations over the entire dataset - and you absolutely must not run them on the production database.

Warning

An analytics query scanning hundreds of millions of rows can crash production in several ways: (1) it holds locks and buffer pages long enough to stall normal OLTP queries; (2) it causes buffer pool thrashing, evicting OLTP’s hot pages from RAM; (3) it saturates I/O and CPU. The result: the entire user-facing app slows or falls over because one analyst ran a report.

The deeper problem: an engine great at small queries (returning one user’s listings) is often terrible at enormous ones. The two workloads have different physical shapes:

  • OLTP (row-oriented): touches a few rows, many columns, high selectivity, continuous writes. Optimized for point lookups.
  • OLAP (column-oriented): scans a few columns across an entire table, big aggregations, batch writes. Optimized for column scans.

The answer is a data warehouse built on columnar storage: ClickHouse, Snowflake, BigQuery, Redshift. We replicate data from the production SQL database into the warehouse on a schedule (say nightly) via an ETL/ELT pipeline (Extract-Transform-Load, or Extract-Load-Transform).

graph LR
    subgraph "OLTP (Row-oriented)"
        App["App Servers"] -->|"Point reads/writes
high selectivity"| PG[("`PostgreSQL Production`")] end subgraph "Pipeline" PG -->|"ETL/ELT
(nightly batch or CDC)"| DW end subgraph "OLAP (Column-oriented)" DW[("`ClickHouse / BigQuery Columnar Storage`")] -->|"Full-table scans
aggregations"| BI["BI / Analysts"] end

The result: analysts query giant datasets on the warehouse while production stays smooth. This is the definition of “the right tool for the job” - and proof that no single database is “best” at everything.

Summary: architectural lessons and the trade-offs

Kate (the character from the original essay) makes an important point: no design is absolutely “right.” The real reason for many technology choices is “we chose X because Sara knows a lot about X” and “we chose Y on the spur of the moment when it didn’t seem like a big decision, and we never found time to re-evaluate.” Admitting this keeps you humble in front of any architecture you meet.

The CAP theorem in practice

The CAP theorem says a distributed system can guarantee only two of three: Consistency, Availability, Partition tolerance. But in real systems, partitions (network splits) are an unavoidable property of any distributed system - cut cables, failed switches, disconnected data centers. So the real choice is not “two of three” but CP or AP when a partition happens: do you refuse to serve (keep consistency) or serve possibly-stale data (keep availability)?

The practical consequence: the whole system does not need one answer. A webhook delivery should be AP (better to double-send than to lose). A payment transaction should be CP (better to reject than to charge wrong). This is why you need explicit awareness of what each subsystem requires.

Conway’s Law

“Organizations design systems that mirror their own communication structure.” (Melvin Conway, 1967)

This is not just philosophy. It is a prediction tool: if you have 4 teams, you will get 4 service boundaries - whether you want them or not. A monolith edited by 20 teams naturally fractures into overlapping modules. Conversely, if you want to split out a service, the fastest way is to stand up a team that owns it end-to-end.

The consequence: architecture and team structure are two sides of one coin. You cannot design microservices with a single 6-person team operating 40 services - you will die from cognitive load and on-call rotation.

The final trade-off: operational complexity vs real scaling needs

Every architectural component you add is paid for with operational complexity: one more thing that can break at 3 AM, one more dashboard to build, one more runbook to write, one more failure class to learn. In exchange, you buy scalability and fault isolation.

The right question is not “which architecture is most modern?” but “what real problem am I hitting, and does this component solve it?”

  • If you do not have millions of rows yet, sharding is over-engineering.
  • If you have no full-text search need yet, you do not need Elasticsearch.
  • If you have no heavy analytics workload yet, you do not need a warehouse.
  • If your team is small, a modular monolith beats microservices.

Great systems evolve from simple systems that worked. A complex system designed from scratch never works properly and cannot be patched into working (Gall’s Law). Start with a simple monolith, measure honestly, and split only when you have concrete evidence of a bottleneck. Sound boring? Correct - good systems look boring on paper, but they run reliably in production.

Share :

Related Posts

Lore: Epic Games Open-Sources a VCS Built for Massive Binaries

Lore: Epic Games Open-Sources a VCS Built for Massive Binaries

Game repositories are not web repositories. In an engine the size of Unreal, most of the weight is binary assets, textures, meshes, world partitions, hundreds of MB to several GB each, churned by thousands of artists and developers. Git handles that badly: LFS is a bolt-on dragging whole 2 GB files, dedup is capped by pack deltas, partial clone plus sparse checkout is experimental, and offline means fail.

Read More
DevOps Is Bullshit: The Illusion of 'You build it, you run it' and the Rise of Platform Engineering

DevOps Is Bullshit: The Illusion of 'You build it, you run it' and the Rise of Platform Engineering

A backend engineer needs an IAM role so their service can read the S3 bucket holding customer invoices. They open Jira, pick the “Infrastructure Request” template, fill in twelve fields, attach the exact labels the README tells them to attach, and hit Create. Three days later, the role exists.

Read More