System Design in Depth

October 3, 2026

System Design in Depth is a free and open-source learning platform for developers who want to understand the engineering mechanics behind scalable and distributed systems.

Why I Built This

Most system design resources explain what components to use without explaining how they work internally. A diagram with a load balancer, cache, and message queue doesn't explain:

  • How consistent hashing distributes data
  • Why databases use B-Trees or LSM-Trees
  • How a write-ahead log provides durability
  • What happens during a Raft leader election
  • How distributed systems handle partial failures
  • What trade-offs emerge as a system grows

System Design in Depth is my attempt to connect high-level architecture decisions with the lower-level mechanisms that make them work.

Screenshot_2026-10-03_at_9.26.27_PM.png

What's Inside

The curriculum covers 200 learning units organized into 6 tracks and 18 modules. It also includes 32 interactive simulators, 15 from-scratch builds, 470 hand-picked videos, and 600 self-check questions.

Every lesson follows the same pattern: intuition first, then videos, diagrams and notes, audio narration, questions to check yourself, and code you can run. Content is backed by authoritative source citations.

The 6 Tracks

  1. Architectural Foundations and Workload Modeling: system invariants, requirements clarification, and how to reason about workloads.
  2. Data Architecture, Storage Engines and State Persistence: relational modeling and query execution, NoSQL architectures, partitioning and IDs, storage engine internals, and blob systems.
  3. Distributed Systems, Consensus and Coordination: network protocols, service communication, distributed coordination, consensus, and leases.
  4. Asynchronous Execution, Queues and Real-Time Processing: caching strategies, asynchronous workflows and partitioned logs, probabilistic sketches and stream analytics, and real-time messaging and feeds.
  5. Production Operations, Resilience and System Hardening: reliability and SLOs, media processing and CDNs, geospatial indexing, and information retrieval and search engines.
  6. Systems Design Labs and Architectural Evolution: full design problems such as a URL shortener, distributed rate limiter, collaborative editor, payment processing, ride matching, and a video-transcoding pipeline. The track ends with real-world architecture evolution and incident case studies:
    • Instagram's early architecture
    • Stripe's idempotency keys
    • Discord's message-storage migration
    • The Amazon Dynamo paper
    • GitLab's database incident

Beyond Static Content

Screenshot_2026-10-03_at_9.28.38_PM.png

Architecture Diagrams

Mermaid diagrams covering CQRS pipelines, Redis Cluster topology, LSM-Tree write paths, Raft election sequences, replication, messaging workflows, and failure recovery, placed alongside the relevant explanations. You can pan and zoom into any of them.

Interactive Simulators

32 simulators, including quorums, load balancing, Raft, cache stampede, MVCC, consistent hashing, token-bucket rate limiting, and cache-eviction policies, so you can change inputs and watch how the algorithms behave.

Videos, Audio, and Self-Checks

Each lesson includes hand-picked video explainers (470 across the curriculum), audio narration, and questions (600 in total) so you can test whether the concept stuck. A built-in glossary and keyboard-first search (⌘K) across all 200 topics make it easy to jump around.

From-Scratch Implementations

15 build exercises:

  • Layer 7 load balancer
  • Consistent hashing ring with virtual nodes
  • Sliding-window rate limiter
  • LRU cache
  • TTL cache with a Redis-style adaptive reaper
  • Bloom filter membership testing
  • Snowflake 64-bit ID generator
  • Keyset (cursor) pagination vs SQL OFFSET
  • Write-ahead log and crash recovery
  • LSM-Tree key/value store
  • Lease-based leader election with fencing tokens
  • Gossip protocol (push-pull anti-entropy and SWIM failure detection)
  • Merkle-tree anti-entropy repair
  • HyperLogLog cardinality sketch
  • Tiny search engine

Who It's For

  • Backend developers learning distributed systems
  • Engineers preparing for system design interviews
  • Developers moving toward senior roles
  • Senior engineers strengthening fundamentals
  • Students exploring software architecture
  • Contributors interested in educational open-source tooling

Key Lessons from Building It

Good diagrams aren't enough. Real complexity lives in failure handling, consistency, concurrency, and operational trade-offs - not in clean architecture boxes.

Every design decision has a cost. Caching adds invalidation problems. Replication creates consistency challenges. Partitioning complicates queries. Async systems introduce ordering, duplication, and observability concerns. Understanding trade-offs matters more than memorizing a "correct" architecture.

Implementation improves understanding. Building a simplified Bloom filter or consistent-hashing ring makes the concept easier to retain and explain than any definition alone.

Educational projects need product thinking. Navigation, search, curriculum order, progress tracking, and visual hierarchy all determine whether learners can actually use the content effectively.

Try It

It's free and open source. Start learning at system-design-in-depth.pages.dev. If you find a mistake or want a topic added, open an issue or PR on GitHub. A star helps others find it.