Subscribe
AWS Well-Architected Framework for AI/ML

AI & ML System Design Architecture Center

Real-world production engineering reference architectures. Design distributed RAG systems, sub-45ms recommendation funnels, and continuous MLOps pipelines with interactive AWS architectural blueprints.

Flight Simulator (Problem Challenges)Interactive

Drag components, wire dataflows, simulate live QPS, and pass interview tests

Challenge Levels (10)
01 • GeneralEasy

URL Shortener (TinyURL)

Design a high-throughput, low-latency URL shortening service like TinyURL or bit.ly. The system must handle 100,000 requests/sec with a 100:1 read-to-write ratio, ensuring sub-25ms redirection latency and persistent long-term storage.

100k QPS•<25ms
02 • GeneralMedium

Distributed Cache Cluster (Redis / Memcached)

Architect a horizontally scalable, in-memory distributed caching tier capable of sustaining 200,000 read QPS with sub-8ms P99 latency. Protect backend persistent databases from cache stampede and hotspot overload.

200k QPS•<8ms
03 • GeneralEasy

Distributed Rate Limiter

Design a distributed API rate limiter capable of throttling abusive clients and enforcing multi-tier tenant quotas across 75,000 requests/sec with strict sub-10ms overhead.

75k QPS•<10ms
04 • SocialMedium

Real-Time Chat & Messaging (WhatsApp / Slack)

Design a high-scale real-time messaging system supporting 1-on-1 and group chats. Support persistent bi-directional WebSocket connections, message ordering, offline storage, and push notifications.

150k QPS•<35ms
05 • SocialHard

Twitter / X News Feed System

Design a scalable social timeline news feed like Twitter / X. The system must support posting tweets, following users, fan-out delivery to millions of followers, and sub-40ms home timeline generation under 150,000 requests/sec.

150k QPS•<40ms
06 • SocialHard

Real-Time Chat System (WhatsApp / Discord)

Design a real-time messaging architecture supporting 1-on-1 private messaging, group chat channels, typing indicators, user online presence, and offline push delivery.

80k QPS•<30ms
07 • MediaHard

Video Streaming Platform (Netflix / YouTube)

Design a planetary-scale video streaming architecture supporting smooth playback, multi-bitrate HLS/DASH transcoding, metadata search, and multi-region CDN caching for millions of concurrent viewers.

200k QPS•<40ms
08 • AI/MLHard

Enterprise Multi-Tenant RAG & LLM Serving

Design an enterprise-scale Retrieval-Augmented Generation (RAG) system capable of serving sub-second grounded answers over 20M enterprise document chunks with strict latency budgets and tenant data isolation.

30k QPS•<320ms
09 • AI/MLHard

Enterprise Generative RAG & Semantic Search

Design an enterprise-grade Retrieval Augmented Generation (RAG) platform. Ingest millions of documentation embeddings, query a distributed vector database, rerank top-K candidates with a cross-encoder, and stream generated responses from vLLM.

25k QPS•<65ms
010 • AI/MLHard

Real-Time Recommendation Engine

Design a sub-45ms real-time recommendation funnel (like TikTok, Netflix, or Spotify) using a two-tower deep learning retrieval architecture, real-time feature stores, and offline continuous training feedback loops.

120k QPS•<45ms
Problem:
Load:100k QPS
Client Traffic
x1
0ms100.0k QPS
Global Route53 DNS
x1
2ms100.0k QPS
Application Load Balancer
x2
1ms100.0k QPS
URL Gateway Cluster
x4
95.7ms100.0k QPS
Redis Hot URL Cache
x2
1ms32.0k QPS
PostgreSQL Primary (Writes)
x1
390.4ms32.0k QPS
PostgreSQL Read Replica
x2
10ms8.3k QPS
100%

Real-Time Traffic Engine

Topological load propagation powered by Kahn's algorithm. Adjust offered QPS from 1,000 to 250,000 requests/sec with smart load balancer splitting and cache hit deductions.

5-Pillar Interview Scoring

Evaluates your architecture like a Principal Engineer across Scalability, Reliability / SPOF detection, Latency Budget SLA, Cost Efficiency, and Layered Decoupling.

35+ Infrastructure Blocks

Stateless app servers, Redis clusters, read replicas, Kafka message queues, vector databases, and vLLM GPU inference nodes with real latency specs.

Continue Learning AI Engineering

Dive into step-by-step algorithms, loss surface mathematics, and complete code walkthroughs.