40 System Design Interview Questions for experienced engineers
40 System Design Interview Questions Explained — From Scaling APIs to Distributed Systems
System-design interviews are rarely about naming the “right” technology. Interviewers usually want to understand whether you can reason about scale, failure, consistency, latency, reliability, security, and trade-offs.
A strong answer normally follows this sequence:
Requirements → bottlenecks → high-level architecture → data flow → failure handling → scaling → trade-offs.
Below is a complete interview-ready guide to all 40 questions from the screenshots you shared.
Part 1 — Scaling Your Backend
Q1. Your backend API must handle 1 million requests per second. How would you design it?
You cannot solve 1M RPS simply by adding a huge server. The system needs to be horizontally scalable.
A possible architecture:
Users
↓
CDN / Edge Network
↓
WAF + Rate Limiter
↓
Global Load Balancer
↓
Regional Load Balancers
↓
Stateless API Servers
↓
Cache / Queue
↓
Database Shards
First, determine what percentage of requests actually need to reach your application. Static responses should be served from a CDN, frequently requested data from Redis, and only necessary operations should hit databases.
API servers should remain stateless so that any request can go to any instance.
For writes that don't need immediate completion, introduce Kafka, Azure Service Bus, AWS SQS, etc.
For example:
POST /analytics-event
↓
API
↓
Kafka
↓
Consumers
↓
Database
Instead of making users wait while analytics processing occurs, the API acknowledges the message quickly.
At this scale, you would also need database partitioning, read replicas, connection pooling, rate limiting, autoscaling, circuit breakers, observability and multi-region deployment.
Interview takeaway: Reduce the amount of work performed synchronously rather than simply adding servers.
Q2. Service A depends on B, C and D. They autoscale and their IP addresses keep changing. Why doesn't the system break?
Because modern distributed systems normally don't communicate using hard-coded IP addresses.
They use service discovery.
For example in Kubernetes:
Service A
↓
http://payment-service
↓
Kubernetes Service
↓
Pod 1
Pod 2
Pod 3
Pods may continuously appear and disappear, but the Kubernetes Service maintains a stable DNS endpoint.
Other options include:
- Kubernetes DNS
- Consul
- Eureka
- AWS Cloud Map
- Azure Service Discovery
- service meshes such as Istio
For example:
payment-service.default.svc.cluster.local
Service A only knows this DNS name.
The infrastructure decides which healthy instance should receive the request.
Q3. Your database handles 1 million reads and writes. Both are slow. Scaling is not the answer. How do you fix it?
Start by identifying why the database is slow.
Possible causes include:
- missing indexes
- inefficient queries
- N+1 queries
- lock contention
- excessive joins
- large table scans
- oversized transactions
- connection pool exhaustion
- unsuitable schema design
Suppose this query runs constantly:
SELECT *
FROM Orders
WHERE CustomerId = 123
AND Status = 'Pending';
If CustomerId and Status aren't indexed, the database might scan millions of rows.
An appropriate composite index could be:
CREATE INDEX IX_Orders_Customer_Status
ON Orders(CustomerId, Status);
Also consider separating workloads:
Writes → Primary DB
Reads → Read Replicas
↑
Cache
For extremely high volume:
Database
↓
Partition / Shard
↓
Shard A
Shard B
Shard C
Database optimization should generally happen before blindly scaling infrastructure.
Q4. Two microservices need to communicate. What options are available?
There are two major approaches.
Synchronous communication
Examples:
REST
gRPC
GraphQL
Flow:
Order Service
↓ HTTP
Payment Service
Use synchronous communication when the caller requires an immediate response.
Example:
Check inventory → response required immediately
Asynchronous communication
Examples:
Kafka
RabbitMQ
Azure Service Bus
AWS SNS/SQS
Flow:
Order Service
↓
OrderCreated event
↓
Kafka
↙ ↓ ↘
Email Inventory Analytics
Use asynchronous communication when services don't need immediate results.
The trade-off is:
Synchronous
+ simple
+ immediate
- tighter coupling
- cascading failures
Asynchronous
+ resilient
+ scalable
+ loosely coupled
- eventual consistency
- operational complexity
Q5. Your API returns one million records. How do you avoid pagination?
Usually, you shouldn't return one million records to a browser at all.
Instead ask what the consumer is trying to accomplish.
For normal UI use:
Cursor pagination
Virtual scrolling
Filtering
Searching
For data exports:
Client
↓
Request export
↓
API
↓
Background Worker
↓
Generate CSV/Excel
↓
Blob Storage
↓
Download URL
The API could return:
{
"jobId": "exp-78231",
"status": "processing"
}
Later:
{
"status": "completed",
"downloadUrl": "..."
}
This is better than keeping an HTTP request open while serializing millions of objects.
For machine-to-machine communication, streaming is another option.
Q6. Your monolith was fast. You moved to microservices and performance collapsed. Why?
Because function calls became network calls.
In a monolith:
OrderService → PaymentService
might simply be an in-process method call taking microseconds.
After migration:
Order Service
↓ HTTP
API Gateway
↓
Payment Service
↓ HTTP
Inventory Service
↓ HTTP
Pricing Service
Every network call introduces:
- serialization
- TLS
- network latency
- connection overhead
- retries
- load-balancer overhead
The bigger problem is often chatty architecture.
Example:
UI
↓
API
↓
Service A
↓
Service B
↓
Service C
↓
Service D
If every service takes 80 ms, latency quickly grows.
Solutions include:
- coarse-grained APIs
- asynchronous events
- caching
- gRPC for internal communication
- parallel service calls
- local aggregation
- avoiding unnecessary service boundaries
Microservices improve organizational and scalability characteristics, but they do not automatically improve performance.
Q7. You sharded the database using User ID. One celebrity joins and one shard freezes. How do you fix it without downtime?
This is a hot shard / hot partition problem.
Normal users may generate 10 requests per minute.
A celebrity could generate millions.
If all celebrity-related data maps to:
Shard 7
Shard 7 becomes overloaded.
Instead of simple partitioning such as:
hash(userId)
you may introduce finer partitioning:
hash(userId + bucketId)
For extremely hot users:
Celebrity
↓
Partition A
Partition B
Partition C
Partition D
Other mitigation techniques:
- cache popular data
- separate celebrity/fan-out workloads
- use consistent hashing
- split the hot partition
- use virtual shards
- rebalance incrementally
To avoid downtime, migrate gradually:
Old shard
↓ dual write
Old + New partitions
↓
Backfill historical data
↓
Switch reads
↓
Stop old writes
Q8. 100 million users refresh a cricket score simultaneously. How do you stop 100 million database hits?
Never make the database answer identical requests independently.
The score should be processed once and distributed many times.
Architecture:
Score Provider
↓
Ingestion Service
↓
Redis / Event Stream
↓
WebSocket Servers
↓
Millions of Users
For polling clients:
Users
↓
CDN
↓
Edge cache
↓
Application cache
Cache:
IND vs AUS score
TTL = 1 second
One backend fetch may therefore serve millions of users.
Even better, use push:
Score changed
↓
Publish event
↓
WebSockets / SSE
↓
Connected clients
This prevents the classic thundering herd problem.
Q9. Two users generate a TinyURL simultaneously and get the same short code. How do you prevent collisions?
Use globally unique ID generation.
Options include:
- database sequences
- UUIDs
- Snowflake IDs
- distributed ID generators
- atomic counters
Example:
Database ID = 987654321
Convert to Base62:
987654321 → 14q60P
Then URL becomes:
short.ly/14q60P
Because the numerical ID is unique, the generated code is also unique.
Additionally enforce:
UNIQUE(short_code)
Even if an application bug creates a collision, the database prevents duplicate entries.
Q10. Autocomplete takes 600 ms with 10 million simultaneous users. How do you improve it?
Autocomplete requires extremely fast prefix lookup.
Don't execute:
SELECT *
FROM Products
WHERE Name LIKE 'iph%';
against a huge database for every keystroke.
Better options:
- Trie
- Elasticsearch/OpenSearch
- precomputed popular queries
- Redis
- edge caching
Also implement frontend debouncing:
User types:
i
ip
iph
ipho
iphone
Without debounce:
5 API calls
With a 200-ms debounce:
1 API call
Architecture:
Browser debounce
↓
CDN
↓
Autocomplete Service
↓
Redis
↓
Search Index
Target latency would generally be closer to tens of milliseconds than hundreds.
Q11. Why can one ticketing system crash during a booking rush while an e-commerce system survives similar traffic?
The critical issue isn't simply traffic volume. It is the shape of traffic and resource contention.
A ticket sale may involve millions of users attempting to acquire the same limited inventory simultaneously.
Example:
Train seats = 1,000
Users = 2,000,000
This creates massive contention.
A resilient design could introduce:
Users
↓
Virtual waiting room
↓
Rate limiter
↓
Booking service
↓
Inventory reservation
↓
Payment
Only a controlled number of requests reach inventory at once.
Use:
- token-based queues
- short-lived reservations
- atomic inventory operations
- idempotency
- asynchronous processing
- aggressive caching
- backpressure
A large traffic number itself is manageable. Uncontrolled synchronized traffic is the bigger danger.
Q12. Ten servers sit behind a load balancer. One reaches 95% CPU while the others are idle. Why?
Possible causes include:
- sticky sessions
- long-lived WebSocket connections
- unequal request complexity
- poor load-balancer algorithm
- stale health data
- uneven connection reuse
For example, round robin sends the same number of requests:
Server A → 100 requests
Server B → 100 requests
But Server A might receive 100 expensive report-generation requests while Server B receives simple health checks.
Better algorithms include:
Least Connections
Weighted Least Connections
Least Response Time
For stateful workloads, also investigate session affinity.
Part 2 — Microservices & APIs
Q13. Server A crashes, traffic moves to B, then B crashes, and eventually all servers fail. What is happening?
This is a cascading failure.
Imagine a dependency becomes slow.
Server A accumulates threads waiting for that dependency and crashes.
Its traffic moves to B:
A fails
↓
B receives extra traffic
↓
B fails
↓
C receives more
↓
C fails
Solutions:
Timeouts
Circuit breakers
Bulkheads
Rate limiting
Backpressure
Load shedding
Retries with exponential backoff
Circuit breaker:
Payment service failing
↓
Stop calling temporarily
↓
Return fallback / failure quickly
A system that fails quickly can sometimes survive better than one that endlessly waits.
Q14. Redis goes down for 30 seconds. It returns, but one million requests hit the database and the DB collapses. How do you prevent it?
This is a cache stampede.
Normally:
Requests → Redis → DB only occasionally
When cache disappears:
1,000,000 requests → DB
Solutions include:
Request coalescing
Only one request reloads the missing key:
1000 requests
↓
lock
↓
1 request → DB
↓
others wait
TTL jitter
Don't make millions of keys expire at exactly:
12:00:00
Instead:
59m 45s
60m 03s
60m 18s
59m 51s
Stale-while-revalidate
Serve slightly stale data while refreshing in the background.
Also use:
- local in-memory caches
- database rate limiting
- circuit breakers
- gradual cache warming
Q15. You use an LRU cache independently on three servers while the load balancer routes users randomly. Does caching work?
It works, but inefficiently.
Suppose:
Request 1 → Server A → cache populated
Request 2 → Server B → cache miss
Request 3 → Server C → cache miss
The same data gets duplicated across three local caches.
Better options:
Shared Redis cache
or:
L1 local cache
↓
L2 Redis cache
↓
Database
A two-layer cache provides extremely fast local access while maintaining a shared distributed cache.
Q16. Hundreds of millions of users produce watch events. How can a recommendation system process them?
You do not synchronously recalculate a user's entire recommendation model after every event.
A typical event-driven design is:
User watches content
↓
WatchEvent
↓
Kafka/Event Stream
↓
Stream Processing
↓
Feature Store
↓
Recommendation Pipeline
Different consumers may process the same event:
WatchEvent
├── analytics
├── recommendation features
├── trending calculation
└── personalization
Online and offline processing may be combined.
Real-time features
+
Batch-trained models
=
Recommendations
Important interview point: large consumer platforms generally use layered recommendation architectures; don't claim a specific proprietary internal design unless you have verified evidence.
Q17. PhonePe transfers money to Google Pay but doesn't access Google's database. How can the payment be verified?
Independent systems communicate through payment network protocols and APIs, not each other's databases.
Conceptually:
App A
↓
Payment Service / PSP
↓
UPI/payment network
↓
Recipient PSP/bank
↓
Confirmation
Each participant maintains its own ledger.
A transaction includes identifiers such as:
Transaction ID
Sender
Recipient
Amount
Timestamp
Status
Systems reconcile based on trusted network messages and transaction IDs.
The core principle is:
Distributed organizations exchange messages/contracts; they do not share internal databases.
Q18. A request goes through 20 microservices. One fails. How do you find where?
Use distributed tracing.
Generate a Trace ID when the request enters the system:
TraceId = ABC123
Propagate it:
API Gateway ABC123
Order Service ABC123
Inventory ABC123
Payment ABC123
Email ABC123
Each service generates spans.
Example:
Order 800 ms
├─ Inventory 50 ms
├─ Pricing 60 ms
└─ Payment 620 ms ← problem
Popular tooling:
- OpenTelemetry
- Jaeger
- Zipkin
- Azure Application Insights
- Datadog
- New Relic
Combine:
Logs + Metrics + Traces
for observability.
Part 3 — Data Consistency & Replication
Q19. Two users edit the same record simultaneously and one overwrites the other. How do you prevent data loss?
This is a concurrency-control problem.
A common solution is optimistic concurrency control.
Record:
Document
Version = 5
User A and B both load version 5.
User A updates:
Version 5 → 6
User B submits version 5.
Database executes:
UPDATE Document
SET Content = @content,
Version = Version + 1
WHERE Id = @id
AND Version = 5;
Zero rows affected means someone already changed it.
The system tells User B:
Document changed. Reload or merge changes.
Other solutions include:
- pessimistic locking
- ETags
- CRDTs
- operational transformation for collaborative editors
Q20. Git can tell that one branch is 17 commits behind another even with millions of commits. How?
Git represents commit history as a directed acyclic graph — DAG.
Each commit contains references to its parent commit(s).
Example:
A → B → C → D → E main
\
F → G feature
Git can identify the merge base and determine commits reachable from one branch but not another.
Implementations optimize graph traversal using indexes and cached metadata, so the system doesn't naively inspect repository contents every time.
Interview concept:
Commit graph + graph traversal + indexing
Q21. A billing job crashes halfway through and runs again. 1,000 customers get charged twice. How do you prevent it?
Make the operation idempotent.
Generate:
BillingOperationId =
Customer123-August2026
Before charging:
Does this operation ID already exist?
YES → do not charge
NO → process charge
Database:
PaymentOperations
OperationId | Customer | Status
Add:
UNIQUE(OperationId)
Payment APIs should also support idempotency keys:
Idempotency-Key: invoice-12345
Retrying then returns the existing result instead of creating another payment.
Q22. The user gets “Payment succeeded” and later “Payment failed.” How do distributed systems avoid contradictory states?
Use an explicit state machine.
For example:
CREATED
↓
PROCESSING
↓
AUTHORIZED
↓
CAPTURED
or:
PROCESSING
↓
FAILED
Define valid transitions.
For example:
CAPTURED → FAILED
may not be allowed directly.
If later compensation occurs:
CAPTURED → REFUND_PENDING → REFUNDED
Also maintain:
- one authoritative source for payment state
- version numbers
- ordered events
- idempotency
- transactional outbox
- reconciliation
Don't let every microservice independently invent a payment status.
Q23. Service A generates a JWT, B accepts it but C rejects it during a network partition. How do you ensure consistent JWT validation?
JWT verification should not require calling the issuing service for every request.
Use asymmetric signing:
Auth Service
Private Key → signs token
Services B/C
Public Key → verify token
Publish keys through JWKS.
Services cache:
kid → public key
During key rotation:
Old key + New key
should overlap for a defined period so already-issued tokens remain verifiable.
Important things to keep consistent:
issuer
audience
algorithm
clock skew
token expiration
public keys
Centralizing authentication policy through an API gateway or service mesh can further reduce inconsistency.
Q24. How do you build a persistent shopping cart that scales?
Don't store the cart only in server memory.
Bad:
Server A memory
If Server A dies, the cart disappears.
Use:
Client
↓
Cart API
↓
Redis / Distributed Store
↓
Persistent Database
Redis gives low latency.
The database provides durability.
Example:
cart:user123
{
product17: 2,
product51: 1
}
Use TTL for abandoned carts, but persist valuable data if users expect carts to survive across devices and long periods.
For anonymous users, create a temporary cart ID and merge it after login.
Part 4 — System Design Patterns
Q25. Your app is used in 100 countries but data exists only in the US. Users expect <100 ms latency. What do you do?
Physics becomes the limiting factor.
A user in India cannot repeatedly communicate with a US database and still reliably achieve tiny latency.
Move data and compute closer to users.
Global DNS / Anycast
↓
Nearest Region
India
Europe
US
Singapore
Australia
Use:
- CDN for static content
- edge caching
- regional application servers
- replicated databases
- multi-region deployment
Architecture:
India Users → Mumbai Region
EU Users → Frankfurt Region
US Users → Virginia Region
The difficult part is global writes.
You need to choose between:
Strong consistency
vs
lower latency / higher availability
depending on the domain.
Q26. How would you efficiently find nearby Uber-like drivers?
Use geospatial indexing.
Store driver locations as geographic coordinates:
Driver123
lat = ...
lng = ...
Instead of comparing every driver against every user, divide the earth into cells using approaches such as:
- Geohash
- H3
- S2
- R-tree-based spatial indexes
Example:
Current user → cell XYZ
Search:
XYZ
+ neighbouring cells
Driver updates:
Driver phone
↓ every few seconds
Location Service
↓
Geospatial store
Nearby search:
Find drivers within 3 km
Redis GEO or databases/search systems with geospatial indexing can support such patterns.
Q27. How do you design a scalable rate limiter for Free and Pro API plans?
A common algorithm is token bucket.
Suppose:
Free:
100 requests/minute
Pro:
10,000 requests/minute
Each customer has a bucket.
capacity = 100
refill = 100/minute
Every request consumes one token.
If:
tokens > 0 → allow
tokens = 0 → HTTP 429
In distributed systems, store counters in Redis and update atomically.
Other algorithms:
- fixed window
- sliding window
- leaky bucket
- token bucket
Token bucket is popular because it supports controlled bursts.
Place rate limiting near the system boundary:
Client
↓
API Gateway
↓
Rate Limiter
↓
Services
This rejects abusive traffic before consuming expensive backend resources.
Q28. Reporting queries lock your transactional SQL database. How do you separate OLTP and OLAP?
Separate workloads.
OLTP
Optimized for:
INSERT
UPDATE
DELETE
small point queries
Examples:
Orders
Payments
Customers
OLAP
Optimized for:
aggregation
analytics
reports
large scans
Architecture:
Transactional DB
↓
CDC / ETL / Streaming
↓
Data Warehouse
↓
BI / Reports
Example technologies:
SQL Server/PostgreSQL
↓
CDC
↓
Snowflake / BigQuery / Synapse / Databricks
Now:
Customer transactions → OLTP
Power BI reports → OLAP
Heavy reports no longer lock the operational database.
Q29. “Place Order” triggers inventory, payment, email and shipping. How do you ensure all steps complete or none do?
A distributed transaction across independent microservices is usually better handled with a Saga pattern than a giant database transaction.
Example:
Create Order
↓
Reserve Inventory
↓
Charge Payment
↓
Create Shipment
Suppose shipping fails.
Compensating actions may be:
Cancel Shipment
Refund Payment
Release Inventory
Cancel Order
There are two Saga styles.
Orchestration
Saga Orchestrator
↓
Inventory
↓
Payment
↓
Shipping
Choreography
OrderCreated
↓
InventoryReserved
↓
PaymentCompleted
↓
ShipmentCreated
Services react to events.
For complex business processes, orchestration is often easier to understand and troubleshoot.
Q30. How would you design an Instagram-like feed?
Two major approaches exist.
Fan-out on write
When a creator posts:
Post
↓
Find followers
↓
Push post ID into follower feeds
Reading is extremely fast.
Problem:
A creator with 100 million followers creates an enormous write workload.
Fan-out on read
When a user opens the app:
Find followed accounts
↓
Read recent posts
↓
Rank
↓
Return feed
Writes are cheap, but reads are expensive.
Large systems often use a hybrid strategy.
For normal accounts:
Fan-out on write
For celebrities:
Fan-out on read
Feed generation may involve:
Candidate generation
↓
Filtering
↓
Ranking model
↓
Caching
↓
User feed
Part 5 — Advanced Topics
Q31. Your system processes 10 TB/day and needs real-time analytics. How do you design the pipeline?
Use a streaming architecture.
Applications / Devices
↓
Kafka / Event Hub
↓
Stream Processing
↓
┌─────────────┬──────────────┐
↓ ↓ ↓
Real-time DB Data Lake Alerts
↓
Dashboards
Possible stack:
Kafka
↓
Flink / Spark Streaming
↓
ClickHouse / Elasticsearch
↓
Dashboard
Raw events should also be stored:
Kafka
↓
Object Storage/Data Lake
This enables historical processing and model training.
Partition streams by a suitable key:
customerId
deviceId
region
to distribute processing across workers.
Q32. Design real-time multiplayer state synchronization for thousands of players.
Use long-lived connections such as WebSockets.
Architecture:
Player
↓
WebSocket Gateway
↓
Game Server
↓
Game State
Each game room should generally have one authoritative state owner.
Example:
Room 123
Players A, B, C, D
↓
Game Server 7
When Player A moves:
A → Game Server
→ validate move
→ update authoritative state
→ broadcast event
→ B,C,D
Do not trust client state.
The server validates:
Is this player's turn?
Is the move legal?
What is the new state?
For large scale:
GameId hash → server/shard
Redis or similar systems can maintain ephemeral shared state, while durable events/results can be persisted separately.
Q33. How do you implement online/typing presence in chat?
WebSockets are well suited.
When a user connects:
user:123 → online
Store presence in Redis:
presence:123 = online
TTL = 60 seconds
Clients send heartbeats:
every 20 seconds
If heartbeat stops and TTL expires:
online → offline
Typing indicators should generally be ephemeral.
User started typing
↓
WebSocket event
↓
Other participant
Do not persist every typing event in a relational database.
Q34. Your public API receives 100× normal traffic due to DDoS. How do you protect it?
Protection should occur before traffic reaches your application servers.
Architecture:
Internet
↓
Anycast / DDoS Protection
↓
CDN
↓
WAF
↓
Rate Limiter
↓
Load Balancer
↓
Applications
Strategies include:
- network-level DDoS mitigation
- WAF rules
- per-IP throttling
- account-level limits
- bot detection
- CDN caching
- request-size limits
- connection limits
- challenge mechanisms
- autoscaling
- load shedding
Protect downstream systems independently too.
Even if APIs survive, attackers shouldn't be allowed to create:
500,000 DB queries/second
Q35. How would you design a CDN?
A CDN replicates/cacheable content close to users.
Origin Server
↓
Edge POPs worldwide
A user in Delhi requests:
/image.jpg
DNS/Anycast directs the request to a nearby edge location.
If cached:
Edge → user
If not:
Edge → Origin
← object
Edge caches object
→ user
Important concepts:
TTL
Cache-Control
ETag
Cache invalidation
Origin shielding
Consistent hashing
Geo-routing
Versioned asset URLs make invalidation easier:
app.v123.js
instead of repeatedly purging:
app.js
Q36. What is CAP theorem?
CAP states that when a distributed-system network partition occurs, the system cannot simultaneously guarantee both:
C — Consistency
Every read sees the latest successful write.
A — Availability
Every request receives a non-error response.
P — Partition tolerance
The system continues operating despite network partitions between nodes.
Because distributed systems must generally tolerate partitions, the meaningful trade-off during a partition is typically:
CP
or
AP
CP example
Bank balance:
Better to temporarily reject a transaction
than show an incorrect balance.
AP example
Some social features:
Better to show slightly stale like counts
than make the application unavailable.
CAP does not mean “choose any two forever.” The trade-off specifically matters under partition conditions.
Q37. How do you guarantee exactly-once notifications despite failures?
True end-to-end exactly-once delivery is difficult because failures can occur after an external side effect but before acknowledgement.
A more practical design is:
At-least-once delivery
+
Idempotent processing
=
Effectively once
Assign each notification:
NotificationId = N123
Consumer processing:
Have I processed N123?
YES → ignore
NO → process and store N123
Use:
- message IDs
- deduplication
- idempotency keys
- durable queues
- acknowledgements
- transactional outbox
Example:
Database Transaction
├─ save business change
└─ save Outbox event
A background publisher sends the outbox event.
This avoids:
DB saved
but
message never published
Q38. Design a secure password-reset flow.
A proper flow is:
User enters email
↓
POST /forgot-password
↓
Server generates random token
↓
Store HASH(token)
with expiration
↓
Send link through email
Link:
https://example.com/reset?token=XYZ
Important: don't store the raw reset token if avoidable.
Store:
SHA256(token)
expiry
userId
used=false
When the user submits a new password:
Hash received token
↓
Find record
↓
Check expiry
↓
Check used=false
↓
Update password
↓
Mark token used
Password itself should be stored using a password hashing algorithm such as:
Argon2id
bcrypt
PBKDF2
Additional protections:
- short token expiry
- single use
- rate limiting
- generic responses such as “If an account exists...”
- revoke sessions where appropriate
- audit logging
Q39. What is a reverse proxy, and how is it different from a load balancer?
A reverse proxy sits in front of backend servers.
Client
↓
Reverse Proxy
↓
Backend
It can provide:
- TLS termination
- authentication
- routing
- caching
- compression
- header manipulation
- logging
- rate limiting
Examples include:
NGINX
Envoy
HAProxy
Traefik
A load balancer has a more specific primary responsibility:
Distribute traffic among multiple instances
Example:
Client
↓
Load Balancer
├─ Server A
├─ Server B
└─ Server C
There is considerable overlap.
Many products perform both roles.
You can say in an interview:
A reverse proxy represents backend services to clients and can perform routing, security and protocol functions. Load balancing is one capability that distributes traffic across backend instances.
Q40. How do you design reliable long-running background jobs such as video processing?
Never make the user keep an HTTP request open while processing a 30-minute job.
Architecture:
Client
↓
Upload Video
↓
Object Storage
↓
Create Job
↓
Queue
↓
Worker Pool
↓
Processing
↓
Output Storage
API immediately returns:
{
"jobId": "job-91827",
"status": "queued"
}
Worker:
Queue
↓
Download video
↓
Process
↓
Upload output
↓
Update job status
To survive failures, maintain states such as:
QUEUED
PROCESSING
COMPLETED
FAILED
RETRYING
Use:
- checkpoints
- retries
- exponential backoff
- visibility timeouts/message locks
- dead-letter queues
- idempotent workers
- heartbeat/lease renewal
- autoscaling based on queue depth
Suppose a worker crashes at 70%.
Instead of blindly starting again, a sophisticated workflow might checkpoint stages:
Upload ✓
Extract audio ✓
Transcode 1080p ✓
Transcode 720p ← crashed
Thumbnail pending
Another worker resumes from the unfinished stage.
For complex pipelines, a workflow/orchestration engine can help:
Temporal
Azure Durable Functions
AWS Step Functions
Airflow
Bringing Everything Together
If you study these 40 questions individually, they may look unrelated. In reality, most system-design problems reduce to a fairly small collection of principles.
| Problem | Common solution | |---|---| | Too many reads | Cache, CDN, replicas | | Too many writes | Partitioning, queues, batching | | Traffic spikes | Queue, rate limit, backpressure | | Service failures | Circuit breaker, timeout, retry | | Duplicate operations | Idempotency | | Cross-service transaction | Saga | | Lost events | Transactional outbox | | Slow synchronous work | Background processing | | Huge datasets | Partition/shard | | Global latency | CDN + multi-region | | Hot partition | Repartitioning / virtual shards | | Concurrent updates | Optimistic locking | | Cache overload | Stampede protection | | Debugging microservices | Distributed tracing | | Real-time updates | WebSockets/SSE | | Search/autocomplete | Specialized indexes | | Analytics hurting DB | OLTP/OLAP separation | | Event processing | Kafka/event streaming |
A System-Design Interview Framework You Can Reuse
For your interviews, don't immediately answer:
“I will use Kafka, Redis and Kubernetes.”
That often sounds tool-driven.
Instead, structure your response.
1. Clarify requirements
Ask:
How many users?
Read/write ratio?
Latency expectation?
Availability requirement?
Data size?
Consistency requirement?
2. Estimate scale
For example:
100M users
10M DAU
5 requests/user/minute
≈ 830K requests/minute
≈ 14K requests/sec average
Then estimate peak load.
3. Draw the high-level architecture
Start simple:
Client
↓
Load Balancer
↓
API
↓
Database
Then evolve it:
Client
↓
CDN
↓
API Gateway
↓
Load Balancer
↓
Services
↓
Cache
↓
DB
4. Identify the bottleneck
Ask yourself:
What breaks first?
CPU?
Database?
Network?
Storage?
Lock contention?
A hot partition?
5. Handle failure
Interviewers love this question:
“What happens if this component fails?”
Discuss:
Retries
Timeouts
Circuit breakers
Dead-letter queues
Replication
Idempotency
Fallbacks
6. Discuss consistency
Explain whether you require:
Strong consistency
or
Eventual consistency
Example:
Money → strong consistency
Instagram like count → eventual consistency may be acceptable
7. Discuss observability
Production architecture isn't complete without:
Logs
Metrics
Traces
Alerts
Dashboards
The Most Important Patterns to Memorize
If you're preparing for Staff Engineer / Tech Lead system-design rounds, I would particularly master these concepts:
Idempotency → prevents duplicate operations.
Saga → manages distributed business transactions.
Transactional Outbox → prevents DB/message inconsistency.
Circuit Breaker → prevents cascading failures.
Bulkhead → isolates resource failures.
Retry + Exponential Backoff + Jitter → handles transient failures.
Rate Limiting → protects services.
Cache-Aside → reduces database pressure.
Cache Stampede Protection → prevents DB collapse.
CQRS → separates read/write models when useful.
Event-driven architecture → decouples services.
Optimistic concurrency → prevents lost updates.
Distributed tracing → diagnoses microservice chains.
Sharding → distributes large datasets.
Consistent hashing → distributes keys while minimizing movement.
WebSockets/SSE → provides real-time communication.
CDC → moves operational data into analytics systems.
DLQ → isolates repeatedly failing messages.
Backpressure → prevents downstream overload.
One Final Interview Example
Suppose an interviewer says:
“Design an invoice-processing platform that receives millions of invoices.”
A strong answer might evolve like this:
┌─────────────┐
│ Client │
└──────┬──────┘
↓
┌─────────────┐
│ API Gateway │
└──────┬──────┘
↓
┌──────────────┐
│ Upload API │
└──────┬───────┘
↓
┌───────────────────────┐
│ Blob Storage │
└───────────┬───────────┘
↓
┌───────────────┐
│ Service Bus │
└───────┬───────┘
↓
┌───────────────┐
│ Worker Pool │
└───────┬───────┘
↓
┌─────────────────┐
│ Document AI/OCR │
└───────┬─────────┘
↓
┌─────────────┐
│ LLM / Rules │
└──────┬──────┘
↓
┌──────────────┐
│ Database │
└──────┬───────┘
↓
Webhook / Result
Then explain:
Idempotency → prevent duplicate invoices
Queue → absorb traffic spikes
DLQ → capture failed invoices
Retry → transient failures
Checkpoint → resume processing
Blob → large document storage
Database → metadata/state
Tracing → follow invoice end-to-end
Autoscaling → based on queue depth
That is the type of thinking interviewers usually want to see.
Closing Thoughts
The biggest shift when moving from a Senior Engineer mindset toward Tech Lead / Staff Engineer system-design interviews is this:
Don't only explain how something works.
Explain:
Why you chose it, what problem it solves, what can fail, how it scales, and what trade-off you're accepting.
For example, instead of saying:
“I'll use Kafka.”
Say:
“I don't need the caller to wait for this processing, and traffic can spike significantly, so I'll decouple ingestion from processing with a durable queue/event stream. That also gives me buffering and independent worker scaling. The trade-off is eventual consistency and additional operational complexity.”
That single style change makes a system-design answer considerably stronger.
For interview preparation, Q13, Q14, Q18, Q19, Q21, Q22, Q27, Q29, Q36, Q37 and Q40 are especially important, because the principles behind them—cascading failure, cache stampede, observability, concurrency, idempotency, distributed state, rate limiting, Saga, CAP, reliable messaging and background processing—reappear in many Staff/Lead-level interviews.