What is a Load Balancer?
A load balancer distributes incoming traffic across multiple servers.
Instead of sending every request to one server:
Users ↓ Single Server
Traffic is distributed like this:
Users ↓ Load Balancer ↓ Server 1 Server 2 Server 3
This improves availability, performance, and scalability.
Short answer: A load balancer distributes traffic across multiple healthy servers so that no single server becomes overloaded
Why Do We Need a Load Balancer?
Suppose an application receives thousands of requests at the same time.
If all requests go to one server, that server may become slow or crash.
A load balancer distributes the requests across multiple servers.
This helps with:
- High traffic
- Better performance
- High availability
- Fault tolerance
- Horizontal scaling
What Happens If One Server Fails?
Load balancers regularly perform health checks on backend servers.
Example:
Load Balancer ↓ Server 1 ✓ Server 2 ✕ Server 3 ✓
If Server 2 becomes unhealthy, the load balancer stops sending normal traffic to it and routes new requests to healthy servers.
Interview point: Load balancing also improves availability by removing unhealthy servers from traffic rotation.
What Is Round Robin Load Balancing?
Round Robin sends requests to servers one after another.
Request 1 → Server 1 Request 2 → Server 2 Request 3 → Server 3 Request 4 → Server 1
It works well when all servers have similar capacity.
Client
Load Balancer
Backend Server 1
Backend Server 2
Backend Server 3
What Is Layer 4 Load Balancing?
Layer 4 load balancing works at the transport layer.
It makes routing decisions mainly using information such as:
- IP address
- Port
- TCP
- UDP
It does not need to understand application URLs such as /orders or /products.
AWS Network Load Balancer is a common Layer 4 example.
What Is Least Connections?
Least Connections sends a new request to the server currently handling the fewest active connections.
For example:
Server 1 → 100 connections Server 2 → 30 connections Server 3 → 70 connections
The next connection may be sent to Server 2.
This can be useful when requests take different amounts of time to complete.
Example: Configuring a Simple NGINX Load Balancer
1http {2 upstream backend {3 server backend1.example.com weight=3;4 server backend2.example.com;5 server backend3.example.com;6 }7 8 server {9 listen 80;10 11 location / {12 proxy_pass http://backend;13 }14 }15}This NGINX configuration defines an upstream group named 'backend' with three servers. The first server has a higher weight, receiving more requests. The server block listens on port 80 and proxies incoming requests to the backend group, effectively acting as a load balancer.
What Is Layer 7 Load Balancing?
What Is Layer 7 Load Balancing?
Layer 7 operates at the application layer.
It understands application-level information such as:
- HTTP methods
- URL paths
- Hostnames
- HTTP headers
Example:
/orders → Order Service /products → Product Service /users → User Service
AWS Application Load Balancer is a common Layer 7 example.
What Are Sticky Sessions?
Normally, each request from a user can go to a different server.
Request 1 → Server 1 Request 2 → Server 2 Request 3 → Server 3
With a sticky session, requests from the same user are directed to the same backend server for a period of time.
User A → Server 2 User A → Server 2 User A → Server 2
Sticky sessions can be useful for stateful applications, but designing applications to be stateless is often more scalable.
What Is SSL/TLS Termination?
SSL termination means the load balancer handles HTTPS encryption and decryption.
Client ↓ HTTPS Load Balancer ↓ Backend Servers
This can reduce the encryption workload on backend servers and centralize certificate management.
ALB vs NLB — When Would You Choose Each?
Choose ALB when you need:
- HTTP/HTTPS traffic
- Path-based routing
- Host-based routing
- REST APIs
- Microservices
Choose NLB when you need:
- TCP/UDP traffic
- Very high connection volumes
- Low latency
- Network-level load balancing
- Static IP support
Simple answer: ALB is for intelligent application routing, while NLB is for high-performance network traffic.
What Is the Difference Between a Load Balancer and a Reverse Proxy?
A reverse proxy sits between clients and backend servers and forwards requests.
A load balancer also forwards traffic, but its main goal is to distribute traffic across multiple backend targets.
Tools such as NGINX and HAProxy can perform both roles.
How Does a Load Balancer Work with Auto Scaling?
Imagine your application normally has three servers.
Load Balancer ├── Server 1 ├── Server 2 └── Server 3
Traffic suddenly increases.
Auto Scaling can add more servers:
Load Balancer ├── Server 1 ├── Server 2 ├── Server 3 ├── Server 4 └── Server 5
The load balancer then distributes traffic across the available healthy servers.
When traffic decreases, unnecessary servers can be removed.
When to Use and When Not to Use Load Balancers
Use Load Balancers When: - You need to scale applications horizontally. - High availability and fault tolerance are required. - Traffic must be distributed evenly or based on custom rules. - SSL termination or session persistence is necessary. Avoid Load Balancers When: - The application is simple with minimal traffic and a single server suffices. - Latency sensitivity prohibits additional network hops. - Backend servers cannot handle stateless requests or session affinity is impossible. - Infrastructure complexity must be minimized.
Key Takeaways
- Load balancers distribute client requests to optimize resource utilization, availability, and performance.
- They operate at network or application layers, using algorithms and health checks to route traffic.
- Understanding differences from reverse proxies and API gateways clarifies their role in system architecture.
- Proper configuration, including health checks and session persistence, is critical for reliability.
- Trade-offs exist between complexity, performance, scalability, and fault tolerance in load balancer design.
Frequently Asked Questions
What is the difference between Layer 4 and Layer 7 load balancing?+
Layer 4 load balancing operates at the transport layer, routing traffic based on IP address and TCP/UDP ports without inspecting packet content. Layer 7 load balancing operates at the application layer, making routing decisions based on application data such as HTTP headers, cookies, or URL paths, enabling more granular control.
How do load balancers detect unhealthy backend servers?+
Load balancers perform health checks by periodically sending probes such as TCP connection attempts, HTTP requests, or custom application-level checks to backend servers. If a server fails a configurable number of consecutive checks, it is marked unhealthy and removed from the rotation until it recovers.
What is session persistence and why is it important?+
Session persistence, or sticky sessions, ensures that requests from the same client are consistently routed to the same backend server. This is important for stateful applications where session data is stored locally on servers, preventing session loss and inconsistent behavior.
Can load balancers handle SSL termination?+
Yes, many load balancers can terminate SSL/TLS connections, decrypting incoming encrypted traffic before forwarding it to backend servers. This offloads the computational overhead from backend servers and centralizes certificate management.