Load balancing is a process of distributing incoming requests across multiple servers to improve performance, availability, and reliability of a system. Instead of sending all requests to a single server (causing bottleneck), requests are distributed. Benefits: high availability (if one server fails, traffic goes to others), better performance (distribute load), scalability (add more servers). Load balancer sits between client and servers, decides which server handles each request.
Load Balancing Algorithms (mention 2–3 in interview):
1. Round Robin 👉 Requests distributed one by one in order
Example: Server1 → Server2 → Server3 → Server1 → repeat
Simple, fair distribution, good for servers with similar capacity
2. Least Connections 👉 Sends request to server with least active connections
Useful when requests take different time to complete
Better than round robin for variable request durations
3. IP Hash 👉 User IP decides which server will handle request
Helps maintain session consistency (same client always goes to same server)
Useful for stateful applications where session data is stored locally