Fault Tolerance
Fault tolerance means the system can keep working even when some parts fail.
In Kafka:
- If a broker fails, leaders are re-elected from replicas.
- If a consumer fails, another consumer in the same group takes over its partitions.
- If a producer fails, retries ensure the message is delivered again.
This design makes Kafka a strong system for real-time, always-on data pipelines.
Diagram
Partitions and Replication
graph TD
subgraph Cluster["Kafka Cluster"]
B1[Broker 1]
B2[Broker 2]
B3[Broker 3]
end
subgraph Orders["Topic: orders"]
O0[Partition 0<br/>Leader: B1, Replicas: B2,B3]
O1[Partition 1<br/>Leader: B2, Replicas: B3,B1]
O2[Partition 2<br/>Leader: B3, Replicas: B1,B2]
end
B1 --> Orders
B2 --> Orders
B3 --> OrdersFault Tolerance (Leader Failure and Recovery)
sequenceDiagram
participant P as Producer
participant B1 as Broker 1 (Leader for Partition 0)
participant B2 as Broker 2 (Replica)
participant C as Consumer
P->>B1: Send message to Partition 0 (Leader)
B1->>B2: Replicate message
B1-->>C: Deliver message to consumer
Note over B1: Broker 1 goes down ❌
Note over B2: Broker 2 promoted to Leader ✅
P->>B2: Continue sending messages
B2-->>C: Consumer keeps reading
Quick Notes
- Partitions split data for scalability.
- Replication makes copies for safety.
- Fault tolerance keeps the system alive during failures.
Conclusion
Kafka’s architecture with partitions, replication, and fault tolerance allows it to be:
- Scalable: handle huge data streams.
- Reliable: no single point of failure.
- Available: always ready to serve data.
This is why many companies trust Kafka for critical real-time systems.
Pages: 1 2
Category: Kafka
