Why Your Heartbeat Timeout Should Be 3x Your Network RTT (And How to Calculate It)
In distributed systems, a heartbeat timeout that is too low causes false failure detection, while one that is too high delays the recovery from actual failures. A safe rule of thumb is to set the heartbeat timeout to at least three times the measured network round-trip time…
In distributed systems, a heartbeat timeout that is too low causes false failure detection, while one that is too high delays the recovery from actual failures. A safe rule of thumb is to set the heartbeat timeout to at least three times the measured network round-trip time (RTT), plus an allowance for jitter. This article explains the math behind that rule, how to measure RTT in your environment, and how to adjust the multiplier when your network is unstable.
Key takeaways
- Heartbeat timeout must account for RTT and jitter, not just a fixed constant.
- A 3x RTT baseline is a minimum; increase it when packet delay variation is high.
- Measure RTT under real load, not during idle periods.
- False positives trigger unnecessary leader elections and service disruption.
- Use exponential backoff or adaptive timers if jitter exceeds 50% of RTT.
Why heartbeats fail even when a node is alive
A heartbeat is a small message that a node sends to a coordinator or to its peers at fixed intervals. The receiver expects the message to arrive before a deadline. If the message does not arrive in time, the receiver marks the sender as failed and starts recovery.
Two common reasons cause a heartbeat to miss its deadline:
- Network queuing. A router or switch holds the packet while it processes other traffic. The packet is not lost; it is only delayed.
- Jitter. The delay between two consecutive packets from the same sender varies. One packet takes 10 ms, the next takes 40 ms, even though the average is 20 ms.
If the timeout is set to the average RTT, any jitter above zero will eventually cause a false failure. The system then re-elects a leader, re-partitions data, or restarts services, all while the original node is still running. This is called a false positive, and it is the main source of instability in production clusters.
The math: timeout = RTT + k × jitter
Let:
RTT= measured round-trip time between two nodes, in milliseconds.J= standard deviation of RTT samples, also in milliseconds.k= safety multiplier, usually between 2 and 4.
A conservative timeout is:
1timeout = RTT + k × JWhen you do not have jitter measurements, use a proxy: assume J is about 30% of RTT. Then:
1timeout = RTT + 3 × 0.3 × RTT
2timeout = RTT + 0.9 × RTT
3timeout ≈ 1.9 × RTTIn practice, networks under load can show jitter much higher than 30%. A multiplier of 3 on the raw RTT (without subtracting jitter) is a simple safe default:
1timeout = 3 × RTTThis covers most cases where jitter is less than 100% of RTT. If your network has bursty congestion or wireless hops, raise the multiplier to 4 or 5.
How to measure RTT correctly
Do not use ping from a laptop during a quiet evening. Measure under production load, from the actual host that will send heartbeats, to the actual host that will receive them.
- Collect samples. Use
tcpdumporperfto record 100-200 heartbeat packets over a 10-minute window that includes peak traffic. - Compute statistics. Calculate the mean (
RTT) and standard deviation (J) of the samples. - Check the distribution. If more than 5% of samples exceed
RTT + 3 × J, your network has heavy tails; increasek. - Repeat on failover. Measure again after a primary node fails, because backup paths often have different latency.
Practical tuning table
| Network type | Typical RTT | Suggested multiplier | Timeout | |--------------|-------------|----------------------|---------| | Same rack, 1 GbE | 0.1-0.5 ms | 2× | 0.2-1.0 ms | | Same data centre, 10 GbE | 0.5-2 ms | 3× | 1.5-6 ms | | Cross-region, internet | 20-80 ms | 3-4× | 60-320 ms | | Satellite or mobile | 200-600 ms | 5-6× | 1000-3600 ms |
These are starting points. Always validate with your own measurements.
What happens if you ignore this rule
When the timeout is too low:
- Spurious leader election. Raft and similar algorithms start a new election, which resets commit indices and can lose committed entries if not handled carefully.
- Session invalidation. In key-value stores, a false failure moves partitions to another node, causing cache stampedes.
- Increased latency for clients. Clients are redirected to the new leader while the old one is still serving reads, leading to inconsistent results.
When the timeout is too high:
- Slow failure detection. A real crash takes
timeoutmilliseconds to notice, during which clients see errors or time out. - Resource leaks. Files, locks, or connections held by the dead node are not released until the detector fires.
Adaptive approaches
If jitter is unpredictable, use adaptive timers:
- Exponential backoff. Double the timeout after each missed heartbeat, up to a cap.
- EWMA filter. Update
RTTandJwith each new sample using an exponentially weighted moving average. Recalculate the timeout every interval. - Delta encoding. Send only the change in RTT, not the absolute value, to save bandwidth.
These methods add complexity but remove the need to guess a static multiplier.
Checklist before you change production
- Measure RTT and jitter on the actual heartbeat path.
- Set initial timeout to
3 × RTT. - Enable logging of missed heartbeats and false-positive events.
- Run a chaos test: kill a node for 30 seconds and observe recovery time.
- If false positives occur, increase the multiplier by 0.5 and repeat.
- Document the final timeout and the measurement date so the next engineer can re-validate later.
Bottom line
A heartbeat timeout is not a magic constant. It is a statistical guard against network jitter. Measure your real RTT, estimate jitter, and set the timeout to RTT + k × J with k between 2 and 4. The 3× rule is a safe default for most data-centre networks, but always verify with your own data. False positives are worse than slow detection, because they break consistency and force unnecessary recovery.