BFD: Bidirectional Forwarding Detection
Status: the plugin is live, the production transport is hardened
(GTSM, IP_TTL=255 outbound, SO_BINDTODEVICE for single-hop and multi-VRF,
RFC 5880 §6.8.7 TX jitter), the BGP peer opt-in is wired through the
reactor lifecycle, and the Stage 4 operator surface (show bfd sessions, show bfd session <peer>, show bfd profile, Prometheus
ze_bfd_* metrics) is available. Adding bfd { ... } inside a
bgp peer connection block opens a per-peer BFD session on
Established and tears the BGP session down with RFC 9384 Cease
subcode 10 ("BFD Down") when BFD reports the forwarding path lost.
Adding strict true to that block runs
draft-ietf-idr-bgp-bfd-strict-mode instead: capability 74 is
negotiated, the BFD session opens before the BGP FSM starts, and the
BGP session is held out of Established until BFD is Up.
BFD (RFC 5880) is a low-overhead liveness protocol that detects forwarding failures between a pair of systems in tens to hundreds of milliseconds, one or two orders of magnitude faster than BGP keepalives or OSPF hellos. In ze, BFD runs as a plugin that exposes a session service to other components: BGP, OSPF (when added), and static-route monitors call it to track peer liveness, and tear down their own adjacencies when BFD reports a session down.
When to use it
| Scenario | BFD value |
|---|---|
| eBGP transit session | Sub-second failure detection without waiting for BGP hold time (default 90 s) |
| iBGP over a loopback | Multi-hop BFD can accelerate route reconvergence if the IGP alone is too slow |
| Static route to a gateway | Withdraw the route when the gateway stops forwarding, even if the physical link is up |
| OSPF adjacency (when ze has OSPF) | Collapse the hello-based timeout from 40 s to 150 ms |
Do not use BFD between two routers where the IGP already runs fast BFD; the redundant sessions double the packet rate and can amplify instability.
Basic configuration
Profiles
Profiles are named bundles of timer and feature parameters. Sessions reference a profile by name and inherit every leaf from it. Profiles are optional: sessions can carry their parameters inline.
bfd {
profile fast-link {
desired-min-tx-us 50000; # 50 ms
required-min-rx-us 50000; # 50 ms
detect-multiplier 3; # 150 ms detection time
}
profile lan-default {
desired-min-tx-us 300000; # 300 ms
required-min-rx-us 300000;
detect-multiplier 3;
}
profile ibgp-loopback {
desired-min-tx-us 300000;
required-min-rx-us 300000;
detect-multiplier 5;
}
}
All interval fields are in microseconds because RFC 5880 expresses
every interval in microseconds. A future release may add -ms aliases.
| Leaf | Range | Default | Meaning |
|---|---|---|---|
desired-min-tx-us |
1 – 4 294 967 295 | 300 000 | Local target transmit rate. RFC §6.8.3 enforces a 1 000 000 µs floor while the session is not Up. |
required-min-rx-us |
0 – 4 294 967 295 | 300 000 | Minimum inter-packet gap the local end can handle. Zero means "do not send me periodic control packets." |
detect-multiplier |
1 – 255 | 3 | Number of consecutive missed packets that trigger a Down transition. |
passive |
boolean | false | Active sessions transmit from creation. Passive waits for the peer's first packet. See RFC 5883 §4.3 for the unidirectional-link use case. |
auth { type key-id secret } |
see below | none | RFC 5880 §6.7 cryptographic authentication (Stage 5). |
echo { desired-min-echo-tx-us } |
see below | none | RFC 5880 §6.4 Echo mode (single-hop only). |
Echo mode
Profiles may carry an echo { desired-min-echo-tx-us N } block
that opts sessions into RFC 5880 §6.4 Echo mode. The block is
valid only on single-hop sessions -- RFC 5883 §4 prohibits
multi-hop echo, and the parser rejects the combination with a
descriptive error.
bfd {
profile fast-echo {
desired-min-tx-us 100000;
required-min-rx-us 100000;
echo {
desired-min-echo-tx-us 50000; # 50 ms
}
}
single-hop-session 203.0.113.9 {
profile fast-echo;
}
}
When a session inherits this profile, every outgoing Control
packet sets RequiredMinEchoRxInterval to the configured rate.
Peers that see a non-zero advertisement learn that the local end
is willing to reflect echo packets.
Current coverage: echo mode is fully implemented including the
YANG surface, wire advertisement, per-session TX scheduler, RX
demux with RTT measurement, echo detection timer, and async
slow-down (RFC 5880 §6.8.9). The ze_bfd_echo_tx_packets_total
and ze_bfd_echo_rx_packets_total Prometheus metrics are live.
Authentication
A profile can carry an auth { ... } block that enables RFC 5880 §6.7
authentication for every session that inherits the profile. Ze supports
the five types RFC 5880 defines: simple-password, keyed-md5,
meticulous-keyed-md5, keyed-sha1, and meticulous-keyed-sha1. There
is no default, so authentication runs only on the type you name.
simple-password (RFC 5880 §6.7.2) puts the password in clear in every
Control packet, and the authentication section is the same in each one.
It detects a misconfigured neighbor. It does not stop a listener on the
path, who reads the password from one packet, and it carries no Sequence
Number, so a captured packet can be replayed. Use a keyed type wherever
the path is not already trusted. The password must be 1 to 16 bytes: the
Auth Len field holds the password length plus three, so a longer password
has no wire encoding and the parser refuses it.
bfd {
persist-dir "/var/lib/ze/bfd";
profile authenticated-fast {
desired-min-tx-us 50000;
required-min-rx-us 50000;
detect-multiplier 3;
auth {
type meticulous-keyed-sha1;
key-id 7;
secret "BFD-SHARED-SECRET-V1";
}
}
}
The meticulous variants enforce strict monotonic sequence numbers;
non-meticulous variants allow a receiver to accept equal sequence
numbers across retransmits. The persist-dir leaf (top-level on
bfd { }) names a directory where ze stores the last TX sequence
per session so a Meticulous session resumes above the peer's replay
window after a process restart. Without persist-dir, Meticulous
sessions still work at runtime but briefly re-synchronize after a
restart while the peer's replay window slides forward.
Authentication failures increment ze_bfd_auth_failures_total{mode}.
Enabling BFD on a BGP peer
The common path. Adding a bfd { ... } container inside a BGP peer's
connection { ... } block opts that peer into BFD. On session
Established, the reactor calls the BFD plugin's api.Service.EnsureSession
with the peer's local / remote address and the configured mode; on
session teardown the handle is released. When BFD reports the
forwarding path Down, the reactor tears the BGP session with RFC 9384
Cease NOTIFICATION (subcode 10, "BFD Down") without waiting for the
hold timer.
bgp {
peer peer1 {
connection {
local {
ip 192.0.2.1
}
remote {
ip 192.0.2.2
}
bfd {
enabled true
mode single-hop
profile fast-link
}
}
session {
asn {
local 65001
remote 65002
}
router-id 192.0.2.1
family {
ipv4/unicast
}
}
}
}
With no bfd container, BGP behaves exactly as today. The container
has YANG presence: its mere existence opts in, and enabled false
suspends the opt-in without removing the config (useful during
maintenance). The profile name references a profile defined under the
top-level bfd { profile ... } block; the BFD plugin resolves it when
it receives EnsureSession. If the BFD plugin is not loaded at all,
the BGP peer starts without BFD and logs a warning -- BGP is not
blocked by a missing BFD plugin.
Multi-hop
For iBGP between loopbacks or any peering that crosses more than one IP hop:
bgp {
peer loop4 {
connection {
local {
ip 10.255.255.1
}
remote {
ip 10.255.255.4
}
bfd {
enabled true
mode multi-hop
min-ttl 250
profile ibgp-loopback
}
}
session {
asn {
local 65000
remote 65000
}
router-id 10.255.255.1
family {
ipv4/unicast
}
}
}
}
mode multi-hop tells the plugin to use UDP port 4784 and skip GTSM.
min-ttl is a weaker replacement for GTSM. Packets arriving with a
TTL below the configured value are discarded. Choose min-ttl to be
256 - max-hops; for example, min-ttl 250 allows the packet to
cross up to 5 hops. min-ttl must be non-zero for multi-hop; the
parser rejects zero at config-validate time.
Strict mode
Plain BFD is a failure detector. It opens after the BGP session is up, so a peer whose control plane answers while its forwarding path is broken still establishes BGP and blackholes traffic until something else notices.
Strict mode closes that window. It is
draft-ietf-idr-bgp-bfd-strict-mode, and it holds the BGP session out of
Established until the BFD session to that neighbour is Up.
bgp {
peer peer1 {
connection {
local { ip 192.0.2.1 }
remote { ip 192.0.2.2 }
bfd {
enabled true
mode single-hop
profile fast-link
strict true
hold-time 30
}
}
}
}
strict true does three things.
| What | Detail |
|---|---|
| Advertises capability 74 | The BFD Strict-Mode Capability, code 74, length 0, in every OPEN. Draft Section 6 makes this a MUST for a speaker with strict mode enabled |
| Opens the BFD session early | Before the BGP FSM starts, and keeps it open while BGP is down. Draft Section 7 asks for both, so the BFD handshake is not serialized behind the BGP one |
| Holds the session in OpenSent | When the peer's OPEN arrives and BFD is not yet Up, Ze withholds its KEEPALIVE and waits. show bgp peer list and show bgp peer detail carry a bfd-sub-state field while it does |
A fourth leaf, hold-down, is optional and described below.
Both speakers must advertise it. Against a peer whose OPEN carries no capability 74, Ze establishes on the normal path and BFD stays a failure detector. That is deliberate: a one-sided strict mode is a session that never comes up.
hold-time bounds the wait, and only when the negotiated BGP hold time is
zero. A non-zero BGP hold time already bounds it, so the timer does not run.
On expiry Ze sends a NOTIFICATION with Cease and subcode 10 ("BFD Down"), and
the peer returns to Idle and retries. The default is 30 seconds, which is the
draft's own.
Nothing waits forever. Where the negotiated BGP hold time is non-zero, the ordinary RFC 4271 hold timer is what ends a wait for a BFD session that never comes up: Ze arms it at the negotiated value on the way into the wait, and its expiry sends NOTIFICATION code 4 (Hold Timer Expired) and returns the peer to Idle. That is why the draft arms its own BfdHoldTimer only for a negotiated hold time of zero.
hold-down damps a flapping link. It is how long the BFD session must stay Up
before Ze lets the BGP session establish, in milliseconds, and it is zero by
default: without it Ze establishes on the first BFD Up, which is what a peer
did before the leaf existed. A BFD session that goes down again inside the
interval never completes it, so a link that flaps carries no BGP session.
draft-ietf-idr-bgp-bfd-strict-mode Section 10 recommends the BFD hold-down and
BGP hold time "use similar values"; set the BGP hold time long enough to cover
the interval, or the far end times out waiting for you.
Two operational notes, both from the draft.
- Give the BGP hold time room (Section 10). It has to cover the time between OpenConfirm, the BFD hold-down interval, and the delay before the BFD session starts. Too short and the speaker that reached Established first times out waiting for the other to follow.
- Authenticate BFD (Section 12). A BGP session now depends on a BFD
session, so anything that can stop BFD coming Up can stop BGP. The
authblock above is the answer.
The BFD plugin must be loaded. A strict peer with no BFD plugin is held down rather than run without the check: establishing would deliver the opposite of what was configured. Ze logs an error naming the peer.
Standalone sessions
Some paths have no protocol client. For those, pin the session explicitly
in the top-level bfd block. It exists for the lifetime of the config.
bfd {
profile gateway {
desired-min-tx-us 200000;
required-min-rx-us 200000;
detect-multiplier 3;
}
single-hop-session 192.0.2.254 {
local 192.0.2.1;
interface eth0;
profile gateway;
}
multi-hop-session 198.51.100.7 {
local 10.0.0.1;
profile gateway;
min-ttl 254;
}
}
Standalone sessions are useful for monitoring a default gateway, a tunnel endpoint, or any path whose liveness a script or external tool wants to react to via the event stream.
Administrative shutdown
shutdown; on a session puts it into AdminDown (RFC 5880 §6.8.16)
without removing the config. The peer sees AdminDown via the State field
and, per RFC 5882, clients with their own liveness signal keep running
while clients without one treat it as Down.
bfd {
single-hop-session 192.0.2.254 {
local 192.0.2.1;
interface eth0;
profile gateway;
shutdown;
}
}
Remove shutdown; (or set it to false in a tool-driven edit) to
re-enable the session; it transitions through Down and completes the
three-way handshake normally.
Session sharing
When multiple clients ask for the same path, the BFD plugin creates one underlying session and refcounts subscribers. Timer parameters are chosen as the most aggressive (smallest) value across requesters. For example, if BGP asks for a 50 ms session and OSPF later asks for a 300 ms session to the same peer, they share one 50 ms session; if the BGP subscriber goes away first, the session drops to 300 ms via Poll/Final.
Editing live
BFD configuration is reachable through the standard ze config edit CLI:
operator@router> configure
operator@router# edit bfd profile fast-link
operator@router# set desired-min-tx-us 100000
operator@router# show
profile fast-link {
desired-min-tx-us 100000;
required-min-rx-us 50000;
detect-multiplier 3;
}
operator@router# commit
On commit, the plugin receives a diff section via OnConfigure and
reconfigures affected sessions in place. Profile changes propagate to
every session that references the profile via an in-band Poll/Final
sequence without a session flap.
Observing state
Stage 4 adds three operator-facing commands served by a snapshot of
the engine's live session state. The handlers live in
internal/component/bfd/cmd/bfd.go and publish JSON payloads so
scripts can parse the output while the interactive CLI renders them.
Let BFD protect a live BGP session
Establish BFD and BGP with a local FRR peer, cut the peer link, and verify BFD drives BGP down before protocol timers expire.
Read the demonstration transcript
An operator needs to verify that BFD, not the 300-second BGP hold timer, protects an edge session.
$ ze config show demos/terminal/bfd-failover/ze.conf bfd
$ ze config show demos/terminal/bfd-failover/ze.conf bgp peer edge-peer connection
$ ze config show demos/terminal/bfd-failover/ze.conf bgp peer edge-peer timer
The daemon configuration shows the 300 ms BFD profile, multiplier 3, single-hop binding, and 300-second BGP hold time.
$ ze cli -c 'show bfd sessions'
The running control plane shows the complete Up BFD session.
$ date -u +%T; ip link set bfd-p down
$ ze cli -c 'show bfd sessions'
$ ze cli -c 'show bgp peer list'
Five seconds after the kernel link is cut, the full command output shows no live BFD session and BGP has left Established.
$ ip link set bfd-p up
$ ze cli -c 'show bgp peer list'
The same peer returns to Established after the link is restored.
Every protocol result comes directly from `ze cli`; the lab helper is used only to create and reset the isolated FRR peer.
List all sessions
operator@router> show bfd sessions
[
{"peer":"192.0.2.2","vrf":"default","mode":"single-hop","state":"up","diag":"no-diagnostic","local-discriminator":1,"remote-discriminator":2147518038,"tx-interval":50000000,"rx-interval":50000000,"detection-interval":150000000,"detect-multiplier":3,"profile":"fast-link","tx-packets":2312,"rx-packets":2310,"refcount":1,...},
{"peer":"198.51.100.7","vrf":"default","mode":"multi-hop","state":"down","diag":"control-detection-time-expired","tx-interval":200000000,"rx-interval":200000000,"detection-interval":600000000,"detect-multiplier":3,"profile":"gateway",...}
]
Fields are sorted stably by (mode, vrf, peer) so successive scrapes
produce diff-able output.
Session detail
show bfd session <peer> returns the same struct for one session
plus the most recent state transitions kept in a bounded in-memory
ring (eight entries by default -- see api.TransitionHistoryDepth):
operator@router> show bfd session 192.0.2.2
{
"peer": "192.0.2.2",
"local": "192.0.2.1",
"interface": "eth0",
"vrf": "default",
"mode": "single-hop",
"state": "up",
"diag": "no-diagnostic",
"profile": "fast-link",
"local-discriminator": 1,
"remote-discriminator": 2147518038,
"tx-interval": 50000000,
"rx-interval": 50000000,
"detection-interval": 150000000,
"detect-multiplier": 3,
"tx-packets": 2312,
"rx-packets": 2310,
"refcount": 1,
"transitions": [
{"when":"2026-04-11T09:14:22Z","from":"down","to":"init","diag":"no-diagnostic"},
{"when":"2026-04-11T09:14:23Z","from":"init","to":"up","diag":"no-diagnostic"}
]
}
List profiles
show bfd profile returns every resolved (post-default) profile
stored by the plugin. Passing a name filters to one entry; an unknown
name returns an error:
operator@router> show bfd profile fast-link
{"name":"fast-link","detect-multiplier":3,"desired-min-tx-us":50000,"required-min-rx-us":50000,"passive":false}
Prometheus metrics
The plugin registers five metric families via
internal/component/bfd/metrics.go. The families appear on the
telemetry endpoint when telemetry { prometheus { enabled true } }
is set.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
ze_bfd_sessions |
gauge | state, mode, vrf | Live session count, updated on every show bfd sessions scrape |
ze_bfd_transitions_total |
counter | from, to, diag, mode | Every session state change |
ze_bfd_detection_expired_total |
counter | mode | Detection-timer expirations (RFC 5880 §6.8.4) |
ze_bfd_tx_packets_total |
counter | mode | Control packets transmitted |
ze_bfd_rx_packets_total |
counter | mode | Control packets received (after the TTL gate) |
Operational guidance
Picking timers
| Situation | TX / RX | Mult | Detection |
|---|---|---|---|
| LAN between routers | 50 ms / 50 ms | 3 | 150 ms |
| Metro eBGP | 100 ms / 100 ms | 3 | 300 ms |
| iBGP over IGP | 300 ms / 300 ms | 5 | 1.5 s |
| WAN over uncertain media | 500 ms / 500 ms | 5 | 2.5 s |
Err on the slow side. A session that runs too fast on a link with jitter produces false-positive failures that tear down routing state. BFD is worse than useless in that mode: it amplifies instability instead of damping it.
TTL security on single-hop
Single-hop BFD enforces GTSM (RFC 5082): packets are sent with TTL=255
(IP_TTL socket option applied at Start) and received packets with any
other TTL are silently discarded by the engine. There is nothing for the
operator to configure -- ze sets the TTL automatically and the engine's
TTL gate fails closed when the kernel cannot extract the received TTL
(for example, when running on a platform without IP_RECVTTL).
Multi-hop sessions have no GTSM equivalent; instead each multi-hop-session
carries a min-ttl leaf (default 254). Packets arriving with a TTL below
that floor are discarded before reaching the FSM. Pick min-ttl to be
256 - max-hops.
Multi-VRF and interface binding
Single-hop pinned sessions may specify an interface leaf. When every
pinned session in the same (vrf, single-hop) pair names the same
interface, the plugin binds the socket to that interface via
SO_BINDTODEVICE so egress and ingress are guaranteed to traverse the
named device. If a pinned set mixes interfaces in the same VRF, the
bind-to-device fallback is disabled (a warning is logged) and the
engine-side TTL gate remains the sole protection.
Non-default VRFs bind the socket to the VRF device name (Linux's
SO_BINDTODEVICE accepts VRF masters as device names). Both paths
require CAP_NET_RAW at daemon startup on Linux.
SO_BINDTODEVICE can only name one device per socket, so under a
non-default VRF the session interface leaf is ignored -- the socket
binds to the VRF master and receives packets from every slave device
in that VRF. Ze logs an Info line naming the dropped interface leaves
whenever this override fires so the operator can correlate a reload
with the behaviour change. If you need per-interface pinning in a
non-default VRF, stand up separate sessions per slave device outside
the shared VRF, or wait for the Stage 3 BGP peer opt-in which can
drive one session per peer.
Transmit jitter
RFC 5880 §6.8.7 requires the TX interval to be reduced by 0-25% on each
packet, and the reduction must be at least 10% when detect-multiplier
is 1 so the receiver cannot detect before the next packet arrives. Ze
implements both bands via engine.Loop.applyJitter; operators do not
configure it.
GC pause sensitivity
At 50 ms intervals with mult=3, a 150 ms GC pause looks indistinguishable
from a real failure. Ze's BFD plugin uses pool-backed buffers and runs
every session on a dedicated goroutine (the "express loop") to minimise GC
pressure on the session-driving thread. On a heavily loaded ze instance,
watch for detection-timer expirations that coincide with high allocation
rates elsewhere in the daemon. The metric surface (coming with the wiring
commit) will expose bfd_control_detection_time_expired_total per session
so the correlation is visible in Prometheus.
Interop
ze's wire format matches FRR bfdd, BIRD 3.x, and Junos. Interop tests
against FRR are the primary validation; test/plugin/bfd/ will carry a
namespace-based scenario once the functional tests land.
What is not yet supported
| Feature | RFC | Status |
|---|---|---|
| Demand mode | 5880 §6.6 | Not implemented; rarely deployed. |
| Seamless BFD (S-BFD) | 7880, 7881 | Not implemented. |
| Micro-BFD on LAG | 7130 | Not implemented. |
| Multipoint BFD | 8562 | Not implemented. |
| Data-plane offload | FRR bfddp_packet.h |
Future work. |
The skeleton is intentionally minimum-viable. See the architecture document for the full gap list and the intended follow-up order.
Reference
- Protocol details:
rfc/short/rfc5880.md,rfc5881.md,rfc5882.md,rfc5883.md - Internal architecture:
docs/architecture/bfd.md - Implementation research:
docs/research/bfd-implementation-guide.md