Skip to content

daemon: skip kernel FIB sync for kernel/local-sourced best paths - #629

Open
jakeuibn wants to merge 1 commit into
osrg:masterfrom
jakeuibn:fix/kernel-fib-self-install-loop
Open

jakeuibn wants to merge 1 commit into
osrg:masterfrom
jakeuibn:fix/kernel-fib-self-install-loop

Conversation

@jakeuibn

Copy link
Copy Markdown

Problem

Enabling the kernel integration ([zebra.config] enabled = true) on a
topology with directly-connected eBGP peers made the daemon withdraw
routes it had just learned and forwarded. Reproduced with a minimal
three-node lab (sender → rustybgpd → receiver, all on-link, peers inside
the connected subnets of the daemon):

  1. Address monitoring injects each connected prefix into the BGP RIB as a
    kernel-sourced path (KernelEvent::Address → inject_kernel_route).
  2. TableShard::distribute_update unconditionally syncs the rank-1 best
    path into the kernel FIB — including kernel-sourced paths. The
    connected route is re-installed as
    10.0.0.0/24 via 10.0.0.1 dev eth1 proto bgp
    (via its own interface address) with .replace(), clobbering the
    kernel's connected entry
    (10.0.0.0/24 dev eth1 proto kernel scope link src 10.0.0.1).
  3. Installing a BGP-learned route whose nexthop lives in that subnet now
    fails: the gateway is resolved through the poisoned entry, and netlink
    returns Network unreachable:
    ERROR rustybgp_kernel] kernel route update failed: netlink: Received a netlink error message Network unreachable (os error 101)
  4. NHT independently marks the nexthop unreachable:
    Handle::lookup_route rejects fib-match routes with
    protocol == Bgp, and the poisoned entry is exactly that
    (RouteProtocol::Bgp). The daemon therefore withdraws the routes the
    peers announced through it — the downstream peer observes
    announce followed by withdraw for a route that never changed.
  5. When the learned route is later withdrawn, RTM_DELROUTE hits a
    non-existent entry and logs
    kernel route update failed: netlink: ... No such process (os error 3).

The same topology without the kernel integration (or with a v0.1-style
setup) is stable: the learned route stays installed and advertised.

Root cause

distribute_update treats every best-path change as BGP-owned work. For
kernel-sourced (connected) and local-sourced (gRPC/CLI injected) paths
the FIB write is at best redundant — the kernel already owns the route —
and in the connected case actively harmful, because the path's nexthop is
the local interface address and reinstalling it via .replace()
destroys the connected route that makes on-link nexthops resolvable.

Fix

  • Add fib_sync_required() and call it from distribute_update: skip
    the FIB sync when the new best path is kernel- or local-sourced. The
    skip is a no-op for the kernel (it already owns those routes) and does
    not affect BGP-side behavior: redistribution to peers, NHT registration
    (nht_register already skips these sources), and Loc-RIB visibility
    are unchanged. A prefix whose best path vanished (new_best() == None)
    still syncs, so a stale BGP-installed route is withdrawn when the
    connected route wins back the prefix.
  • Treat ESRCH from RTM_DELROUTE as success in
    rustybgp_kernel::Handle::withdraw: a withdraw for a never-installed
    or already-removed prefix is routine (e.g. the kernel removed its
    connected route first) and should not surface as an error log.

Verification

  • cargo test -p rustybgpd: 581 passed (7 new: fib_sync_required_*
    decision table for BGP / kernel / local / emptied / BGP-take-over /
    kernel-restore cases, plus a Loc-RIB integration test).

  • cargo clippy -p rustybgp-kernel -p rustybgpd and cargo fmt --check: clean.

  • End-to-end on a containerlab sender → rustybgpd → receiver eBGP lab
    with [zebra.config] enabled = true, before vs after:

    observation before after
    kernel connected routes replaced by via <own addr> proto bgp intact (proto kernel scope link)
    learned route in FIB install fails (Network unreachable) installed (<learned prefix> via <peer addr> dev ethX proto bgp)
    downstream view announce then withdraw announce only, stable
    daemon log 2× netlink errors clean
    gobgp global rib entries flap connected + learned entries stable

Enabling the kernel integration ([zebra.config] enabled = true) on a
topology with directly-connected eBGP peers made the daemon withdraw
routes it had just learned and forwarded. Sequence observed with two
on-link peers (sender -> rustybgpd -> receiver):

1. Address monitoring injects each connected prefix into the BGP RIB as
   a kernel-sourced path (Handle::Address -> inject_kernel_route).
2. distribute_update unconditionally syncs the rank-1 best into the FIB,
   so the connected route is re-installed as 'via <own address> ...
   proto bgp' with .replace(), clobbering the kernel's connected
   'dev ethX proto kernel scope link' entry.
3. Installing a BGP-learned route via that nexthop now fails with
   'Network unreachable' (gateway resolution goes through the poisoned
   entry), and NHT marks the nexthop unreachable: lookup_route() rejects
   fib-match routes with protocol == Bgp, and the poisoned entry is
   exactly that. The daemon then withdraws the routes peers announced
   through it (receiver observes announce followed by withdraw).

Fix: skip the FIB write when the new best path is kernel- or local-
sourced. The kernel already owns these routes (connected entries and
gRPC/CLI-injected static routes were never ours to install), so the
skip is a no-op there, while export to peers and NHT registration are
unchanged. A best path that vanished (new_best() == None) still syncs,
so stale BGP-installed routes are withdrawn. Additionally, treat ESRCH
from RTM_DELROUTE as success: a withdraw for a never-installed or
already-removed prefix is routine and should not surface as
'kernel route update failed: No such process'.

With the fix and [zebra.config] enabled, the same topology shows:
connected entries intact, the learned route installed in the FIB
(proto bgp), no netlink errors, and a stable announce on the downstream
peer.

Assisted-by: GLM (ZCode CLI)
Signed-off-by: qfr23 <qfr23@mails.tsinghua.edu.cn>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant