BGP EVPN in SONiC – Part II: BGP EVPN – Inter-Subnet Forwarding

Introduction

 

Though inter-subnet communication in a BGP EVPN-based network between hosts happens through the anycast default gateway of the subnet, it can use two different forwarding models: Asymmetric and Symmetric Integrated Routing and Bridging (IRB). The basic difference between these two IRB models is how the packet is forwarded between the ingress and egress VTEPs.

When an ingress VTEP receives a frame whose destination is in a different subnet from the source, the initial forwarding operations are common to both IRB models. The VTEP performs a MAC lookup in the local VLAN's MAC address table and resolves the anycast gateway (AGW) MAC address. The packet is then routed toward the tenant IP-VRF, where an IP lookup determines the route toward the destination IP address. The routing information identifies the remote VTEP as the overlay next hop. The remote VTEP's underlay IP address is then resolved for use as the outer VXLAN destination IP address. This resolution can involve recursive lookup through the underlay routing table.

The forwarding paths diverge after the ingress VTEP performs the IP lookup. In asymmetric IRB, the ingress VTEP performs an additional MAC lookup in the bridge table associated with the destination subnet. The resulting Continue reading

BGP EVPN in SONiC – Part I: L2VNI for Inter-Switch VLAN Extension

 

Introduction

 

The previous chapter explained how to build an IPv6-only multi-ASN Clos fabric where IPv4 prefixes can be advertised with IPv6 link-local next hops.

In this chapter, we extend VLAN 10 between Leaf-101 and Leaf-102 using BGP EVPN as the control plane and VXLAN as the data plane.

Figure 7-1 shows the main building blocks required to provide this functionality. The IP underlay, where leaf and spine switches use IPv6 link-local addresses for eBGP peering, has already been configured in the previous chapter. Loopback0 is used only as the BGP Router ID and is not advertised to any peer.

The IP address of Loopback1, in turn, is used as the VXLAN Tunnel Endpoint (VTEP) address. Each switch therefore advertises its Loopback1 IP address using BGP IPv4 Unicast so that the VTEP addresses are reachable across the fabric. This address identifies the local VTEP in the EVPN control plane and is used as the BGP next hop for EVPN routes. A VXLAN packet sent to a remote VTEP uses the remote VTEP address as the destination address in the outer IP header.

Each extended VLAN is mapped to an L2VNI (Layer-2 VXLAN Network Identifier). In Figure 7-1, VLAN 10 Continue reading

SONiC Deep Dive: BGP Update Message Processing

 

BGP Route Advertisement

The next step after configuring BGP peering using BGP Unnumbered with IPv6 link-local addresses is to advertise the host networks. The lower part of Figure 6-10 shows the sonic-cli configuration commands used to advertise the local VLAN subnet on each leaf. On Leaf-101, we advertise 10.0.10.0/24, which is the subnet for VLAN 10. On Leaf-102, we advertise 10.0.20.0/24, which is the subnet for VLAN 20.

By examining the BGP Loc-RIB on Leaf-102, for example, we can see that the next hop for 10.0.10.0/24, originated on Leaf-101, is the IPv6 link-local address of Spine-11's Ethernet0 interface.

Figure 6-10: IPv4 Routes Installed in the BGP Loc-RIB Table.

Figure 6-11 verifies that the host subnets learned through BGP have been installed in the IP routing table on all switches. For example, the destination prefix 10.0.10.0/24 is installed on Leaf-102 with the next hop fe80::e22:34ff:feb6:a and the outgoing interface Ethernet0.

This also illustrates the separation between the BGP control plane and the data plane. BGP uses IPv6 link-local addresses to establish the BGP session and exchange BGP messages carrying IPv4 NLRI, with an IPv6 link-local address specified as Continue reading

SONiC Deep Dive: BGP Peer Configuration – BGP Unnumbered

 

After enabling IPv6 on the Ethernet0 interface, we can configure BGP peering. Instead of statically defining a peer IP address and AS number, the BGP neighbor is defined at the interface level. This tells BGP to expect a neighbor and initialize the BGP peering over Ethernet0. Because our intent is to transport IPv4 traffic over the IPv6 network, we activate both the IPv4 and IPv6 address families. We also configure the peer as an external BGP (eBGP) neighbor by using a command remote-as external. This tells BGP that the peer must use an AS number different from the local AS number.

Because the network uses IPv6 as the transport for IPv4 traffic, IPv4 Network Layer Reachability Information (NLRI) exchanged through BGP use IPv6 next-hop addresses. BGP therefore needs the Extended Next Hop Encoding capability to support IPv4 routes with IPv6 next hops. In this configuration, the capability is explicitly enabled with the capability extended-nexthop command.

In FRRouting, the v6only option controls which IP version is used for interface-based BGP Unnumbered peering. Without v6only, FRR uses the interface's IPv6 link-local address for peering only when no suitable IPv4 address is configured on the interface. With v6only, FRR skips Continue reading

Introducing Adaptive Intelligence: Undermining the economics of every bot attack

Modern bot threats are increasingly driven by determined, sophisticated attackers. Often it is not even one person, but a group trading techniques with each other or a commercial service sold to anyone willing to pay. For many of them, getting past bot detection is a full-time job they genuinely enjoy. Block them and they get to work, finding a workaround. AI has simplified this further, making it even easier to set up complex configurations for attackers, lowering the overhead of an attack. 

This shift puts defenders at an economic disadvantage. Responding and adapting to new attacks takes care, evidence, and effort to ensure efforts to block attackers don’t impact real users on the way. Attackers have no such concerns and are primarily constrained by their time and their pool of proxies, and ensuring their infrastructure providers don’t shut down their accounts.

Their advantage is the cost of adaptation. Attackers can adapt as often and continuously as they need, while most defenses are deployed in discrete, managed releases. Cloudflare analyzes more than a trillion requests a day for signs of automated abuse, so we see how fast attackers change tactics. That gap in responsiveness is widening.

The inconvenient truth: bot Continue reading

Why Is My Switch Still Red? The Future of Network Switches

Rethinking Network Switches The rate of technological change today can seem overwhelming but if you are old enough, you start to realize its pretty much more of the same. My biggest fear with tech change is that I don't want it to happen TO me. I may not be in a position to be in READ MORE

The post Why Is My Switch Still Red? The Future of Network Switches appeared first on The Gratuitous Arp.

BotBase for Operators: A clearer path to joining Cloudflare’s directory of bots and agents

Last month, on our second Content Independence Day, we announced a couple of features designed to give website owners more visibility and control over automated traffic: BotBase added a searchable directory of known bots to the Cloudflare dashboard, while Business Insights helped owners understand how crawlers interact with their content. We know that the ecosystem of bots is vast, making it all the more important for site owners to be able to manage bot traffic sustainably.

But this ecosystem goes both ways. While website owners need to decide which automated traffic they allow, bot operators need a clear way to identify themselves, explain what their bots do, and keep that information current. BotBase works best when both sides can participate.

When we launched BotBase, we said we would build tools to bring bot operators into this ecosystem. Until now, their experience largely ended at submission. After pressing submit, an operator had no easy way to check the submission's status, understand why it was rejected, or update an existing entry. Today, we start to change that with the launch of BotBase for Operators, tackling what bot operators need first: transparency.

A new home for bot submissions

Imagine you’re a bot operator Continue reading

The Safest Place to Run an AI Agent Is On a Cluster That Doesn’t Trust It

Every organization running AI agents has already made a hosting decision. Most made it by accident.

The sales team switched on the agent built into their CRM. Engineering is piloting a coding agent in a vendor’s cloud. Someone on the data team deployed a LangGraph service to a VM with a database key in an environment variable, and someone else is running an agent framework on a laptop with production credentials in a dotfile. Each of these is a hosting decision. Each one quietly settled who holds the agent’s credentials, what network paths it can reach, what gets recorded when it acts, and who can stop it. Nobody ran an architecture review, because no single deployment looked big enough to deserve one.

The scale says otherwise. By May 2025, 82% of organizations surveyed by SailPoint were already using AI agents. Only 44% had policies for securing them, 80% said their agents had already taken unintended actions, and 23% had watched an agent get tricked into revealing credentials. A year later the bill arrived: IBM’s 2026 Cost of a Data Breach report found that one in four malicious breaches is now AI-enabled, up 56% in a single year, and that those Continue reading

How we saved 100 terabytes of memory by optimizing 1.1.1.1’s DNS cache

Big Pineapple, the platform behind 1.1.1.1, Gateway DNS, DNS Firewall, AS112, and several other Cloudflare DNS services, stores over 250 billion DNS cache entries at any given time. At that scale, wasting a single byte per entry costs more than 250 gigabytes of memory across our fleet.

Five successive changes to how cache entries are stored in memory cut the per-entry footprint by over 50%. Across our fleet, these changes freed up roughly 100 terabytes of memory, equivalent to the amount of RAM in 130 of our Gen 13 servers. The cache also got faster. Insert throughput rose 43% and lookup latency dropped 19%, as fewer allocations and better memory locality meant we did not trade speed for space.

What we cache

On cold start, Big Pineapple starts out with an empty cache. As DNS queries arrive, the cache fills until it hits its maximum entry count, at which point we evict older or less popular items to make room.

The exact cache size varies by data center. When EDNS Client Subnet (ECS) is in use, authoritative servers return different answers depending on the client's network, so we cache multiple versions of Continue reading

Dear Junos, Tunnels Are Not Virtual Links

In late June, we added GRE tunnels to netlab, including a Junos implementation. It looked great (as in “everything worked”) until I changed the integration tests to have GRE tunnels between a tested device and a pair of FRR containers. All other implementations worked as before, but Junos failed to establish an OSPFv3 adjacency over the GRE tunnel with FRR.

Stefano Sasso quickly identified the culprit: Junos OSPFv3 process thinks it should send the DBD packets over GRE tunnels with MTU set to zero (the behavior reserved for virtual links)1.

Testing DDoS mitigation software

The recently released open source sflowgen tool is a synthetic sFlow generator intended for testing, demonstrations, dashboards, and attack-detection validation. This article uses sflowgen to test the DDoS Protect application.

The screen capture shows how sflowgen DDoS attack telemetry appears in the DDoS Protect dashboard. Charts in the dashboard show different types of DDoS attack. You can see normal low level background activity below the red horizontal threshold line. DDoS traffic can be seen ramping up on the udp_flood and ip_flood charts and leveling off once the attack reaches maximum intensity. Notice that as soon as the attack traffic crosses the threshold, the Controls chart indicates that a Pending control has been created to mitigate the attack.

The Controls table shows the udp_flood attack against host 10.10.0.42 using port 443 as the attack vector. In this case the control action is set to drop, i.e. use a Remotely Triggered Black Hole (RTBH) as the mitigation action. The pending status indicates that it is waiting for user confirmation before being applied.

The Settings tab has been configured to add a local address group containing the 10.10.0.0/24 and 2001:db8:10::/64 CIDRs used in the default Continue reading

1 2 3 … 3,900