Collective Communications: The Operations Behind Every AI Network Flow

This post, part of a series on AI data center networking, clarifies the traffic patterns arising from collective communications during GPU operations. It highlights key operations like broadcast, scatter, gather, and reduce, explaining their significance for both training and inference, especially in Mixture-of-Experts models. Understanding these patterns aids in optimizing network performance.

The post Collective Communications: The Operations Behind Every AI Network Flow appeared first on /overlaid.

NAN132: The AI-Augmented Engineer

Garrett Masters talks with Eric Chou about the concept of the AI-Augmented Engineer. Together they discuss how IT professionals can use Python and NetMiko to automate tasks and even connect AI directly to lab environments. Garrett also shares practical advice on prompt engineering, the importance of context management, and how engineers can safely adopt AI... Read more »

TCG085: Enterprise Salescraft with Drew Meli

William Collins and guest Drew Meli explore the realities of buying enterprise SaaS and infrastructure. They cover the common engineering misconceptions about enterprise sales, navigating rigorous security reviews, establishing clear proof of value criteria, and how AI is reshaping buying behaviors. AdSpot Sponsor: Gartner AI is rewriting your cloud architecture, your operating model, and your... Read more »

Managing Claude Code Sessions Through Lynx

At Tigera, we spend a lot of time thinking about agent security: identity, policy, runtime controls, and the record left behind after an agent acts.

Coding agents create an interesting problem because, in most organizations, they didn’t arrive through the front door.

Few companies ran a platform evaluation and rolled Claude Code out to 500 developers. Developers installed it themselves. By the time security and platform teams started asking how coding agents should be governed, they were already running on laptops with access to source code, credentials, SSH keys, kubeconfigs, internal services, and whatever else the developer could reach.

The long-term answer is increasingly clear I think: move coding agents into isolated environments you control.

Anthropic’s sandboxing work draws filesystem and network boundaries using OS primitives such as bubblewrap and seatbelt. Its reference devcontainer includes an egress firewall. Docker has introduced sandboxes for running coding agents, and Kubernetes-based approaches can add stronger workload isolation, network policy, and disposable development environments.

That direction makes sense.

Isolation governs what an agent can do. A gateway governs what it can send.

And unlike a complete move to remote development environments, the second boundary is something you can introduce today.

Start with one environment Continue reading

HS144: Next-Generation Data Centers: The Rules Are Changing, How Should Your Strategy Change?

In-house data centers are a hot topic again! Companies of all sorts are repatriating cloud workloads. Many want to ramp up in-house AI development, training, and production deployment. However much (or little) you know about building a data center, the rules have all changed in the last few years. From political/regulatory concerns to economics to... Read more »

BGP EVPN in SONiC – Part II: BGP EVPN – Inter-Subnet Forwarding

Introduction

 

Though inter-subnet communication in a BGP EVPN-based network between hosts happens through the anycast default gateway of the subnet, it can use two different forwarding models: Asymmetric and Symmetric Integrated Routing and Bridging (IRB). The basic difference between these two IRB models is how the packet is forwarded between the ingress and egress VTEPs.

When an ingress VTEP receives a frame whose destination is in a different subnet from the source, the initial forwarding operations are common to both IRB models. The VTEP performs a MAC lookup in the local VLAN's MAC address table and resolves the anycast gateway (AGW) MAC address. The packet is then routed toward the tenant IP-VRF, where an IP lookup determines the route toward the destination IP address. The routing information identifies the remote VTEP as the overlay next hop. The remote VTEP's underlay IP address is then resolved for use as the outer VXLAN destination IP address. This resolution can involve recursive lookup through the underlay routing table.

The forwarding paths diverge after the ingress VTEP performs the IP lookup. In asymmetric IRB, the ingress VTEP performs an additional MAC lookup in the bridge table associated with the destination subnet. The resulting Continue reading

Prefix Sets: Simplifying netlab ACLs and Prefix Filters

When I implemented prefix filters in netlab (release 1.9.0), I wanted them to match the final device configurations as closely as possible, including the one prefix per entry rule to match the device configuration sequence numbers. That worked well, but resulted in YAML bloat, since each matching prefix requires several lines of YAML.

We tried to use the same approach with ACLs but quickly gave up when we realized the port not in range operation requires multiple ACL entries on most platforms. With the matching sequence numbers paradigm gone with the wind, it made no sense to limit ourselves to one prefix per ACL entry – the initial ACL entry implementation already accepted lists of prefixes and expanded them into ACL entries.

QO-100 GPS locked

Previously I successfully did some receiving and some transmitting over the amateur radio transponder on the satellite QO-100. That is to say, my FT8 was heard by others (but not me, due to my small dish), and I successfully decoded FT8 from some others.

There are three problems with my QO-100 ground station:

  1. It’s not permanent. It’d need some weather sealing, both for the dish and other active equipment near it.
  2. The dish is too small. It works well enough to prove that the satellite is there and everything else works, but a permanent install should have a bigger dish.
  3. The frequency stability was not good enough for decoding everything. In this screenshot we can see all signals skewing the same way, which means it’s my receiver that’s skewing.

I really appreciate the constant FT8 activity when tuning my station. Sure, I could look at the band edge beacons, but I like this too. And strangely, I can not actually find the edge beacons.

Today is about the third issue: Frequency stability. Because the B200 was GPS locked last time, the main suspect was the LNB. The Othernet Bullseye LNB TCXO from Passion-Radio.com is a fine LNB, but it’s Continue reading

HN843: From Network Engineer to Network Architect with Kevin Nanns

Drew and Ethan talk with Kevin Nanns about his journey from network engineer to network architect. Kevin shares his experience grappling with the shift in responsibilities, increased pressure, and communication and management challenges associated with his new role. They also discuss his content creation work on TikTok and his strategies for translating complex technical concepts... Read more »
1 2 3 … 3,906