rvachev.orgEN / RU / 🤖
← Back to essays
· Essay · 2 min

AI Takes Jobs Not Only from People but Also from the TCP Protocol.

AI is changing familiar things even at the most basic level, and the TCP protocol has proven unprepared for new loads.

🌐 AI takes jobs not only from people but also from the TCP protocol.

Professor Stanford John Ousterhout (now emeritus) built an entire lecture on this idea: AI is changing familiar things even at the most basic level, and the protocol that has supported the entire internet since the 1970s has proven simply unprepared for new loads.

The fact is that training models involves massive data transfers where only bandwidth matters, and here TCP and RDMA handle it fine. However, inference and especially agentic workloads involve constant exchange of small messages: checking the KV-cache, synchronization between GPUs after a computation round. Here, it's not the average latency that matters but the tail latency (P99) - if even one such exchange out of a thousand slows down, all GPUs waiting for synchronization are idled. Previously, computations took seconds, and synchronization took milliseconds, so the difference was insignificant. Now, agentic workloads generate tokens at millisecond intervals, and synchronization of the same length consumes a significant portion of GPU time wastefully.

TCP struggles with this for two reasons. The first is incast: when several nodes simultaneously send data to one receiver, packets accumulate in the switch queue, and a short message gets stuck behind long ones. Congestion control in TCP is handled by the sender, although congestion occurs on the receiver's side - it takes several rounds through ECN marking for the sender to understand what's happening and adjust the speed, causing the system to constantly oscillate between "too fast" and "too slow." This problem has been known for over 20 years and remains unsolved. The second reason is that TCP has no concept of message boundaries; it's just a byte stream, so a short message cannot be prioritized ahead of long ones - it physically stands in line behind them.

Ousterhout and his research group at Stanford created Homa - a transport protocol from scratch for these workloads. Homa works with messages, not a byte stream, so it knows the message length from the first packet and prioritizes short ones through SRPT (shortest remaining processing time). Congestion control is given to the receiver - it sees the whole picture and sends "grant" packets, allowing senders to send the next piece of data. On top of this, Homa uses priority queues in switches so that short messages physically overtake long ones.

In benchmarks on 40 nodes, Homa provides tail latency (P99) for short messages 7-83 times lower than TCP and DCTCP, and even transmits long messages almost twice as fast. The implementation is a kernel module for Linux, available on GitHub, and Ousterhout is currently trying to push Homa upstream into the kernel.

https://youtu.be/eZ8WWZzoaR0
https://github.com/PlatformLab/HomaModule

#ai@rvnikita_blog #networking@rvnikita_blog #stanford@rvnikita_blog #john_ousterhout@rvnikita_blog

Newsletter

Get new posts by email

A weekly digest of what I write about AI, software and products. Unsubscribe anytime.