Homa:TCP 为何在 AI 集群走向终结
Homa: The End of TCP for AI Clusters [video]
Stanford 教授 John Ousterhout 在 AI Engineer 频道发表演讲,阐述 AI 网络负载正从吞吐密集型转向延迟敏感型,TCP 与 RDMA 在数据中心中已成瓶颈。他提出 Homa 传输协议作为重新设计的方案,针对消息边界缺失与 incast 拥塞等机制问题,解决单个慢速交换拖垮整组 GPU 的难题。
Skip navigation
Search
Search with your voice
Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford
Tap to unmute
2x
Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford
AI Engineer 236,442 views 2 weeks ago
Copy link
Info
Shopping
If playback doesn't begin shortly, try restarting your device.
•
You're signed out
Videos you watch may be added to the TV's watch history and influence TV recommendations. To avoid this, cancel and sign in to YouTube on your computer.
Cancel Confirm
Share
- [x] Include playlist
An error occurred while retrieving sharing information. Please try again later.
0:00
0:00 / 0:00
Live
•Watch full video
•
Why latency is becoming the metric that matters
•
57:14 Stevie Nicks - Stevie Nicks: Live At Red Rocks Music • 2020 1y ago Live Playlist ()Mix (50+)19:28 How AI Agents Actually Work (Every Piece Explained & Built)Tech With Tim 180K • 1mo ago Live Playlist ()Mix (50+)19:18 Why Software Factories Fail AI Engineer 442K • 2mo ago Live Playlist ()Mix (50+)51:53 Kent Beck: Software Engineering in the Age of AI | Prodacity 2026 Rise8 114K • 5d ago Live Playlist ()Mix (50+)19:06 Deep Reinforcement Learning for Competitive Agents in MicroRTS | Master's Thesis Presentation (2026)Mathis Delsart 19 • 1d ago Live Playlist ()Mix (50+)15:03 What Is Jev? The AI Model That Doesn't Generate Text IBM Technology 634K • 3d ago Live Playlist ()Mix (50+)21:18 Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley AI Engineer 434K • 2mo ago Live Playlist ()Mix (50+)14:03 Need to scale postgres? This is how The SF Database Meetup and PlanetScale 12K • 4d ago Live Playlist ()Mix (50+)46:05 The Buckmaster Interview (Navier-Stokes) - Numberphile Numberphile 326K • 4d ago Live Playlist ()Mix (50+)19:09 The State of AI in Software Development: Data from 400+ Orgs — Justin Reock, DX AI Engineer 36K • 4d ago Live Playlist ()Mix (50+)55:45 Parallax: Building the Decentralized "Linux of AI" With No Off Switch | Jon Durbin at Exploit The Opentensor Foundation | Bittensor TAO 2.5K • 4d ago Live Playlist ()Mix (50+)49:10 OOP vs. Data-Oriented Programming: Which One to Choose?Java 10K • 10h ago Live Playlist ()Mix (50+)
Sign in to confirm you’re not a bot
This helps protect our community
Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford
AI Engineer
654K subscribers
Subscribe
Subscribed
2.5K
Share
Download
Download
Save
236K views 2 weeks ago
236,442 views • Sep 17, 2026
John Ousterhout, Stanford Professor and author of A Philosophy of Software Design, turns our attention to the evolving nature of AI networking workloads and why traditional protocols like TCP and RDMA are becoming bottlenecks in modern data center environments! …...more
...more
How this was made
Auto-dubbed
Audio tracks for some languages were automatically generated. Learn more
Chapters
View all
Why latency is becoming the metric that matters
0:00
### The old workload: gigabytes and throughput ### The old workload: gigabytes and throughput 2:19
The old workload: gigabytes and throughput
2:19
### The new workload: metadata and coordination ### The new workload: metadata and coordination 3:09
The new workload: metadata and coordination
3:09
### How one slow exchange stalls every GPU ### How one slow exchange stalls every GPU 4:24
How one slow exchange stalls every GPU
4:24
### Incast, and where the queue actually builds ### Incast, and where the queue actually builds 6:07
Incast, and where the queue actually builds
6:07
Why congestion control lives on the wrong end
7:10
### A byte stream has no message boundaries ### A byte stream has no message boundaries 10:05
A byte stream has no message boundaries
10:05
### Homa, and a clean slate redesign ### Homa, and a clean slate redesign 11:08
Homa, and a clean slate redesign
11:08
Transcript
Follow along using the transcript.
Show transcript
### AI Engineer 654K subscribers
VideosAbout Linkedin — Work with us!
Show less
Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford
236,442 views 236K views
Sep 17, 2026
2.5K
Share
Download
Download
Save
192 Comments
Sort comments
Sort by
Top Show featured commentsNewest Show recent comments, including potential spam
Add a comment...
Pinned by @aiDotEngineer
@aiDotEngineer
@aiDotEngineer @aiDotEngineer2 weeks ago
Join us for the next AIEs: NYC Oct 12-14 https://ai.engineer/nyc/2026 and SF Nov 10-12 https://ai.engineer/code/2026
Show less Read more
Like
2
Dislike
Reply
@twoplustwo5
TCP also has lost it's job to AI
Show less Read more
Like
401
Dislike
Reply
10 replies
Hide replies
10 replies
@lagrangian_density
High performance computing left TCP two decades ago.
Show less Read more
Like
35
Dislike
Reply
@mehrdadfeller
I wish there was also a comparison against QUIC which addressing a lot of the downsides of TCP
Show less Read more
Like
39
Dislike
Reply
4 replies
Hide replies
4 replies
@blargathon
TCP is a legacy protocol and I’m not sad to see it go. Modern protocols like QUIC and others will serve us better across the board.
Show less Read more
Like
Dislike
Reply
@DanMcGrath77
Everyone on Earth: "We need to slow down with AI!" John Ousterhout: "Wait wait wait wait, I found how to make AI FASTER!"
Show less Read more
Like
2
Dislike
Reply
@JackieSoares-c7d
Homa's approach to transmitting datagrams closely mirrors Demand Assigned Multiple Access (DAMA) and Priority-Oriented Demand Assignment (PODA) protocols used in satellite communications, since both solve the same core problem: allocating a contested transmission medium efficiently without wasting time on handshakes. For short data, both systems favor a "send first" approach — satellite terminals inject brief messages directly into shared contention slots (like Slotted ALOHA), while Homa lets senders push out an initial unscheduled burst (roughly the RTT-bytes window) without waiting for permission, so small transfers complete in a single network round trip. For long transfers, both switch to explicit coordination: satellite terminals request dedicated space segments from a master station and wait for assignment before streaming, while Homa senders pause after their unscheduled burst and wait for the receiver to issue GRANT packets specifying how much more they can send. The key structural difference is where control lives — satellite systems centralize scheduling in a master earth station or transponder, whereas Homa is fully decentralized, with each receiving host acting as its own local scheduler, using shortest-remaining-processing-time logic to grant capacity to whichever sender it chooses.
Show less Read more
Like
6
Dislike
Reply
@EldonFreeman-v2k
Unbelievable how this guy can be so relevant and inspirational still! He was a role model for open source projects when he invented Tcl/Tk way back in the 80's/90's and he apparently hasn't got stuck...
Show less Read more
Like
25
Dislike
Reply
2 replies
Hide replies
2 replies
@selub1058
Excellent presentation. I hope Homa will be ready enterprise solution and will improve LLM generation.
Show less Read more
Like
10
Dislike
Reply
@vicaya
Neat. Application aware QoS finally found a viable use case.
Show less Read more
Like
10
Dislike
Reply
@johngregor6743
Homa's been kicking around since 2018. As of yet, it has not set the world on fire.
Show less Read more
Like
30
Dislike
Reply
6 replies
Hide replies
6 replies
@ScarlettDanger
Thanks for sharing!!
Show less Read more
Like
Dislike
Reply
@darkbit1001
This is specific to high message count, low-latency throughput (dispersed small/large packets) vs infiniband which is high bandwidth throughput ( features very large packets )
Show less Read more
Like
7
Dislike
Reply
@alimurreza
Excellent talk!
Show less Read more
Like
Dislike
Reply
@NyxSalazar
This deserves way more views.
Show less Read more
Like
Dislike
Reply
@dastin7276
Great to see the modeling and results.
Show less Read more
Like
Dislike
Reply
@SinergiasHolisticas
Perfect!!!
Show less Read more
Like
2
Dislike
Reply
@erionalite
great presentation thank you
Show less Read more
Like
Dislike
Reply
@Imaginasercasilisto
Homa looks very promising. Thank you!
Show less Read more
Like
1
Dislike
Reply
@avataros111
An interesting experiment would be to watch how it behaves over unstable WiFi, where the packets don't arrive in a particular order. :)
Show less Read more
Like
Dislike
Reply
Comments 192
Top Show featured commentsNewest Show recent comments, including potential spam
In this video
Chapters
Transcript
Chapters
Why latency is becoming the metric that matters
0:00
### The old workload: gigabytes and throughput ### The old workload: gigabytes and throughput 2:19
The old workload: gigabytes and throughput
2:19
### The new workload: metadata and coordination ### The new workload: metadata and coordination 3:09
The new workload: metadata and coordination
3:09
### How one slow exchange stalls every GPU ### How one slow exchange stalls every GPU 4:24
How one slow exchange stalls every GPU
4:24
### Incast, and where the queue actually builds ### Incast, and where the queue actually builds 6:07
Incast, and where the queue actually builds
6:07
Why congestion control lives on the wrong end
7:10
### A byte stream has no message boundaries ### A byte stream has no message boundaries 10:05
A byte stream has no message boundaries
10:05
### Homa, and a clean slate redesign ### Homa, and a clean slate redesign 11:08
Homa, and a clean slate redesign
11:08
### Messages, not streams ### Messages, not streams 12:12
Messages, not streams
12:12
### Controlling congestion from the receiver ### Controlling congestion from the receiver 13:30
Controlling congestion from the receiver
13:30
Using the priority queues already in the switch
14:59
### The benchmark against TCP ### The benchmark against TCP 15:52
The benchmark against TCP
15:52
Sync to video time
Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford
AI Engineer
AI Engineer
2.5K Likes
236K Views
Sep 17 2026
John Ousterhout, Stanford Professor and author of A Philosophy of Software Design, turns our attention to the evolving nature of AI networking workloads and why traditional protocols like TCP and RDMA are becoming bottlenecks in modern data center environments! The Shift in Workloads Historical context: AI traffic was dominated by massive, long-running transfers (gigabytes of gradients), where throughput was the primary metric (2:19-2:42). Modern AI: Workloads, especially inference and agentic applications, now rely on frequent, small coordination messages (e.g., KV cache lookups, barrier synchronization). These small messages are highly sensitive to latency (3:09-4:02). The Bottleneck: When small synchronization messages are mixed with large traffic, they get trapped in queues (caused by incast), significantly increasing 99th percentile (tail) latency. This causes GPUs to sit idle, wasting expensive compute resources (4:24-5:34). Why Legacy Protocols Struggle Sender-Driven Congestion Control: TCP and RDMA rely on the sender to detect congestion, often via packet drops or delayed signals from switches. This process is inherently reactive and oscillates, leading to unstable performance (7:10-9:58). Byte Stream Model: These protocols view data as an opaque stream of bytes rather than discrete messages, making it difficult to prioritize short, critical tasks (10:05-11:05). The Homa Solution John introduces Homa, a clean-slate transport protocol designed for data centers (11:15-12:15): Message-Based: Unlike byte streams, Homa understands message boundaries, allowing it to predict traffic and prioritize short messages using Shortest Remaining Processing Time (SRPT) (12:22-13:28). Receiver-Driven: The receiver controls the flow by issuing grants to senders, effectively managing congestion before it occurs at the switch (13:30-14:58). Priority Queues: Homa leverages the multiple hardware queues already present in modern switches to bypass long, queued traffic with low-latency short messages (14:59-15:51). Performance Results: Benchmarks show Homa can reduce tail latency for short messages by over 10x compared to TCP, while simultaneously improving performance for large messages (15:52-17:27). Speaker info:
Timestamps: 0:00 - Why latency is becoming the metric that matters 2:19 - The old workload: gigabytes and throughput 3:09 - The new workload: metadata and coordination 4:24 - How one slow exchange stalls every GPU 6:07 - Incast, and where the queue actually builds 7:10 - Why congestion control lives on the wrong end 10:05 - A byte stream has no message boundaries 11:08 - Homa, and a clean slate redesign 12:12 - Messages, not streams 13:30 - Controlling congestion from the receiver 14:59 - Using the priority queues already in the switch 15:52 - The benchmark against TCP…...more
...more Show less
How this was made
Auto-dubbed
Audio tracks for some languages were automatically generated. Learn more
Chapters
View all
Why latency is becoming the metric that matters
0:00
### The old workload: gigabytes and throughput ### The old workload: gigabytes and throughput 2:19
The old workload: gigabytes and throughput
2:19
### The new workload: metadata and coordination ### The new workload: metadata and coordination 3:09
The new workload: metadata and coordination
3:09
### How one slow exchange stalls every GPU ### How one slow exchange stalls every GPU 4:24
How one slow exchange stalls every GPU
4:24
Transcript
Follow along using the transcript.
Show transcript
### AI Engineer 654K subscribers
VideosAbout Linkedin — Work with us!
Transcript
NaN / NaN
Stevie Nicks - Stevie Nicks: Live At Red Rocks
Music 2020
Free with ads
G
How AI Agents Actually Work (Every Piece Explained & Built)
Tech With Tim
180K 1mo ago
Why Software Factories Fail
AI Engineer
442K 2mo ago
Kent Beck: Software Engineering in the Age of AI | Prodacity 2026
Rise8
114K 5d ago
Deep Reinforcement Learning for Competitive Agents in MicroRTS | Master's Thesis Presentation (2026)
Mathis Delsart
19 1d ago
What Is Jev? The AI Model That Doesn't Generate Text
IBM Technology
634K 3d ago
Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley
AI Engineer
434K 2mo ago
Need to scale postgres? This is how
The SF Database Meetup and PlanetScale
12K 4d ago
The Buckmaster Interview (Navier-Stokes) - Numberphile
Numberphile
326K 4d ago
The State of AI in Software Development: Data from 400+ Orgs — Justin Reock, DX
AI Engineer
36K 4d ago
Parallax: Building the Decentralized "Linux of AI" With No Off Switch | Jon Durbin at Exploit
The Opentensor Foundation | Bittensor TAO
2.5K 4d ago
OOP vs. Data-Oriented Programming: Which One to Choose?
Java
10K 10h ago
Jev CEO: I made ChatGPT, now I'm building what's next
AI Engineer
492K 2mo ago
Fixing the PR Bottleneck — Matt Pocock, AIHero
AI Engineer
152K 8d ago
Harnesses in AI: A Deep Dive — Tejas Kumar, IBM
AI Engineer
301K 4mo ago
A total disaster
The PrimeTime
581K 1d ago
DHH has gone completely off the rails...
Fireship
1.9M 6d ago
AI Infrastructure Explained (GPUs, vLLM, and LLM-D)
KodeKloud
337K 1mo ago
The Death of the Code Review: What the Data Actually Says — Laurie Voss, Arize AI
AI Engineer
47K 4d ago
Which is The Best Qwen3.8-27B?
Sam Witteveen
5.6K 3h ago
Show more
Show more
来源:Hacker News · youtube.com
57:14
19:28
19:18
51:53 New
19:06 New
15:03 New
21:18
14:03 New
46:05 New
19:09 New
55:45 New