跳到正文
原文
Hacker News· signa11·· 3 小时前AI 评分31

Homa:TCP 为何在 AI 集群走向终结

Homa: The End of TCP for AI Clusters [video]

AI 导读

Stanford 教授 John Ousterhout 在 AI Engineer 频道发表演讲,阐述 AI 网络负载正从吞吐密集型转向延迟敏感型,TCP 与 RDMA 在数据中心中已成瓶颈。他提出 Homa 传输协议作为重新设计的方案,针对消息边界缺失与 incast 拥塞等机制问题,解决单个慢速交换拖垮整组 GPU 的难题。

正文

Back Image 1

Skip navigation

Search

Search with your voice

Sign in

Image 2

Video 1

Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford

Tap to unmute

2x

Image 3

Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford

AI Engineer 236,442 views 2 weeks ago

Copy link

Info

Shopping

Image 4

Image 5

If playback doesn't begin shortly, try restarting your device.

•

You're signed out

Videos you watch may be added to the TV's watch history and influence TV recommendations. To avoid this, cancel and sign in to YouTube on your computer.

Cancel Confirm

Share

- [x] Include playlist

An error occurred while retrieving sharing information. Please try again later.

Image 6

0:00

0:00 / 0:00

Live

•Watch full video

•

Why latency is becoming the metric that matters

•

57:14 Stevie Nicks - Stevie Nicks: Live At Red Rocks Music • 2020 1y ago Live Playlist ()Mix (50+)19:28 How AI Agents Actually Work (Every Piece Explained & Built)Tech With Tim 180K • 1mo ago Live Playlist ()Mix (50+)19:18 Why Software Factories Fail AI Engineer 442K • 2mo ago Live Playlist ()Mix (50+)51:53 Kent Beck: Software Engineering in the Age of AI | Prodacity 2026 Rise8 114K • 5d ago Live Playlist ()Mix (50+)19:06 Deep Reinforcement Learning for Competitive Agents in MicroRTS | Master's Thesis Presentation (2026)Mathis Delsart 19 • 1d ago Live Playlist ()Mix (50+)15:03 What Is Jev? The AI Model That Doesn't Generate Text IBM Technology 634K • 3d ago Live Playlist ()Mix (50+)21:18 Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley AI Engineer 434K • 2mo ago Live Playlist ()Mix (50+)14:03 Need to scale postgres? This is how The SF Database Meetup and PlanetScale 12K • 4d ago Live Playlist ()Mix (50+)46:05 The Buckmaster Interview (Navier-Stokes) - Numberphile Numberphile 326K • 4d ago Live Playlist ()Mix (50+)19:09 The State of AI in Software Development: Data from 400+ Orgs — Justin Reock, DX AI Engineer 36K • 4d ago Live Playlist ()Mix (50+)55:45 Parallax: Building the Decentralized "Linux of AI" With No Off Switch | Jon Durbin at Exploit The Opentensor Foundation | Bittensor TAO 2.5K • 4d ago Live Playlist ()Mix (50+)49:10 OOP vs. Data-Oriented Programming: Which One to Choose?Java 10K • 10h ago Live Playlist ()Mix (50+)

Sign in to confirm you’re not a bot

This helps protect our community

Sign inLearn more

Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford

Image 7: AI Engineer

AI Engineer

AI Engineer

654K subscribers

Subscribe

Subscribed

2.5K

Share

Download

Download

Save

236K views 2 weeks ago

236,442 views • Sep 17, 2026

John Ousterhout, Stanford Professor and author of A Philosophy of Software Design, turns our attention to the evolving nature of AI networking workloads and why traditional protocols like TCP and RDMA are becoming bottlenecks in modern data center environments! …...more

...more

How this was made

Auto-dubbed

Audio tracks for some languages were automatically generated. Learn more

Chapters

View all

Image 8 ### Why latency is becoming the metric that matters ### Why latency is becoming the metric that matters 0:00

Why latency is becoming the metric that matters

0:00

Image 9 ### The old workload: gigabytes and throughput ### The old workload: gigabytes and throughput 2:19

The old workload: gigabytes and throughput

2:19

Image 10 ### The new workload: metadata and coordination ### The new workload: metadata and coordination 3:09

The new workload: metadata and coordination

3:09

Image 11 ### How one slow exchange stalls every GPU ### How one slow exchange stalls every GPU 4:24

How one slow exchange stalls every GPU

4:24

Image 12 ### Incast, and where the queue actually builds ### Incast, and where the queue actually builds 6:07

Incast, and where the queue actually builds

6:07

Image 13 ### Why congestion control lives on the wrong end ### Why congestion control lives on the wrong end 7:10

Why congestion control lives on the wrong end

7:10

Image 14 ### A byte stream has no message boundaries ### A byte stream has no message boundaries 10:05

A byte stream has no message boundaries

10:05

Image 15 ### Homa, and a clean slate redesign ### Homa, and a clean slate redesign 11:08

Homa, and a clean slate redesign

11:08

Transcript

Follow along using the transcript.

Show transcript

Image 16 ### AI Engineer 654K subscribers

VideosAboutImage 17 Linkedin — Work with us!

Show less

Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford

236,442 views 236K views

Sep 17, 2026

2.5K

Share

Download

Download

Save

192 Comments

Sort comments

Sort by

Top Show featured commentsNewest Show recent comments, including potential spam

Image 18: Default profile photo

Add a comment...

Image 19

Pinned by @aiDotEngineer

@aiDotEngineer

@aiDotEngineer @aiDotEngineer2 weeks ago

Join us for the next AIEs: NYC Oct 12-14 https://ai.engineer/nyc/2026 and SF Nov 10-12 https://ai.engineer/code/2026

Show less Read more

Like

2

Dislike

Reply

Image 20

@twoplustwo5

2 weeks ago

TCP also has lost it's job to AI

Show less Read more

Like

401

Dislike

Reply

10 replies

Hide replies

10 replies

Image 21

@lagrangian_density

2 weeks ago

High performance computing left TCP two decades ago.

Show less Read more

Like

35

Dislike

Reply

Image 22

@mehrdadfeller

12 days ago

I wish there was also a comparison against QUIC which addressing a lot of the downsides of TCP

Show less Read more

Like

39

Dislike

Reply

4 replies

Hide replies

4 replies

Image 23

@blargathon

1 hour ago

TCP is a legacy protocol and I’m not sad to see it go. Modern protocols like QUIC and others will serve us better across the board.

Show less Read more

Like

Dislike

Reply

Image 24

@DanMcGrath77

13 days ago

Everyone on Earth: "We need to slow down with AI!" John Ousterhout: "Wait wait wait wait, I found how to make AI FASTER!"

Show less Read more

Like

2

Dislike

Reply

Image 25

@JackieSoares-c7d

8 days ago

Homa's approach to transmitting datagrams closely mirrors Demand Assigned Multiple Access (DAMA) and Priority-Oriented Demand Assignment (PODA) protocols used in satellite communications, since both solve the same core problem: allocating a contested transmission medium efficiently without wasting time on handshakes. For short data, both systems favor a "send first" approach — satellite terminals inject brief messages directly into shared contention slots (like Slotted ALOHA), while Homa lets senders push out an initial unscheduled burst (roughly the RTT-bytes window) without waiting for permission, so small transfers complete in a single network round trip. For long transfers, both switch to explicit coordination: satellite terminals request dedicated space segments from a master station and wait for assignment before streaming, while Homa senders pause after their unscheduled burst and wait for the receiver to issue GRANT packets specifying how much more they can send. The key structural difference is where control lives — satellite systems centralize scheduling in a master earth station or transponder, whereas Homa is fully decentralized, with each receiving host acting as its own local scheduler, using shortest-remaining-processing-time logic to grant capacity to whichever sender it chooses.

Show less Read more

Like

6

Dislike

Reply

Image 26

@EldonFreeman-v2k

12 days ago

Unbelievable how this guy can be so relevant and inspirational still! He was a role model for open source projects when he invented Tcl/Tk way back in the 80's/90's and he apparently hasn't got stuck...

Show less Read more

Like

25

Dislike

Reply

2 replies

Hide replies

2 replies

Image 27

@selub1058

2 weeks ago

Excellent presentation. I hope Homa will be ready enterprise solution and will improve LLM generation.

Show less Read more

Like

10

Dislike

Reply

Image 28

@vicaya

2 weeks ago

Neat. Application aware QoS finally found a viable use case.

Show less Read more

Like

10

Dislike

Reply

Image 29

@johngregor6743

2 weeks ago

Homa's been kicking around since 2018. As of yet, it has not set the world on fire.

Show less Read more

Like

30

Dislike

Reply

6 replies

Hide replies

6 replies

Image 30

@ScarlettDanger

1 day ago

Thanks for sharing!!

Show less Read more

Like

Dislike

Reply

Image 31

@darkbit1001

2 weeks ago

This is specific to high message count, low-latency throughput (dispersed small/large packets) vs infiniband which is high bandwidth throughput ( features very large packets )

Show less Read more

Like

7

Dislike

Reply

Image 32

@alimurreza

10 days ago

Excellent talk!

Show less Read more

Like

Dislike

Reply

Image 33

@NyxSalazar

11 days ago

This deserves way more views.

Show less Read more

Like

Dislike

Reply

Image 34

@dastin7276

7 days ago

Image 35: 👍 Great to see the modeling and results.

Show less Read more

Like

Dislike

Reply

Image 36

@SinergiasHolisticas

2 weeks ago

Perfect!!!

Show less Read more

Like

2

Dislike

Reply

Image 37

@erionalite

6 days ago

great presentation thank you

Show less Read more

Like

Dislike

Reply

Image 38

@Imaginasercasilisto

13 days ago

Homa looks very promising. Thank you!

Show less Read more

Like

1

Dislike

Reply

Image 39

@avataros111

2 days ago (edited)

An interesting experiment would be to watch how it behaves over unstable WiFi, where the packets don't arrive in a particular order. :)

Show less Read more

Like

Dislike

Reply

Comments 192

Top Show featured commentsNewest Show recent comments, including potential spam

In this video

Chapters

Transcript

Chapters

Image 40 ### Why latency is becoming the metric that matters ### Why latency is becoming the metric that matters 0:00

Why latency is becoming the metric that matters

0:00

Image 41 ### The old workload: gigabytes and throughput ### The old workload: gigabytes and throughput 2:19

The old workload: gigabytes and throughput

2:19

Image 42 ### The new workload: metadata and coordination ### The new workload: metadata and coordination 3:09

The new workload: metadata and coordination

3:09

Image 43 ### How one slow exchange stalls every GPU ### How one slow exchange stalls every GPU 4:24

How one slow exchange stalls every GPU

4:24

Image 44 ### Incast, and where the queue actually builds ### Incast, and where the queue actually builds 6:07

Incast, and where the queue actually builds

6:07

Image 45 ### Why congestion control lives on the wrong end ### Why congestion control lives on the wrong end 7:10

Why congestion control lives on the wrong end

7:10

Image 46 ### A byte stream has no message boundaries ### A byte stream has no message boundaries 10:05

A byte stream has no message boundaries

10:05

Image 47 ### Homa, and a clean slate redesign ### Homa, and a clean slate redesign 11:08

Homa, and a clean slate redesign

11:08

Image 48 ### Messages, not streams ### Messages, not streams 12:12

Messages, not streams

12:12

Image 49 ### Controlling congestion from the receiver ### Controlling congestion from the receiver 13:30

Controlling congestion from the receiver

13:30

Image 50 ### Using the priority queues already in the switch ### Using the priority queues already in the switch 14:59

Using the priority queues already in the switch

14:59

Image 51 ### The benchmark against TCP ### The benchmark against TCP 15:52

The benchmark against TCP

15:52

Sync to video time

Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford

Image 52: AI Engineer

AI Engineer

AI Engineer

2.5K Likes

236K Views

Sep 17 2026

John Ousterhout, Stanford Professor and author of A Philosophy of Software Design, turns our attention to the evolving nature of AI networking workloads and why traditional protocols like TCP and RDMA are becoming bottlenecks in modern data center environments! The Shift in Workloads Historical context: AI traffic was dominated by massive, long-running transfers (gigabytes of gradients), where throughput was the primary metric (2:19-2:42). Modern AI: Workloads, especially inference and agentic applications, now rely on frequent, small coordination messages (e.g., KV cache lookups, barrier synchronization). These small messages are highly sensitive to latency (3:09-4:02). The Bottleneck: When small synchronization messages are mixed with large traffic, they get trapped in queues (caused by incast), significantly increasing 99th percentile (tail) latency. This causes GPUs to sit idle, wasting expensive compute resources (4:24-5:34). Why Legacy Protocols Struggle Sender-Driven Congestion Control: TCP and RDMA rely on the sender to detect congestion, often via packet drops or delayed signals from switches. This process is inherently reactive and oscillates, leading to unstable performance (7:10-9:58). Byte Stream Model: These protocols view data as an opaque stream of bytes rather than discrete messages, making it difficult to prioritize short, critical tasks (10:05-11:05). The Homa Solution John introduces Homa, a clean-slate transport protocol designed for data centers (11:15-12:15): Message-Based: Unlike byte streams, Homa understands message boundaries, allowing it to predict traffic and prioritize short messages using Shortest Remaining Processing Time (SRPT) (12:22-13:28). Receiver-Driven: The receiver controls the flow by issuing grants to senders, effectively managing congestion before it occurs at the switch (13:30-14:58). Priority Queues: Homa leverages the multiple hardware queues already present in modern switches to bypass long, queued traffic with low-latency short messages (14:59-15:51). Performance Results: Benchmarks show Homa can reduce tail latency for short messages by over 10x compared to TCP, while simultaneously improving performance for large messages (15:52-17:27). Speaker info:

Timestamps: 0:00 - Why latency is becoming the metric that matters 2:19 - The old workload: gigabytes and throughput 3:09 - The new workload: metadata and coordination 4:24 - How one slow exchange stalls every GPU 6:07 - Incast, and where the queue actually builds 7:10 - Why congestion control lives on the wrong end 10:05 - A byte stream has no message boundaries 11:08 - Homa, and a clean slate redesign 12:12 - Messages, not streams 13:30 - Controlling congestion from the receiver 14:59 - Using the priority queues already in the switch 15:52 - The benchmark against TCP…...more

...more Show less

How this was made

Auto-dubbed

Audio tracks for some languages were automatically generated. Learn more

Chapters

View all

Image 53 ### Why latency is becoming the metric that matters ### Why latency is becoming the metric that matters 0:00

Why latency is becoming the metric that matters

0:00

Image 54 ### The old workload: gigabytes and throughput ### The old workload: gigabytes and throughput 2:19

The old workload: gigabytes and throughput

2:19

Image 55 ### The new workload: metadata and coordination ### The new workload: metadata and coordination 3:09

The new workload: metadata and coordination

3:09

Image 56 ### How one slow exchange stalls every GPU ### How one slow exchange stalls every GPU 4:24

How one slow exchange stalls every GPU

4:24

Transcript

Follow along using the transcript.

Show transcript

Image 57 ### AI Engineer 654K subscribers

VideosAboutImage 58 Linkedin — Work with us!

Transcript

NaN / NaN

Image 59 57:14

Image 60

Stevie Nicks - Stevie Nicks: Live At Red Rocks

Music 2020

Free with ads

G

Image 61 19:28

Image 62

How AI Agents Actually Work (Every Piece Explained & Built)

Tech With Tim

180K 1mo ago

Image 63 19:18

Image 64

Why Software Factories Fail

AI Engineer

442K 2mo ago

Image 65 51:53 New

Image 66

Kent Beck: Software Engineering in the Age of AI | Prodacity 2026

Rise8

114K 5d ago

Image 67 19:06 New

Image 68

Deep Reinforcement Learning for Competitive Agents in MicroRTS | Master's Thesis Presentation (2026)

Mathis Delsart

19 1d ago

Image 69 15:03 New

Image 70

What Is Jev? The AI Model That Doesn't Generate Text

IBM Technology

634K 3d ago

Image 71 21:18

Image 72

Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley

AI Engineer

434K 2mo ago

Image 73 14:03 New

Image 74

Image 75

Need to scale postgres? This is how

The SF Database Meetup and PlanetScale

12K 4d ago

Image 76 46:05 New

Image 77

The Buckmaster Interview (Navier-Stokes) - Numberphile

Numberphile

326K 4d ago

Image 78 19:09 New

Image 79

The State of AI in Software Development: Data from 400+ Orgs — Justin Reock, DX

AI Engineer

36K 4d ago

Image 80 55:45 New

Image 81

Parallax: Building the Decentralized "Linux of AI" With No Off Switch | Jon Durbin at Exploit

The Opentensor Foundation | Bittensor TAO

2.5K 4d ago

Image 82 49:10 New

Image 83

OOP vs. Data-Oriented Programming: Which One to Choose?

Java

10K 10h ago

Image 84 18:05

Image 85

Jev CEO: I made ChatGPT, now I'm building what's next

AI Engineer

492K 2mo ago

Image 86 22:35

Image 87

Fixing the PR Bottleneck — Matt Pocock, AIHero

AI Engineer

152K 8d ago

Image 88 20:27

Image 89

Harnesses in AI: A Deep Dive — Tejas Kumar, IBM

AI Engineer

301K 4mo ago

Image 90 13:19 New

Image 91

A total disaster

The PrimeTime

581K 1d ago

Image 92 6:44 New

Image 93

DHH has gone completely off the rails...

Fireship

1.9M 6d ago

Image 94 51:00

Image 95

AI Infrastructure Explained (GPUs, vLLM, and LLM-D)

KodeKloud

337K 1mo ago

Image 96 24:41 New

Image 97

The Death of the Code Review: What the Data Actually Says — Laurie Voss, Arize AI

AI Engineer

47K 4d ago

Image 98 21:19 New

Image 99

Which is The Best Qwen3.8-27B?

Sam Witteveen

5.6K 3h ago

Show more

Show more

来源:Hacker News · youtube.com