跳到正文
原文
Hacker News· zed_labs_dev·· 7 小时前AI 评分40

Claude Code 建议消息功能解析:真正的客户是模型本身

Claude Code’s suggested message feature: I think the real customer is the model

AI 导读

Claude Code 在完成任务后会预填一条建议消息,工程师 Zohaib Ansari 认为该功能表面是为用户节省输入,实际上真正的"客户"是模型本身。每条建议都是对用户下一步动作的预测,用户直接发送即为正向反馈,编辑后则形成偏好对,这些来自真实仓库的数据可作为强化学习人类反馈(RLHF)的训练素材。作者承认这只是猜测,会话是否被用于训练取决于用户的隐私设置。

正文

Zohaib AnsariGitHubLinkedInEmail

  • About 01

Software engineer in Waterloo, ON. I started with Android apps and now work mostly on full-stack systems and AI agents. I care about making products that make life easier, feel fast, and stay fun to use.

Waterloo, ON 17:26 EDT 01 / 08

2026-10-06 Close

The smartest Claude code feature is not for its users

Claude Code recently started filling in the prompt box for me. After it finishes a task, a suggested next message is already sitting in the text field, something like "run the tests" or "commit this". I can send it as is or edit it first.

its off, can u verify↵

  • Bypass permissions Opus 5.5 Medium

A suggested next message, pre-filled in the Claude Code prompt box.

As a user feature it is minor. It saves a few seconds of typing, and I often write my own message anyway.

I think the real customer is the model.

Getting useful feedback out of users is hard. Every AI product has thumbs up and thumbs down buttons and few people click them. The ones who do tend to be annoyed, so the labels are sparse and carry a selection bias. Paying annotators works, but it is expensive, and an annotator reading someone else's codebase is guessing at what the developer wanted.

The suggestion box gets around both problems. Each suggestion is a prediction of my next turn, conditioned on the whole session up to that point. I then grade it without thinking of it as grading. Sending it untouched is a positive label. Editing it is worth more, because the original and my edited version form a preference pair, and the diff between them shows where the prediction went wrong. If the suggestion says "run the tests" and I change it to "run only the auth tests, the full suite takes ten minutes", that is a correction from someone who knows the project, written at the moment they care about getting it right.

This is the raw material for reinforcement learning from human feedback. The standard recipe is to collect human preferences between model outputs and train a reward model on them, which the policy is then optimised against. Collecting the preferences has always been the expensive part. Here they fall out of people doing their jobs, and they come from real repositories instead of a synthetic benchmark, so they are in distribution for the work the model will be asked to do.

There is a second benefit. Predicting the next user turn is a useful training objective on its own. A model that can guess what a competent developer asks for next has learned something about how work is sequenced. After a refactor you run the tests. Once they pass you commit. That is a short step from an agent that takes the next action without being asked.

I should say that I am guessing. I have no inside knowledge of what Anthropic does with this signal, and whether your sessions are used for training at all depends on your plan and privacy settings. But if I had to design a way to collect preference data from thousands of working engineers, I would struggle to come up with something cheaper. It costs one line of text in an input box, and it ships as a convenience.

来源:Hacker News · zohaib.cc