ESC
↑↓ 选择↵ 打开esc 关闭⌘K 唤起
← 返回速读报告 回声编辑部 · NO.93 · 全文

Harness Engineering is not Enough: Why Software Factories Fail — Dex Horthy, HumanLayer

频道: AI Engineer
视频: https://www.youtube.com/watch?v=Ib5GBkD555M
原文语言: en
统计: 共 55 轮


[0:01]

[music] What's up everybody? How we doing guys? Give it up for all the great speakers today so far. Um, [applause] all right. This is uh harness engineering is not enough and why software factories fail and we're going to click maybe. Oh, that's way too many slides. Hold on,

[音乐] 大家好啊?各位感觉怎么样?咱们给今天到目前为止所有精彩的演讲者鼓个掌吧。呃,[掌声] 好。这场演讲叫“harness engineering 还不够,以及软件工厂为什么会失败”,我们来点一下,可能……哦,这幻灯片也太多了。等一下,


[0:49]

guys. Okay. Um, so we're all racing to put AI coding into production and uh there's been lots been said about loop engineering and uh we should probably write more loops and uh yeah, I don't know. I guess we're doing loops now. Uh, strong DM built a lights off software factory where nobody even reads the code and the prevailing narrative is we should just spend more tokens. You are the bottleneck. The models are good

各位。好。呃,现在我们都在争先恐后地想把 AI 编程搬进生产环境,关于 loop engineering 也已经有很多讨论了,说我们大概应该多写点 loop,呃,是啊,我也说不好。我猜我们现在都在搞 loop 了吧。呃,StrongDM 搭了一个熄灯软件工厂(lights off),甚至都没人去读代码,而现在主流的说法就是我们只要多花点 token 就行。你才是瓶颈,模型已经足够好了


[1:18]

enough. Code is free. just ship more stuff. But at the same time, we are starting to see the cracks. Our friend uh Mario at AI Engineer Europe begged us to slow down because companies that should not be having outages because of coding agents are having outages due to coding agent mishaps. Um code bases are falling apart faster than they ever have before. And our friends at Farosai actually even did a report since we all

,代码不要钱,只管多发布就是了。但与此同时,我们也开始看到裂缝了。我们的朋友 Mario 在 AI Engineer Europe 上恳求大家慢下来,因为一些本不该因为 coding agent 而出故障的公司,正因为 coding agent 的失误而频频宕机。呃,代码库崩坏的速度比以往任何时候都快。我们在 Faros AI 的朋友们甚至还做了一份报告,因为自从我们大家


[1:44]

picked up all these AI coding tools in January, maybe February. Um, pull request code review quality is way down. We're having more comments, longer comments, and tons of PRs being merged without any review at all. Incidents are way up, bugs per developer are way up.

在一月、大概二月开始用上这些 AI 编程工具之后……呃,pull request 的代码审查质量大幅下滑。评论变多了,评论也变长了,还有大量 PR 根本没经过任何审查就被合并了。故障(incident)大幅上升,每个开发者身上的 bug 数也大幅上升。


[2:01]

And uh, many people will tell you that you're holding it wrong. That's the only reason. You're not. Well, maybe you are, but that's not the point. Um, I've spoken a lot about how to hold it better when it comes to working with AI. uh probably a million views on YouTube at this point across a bunch of different talks. Um and the basic thing is like as engineers we've been told that if token maxing isn't working then it's a skill issue. You

然后呢,很多人会告诉你,是你用错方法了。就这一个原因。其实不怪你。嗯,也许真怪你,但那不是重点。呃,关于怎么更好地跟 AI 协作,我已经讲过很多了,散在一堆不同的演讲里,现在 YouTube 上的播放量加起来大概有一百万了吧。嗯,基本意思是,作为工程师,我们一直被灌输:如果拼命堆 token 还不管用,那就是你的技术问题。你


[2:27]

just need to spend more tokens. Uh let go of reading the code. That with enough harness engineering if we maybe sprinkle some magic words adversarial review on enough of our PR bots that we can get the best of both worlds. 10 to 100x faster, high quality, and nobody has to do that thing we all hate called code review. Uh, I'm here to convince you today that this is in fact not a scale issue. That no amount of harness

只要多花点 token 就行了。呃,别再执着于读代码了。只要 harness engineering 做得够足,也许再往足够多的 PR 机器人上撒点像 adversarial review 这样的魔法咒语,我们就能两全其美:快上 10 到 100 倍、质量还高,而且再没人需要去做那件我们都讨厌的事——code review。呃,我今天要说服你们的是:这其实根本不是一个规模(scale)问题。再多的 harness


[2:54]

engineering or loops maxing can solve what is fundamentally a model training issue. That's why we say the harness is not enough. Um, and to understand this, we kind of have to grapple and dig into how coding models are trained. I'm going to talk about what I think the shortcomings are with some of the current benchmarks and what better ones might look like. And we'll talk about how to move faster safely in the meantime. Um, it's going to sound like a

engineering、再怎么疯狂加 loop,都解决不了一个本质上属于模型训练的问题。这就是为什么我们说 harness 还不够。嗯,要理解这一点,我们多少得去啃一啃、深挖一下 coding 模型到底是怎么训练出来的。我会讲讲我认为现在一些 benchmark 存在哪些不足,以及更好的 benchmark 可能长什么样。我们也会聊聊在这中间,怎么既跑得更快又足够安全。嗯,这听起来会有点像一场


[3:18]

rant. Uh, but there is hope here. Uh, I'm going to talk about our journey and a bunch of the landmines we've hit, uh, building in this world, a bunch of exciting new techniques that we've been working with, uh, a lot of our users and customers to develop, and I think how we all as a community get to the next chapter of agentic engineering after whatever this thing that we're in. Um,

吐槽。呃,但这里面是有希望的。呃,我会讲讲我们这一路走来的经历,以及我们在这个世界里做东西时踩过的一堆雷,还有一堆令人兴奋的新技术——都是我们和很多用户、客户一起摸索出来的,以及我认为我们整个社区要怎么走出眼下这个阶段(不管它到底叫什么),进入 agentic engineering 的下一个篇章。嗯,


[3:36]

so we use a lot of words here. I'm going to zoom out a little bit. I want to give you kind of like a brief history of the software factory. Um, and it was actually I don't I I I just learned this last week. It was the term software factory was defined at a NATO conference in 1968.

我们这儿用了很多词儿。我把镜头往外拉一点。我想给大家讲一段软件工厂的简史。嗯,其实——我也是上周才知道的——“软件工厂”这个词最早是在 1968 年的一次 NATO 会议上被定义出来的。


[3:49]

Uh, we're going to start around 2022, like right before AI started coming around. Um, and basically in a typical 2022 software factory, you will have some people building stuff. You'll have engineers, you'll have PMs, maybe you have some sort of leadership team that is driving the vision here. and they all decide that stuff needs to get done. And so you put it in a tracker, a linear, a Jira, a beads, some sort of state

呃,我们就从 2022 年前后讲起,差不多就是 AI 火起来之前那会儿。嗯,基本上,在一个典型的 2022 年软件工厂里,你会有一些人在做东西。你有工程师,有 PM,可能还有某种领导团队在这儿把控愿景。他们一起决定哪些东西该做。于是你把它丢进一个 tracker,可能是 Linear、Jira、Beads,某种用来追踪待办事项的状态


[4:11]

machine that tracks what needs to be done. And then someone goes and grabs something off of there and they build the thing. And there may be some automated testing in that process, maybe some manual testing in that process. At a certain point, we make this pull request thing says, "Okay, cool. We got to run a bunch of checks, automated stuff. A human's going to review the change and review the code. And perhaps we might even have uh a human pull it

机(state machine),用来追踪有哪些事情要做。然后有人从里面领一个任务,把东西做出来。这个过程里可能会有一些自动化测试,也可能会有一些手动测试。到某个节点,我们会发起 pull request,说:“好,很好。我们得跑一堆检查,都是自动化的。会有一个人来审查这次改动、审查代码。也许我们甚至还会让一个人把它拉


[4:32]

down and test it somehow. And if anything goes wrong here, we loop back to someone builds the thing. Uh, and eventually we're ready for prod. And so we ship it to production. And once it's in prod, it makes contact with our users. And users do, uh, a thing that,

下来,用某种方式测一测。如果这里出了任何岔子,我们就绕回到“有人来做东西”那一步。呃,最终我们准备好上生产环境了。于是我们把它发布到生产环境。一旦它进了生产环境,就要跟我们的用户接触了。而用户会做一件,呃,一件


[4:47]

uh, we all love. Uh, users love to complain. I I love our users. Uh, but yeah, they're going to ask for things. They're going to find bugs. They're going to file feature requests. Uh, and that goes back to your team. You might also add monitoring. And so, uh, you know, what do we want more than anything else? We want to wake up engineers at 3 in the morning when something breaks so they can get dragged out of bud to try to go

呃,我们都特别“喜欢”的事。呃,用户超爱抱怨。我——我是很爱我们的用户啦。呃,但没错,他们会提各种需求,会找出 bug,会提交功能请求。呃,这些又都回到你的团队这边。你可能还会加上监控。所以,呃,你懂的,我们最最想要的是什么?我们最想要的就是:一出问题,凌晨三点把工程师叫醒,好让他们被从被窝里拽出来,去试着


[5:08]

fix it. Uh and we go on and on in this loop uh and we ship a bunch of code. Um and one thing that we noticed here is that uh teams figured this out decades ago is that this someone builds the thing step is usually going to take hours or days in most cases. And the review part will also take hours or days for large things. And so teams started doing these upfront planning,

修好它。呃,我们就这样在这个循环里一圈又一圈,发布出一大堆代码。嗯,我们注意到一件事——其实各个团队几十年前就想明白了——就是“有人来做东西”这一步,大多数情况下通常要花上几个小时甚至几天。而审查那部分,碰上大改动同样要花几个小时甚至几天。于是团队们开始做这些前期的规划、


[5:30]

architecture, proposals, sprint planning, and they would collaborate this on on these things as a team uh with the hopes that we might decrease the percent chance that something would need to be reworked that we would be able to reduce the time spent in reviewing every line of code because we aligned on everything ahead of time.

architecture、方案提议、sprint planning,他们会作为一个团队一起来协作搞这些东西,呃,希望借此降低某个东西需要返工的概率,也希望能减少逐行审查代码所花的时间,因为我们提前就把所有事情都对齐好了。


[5:47]

This brings us to the agentic software factory. Uh every company and their mother is talking about how they built a coding agent factory that ships 75% of their code. Now uh literally everybody uh and so if we look at the software factory from 2022 uh we just replace someone builds the thing with an agent builds the thing and we have an orchestration and a harness and a sandbox and a model and computer use and I'm not going to get into like the

这就把我们带到了 agentic 软件工厂。呃,现在几乎每一家公司都在讲他们怎么搭了一个 coding agent 工厂、由它产出了他们 75% 的代码。真的是人人都在讲。所以,如果我们回看 2022 年的那个软件工厂,呃,我们只要把“有人来做东西”换成“agent 来做东西”,然后我们有 orchestration、有 harness、有 sandbox、有 model、有 computer use——我不打算去细讲那些


[6:12]

details of that. You can watch a hundred talks about that this week I'm sure. Um but now the building part takes minutes or hours but this human part still takes hours or days if you're going to review the code and you're going to test the changes. And so we bring in agentic code review and we bring in agentic regression testing. Uh, and it makes this part faster, but it's probably still the bottleneck. But we can do more

细节了。这周你肯定能看到上百场讲这个的演讲。嗯,但现在“做东西”那部分只要几分钟到几个小时,可“人来做”那部分——如果你要审查代码、要测试改动——仍然得花几个小时甚至几天。于是我们引入了 agentic 代码审查,引入了 agentic 回归测试。呃,这让这部分快了一些,但它很可能仍然是瓶颈。不过我们可以多加点


[6:32]

loops here. Why not? Let's do some more loops. So, we can route all incidents straight into the factory. Why does someone need to get woken up uh and try to fix it when they could just wake up to a pull request and uh maybe that fixes the issue for you? You can take all the user feedback and just stick it straight into the factory so that uh people ask for stuff and it gets built.

loop 嘛。为什么不呢?咱们再多来点 loop。所以,我们可以把所有故障都直接路由进工厂。既然醒来就能看到一个现成的 pull request、也许问题就顺手帮你解决了,那又何必大半夜把某个人叫醒去修呢?你可以把所有用户反馈直接塞进工厂,这样,呃,别人一提需求,东西就被做出来了。


[6:50]

And now your only job is how much things can you stuff into the queue of stuff to do and how fast can you review and test the changes. Which brings us of course to I'm sure you know the lights off software factory where basically Dan Shapiro coined this is we no longer read the code. We say you know what this is going great that code review thing no thanks. We're just not going to do that anymore. Uh and we invest into all these

而现在你唯一的工作,就是你能往待办队列里塞多少东西、以及你审查和测试改动能有多快。这当然就把我们带到了——你们肯定都知道——熄灯软件工厂(lights off),这个说法基本上是 Dan Shapiro 提出来的,意思就是我们不再读代码了。我们说:你猜怎么着,这进展棒极了,code review 那玩意儿——不了谢谢,我们以后就不干那个了。呃,我们把精力都投到所有这些


[7:13]

other parts of the system your your testing your monitoring your rollout everything else. we just write more code and build those systems better. And now our job really is just how how much stuff can we ask the agent to build. I am going to pause that does not work.

系统的其他部分上——你的测试、你的监控、你的 rollout,其他所有东西。我们就只管多写代码、把那些系统做得更好。而现在我们的工作真的就只剩下:我们能让 agent 做多少东西。我要在这儿停一下——那是行不通的。


[7:28]

Uh and this is why software factories fail. Um as as an aside, what I'm going to say has nothing to do with vibe coding. So Addy had this uh great post. I'm going to just go literally take his quote verbatim. a developer vibe coding a side project a dozen people will ever run and a team keeping a 10-year-old enterprise system alive for another quarter share almost no constraints worth naming and most of what you hear

呃,这就是软件工厂失败的原因。嗯,顺带说一句,我接下来要讲的跟 vibe coding 一点关系都没有。Addy 写过一篇特别棒的帖子,我干脆一字不差地引用他的原话:一个开发者 vibe coding 一个总共也就十来个人会用的副业项目,跟一个团队要让一套十年的企业系统再撑一个季度——这两者之间几乎没有任何值得一提的共同约束,而你在网上听到的大多数东西


[7:51]

on the internet is one of these groups of people telling the other group of people how to live their lives if you love vibe coding please go on um at human layer what we care about is how do we help people solve hard problems in complex code bases um we use the word brownfield a lot which historically has meant like some 10-year-old Java thing.

都是这两拨人里的一拨在教另一拨人该怎么过日子。如果你热爱 vibe coding,那请尽管继续。嗯,在 HumanLayer,我们关心的是:怎么帮人们在复杂的代码库里解决那些难题。嗯,我们经常用 brownfield(棕地/存量系统)这个词,它历来指的就是那种十年老 Java 项目之类的东西。


[8:10]

I actually think agents really start to struggle after maybe 3 to 6 months, especially with the pace at which we can ship. Now, um you can ask me how I know this and I will tell you that it is because in July 2025, we tried this. We went full lights off and uh if you have tried this seriously for a number of months, you probably found at least one issue that the agent couldn't solve.

我其实觉得,agent 大概在三到六个月之后就真的开始吃力了,尤其是以我们现在能达到的这种发布节奏。嗯,你可以问我是怎么知道的,我会告诉你:因为在 2025 年 7 月,我们试过这个。我们彻底走了熄灯(lights off)这条路,呃,如果你认真试上好几个月,你多半会碰到至少一个 agent 解决不了的问题。


[8:31]

even with your most advanced prompting. You do research, you do reproductions, you just you have to go and dig into that codebase that you stopped reading three months ago to try to figure out what's broken. And in the meantime, your site was down, your users were pissed,

哪怕用上你最高级的 prompt 技巧也没用。你得去做调研,去做复现,你就是不得不一头扎进那个你三个月前就不再读的代码库,去搞清楚到底哪儿坏了。而与此同时,你的网站已经挂了,你的用户已经气炸了,


[8:46]

and you were, if you were like me, you were probably miserable reading all this slop code that you let slip into your system. And what I want to get to is basically models have a shortcoming. um they can't maintain and improve codebase quality over time, not without a decent amount of human steering. Um and when I say maintainability, I'm basically talking about issues like it becomes really really hard to make a change in

而你呢,如果你跟我一样,那你大概正痛苦地读着这些被你放任溜进系统里的垃圾代码。我想说到的核心是:模型有一个短板。嗯,它们没法随着时间推移去维护和提升代码库的质量——至少在没有相当程度的人为引导(steering)的情况下做不到。嗯,我说的可维护性,基本上指的是这类问题:想在代码库的某一处做改动会变得非常非常难,


[9:06]

one part of the codebase without breaking other parts of the codebase. This is Martin Fowler shotgun surgery textbook code smell. Um I'm not going to say much more about maintainability. There's a bunch of books that you can go read about it. In fact, John Austerhood is actually here speaking this week, so you can go ask him in person about the philosophy of software design if you want to. Um, but it brings us to this

而且改的时候还不能弄坏代码库的其他部分。这就是 Martin Fowler 说的“shotgun surgery”(霰弹式修改),教科书级别的 code smell。嗯,可维护性我就不多说了,有一堆书你可以拿去读。事实上,John Ousterhout 这周也在这儿演讲,所以如果你愿意的话,可以当面去问他关于《软件设计的哲学》的问题。嗯,但这就把我们引向了这个


[9:24]

question of like why can't models do software maintainability. Um, and you may also be saying, "But Dex, you know, surely the models have gotten much better since then." Um, they've gotten better in some ways, but they're still about the same in others. Um, if you want to solve one-off problems or vibe code a new marketing site, yes, they got way better since 2025 and 2024, but as far as improving codebase quality, I

问题:为什么模型做不好软件的可维护性。嗯,你可能还会说:“可是 Dex,你也知道,模型从那会儿到现在肯定进步了一大截啊。”嗯,它们在某些方面确实变好了,但在另一些方面还是老样子。嗯,如果你想解决那种一次性的问题,或者 vibe code 一个新的营销网站,那没错,比起 2024、2025 年,它们强了太多;但要说提升代码库的质量,我


[9:47]

think uh they have not gotten much better. Now, I cannot prove this because there are no good benchmarks for a model's ability to maintain codebased quality. And I'll get into like where we're going with that. Um, but if you've worked with coding agents for a while, a lot of people are posting about this.

觉得,呃,它们并没有强多少。当然,这一点我没法证明,因为根本就没有好的 benchmark 来衡量模型维护代码库质量的能力。这块我们要往哪儿走,我后面会讲。嗯,但只要你跟 coding agent 打过一阵子交道,就会发现很多人都在发帖聊这件事。


[10:01]

It's just like you probably have this vibe that they they generally make things worse over time and make the codebase harder to work in. And to figure out why this happens, I want to zoom out to the first great coding agent. Why did cloud code go from nothing to four billion? And I think now they're at nine billion in revenue in under a year because they were great CLI agents before cloud code. You had ADER,

就是你大概会有这么一种感觉:随着时间推移,它们总体上会把事情越弄越糟,让代码库越来越难上手。为了搞清楚这是为什么,我想把镜头拉回到第一个伟大的 coding agent。为什么 Claude Code 能在不到一年里从零做到 40 亿——我记得他们现在营收已经到 90 亿了?在 Claude Code 之前,也有很棒的 CLI agent。有 Aider,


[10:24]

you had CodeBuff. There was a bunch of tools in this category. They had all the same tools, read, write, edit, grab, bash. Um, so what was the difference? The difference was was that the this was the first time that a model lab trained a model against the harness that they were going to distribute it to users in.

有 Codebuff。这个品类里当时有一大堆工具。它们用的工具都一样——read、write、edit、grep、bash。嗯,那区别到底在哪儿?区别在于,这是头一回有一家模型实验室,针对他们即将分发给用户的那个 harness 去训练模型。


[10:40]

Um, and it got really really good. This is just some of the tools, but it got really really good at calling these sorts of tools in an agentic loop. In fact, the OpenAI team did a talk in November about basically if you are a uh harness builder and you don't own the model weights and you can't RL the model in your harness, you will always be at a disadvantage compared to somebody who owns both the model and the harness. Um,

然后它变得非常非常厉害。这只是其中一部分工具,但它在 agentic loop 里调用这类工具变得非常非常在行。其实 OpenAI 团队十一月做过一个演讲,大意是:如果你是一个 harness 的构建者,但你并不拥有模型权重,也没法在你的 harness 里对模型做 RL,那你相比那些同时拥有模型和 harness 的人,就永远处于劣势。嗯,


[11:05]

and I'm going to site a couple slides from my buddy Calvin French Owen who was a MTS on Codeex during the initial launch. Um, but LM are just next token predictors. Uh, this is a slide from over a year ago where basically as you're doing your agentic loop, context window goes in, next step comes out.

接下来我要引用我朋友 Calvin French Owen 的几张幻灯片,他在 Codex 最初发布的时候是那边的 MTS(技术团队成员)。嗯,不过 LLM 本质上就是 next token 预测器。这是一张一年多以前的幻灯片,基本意思是,当你在跑 agentic loop 的时候,context window 输进去,下一步就输出来。


[11:20]

And, uh, we're going to try to do this. I haven't actually timed this, but we're going to see if we can do coding agent reinforcement learning in 60 seconds. So, what we're going to do if we want to train a model to get better at tool calling, better at solving software problems. We're going to generate a bunch of we're going to give it a problem and we're going to generate a bunch of traces, try to solve the problem a bunch of different times.

然后,呃,我们来试试看。我其实没掐过表,但我们来看看能不能在 60 秒内讲完 coding agent 的 reinforcement learning。所以,如果我们想训练一个模型,让它更擅长 tool calling、更擅长解决软件问题,我们要做的是:我们会生成一堆——我们会给它一个问题,然后生成一堆 trace,把这个问题反复尝试解很多遍。


[11:36]

We're going to score them all on correctness and did the test pass and all this stuff. Uh, and then we're going to reinforce. We're going to make the bad behavior less likely, and we're going to update the weights to make the good behavior more likely. Um, this one of the classic ones here is Swebench Multilingual. Uh, they're about 15-minute tasks. They're from open source repos like Reddus, JQ, and Django and all this stuff. And they have binary

我们会根据正确性、test 有没有通过之类的标准,把它们全部打分。呃,然后我们做 reinforce(强化):让不好的行为出现概率变小,更新权重让好的行为出现概率变大。嗯,这里一个经典的例子是 SWE-bench Multilingual。呃,这些大概是 15 分钟量级的任务,来自像 Redis、jq、Django 这些开源仓库。它们的 reward 是二元的——


[11:56]

one or zero rewards on did you fix the problem you were trying to fix and did you do it without breaking anything else. Um, and we look at actually a real problem from one of these benchmarks. This is Fastlane, which is a Ruby project. Um, basically there was some issue where we weren't checking for nil and we have a stack trace blow up because you have a null pointer exception. And in this um in this benchmark you have a base commit that

reward 要么是 1 要么是 0,看你有没有修好你要修的那个问题,以及有没有在修的过程中没搞坏别的东西。嗯,我们实际来看这些 benchmark 里的一个真实问题。这是 Fastlane,一个 Ruby 项目。嗯,基本上就是有个地方我们没检查 nil,结果因为空指针异常,stack trace 直接炸了。在这个 benchmark 里,你会有一个 base commit,


[12:16]

we're going to check out before the issue was solved by a human in the past. We're going to give it a test patch that says here's what the behavior should be afterwards. We have a golden patch. Both these are hidden from the model. Uh, and so we have the agent go try to solve the problem. We store its patch. we undo all the changes it made to any test files because I'm sure you've seen models comment out tests just to get things

我们会 check out 到过去某个人类解决这个问题之前的那个状态。我们会给它一个 test patch,告诉它修好之后应该表现出什么样的行为。我们还有一个 golden patch。这两个都对模型隐藏起来。呃,然后我们让 agent 去尝试解决这个问题,把它的 patch 存下来。我们会撤销它对任何 test 文件所做的全部改动,因为我相信你们肯定见过,模型为了让东西跑通,直接把 test 注释掉——


[12:37]

working. And then um we're going to apply our golden test patch. Uh and then we're going to run the test old test and did the new test pass. And if they both pass then uh then we get the reward otherwise we don't. Um and so models are trying to get the test to pass. There's no way in this system that we can penalize it for poor program design or for eroding the maintainability of our systems. That's why we get things like

就为了跑通。然后,嗯,我们会应用我们的 golden test patch。呃,接着我们会跑那些 test——旧的 test,以及新的 test 有没有通过。如果两个都过了,那我们就给 reward,否则就不给。嗯,所以模型是在努力让 test 通过。在这个体系里,我们根本没办法因为糟糕的程序设计、或者因为它侵蚀了系统的可维护性,去惩罚它。这就是为什么我们会看到诸如


[12:58]

this try catches around things that probably don't need a try catch or things like this. I think VBOV gave us this example earlier of casting things to other things just so the model can just just it just wants to get the test to pass. Um and so if you can't verify the uh maintainability of the code, it gets way harder to train on this stuff.

这样的东西:在一些其实根本不需要 try catch 的地方套上 try catch,诸如此类。我记得 VBOV 之前给过我们一个例子,就是把一个东西强转成另一个东西,纯粹是为了让模型能——它就是想让 test 通过而已。嗯,所以如果你没法验证代码的可维护性,那在这上面做训练就会难得多。


[13:16]

Um, so you remember this picture. Verifying code quality and maintainability is orders of magnitude harder than the code runs and the test pass because the cost function of bad architecture is measured in months and years if you have a coding episode and then you only find out months later that like somebody vied this a little bit too hard. It's really hard to propagate that reward signal back across the gap. And now the frontier is getting better

嗯,还记得这张图吧。验证代码质量和可维护性,比验证代码能不能跑、test 能不能过要难上好几个数量级,因为糟糕 architecture 的代价函数是以月、以年来衡量的——如果你有一次 coding 的过程,几个月之后你才发现,有人这块 vibe(凭感觉写)得有点太狠了。要把那个 reward 信号跨越这么长的时间间隔传回去,是非常难的。现在前沿虽然在慢慢变好——


[13:40]

slowly. And since I know someone's going to be in the YouTube comments about this, yes, I know benchmarks and verifiers are different and they actually have to be separate data sets, but they're shaped the same and the the structure of these benchmarks is directionally correct. So, we're going to look at these as like what is the future of evaluating code maintainability. Um, there's a really cool one called SWE Marathon from

慢慢地变好。我知道肯定会有人在 YouTube 评论区里揪这一点,是的,我知道 benchmark 和 verifier 是不一样的,它们实际上必须是分开的数据集,但它们的形态是一样的,而且这些 benchmark 的结构在方向上是对的。所以我们就把它们当作评估代码可维护性的未来来看。嗯,有一个特别酷的东西叫 SWE Marathon,来自


[13:58]

Abundant AI where they do like 400hour tasks of like clone all of Microsoft Excel, every single feature. Uh, and they have some sophisticated reward channel stuff. uh deep suite from data curve is also like large tasks on oss repos that are not actually in the training set because they were never actually built in the real world. Uh and then you have frontier code from cognition um which is multipr tasks.

Abundant AI,他们做的是那种 400 小时的任务,比如把整个 Microsoft Excel 克隆出来,每一个功能都要。呃,他们有一些相当精巧的 reward channel 机制。呃,Datacurve 的 DeepSuite 也是在 oss 仓库上做大型任务,而且这些任务其实不在训练集里,因为它们在现实世界里从来就没被真正做出来过。呃,然后还有 Cognition 的 Frontier Code,是那种 multi-PR(多个 PR)的任务。


[14:20]

They do interesting things like hey if the model writes tests that don't fail on the pre- patch code then it gets penalized and we have a judge model that says okay uh did this follow all of our code quality rules. Um, so we're getting better, but I think models judging quality can only go so far. Uh, because if the new model, if the model knew what good code looks like, it would probably write it in the first place. Uh, and

他们会做一些很有意思的事情,比如说,如果模型写的 test 在打 patch 之前的代码上根本不会失败,那它就会被扣分;我们还有一个 judge model 来判断,好,这段代码有没有遵守我们所有的代码质量规则。嗯,所以我们是在进步,但我觉得让模型来评判质量,能走的路是有限的。因为如果那个新模型——如果模型真知道好代码长什么样,那它一开始大概就直接写出来了。呃,还有——


[14:41]

review agents and throwing more tokens at the problem, it can raise the floor. Um, but we're still constrained by what we can teach during RL. Um, and so I will I will posit that for now we're stuck reading the code. Uh, but we can still move pretty fast. And of course there's a world where this is solved uh in the future. And if you want to just keep yoloing prompts until you get to GPT7, you don't have to think about

review agent,以及往这个问题上砸更多的 token,是可以抬高下限的。嗯,但我们仍然受限于在 RL 阶段能教给模型的东西。嗯,所以我要放话:至少现在,我们还是逃不掉要去读代码。呃,但我们仍然可以跑得相当快。当然,也存在一种未来,这个问题被解决了。如果你就是想一直 yolo 各种 prompt,一路等到 GPT7,那你根本不用去操心


[15:00]

this. By all means, please. Uh, but bitter lesson be damned. We've got some problems to solve. So, let's engineer our way out of this. Um, so turning the lights back on, we're going to put the code review back. Uh, we're going to embrace this approach of like how do we plan up front to reduce the chance that we have a long or uh difficult review process. We're going to find leverage.

这些。真的,尽管去。呃,但去他的 bitter lesson(苦涩的教训)。我们还有些问题要解决呢。所以,让我们用工程的办法把自己从这个困境里救出来。嗯,那把灯重新打开,我们要把 code review 放回来。呃,我们要拥抱这样一种思路:怎么在前期就做好规划,来降低我们碰上一个漫长或者困难的 review 过程的概率。我们要去寻找杠杆。


[15:19]

We're going to use AI to help with this. Um, the first thing we're going to do is we're going to do some sort of product review. understanding what problem we're solving, what's the desired behavior, maybe looking at mockups. Here's a product review I was working on yesterday with a mockup of a new feature. Um, once we have our product review, we're going to, and by the way,

我们会用 AI 来帮忙做这件事。嗯,我们要做的第一件事是做某种产品评审(product review):搞清楚我们在解决什么问题、期望的行为是什么,可能还会看看 mockup。这是我昨天在做的一个产品评审,里面有一个新功能的 mockup。嗯,一旦我们有了产品评审,我们就要——顺便说一句,


[15:35]

we don't small stuff still just goes straight to the agent. Um, but once we have the product review, we're going to also do architecture, system architecture. A lot of people have been doing this for a while. Component contracts, data models, constraints. Um,

小的东西我们不会走这一套,还是直接丢给 agent。嗯,但一旦我们有了产品评审,我们还会做 architecture,系统 architecture。很多人已经这么做有一阵子了。组件契约、数据模型、各种约束。嗯,


[15:46]

this is an example of a doc that we build to understand how these systems are going to fit together and what's like the highle picture of it. From there, we do something uh that I think is really underemphasized in uh agentic coding these days, which is program design. I think people assume that once you get the architecture right, the model can just cook. Um but we I we often look into the types and the method signatures, the program layout and the

这是我们写的一份文档的例子,用来理解这些系统之间要怎么拼合到一起、它的高层全景是什么样的。在这之后,我们会做一件我觉得在如今的 agentic coding 里被严重低估的事,那就是程序设计(program design)。我觉得大家都假设,只要 architecture 定对了,模型就可以尽情发挥了。嗯,但我们——我们经常会深入去看类型、方法签名、程序的布局,以及


[16:08]

call stacks. And so here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at is how are we actually going to lay this stuff out and how are these systems going to interact? Uh Dylan Mullroy from Cloudflare talks a lot about how he's using these call graphs as part of his planning process. I think this is exactly right. Um and then once we've done the pro program design, we can do

调用栈(call stack)。这里有一些例子。我估计你们看不清这一张,但这就是我们所处的抽象层级——我们究竟要怎么把这些东西铺排开、这些系统之间又要怎么交互?呃,Cloudflare 的 Dylan Mullroy 经常谈到他是怎么把这些 call graph 用作他规划流程的一部分的。我觉得这完全正确。嗯,然后一旦我们做完了程序设计,我们就可以做


[16:27]

this thing called vertical slices. Um which is the order of implementation, multi-reo coordination. How are we going to build this across our entire system and how are we going to check it along the way? I've talked a little bit about how models have horizontal plans. I won't go too deep into it. If you want to learn more about this, you can go watch our talk from AI engineer Miami.

一件叫做垂直切片(vertical slices)的事。嗯,也就是实现的先后顺序、多仓库(multi-repo)之间的协调。我们要怎么在整个系统范围内把它建起来,又要怎么在过程中一路验证它?我之前稍微讲过模型是怎么做横向规划的。我就不深入展开了。如果你想更多了解这方面,可以去看我们在 AI Engineer Miami 的那场演讲。


[16:45]

um couple shots of a dock like this going through the tests and the steps in between each phase. Um the main idea here is 30 minutes over here in pre-planning and alignment can save you hours in review and so it's actually feasible to still read every line of code. Um we're skip this part. Uh basically the the summary here is like you don't have too many PRs. If you're drowning in PRs, you actually have too many bad PRs.

嗯,几张这样一份文档的截图,会走一遍那些 test,以及每个阶段之间的步骤。嗯,这里的核心思想是:在前期规划和对齐上花的这 30 分钟,可以帮你在 review 上省下好几个小时,所以要逐行读完每一行代码其实是可行的。嗯,这部分我们跳过。呃,基本上这里的总结就是:你的 PR 并不是太多了。如果你被 PR 淹没了,其实是你有太多糟糕的 PR。


[17:12]

Um because a good PR is a joy to to review. It's it's you're just reading through like, "Yep, this is great. This is what we discussed. This is what we talked about." Um but even if a PR needs 20% rework, which is generous for a lot of AI AI vibecoded slop um it's an it's an emotional and intellectual burden on both the reviewer and the submitter. Um and so if you use model assisted planning and alignment, your alignment

嗯,因为一个好的 PR 是一种享受,review 起来很愉快。你就是一路读下去:「嗯,这很棒。这就是我们讨论过的。这就是我们说好的。」嗯,但哪怕一个 PR 只需要返工 20%——对很多 AI 那种 vibe coding 出来的垃圾来说,20% 已经算很客气了——它对 review 的人和提交的人来说,都是一种情绪上和智力上的负担。嗯,所以如果你用模型辅助的规划和对齐,你的对齐


[17:35]

is shorter because you use AI to get all the information at once. your code review is faster because you aligned up front and your coding is faster because AI did it. And so you're now you're actually really moving faster, but you're still reading everything and you're still owning the code. So closing advice, um, it's easy to hear all this and be a little bummed out. Uh, I really like the world where we just yolo everything and we can just like not have

就会更短,因为你用 AI 一次性拿到了所有信息。你的 code review 更快,因为你在前期就对齐了;你的编码也更快,因为是 AI 写的。所以现在你其实是真的在加速,但你依然在读每一样东西,依然对代码负责。那么,作为收尾的建议,嗯,听完这些,很容易让人有点沮丧。呃,我其实很喜欢那种什么都 yolo、然后我们可以再也不用去


[17:58]

to ever read code ever again. But, uh, we're engineers and these are just constraints and models are good at certain things and they're not good at other things. And so go figure out how to solve problems given a set of constraints. Uh use loops. They're great. Go solve hard problems. Seek leverage. Um if you want to help with this, um we're building human layer.

读代码的世界。但是,呃,我们是工程师,这些不过是一些约束条件而已,模型在某些事情上很擅长,在另一些事情上不擅长。所以去搞清楚,在一组给定的约束下该怎么解决问题吧。呃,用好这些循环(loop),它们很棒。去解决难题。去寻找杠杆。嗯,如果你想在这方面出一份力,嗯,我们正在做 HumanLayer。


[18:18]

Human layer is an AI IDE and collaboration platform. It's building blocks for your software factory. Um and soon to be better verifiers for software quality. Um we've got sort of a Figma for cloud code and codec style collaborative workspace. It walks you through the workflows for doing this sort of work. And uh we are talking to design partners. We are hiring founding engineers here in San Francisco. And uh these slides are live. You can go get

HumanLayer 是一个 AI IDE 和协作平台。它是你软件工厂的构建模块。嗯,很快还会有更好的、面向软件质量的 verifier。嗯,我们做了一个类似「Claude Code 和 Codex 版 Figma」的协作式工作空间。它会带着你一步步走完做这类工作的工作流。呃,我们正在和 design partner(设计合作伙伴)洽谈。我们正在旧金山这里招募创始工程师。呃,这些幻灯片是实时公开的,你现在就可以去拿


[18:44]

them right now. You can try human layer at human layer.com. Uh it's free for small teams. Go solve hard problems in complex codebases. Thank you all for your energy. [applause] >> [music]

到它们。你可以在 humanlayer.com 上试用 HumanLayer。呃,它对小团队是免费的。去在复杂的代码库里解决难题吧。谢谢大家的热情。[掌声] >> [音乐]