Realtime multiplayer, automation, and you! — Idan Gazit, GitHub
频道: AI Engineer
视频: https://www.youtube.com/watch?v=iQ5xldZ9StU
原文语言: en
统计: 共 17 轮
[0:01]
[music]
[音乐]
[0:13]
going to start 1 minute early, which gives me one extra minute. And then anybody who came on time is going to miss the super enthralling introduction. Hi, my name is Eitan. Nice to meet you all. Uh I lead GitHub Next, which is the Labs team of GitHub. I like to call us the Department of Fool Around and Find Out, but I usually don't say the word fool. We're the team that created Copilot uh and pioneered a ton of areas since then, right? Uh spec-based programming, natural language to app, lots more. Not everything uh that we do turns into a finished product. Our job is to sort of explore the future and scout it out. Um but our job is to reach for the GitHub that's going to be next year. Maybe not tomorrow's GitHub, but uh the tools that we're all going to use to make software a year from now, 2 years from now. That's pretty hard, cuz my crystal ball barely works into uh next week. Uh and we're really fortunate that we get to do most of our work in the open, so you can check out githubnext.com and our socials, which we occasionally remember to post stuff to. And what we do isn't really research, right? Because the only way to know what's going to be good uh is to make stuff. So, we make a lot of stuff, and the hard part about being an undirected research team is always the question of what's worth our time.
我打算提前一分钟开始,这样就能多出一分钟时间。当然,准时来的人就会错过这段无比精彩的开场介绍了。大家好,我叫 Idan,很高兴见到各位。我负责 GitHub Next,也就是 GitHub 的实验室团队。我喜欢管我们叫「瞎折腾看看会出什么事部」,不过那个词我一般不会说出口。Copilot 就是我们这个团队做出来的,从那之后我们还开拓了一大堆方向——基于 spec 的编程、用自然语言直接生成应用,还有很多。我们做的东西并不是每一样都会变成正式产品,我们的活儿是去探索未来、去探路。我们要够到的是「明年的 GitHub」——也许不是明天的 GitHub,而是一年后、两年后我们所有人用来做软件的那套工具。这挺难的,因为我的水晶球连下周都看不太清。我们也很幸运,大部分工作都能公开地做,所以你们可以去 githubnext.com 看看,也可以关注我们的社交账号——我们偶尔会想起来在上面发点东西。我们做的事情其实算不上研究,因为想知道什么东西会成,唯一的办法就是把它做出来。所以我们做了非常多的东西。而当一个没有既定方向的研究团队,最难的永远是那个问题:什么才值得我们花时间。
[1:34]
Even if you're a token billionaire, uh even if you have 10 terminals running Fable night and day, then opportunity cost is is still there. It's everything. Uh so, if in a in a world where the uh marginal cost of a line of code is approaching zero, uh and AI can help us to think and to make, what do we make, right? How do we even choose what's important uh when the market is super noisy and the tech changes every week? And this isn't even really a next problem anymore. This is an all of us problem now. We're all labs teams now. And the way that next thinks about this stuff is to look for durable themes. Things that will be true no matter what the technology that exists tomorrow. And I think that the theme of this moment is very much an evergreen one, right? It's AI started with a surge of personal productivity, right? The LLMs completed what I type and the agents go fetch me the thing that I need. And now I have many agents helping me to parallelize myself. But the greatest value doesn't come from multiplying me into more me. It comes from enabling groups of people to do more. That's always been true. And we're thinking about how to accomplish that through two lenses. Every industrial revolution came about through automation, right? It's funny to think about our giant software industry as being pre-industrial, but on some level it is because until now the only automations that we had were heuristics like make sure there's a semicolon at the end of every line.
就算你是个 token 亿万富翁,就算你开着 10 个终端让 Fable 日夜不停地跑,机会成本依然在那儿——而且它就是一切。所以,在一个代码的边际成本趋近于零、AI 又能帮我们思考和创造的世界里,我们到底该做什么?在市场噪音这么大、技术每周都在变的情况下,我们怎么判断什么才重要?而且这已经不只是 Next 团队的问题了,现在这是我们所有人的问题。我们每个人都成了实验室团队。Next 思考这类问题的方式,是去找那些「持久的主题」——不管明天出现什么技术,它都依然成立的东西。我觉得当下这个阶段的主题恰恰是个常青主题:AI 的第一波是个人生产力的暴涨——LLM 帮我把打的字补全,agent 帮我去把需要的东西取回来。现在我有一大堆 agent 帮我把自己并行化。但最大的价值并不来自把「我」复制成更多个「我」,而是来自让一群人能做成更多事。这一点从来都成立。我们从两个角度思考怎么实现它。每一次工业革命都是靠自动化发生的。把我们这个庞大的软件行业叫作「前工业时代」听起来有点好笑,但某种程度上确实如此——因为在此之前,我们拥有的自动化只有一些死规则,比如「检查每行末尾有没有分号」。
[3:11]
But now AI can help us to automate things that require some amount of basic judgment and intelligence. And there's no magic trick to making great software, right? It costs time. And we can buy that time by automating away the things that we used to need to do manually. Like the more we automate, the more time we have to spend on craft or on our product or on making it really good or on features, right? Either you hire more people or you automate away part of what your people are currently doing in order to spend that time. And at the same time, how are we going to work together, right? How does collaboration look like in the future? Whoops. Oh well, sorry about that. Yesterday uh, Jeffrey Lit talked about understanding being the bottleneck, and that's very true at a me level. Uh, but my personal understanding was never sufficient for shipping code inside a team, right? Our understanding at an us level can only happen at the end of the process. Um, sorry. Uh, uh, I it can't only happen at the end of the process, uh, when the process happens so much faster. So, going faster means that a small misalignment, uh, can snowball into a ton of wasted work, uh, and that work costs tokens, and tokens cost real money now, so, uh, on top of the time that you're mis spending.
但现在 AI 能帮我们把那些需要一点基本判断力和智能的事情也自动化掉。做出好软件没有什么魔法,它就是要花时间。而我们可以把原本必须手动做的事情自动化掉,用这种方式把时间买回来。自动化得越多,我们能花在手艺上、花在产品上、花在把东西真正做好、花在功能上的时间就越多。要么你多招人,要么你把手下人现在正在做的一部分事情自动化掉,才能腾出这些时间。与此同时还有另一个问题:我们将来要怎么一起工作?未来的协作会是什么样子?哎哟,翻页出问题了,抱歉。昨天 Geoffrey Litt 讲到「理解才是瓶颈」,这话在「我」这个层面上非常对。但我个人的理解,从来都不足以让代码在一个团队里发布出去。「我们」这个层面上的理解——抱歉,我想说的是,它不能只发生在流程的最后一步,尤其当整个流程跑得这么快的时候。跑得越快就意味着,一个小小的错位可能滚雪球滚成一大堆白费的工作,而这些工作要烧 token,token 现在是真金白银——这还没算上你浪费掉的时间。
[4:30]
So, today I'll give you a quick tour of two prototypes that we're working on at GitHub Next in each of these themes. Agentic Workflows is our take Why is that not there? Oh, I had to click again. Uh, Agentic Workflows is our take on how automations should work in an agentic world, and Aces a prototype that explores what real-time multiplayer software development looks like. So, I'll start by showing off Agentic Workflows, and it requires me doing this. Okay, cool. Uh, this is my personal website, not that interesting. I'm showing it to you. This is like Chekhov's gun, we're going to see it again later. Um, and my personal website is built with this framework called Astro. Astro is a great web framework. The greatest part about it is that they release like 50 things a month, which means that I'm constantly on the upgrade treadmill, and there's a great GitHub product called Dependabot, which notifies me when my stuff is out of date. Um, but the problem is is that when I do these upgrades, I frequently need to make code changes. So, what I really want is a kind of super Dependabot that's always there, automatically looking in the background at my dependencies and figuring out how to upgrade me, including the code changes, the breaking changes. Um, and because I'm lazy, and I like not doing work, um, I used Copilot, uh, to create an agentic workflow, and there's this magic line up top where I supply effectively a skill saying like, "Hey, create an agentic workflow. Here's a document that tells you everything you need to know about that.
所以今天我带大家快速过一遍我们 GitHub Next 在这两个主题下做的两个原型。Agentic Workflows 是我们对……咦,怎么没出来?哦,得再点一下。Agentic Workflows 是我们给出的答案:在一个 agentic 的世界里,自动化应该怎么运作。而 Aces 是另一个原型,探索的是实时多人协作的软件开发长什么样。我先演示 Agentic Workflows,得先这么操作一下。好,可以了。这是我的个人网站,没什么意思,就给大家看一眼——这就像契诃夫的枪,后面还会再出现。我的个人网站是用一个叫 Astro 的框架搭的。Astro 是个很棒的 Web 框架,它最棒的一点是每个月能发布大概 50 个东西,这意味着我永远在升级的跑步机上跑。GitHub 有个很好的产品叫 Dependabot,我的依赖过期了它会通知我。但问题是,每次做这些升级,我经常还得改代码。所以我真正想要的,是一个一直待在后台的「超级 Dependabot」,自动盯着我的依赖,自己琢磨怎么帮我升上去,包括那些代码改动、破坏性变更。而因为我懒,我喜欢不干活,所以我用 Copilot 创建了一个 agentic workflow。最上面有神奇的一行,我在那儿基本上就是塞给它一个 skill,说:「嘿,帮我创建一个 agentic workflow,这份文档里写了你需要知道的一切。」
[5:50]
And then what comes below that is something a lot like a Slack message that I'd send to a junior developer on my team. Like, every day I want you to check if there's a new release, look at the change log, look at the docs, come up with a plan for the upgrade, and then create a PR with the thing and here's the links to the docs. Right, this is like a message that I would send to somebody on my team, go write a playbook. And when I went and created this, it did go and create a playbook. In fact, that's what agentic workflows kind of look like. They look like markdown documents. Like, if GitHub Actions and Copilot had a baby together, and it ran on markdown, this is what it is. So, what does this agentic workflow look like? Well, you know, it's an upgrade checker, it's got my tasks, step one, check for new releases. Again, because it sees my codebase, it was able to infer what it even needs to check and it actually found these specific dependencies. Review the change log and the upgrade guide, apply the upgrade, and then create a pull request. Right, I didn't ask for any of this that explicitly, but it turns out that Copilot is pretty good at sussing out my little three-line message into a full playbook. And then at the top, I've got this special section, this is the what we're calling No, don't collapse it. Oh, man.
再往下的内容,就很像我发给团队里某个初级开发的一条 Slack 消息:每天帮我看看有没有新版本发布,去读 changelog,去看文档,想一个升级方案出来,然后开一个 PR 把东西提上来,这儿是文档的链接。对,就是这么一条我会发给同事的消息——去写个 playbook 吧。我这么一说,它还真就去写了一个 playbook。事实上,agentic workflow 大致就长这样:它们看起来就是 Markdown 文档。就像 GitHub Actions 和 Copilot 生了个孩子,而这孩子跑在 Markdown 上——差不多就是这么回事。那这个 agentic workflow 长什么样呢?它是个升级检查器,写着我的任务:第一步,检查有没有新版本发布。同样因为它能看到我的代码库,它自己推断出该去查什么,而且真的找到了这些具体的依赖。然后是查看 changelog 和升级指南、执行升级,最后创建一个 pull request。这些我一条都没有明说,但事实证明,Copilot 挺擅长把我那三行小消息挖出意思、扩写成一整份 playbook。然后在最上面,还有一段特殊的内容,就是我们叫作——别折叠啊,天哪。
[7:08]
Scrolling is wonky when you zoom in. This YAML front matter. This is where we stick the guardrails cuz if we're going to be not supervising agents doing things, then we're going to need much stronger guardrails around what they're allowed to do, what they're allowed to read, what they're allowed to write. And where are we going to specify that? And it's not enough to just prompt the agent and be like, "Listen, bro, I don't want you to buy Bitcoin for me ever." That's not enough cuz somebody else can prompt inject the agent and take it in a direction that you don't expect. So, any of the guardrails, if you're prompting the guardrails at the agent, you're effectively letting the fox loose in the henhouse. It's not actually a guard rail. Um, so here, uh, you can see that I'm specifying deterministically like my permissions are read all, what tools am I allowed to use, uh, what network, uh, requests is it allowed to make? It's not allowed to just go to bitcoin.com or whatever. Uh, in fact, it's only allowed to go to some specified set of default websites, the NPM ecosystem cuz it's got to check for like, you know, what's new, GitHub, and of course the Astro docs which I specified in my original prompt. Uh, and I've got this block called safe outputs which is basically saying these are the only things that the agent is allowed to write. And so I'm saying in this case the agent will is allowed to create pull request. Pull request single uh, because I don't want the agent to get prompt injected to create 500 pull requests. That would be a denial of service. Um, or and this is the other thing, I explicitly said you're allowed to do nothing, right? Which sounds silly, but it actually matters because in a world where I have lots of automations, the last thing I want is noise. I don't want the agents denial of servicing me.
放大之后滚动就很别扭。这段 YAML front matter,就是我们放护栏的地方。因为如果我们不打算全程盯着 agent 干活,那就需要给它们能做什么、能读什么、能写什么,加上强得多的护栏。这些东西写在哪儿呢?光靠给 agent 写提示词是不够的,比如「兄弟听着,你永远不许拿我的钱去买比特币」——这不够,因为别人可以通过 prompt injection 把 agent 带到你意想不到的方向去。所以,如果你的护栏是靠提示词喂给 agent 的,那基本上等于把狐狸放进鸡窝,它根本算不上护栏。所以在这里你可以看到,我是用确定性的方式规定的:我的权限是只读,允许用哪些工具,允许发哪些网络请求。它不能想去 bitcoin.com 就去——事实上,它只能访问一组指定的默认网站:NPM 生态(因为它得去查有什么新东西)、GitHub,当然还有我在原始提示里写的 Astro 文档。我还有一块叫 safe outputs 的配置,意思是:这些是 agent 唯一被允许写入的东西。所以在这个例子里,我写的是允许 agent 创建 pull request——单数的 pull request,因为我不希望 agent 被 prompt injection 之后一口气开 500 个 PR,那就成拒绝服务攻击了。还有另一件事,我明确写了「你也可以什么都不做」。这听起来有点傻,但其实很重要,因为在一个我有一大堆自动化的世界里,我最不想要的就是噪音。我可不想被这些 agent 反过来拒绝服务了。
[8:48]
So, okay, I've created this and I've run it and this is actually my actual automation on my actual personal website. I didn't ask for any of this, but it did a pretty good job of like saying, "Hey, here's the highlights of what you get from going from this version that you're currently on to the version that is the target, right? It's read all of the release notes in the middle. This is normally what I would do as a human. Uh, and it's built me like, you know, sort of like a tailored description. It's figured out there's no breaking changes. It's actually verified this by running and building my project. And because I happen to have this deployed to Cloudflare, um, or whatever, anything with preview deploys, I can click that open and see that nothing has changed in my website, which is exactly what I want, right? Like it's done the upgrade and I see that it still works exactly as it did before. But this was like a minor point release. That doesn't really count. Let's look at a major upgrading change. And actually, I'm lucky that Astro just released Astro 7 because this is actually jumping two major revisions from five to seven. And so, now it's saying like, "Okay, Astro 7 has brought me all of these things. And Astro 6 would have brought me all of that stuff, but I neglected to do the upgrade so I could have a cool demo for you all."
好,我把它建出来跑了一遍——这是我个人网站上真实在跑的自动化。这些内容我一句都没要求过,但它做得相当不错:「嘿,从你现在这个版本升到目标版本,你能拿到的亮点是这些。」它把中间所有的 release notes 都读了一遍——正常情况下这活儿得我这个人类来干。它还给我写了一份量身定制的说明,判断出没有破坏性变更,而且真的跑了一遍构建来验证。因为我这个站正好部署在 Cloudflare 上——其实任何支持预览部署的平台都行——我可以点开那个链接,看到网站没有任何变化,这正是我要的:升级做完了,站还跟原来一模一样。不过这只是个小版本更新,不太算数。我们来看一次大版本升级。而且我运气不错,Astro 刚好发布了 Astro 7,所以这次是从 5 一口气跳两个大版本到 7。它现在说:好,Astro 7 给你带来了这些东西,Astro 6 本来会带来那些东西,但我一直没升级,就是为了给大家留一个够酷的 demo。
[9:57]
And it's found all of the code changes that were broken, and it updated them. It also verified that the build runs. And it also highlighted manual steps that things that I would need to do later. And again, you know, if I go down here and I click on this, I can see, "Hey, still works." So, cool. Now, uh it's just markdown. It's easy to iterate on that markdown, right? If you don't like the way that the automation works, just edit the English. It gets recompiled into an Actions workflow. Like, the markdown is the source code. The YAML is like a compiled artifact. You never look at it. Um but, we've also given you a whole library of Agendic workflows for you to use as a starting point to customize. So, an issue triager. Internally, GitHub has actually used this as the basis for like spiking out our own internal issue triager or for like hunting down N+1 queries in our like monolith or all kinds of things. There's a ton of things that are super helpful there. Repo assist. This is actually a swarm of Agendic workflows that work together to help you maintain your project by finding low-hanging fruit, fixing them, identifying tickets that need nudging or feedback that you need from people who have filed issues, whatever. CI doctor.
它把所有被改坏的代码都找了出来,并且更新掉了,还验证了构建能跑通,同时把那些需要我事后自己动手的手工步骤也标了出来。同样,我往下拉、点开这个链接,就能看到:嘿,站还是好的。挺好。而这东西就是一份 Markdown,所以迭代起来非常容易——你不喜欢这个自动化的行为方式,改英文就行。它会被重新编译成一个 Actions workflow。Markdown 才是源代码,YAML 只是编译产物,你根本不用去看它。另外,我们还给大家准备了一整个 agentic workflows 的库,可以拿来当起点自己改。比如一个 issue 分诊器——GitHub 内部其实就是拿它当基础,快速搭出了我们自己的 issue 分诊器,或者去我们那个大单体里揪 N+1 查询,还有各种各样的用法,里面有一堆非常好用的东西。还有 repo assist,它其实是一群 agentic workflow 协同工作,帮你维护项目:找出容易摘的低垂果实并修掉、找出哪些工单需要推一把、或者哪些提了 issue 的人那边你还需要拿到反馈,等等。还有 CI doctor。
[11:13]
How many times have you responded to a busted CI run by just running it again? All of us. Anybody who hasn't raised their hand is lying. Uh a million more. Like, you know, goals, sure. Daily team status and repo status. If I want this to go do like homework on the internet, I can. So, this is not just for engineers, this is also for product managers whose job it is to look at information over here and summarize those tickets over there, right? We can start to get everybody involved in automation. That's how you actually get industrial scale. Uh So, uh that's Agentic Workflows. Um the security guardrails, we have sort of four principles that we believe uh everybody should burn into their brains. Uh defense in depth, one layer is never enough. Uh that was always true. Never trust agents with secrets. If an agent can know a secret, that secret, you need to treat it as if it's already been compromised. Uh because you have no idea whether or not somebody's injected the agent to reveal that secret somewhere else. So, if an agent can see the secret, um it's bad. In Agentic Workflows, the secrets are all kept outside of the agent's jail, and when the agent wants to use the secret to call something, it needs to ask the warden, "Hey, mother may I please go talk to that service?" Uh stage and vet all rights, just so that it's auditable, and log everything, just so that it's auditable. Uh and when we give this to existing projects like the Home Assistant project, which is a huge open-source project, um the first uh Agentic Workflow they built was something that looks at every submitted issue, walks the Python stack trace to figure out if the bug is in first-party code or third-party code, closes the issue if it's not their issue, right? That's something that was not possible before AI, not possible with heuristics, uh but is possible now.
有多少次,你面对一个跑挂的 CI,处理方式就是再跑一遍?我们所有人都这样。没举手的都在撒谎。还有一百万个别的,比如目标追踪之类的,都有。每日团队状态、仓库状态。如果我想让它上网去做点功课,也可以。所以这不只是给工程师用的,也是给产品经理用的——他们的工作就是看这边的信息、把那边的工单总结出来。我们可以开始让所有人都参与到自动化里来,这才是真正做到工业级规模的方式。这就是 Agentic Workflows。关于安全护栏,我们有四条原则,我觉得所有人都该把它们刻进脑子里。第一,纵深防御,一层永远不够——这一点一直都成立。第二,永远别让 agent 碰密钥。如果一个 agent 能知道某个密钥,你就得当这个密钥已经泄露了来处理,因为你根本不知道有没有人注入了这个 agent、让它把密钥透露到别处去。所以只要 agent 能看见密钥,就是坏事。在 Agentic Workflows 里,所有密钥都放在 agent 的「牢房」外面,当 agent 想用密钥去调某个服务时,它得先问看守:「报告,我能去跟那个服务说句话吗?」第三,所有写操作都要先暂存、再审核,这样才可审计。第四,把所有事情都记录下来,同样是为了可审计。而当我们把这套东西交给已有的项目,比如 Home Assistant 这个巨大的开源项目,他们做的第一个 agentic workflow 就是:看每一个提交上来的 issue,顺着 Python 的调用栈判断这个 bug 到底出在第一方代码还是第三方代码里,如果不是他们的问题就直接把 issue 关掉。这件事在 AI 出现之前是做不到的,用死规则也做不到,但现在可以了。
[13:01]
Agentic Workflows is in public preview today. You can go and kick the tires. So, go ahead, go wild. Uh we actually believe that this is going to be a bigger category than interactive AI because automations that run in the background while you sleep, that's the ballgame. Okay, so let's talk about the collaboration piece. So, this is how we've always built software, right? Because the cost of writing code was so high, uh but that's not true anymore. We would plan and review together, but the building part was done alone. Like, you know, illuminated by the light of my monitor, uh, I would build. But now, none of it is alone, right? Planning isn't before, and review isn't after. We iterate on the direction together, and AI takes a step, and then we iterate more in the direction. So, what's an interface that makes sense for that style of development? I'm only slightly trolling, right? Slack was designed to be better, uh, than email for the average office worker. It was never designed for making software or the needs of everyone involved in that. But, what this is good for is surfacing all the facts that are not in code. Anything that's in code, any fact that's in code, the agents can figure out by reading the code. What's left are the things that are not in code, like political considerations. Like, "Hey, if we do it that way, that VP over there is going to vibe with that direction." Or, like, "We should make it purple because that's their favorite color." Or, "We get a really sweet deal, uh, on infrastructure from Azure. Therefore, we should be building on Azure, not on, uh, GCP or AWS. Whatever."
Agentic Workflows 今天已经开放 public preview 了,大家可以去随便折腾,尽管撒欢。我们其实相信,这会是一个比交互式 AI 更大的品类——因为真正决定胜负的,是那些趁你睡觉时在后台自己跑的自动化。
好,接下来聊协作这部分。我们一直以来就是这么做软件的,对吧?因为写代码的成本太高了。但现在不是这样了。以前我们会一起做规划、一起做 review,但真正动手建的那段是一个人干的——就我一个人,对着显示器的光闷头写。可现在,没有哪一步是一个人干的了。规划不再是「之前」的事,review 也不再是「之后」的事。我们一起把方向捋一遍,AI 往前走一步,然后我们再一起把方向往前推一点。
那什么样的界面配得上这种开发方式?我下面这话只是稍微有点找茬啊——Slack 当年的设计目标是「对普通办公室白领来说比邮件更好用」,它从来就不是为做软件、也不是为做软件这件事里牵涉到的所有人设计的。但它真正擅长的,是把那些不在代码里的事实浮出来。凡是代码里有的事实,agent 自己读代码就能搞明白;剩下的就是代码里没有的东西,比如政治层面的考量——「嘿,要是我们那么干,那边那位 VP 会挺认这个方向的。」或者「我们应该做成紫色,因为那是人家最喜欢的颜色。」再或者「我们从 Azure 拿到的基础设施价格特别香,所以应该建在 Azure 上,而不是 GCP 或者 AWS,随便吧。」
[14:30]
But, the biggest win is the same win that we've already seen over and over, right? I don't email Word documents around anymore. I create and collaborate in the same surface, in the same place. This is coming for code, a trillion percent, right? So, let me show you what we have here. Oops. Here we go. I got to find the tab. All right. Uh, this is Ace. Let's switch to the repository. So, Ace looks an awful lot like Slack, right? And over here on the left, I've got sessions, and I can create new ones. And, you know, so far this kind of looks like every other conductor-like product out there. Um, the difference being is that every one of these is not on my machine. In fact, none of this is running on my machine. It's all micro VMs in the cloud. So, every session is just a branch of my repo checked out to a spot in the cloud. Uh, and I can create them and do stuff in them and talk with my teammates. So, like, uh hey, um uh what's your favorite color, right? Uh and meanwhile, I'm going to like install my dependencies. And then when that's done, I'm going to do like uh bond dev. I'm going to run the dev server. Um and here, like, Russ and I are having a discussion, like, are you sure? Maybe uh maybe green is calmer.
但最大的收益,其实还是我们已经反复见过的那个收益:我现在不会再把 Word 文档用邮件来回传了,我在同一个界面里创作,也在同一个地方协作。这件事一定会发生在代码上,一万个百分点地确定。
那我给大家看看我们手上的东西。哎哟,来了,我得先找到那个标签页。好,这个就是 Ace。我们切到这个 repo。你看,Ace 长得特别像 Slack 对吧。左边这里是 sessions,我可以新建。到这儿为止,它跟市面上那些 Conductor 之类的产品看着差不多。区别在于,这里面没有一个是跑在我这台机器上的——事实上这里没有任何东西跑在我本机,全都是云上的 micro VM。所以每个 session 就是我这个 repo 的一个分支,被 check out 到云端的某个地方。我可以创建它们、在里面干活,同时跟队友聊天。比如我问一句「嘿,你最喜欢什么颜色?」,与此同时我这边先把依赖装上;装完之后再跑一下 dev 命令,把 dev server 起起来。然后你看,我跟 Russ 在这儿讨论上了——「你确定吗?也许绿色更让人平静一点。」
[15:49]
Um Oop, nope. I sent that as a terminal command. Good job, me. Um I do not want that as a thing. Great. I'll do it like this. Uh and I can open up my preview. Oops. Give me a preview. I'd like a browser preview. Okay. So, so far, not that different from developing with any sort of like multiplayer tool. And here, I've got this sort of calm hacker news thing. I've just had a whole discussion with my teammate. I don't want to turn around and now like emit those instructions again. Instead, I just want to be like, yo, Ace, do it. Uh and because it sees the entire backscroll of my conversation with my peers, with my team, it's able to act on that uh uh on on that chat history. And if the Wi-Fi was nice, then it would be doing it faster. Um uh you're going to have to trust me on this because I don't have enough time to wait for this. That it's going to just respond to the fact that we had a discussion about colors. And AI is also really good at fishing out that final state. Like, very frequently, what do engineering conversations sound like? They sound like like, hey, we should try it this way. No, wait, I thought of like an edge case. We should actually do it that way. Let's go back to the first idea, right? But instead of me sort of like figuring teasing out that final state from that long conversation, I can just let AI do it and it'll figure it out. So, I don't need to work for the robots. And sometimes we have things that are a lot more um complicated. Like here, I wanted to add selectable time frames to my app. And so, I asked it to make a plan, and that plan comes as a uh uh markdown document. Uh but, this markdown document is not just for me to look at and edit, it's for us to look at and edit together. So, Russ is somewhere uh here in this document, and like, you know, maybe he thinks that we should add an all time, and I'm going to get rid of the today, and here I can again do like
呃,糟糕,我把它当成终端命令发出去了,干得漂亮啊我。我不想要那个东西。好,这样发。然后我可以打开预览……哎哟,给我个预览,我要一个浏览器预览。好。
到这儿为止,跟用任何一个多人协作工具做开发都没多大区别。这里是我做的一个风格很「静」的 Hacker News。我刚跟队友讨论了一整轮,我不想转过头来把那些指令再重新说一遍。我只想说一句:「Ace,照做。」因为它能看到我跟同事、跟团队之间完整的聊天记录,它就能基于这段聊天历史去行动。要是 Wi-Fi 争气一点,它现在早跑完了。这块你们得信我一次,因为我没时间在这儿干等——它会直接照着我们刚才那段关于颜色的讨论去做。
而且 AI 特别擅长从对话里把最终结论捞出来。工程师之间的讨论通常长什么样?「嘿,我们试试这么干。」「等下,我想到一个边界情况,其实应该那么干。」「算了,还是回到第一个方案吧。」与其让我自己从这么长一段对话里把最终状态一点点抠出来,不如让 AI 去干,它自己能想明白。所以我不用反过来伺候机器人。
有时候事情要复杂得多。比如这儿,我想给我的 app 加上可切换的时间范围,于是我让它出一份计划,这份计划是以一个 Markdown 文档的形式给我的。但这个 Markdown 文档不只是给我一个人看、一个人改的,它是给我们一起看、一起改的。Russ 现在就在这个文档里的某个位置,比如他觉得我们应该加一个「全部时间」,那我就把「今天」去掉。然后我在这儿又可以来一句——
[17:39]
uh we've updated the plan, do it. Um Uh uh and it'll just respond to the plan that we've edited together. And as we see now, we're moving to this future where uh more and more of the work that we're doing with AI results in documents like markdown documents in a docs folder that captures sort of the truth, and maybe more and more in the future we're going to be editing those documents as the way that we do development. Like, in order to change something about my application, I'm going to edit a document, and I'm going to tell AI, "Hey, make the document true." So, this shared document editing is not just like, "Oh, a nice to have." Maybe this is actually sort of the uh interface that we like to work in. But, there's also the uh social coding aspect, right? Like, if I'm working with other people on my team. Um remember when that was a thing that was a tagline under the GitHub logo? Um so, uh how can it help me stay up to date with what everybody else on my team is working on? Like, it's not just enough to have like real-time multiplayer, I also want to be ambiently aware of what everybody's going going on about. So, Kristoff is working on Vian tooling. This is actually work that we're doing on Ace, and Maggie wrote this dashboard and hardcoded her name, and so that's why we're looking at Maggie's name. Um and David worked on whatever. All this stuff to help me stay aligned with my team. And when I look to the future, I'm starting to think about how do automations surface themselves in this?
——「计划我们已经更新过了,照做。」它就会照着我们俩一起改出来的这份计划去执行。
你看,我们正在走向这样一个未来:我们跟 AI 一起干的活儿,越来越多地会沉淀成文档,比如 docs 目录里的 Markdown 文档,那里承载的才是「事实」。也许再往后,我们做开发的方式就是去编辑这些文档:要改我这个应用里的某个东西,我就去改一份文档,然后跟 AI 说「把这份文档变成真的」。所以这种共享文档编辑不只是「有了更好」的锦上添花,它可能真的就是我们愿意待着干活的那个界面。
另外还有社交化编程这一面,对吧?就是我跟团队里其他人一起干活的时候。还记得当年 GitHub logo 底下挂着那句 slogan 吗?那它怎么帮我跟上团队里其他人在做什么?光有实时多人协作还不够,我还想在不经意间就知道大家都在忙活什么。比如 Kristoff 在做工具链方面的东西——这其实就是我们开发 Ace 时候的真实工作;Maggie 写了这个 dashboard,还把她自己的名字写死在了里面,所以我们这儿看到的是 Maggie 的名字;David 做了别的什么。所有这些都是为了帮我跟团队保持在一条线上。
再往未来看,我开始琢磨:自动化要怎么在这里露面?
[19:08]
If I want to talk with my automation, uh there's lots of things that I want to do in this kind of interface, like when an agent wants to tap me on the shoulder and ask me a question, um, that I think are very interesting. So, that's a short Ace demo. We're going through this weird inversion of our relationship with the agents. Like, the better that we get at articulating, uh, our goals to the agents, the less they need us. Uh, and as the models get better, they're also good at spotting like underspecified behaviors and then asking us to clarify. Uh, and then whenever they need a pair of hands, they can ask us to be the pair of hands. But, either way, the interfaces now have the ability to support the ability of agents to listen to everything and invoke us when they need it. Which is a little funny to think about. It's maybe like sort of we're coming at it from this side and like open claws coming at it from this side, but like we're landing in sort of a similar spot. And I'll close with this thought. Um, for the past few years, AI has helped me to type. But, if you look at the science of the matter, it's only about 5% of the job. Like, this was a longitudinal study conducted on like 100 developers over thousands of hours. Turns out that the hands-on keyboard typing part is 5% of the time. Now, AI has to help me with the other 95%. Where is the system that I want to touch? How does it work today?
如果我想跟我的自动化对话——在这种界面里有很多我想做的事,比如当一个 agent 想拍拍我肩膀、问我一个问题的时候,我觉得那些场景都非常有意思。
以上就是 Ace 的一个简短 demo。
我们跟 agent 的关系正在经历一次很奇怪的反转:我们越是能把自己的目标讲清楚,它们就越不需要我们。而随着模型变强,它们也越来越擅长发现哪些地方我们没说明白,然后回头让我们澄清;等它们需要一双手的时候,再叫我们去当那双手。但不管从哪边看,现在的界面已经有能力支撑这种模式了——agent 可以听着所有的动静,需要的时候把我们叫过来。想想还挺好笑的。有点像是我们从这一头往里走,OpenClaw 他们从另一头往里走,最后落到了差不多的位置。
最后我用这个想法收尾。过去这几年,AI 帮我的是「打字」。但你要是看研究数据,打字只占这份工作的大概 5%。有一项纵向研究,跟踪了大约 100 名开发者、上千小时,结果发现真正手放在键盘上敲代码的时间只占 5%。那现在,AI 得来帮我搞定剩下那 95%:我想动的那个系统在哪儿?它今天是怎么跑起来的?
[20:33]
What do other people think about like how we could mutate it or should mutate it? When AI can discover anything in my code base, like, how do we How do we help scale up all those other things, right? Like, not just the 5%, which is what all the tools have been helping us to do so far. So, that's Ace and that's a genetic workflows. Uh, please, uh, come by and talk to us. Uh, we have, uh, a booth down in the Microsoft booth because we're a Microsoft company. Uh, and you can find us on the socials and get at next.com. So, if any of this resonates and you're interested in it and you want to give it a shot, ACE is going to be in technical preview hopefully later this month and genetic workflows is already out there for you to kick the tires and we'd love to hear from you and how you want to use this. Thank you so much.
别人怎么看这个系统能怎么改、该不该改?当 AI 能把我代码库里的任何东西都翻出来的时候,我们该怎么把这些「其他的事」一起放大?而不是只盯着那 5%——到目前为止,所有工具帮我们做的都只是这 5%。
这就是 Ace,也是 agentic workflows。欢迎过来跟我们聊聊,我们在楼下微软展台那边有个位置,因为我们本来就是微软旗下的公司。你也可以在各个社交平台上找到我们,网站是 githubnext.com。所以,如果刚才这些说到你心坎上了、你也想试一把:Ace 争取这个月晚些时候进 technical preview,agentic workflows 已经放出来了,随时可以去折腾。我们非常想听到你们的反馈,以及你们打算拿它来干什么。谢谢大家!
[21:18]
[applause]
[掌声]