How to Build a Self-Improving Company with AI
频道: Y Combinator
视频: https://www.youtube.com/watch?v=X_JsIHUfUjc
原文语言: en
统计: 共 23 轮
[0:00]
This is based a little bit off a talk Diana gave. There's a video up over the weekend which is super cool. Um Jack Dorsey was tweeting some stuff like two or three weeks ago that I thought was super cool and I've kind of um stolen a bunch of those ideas and shove them into here. This talk is like pretty conceptual and high level about thinking about how to build companies. So the Roman legions were designed to project power over two continents or something from Rome at the center to like these people on Hadron's wall up in Scotland. And the idea was um this nested hierarchies with consistent spans of control and you had like named individual with spans of control to pass
这次分享多少有点借鉴了 Diana 之前的一次演讲,她那个视频上周末刚发出来,特别酷。还有 Jack Dorsey,大概两三周前在推上发过一些东西,我觉得特别有意思,所以我就把他那一堆点子顺手「偷」过来塞进了这次分享里。这次分享比较概念化、比较高层,讲的是怎么去思考「公司到底该怎么搭建」。你看罗马军团,它当年的设计目的是把权力从罗马这个中心投射到横跨两个大陆之外——一直到苏格兰北边哈德良长城上戍守的那些人身上。它的核心思路就是这种层层嵌套的等级结构,每一层的管辖幅度都一样,每个岗位都有一个具名的人,负责一定范围的管控,把
[0:42]
orders down and send information back up the hierarchy. And if you think about most companies today, they are organized like a Roman legion where human beings are the conduit for information flowing up and down. And so Jack Dorsey's tweet which I thought was great was it's like this underlying assumption that hierarchically organized companies are the are the way that we should be organizing like our economic units of value. And I think AI basically breaks that. If you talk to people a year ago about how AI was useful, they talked about productivity, like co-pilots, making engineers 20% more productive, adding co-pilots to workflows, shipping more software. But I think that is
命令一层层往下传,再把信息一层层往上报。你想想今天绝大多数公司,其实就是按罗马军团这套组织起来的——人,就是信息上下流动的管道。Jack Dorsey 那条推我觉得特别棒,他说的是:我们默认「公司就该按等级制来组织、把它当成创造价值的经济单元」,这其实只是一个潜在假设而已。而我认为 AI 基本上把这个假设给打破了。一年前你跟人聊 AI 有什么用,大家说的都是「提效」——什么 copilot 啦,让工程师效率提升 20% 啦,给工作流加个 copilot、多出点软件啦。但我觉得这其实是
[1:22]
actually a broken way of thinking about AI. That's like Pete had a great blog post. We're basically just like taking the old way of working and adding like a more powerful engine onto it. And instead of that, I think you can reimagine like what a company is and how it acts. And so as Gary's talking like he I genuinely believe can produce more code than an entire engineering team. The thing that's really stuck with me is this idea of like extracting the domain knowledge from your company and defining it as a as like context or a set of skills or whatever you want to call it.
一种很错误的看待 AI 的方式。Pete 写过一篇很棒的博客,意思是我们基本上就是把旧的工作方式原封不动地拿过来,然后给它装了个更强劲的发动机而已。但我觉得,与其这样,你完全可以重新想象「公司是什么、公司该怎么运转」。就像 Gary 那样——我是真心相信他一个人产出的代码,能比一整个工程团队还多。真正让我印象深刻的,是这个想法:把你公司里的领域知识抽取出来,把它定义成一段 context、一套 skills,或者你想怎么叫都行。
[1:55]
But like this idea that there's domain knowledge or business knowledge or like some knowhow that's inside the heads of people and in Slack messages and in emails and in notion. All of this like information together defines how your company works. And if you can make that legible, you suddenly can can move from this hierarchal organization to a sort of intelligent AI powered organization with AI native software. AI isn't the some it's not something you bolt onto the side of a company. It's not like a tool you give to your engineers to make them more productive. But I think you can reimagine what a company is as a set of recursive self-improving AI loops. I think this is really, really, really
就是这个思路——有那么一些领域知识、业务知识、某种 know-how,它们藏在人的脑子里、藏在 Slack 消息里、藏在邮件里、藏在 Notion 里。所有这些信息合在一起,定义了你公司是怎么运转的。如果你能把它变得「可读」「可被机器理解」,你一下子就能从这种等级制组织,迁移到一种基于 AI 原生软件、智能化、AI 驱动的组织形态。AI 不是你往公司侧面贴上去的一个东西,不是你发给工程师、让他们提效的一个工具。我觉得你完全可以把公司重新想象成一组「递归式自我改进的 AI 循环」。我觉得这一点真的真的真的非常
[2:38]
important because when it gets there, I think the company starts to self-improve even when you're sleeping. So, let me give you an example. Diana's talks about this as well. this AI loop. You start with like a sensor layer, which is like that's a fancy word, but really it might be like emails from your customers. Might be support tickets, code changes, people canceling their subscription, product telemetry. It's like sensor data to get information from the outside world. And then a a policy layer, decision layer, like rules about what you can do, what it has to ask a human permission for, what it must log. A tool layer, that's kind of Gary's skills and code. Like the tool layer is Gary's
重要,因为一旦做到这一步,公司在你睡觉的时候都还在自我改进。我举个例子吧,Diana 也讲过这个 AI 循环。你从一个「传感层」开始——这词儿听着挺高大上,但其实它可能就是你客户发来的邮件、可能是工单、代码变更、用户取消订阅、产品埋点数据。它就是「传感数据」,是从外部世界拿信息进来。然后是一个「策略层」「决策层」——就是一套规则,规定 AI 能做什么、哪些必须先问人类拿许可、哪些必须记录在案。再然后是「工具层」,这差不多就是 Gary 的那些 skills 和代码。工具层就是 Gary 的
[3:19]
code. It's basically deterministic APIs, things like query my database or look at my calendar. Um, a set of tools that the the AI can call a quality gate like that might be evalistic checks, safety filters, human review for high-risk stuff. and then a learning mechanism. It's like your system interacts with the real world, picks up where it doesn't work, and loops back into the top again. And if you can run every single step of that without human intervention, without with minimal human intervention, your system gets better and better and better while you're sleeping. And I can give you actual examples of this that are live right now. We started with an agent that you can ask and it it has
代码,基本上就是一堆确定性的 API——比如查我的数据库、看我的日历,就是这么一组 AI 可以调用的工具。再来是「质量闸门」,可能是 eval 式的检查、安全过滤、对高风险动作做人工审核。最后是一个「学习机制」——你的系统跟真实世界打交道,发现哪儿不灵了,再绕回循环的顶端重新来一遍。如果你能让这一整套的每一步都不需要人工介入、或者只需要极少的人工介入,那你的系统就会在你睡觉的时候越变越好、越变越好。我可以给你举几个现在就真实跑着的例子。我们一开始做了个 agent,你可以问它问题,它有
[3:58]
deterministic tools to query our database. Pretty simple, like when did I last have office hours with this company? Then it got a little bit smarter which was like for this company I'm doing offices hours with right now they need introductions for anyone in petrochemicals or something and it could query the database in different ways and use rag and all sorts of stuff to like come up with five relevant founders for you to meet. But again this is like this is a sidekick right this is an agent this is like the old this is last year's version of how AI is making me better as a group partner. It's making me 20 or 30% more effective. The aha moment for me came when we put a monitoring agent
一组确定性工具去查我们的数据库。挺简单的,比如「我上次跟这家公司开 office hours 是什么时候?」后来它变得稍微聪明了点,比如「我现在正在跟这家公司开 office hours,他们想认识石化行业的任何人」,它就能用不同的方式去查数据库,用上 RAG 之类各种花活,给你凑出五个值得见的相关创始人。但话说回来,这还只是个「跟班」对吧,还只是个 agent,这就是那种老款的、去年那一版「AI 怎么让我这个 group partner 变得更强」的玩法——它让我的效率提升了 20%、30%。真正让我「啊哈」一下的,是我们在这上面加了一个「监控 agent」
[4:33]
on top of that which looked at every single query every single YC employee was doing and saw when it worked and when it did not work and when it did not work it's like oh why not what would have made this query work do we need different deterministic tools do we need to update the skills file do we need a different database view do we need a new index and this happen this literally happens overnight now let's write the code put in a merge request to the YC codebase have an agent review it and merge it and deploy it. So when a human comes the next day to ask the same query, it will now succeed. For me, that was like the holy [ __ ] [ __ ] right?
之后——这个监控 agent 会去看每一个 YC 员工发出的每一条查询,看它什么时候奏效、什么时候不奏效;一旦不奏效,它就琢磨:哦,为啥不行?怎么才能让这条查询跑通?是不是要换一套确定性工具?要不要更新 skills 文件?要不要换个数据库视图?要不要加个新索引?而这事现在真的就一夜之间发生了——它会直接写好代码,往 YC 的代码库提一个 merge request,再让另一个 agent 去 review、合并、部署。这样第二天人类来问同一条查询的时候,它现在就能跑通了。对我来说,那一下真的是「卧槽」对吧?
[5:09]
That's not just AI making you 20 or 30% more valuable. It is the AI going through this loop to figure out how to self-improve. And I think basically if you can identify parts of your company that work like this and eliminate as have the human and kind of a monitoring of supervisory capacity, you can just throw tokens at this problem and your company will get better. And so other examples might be if you have product analytics, having an agent go through your product analytics to to figure out what part of your sales funnel is presenting the highest amount of friction, researching best practices, putting in place an AB test, running it for a week, picking the best version,
这就不只是 AI 让你的价值提升 20%、30% 了。这是 AI 自己走完整个循环,去搞明白「怎么自我改进」。我觉得基本上,如果你能识别出公司里哪些环节是按这个模式运转的,把人类从里头抽出去、只让人当个监控者、起个监督的作用,那你就只管往这个问题上「砸 token」,你的公司就会越来越好。其他例子还有,比如你有产品分析数据,让一个 agent 去把产品分析数据过一遍,找出你销售漏斗里摩擦最大的那一段,去研究最佳实践,上一个 AB test,跑一周,挑出最优的那一版,
[5:47]
and deploying it. Then doing that again and again and again for your product. Just have a self-optimizing like product loop. Or you do it with customer service queries. You have customer suggestions coming in and in and in. you triage it with a kind of you have to have an agent which is like your chief product officer and your chief technology officer who make kind of judgment calls about okay this is a suggestion we just don't want to do we'll discard it but no this is a suggestion which is now in line with our road map um we can do it overnight let's write the code let's deploy it let's ship it to the customer without a human being involved so I think if you can
然后部署上线。然后一遍一遍一遍地这么干下去,给你的产品搭一个自我优化的产品循环。或者你拿客服工单来做也行——用户的建议源源不断地涌进来,你做个分流,你得有一个相当于「首席产品官」加「首席技术官」的 agent,由它来做判断:好,这条建议我们就是不想做,丢掉;不对,这条建议是符合我们路线图的,那我们一夜之间就能搞定,写代码、部署、发布给客户,全程不需要人参与。所以我觉得,如果你能
[6:18]
think about each part of your company as a self-improving like recursive AI loop it becomes very very different to this like hierarchically organized Roman legion from a company so what So like if you want to do this, what are the implications? One is like burn tokens, not headcount. We are seeing companies get to demo day with about 5x more revenue per employee than they did 18 months ago. And I think that's going to continue to series A and series B. And so I think you're going to be constrained on token usage, not on headcount really, really soon. The blunt measure now is just like measuring everyone's token usage, which is obviously like dumb and gameable at the
把公司的每一个环节都想象成一个自我改进的、递归式的 AI 循环,那它跟那种等级森严、像罗马军团一样的公司就变得截然不同了。那么——如果你真想这么干,意味着什么?第一条就是:烧 token,别堆人头。我们现在看到,企业走到 demo day 时的人均营收,比 18 个月前高了大约 5 倍。我觉得这个趋势会一路延续到 A 轮、B 轮。所以我觉得很快很快,你的瓶颈就会是 token 用量,而不再是人头了。现在最粗暴的衡量办法,就是去看每个人的 token 用量——这显然挺蠢的,极端情况下也很容易被钻空子,但
[6:55]
extreme, but directionally I think is correct. We're in the phase of like what is possible right now and so everyone should be experimenting to the max to figure out what we can even do with this crazy new intelligence we have. As soon as you turn it into a leaderboard and people get promoted or fired based on it, obviously it gets gamed, obviously that's dumb. But I think directionally figuring out who in your organization is token maxing, who is not is like a good way to think about which employees you should be spending your time with. I think middle management is done. I just don't think you need middle management for this coordination problem. I think AI should be doing it. And for me, there
方向上我觉得是对的。我们正处在「现在到底能做到什么」的阶段,所以每个人都应该把实验拉满,去搞清楚拿这股我们刚到手的、疯狂的新智能,到底能干出什么名堂来。一旦你把它变成一个排行榜、让人因此被升职或被开除,那它显然会被钻空子,那显然很蠢。但我觉得方向上,搞清楚你组织里谁在「把 token 用到极致」、谁没有,是个不错的视角,能帮你判断该把时间花在哪些员工身上。我觉得中层管理已经没戏了。我就是不觉得为了解决这个协调问题你还需要中层管理,我觉得这事该交给 AI 来做。对我来说,那里
[7:29]
are two roles. Jack Dorsey has three. I actually don't like the third one, so I deleted it. But there are two roles that really, really matter for me. I think everyone just has to be an IC now, a builder, an operator. And I think crucially having directly responsible individuals to get anything done I think you need a named human not a committee not a group of people just a single person and I think you can build companies based on IC's effectively I think just middle management is is over so building this self-improving company that's a dream and by the way I think like people are at the bleeding edge of this right now I'd be interested to see where you all are but it feels like
有两个角色。Jack Dorsey 列了三个,不过第三个我其实不太喜欢,就把它删掉了。但对我来说,真正非常非常重要的就两个角色。我觉得现在每个人都得是个 IC(一线贡献者),一个动手做事的 builder、operator。而且关键是,要有明确的、直接负责的个人——任何事情要推进,你需要一个有名有姓的人,而不是一个委员会,不是一群人,就是单独一个人。我觉得你完全可以靠 IC 把公司搭起来,中层管理真的可以说是终结了。所以打造这种自我改进的公司,是个梦想。顺便说一句,我觉得大家现在都站在这件事的最前沿,我也很想知道你们各自做到了哪一步。但感觉就像……
[8:03]
people are like exploring the boundaries here I'm not sure anyone has a truly self-improving company in every function. I might be wrong. You might prove me wrong. What would I do? First of all, this is really, really important. I would make the entire organization legible to AI. What does that mean? It means you've got to record everything. Simplistically, all of our um partner emails. Now, if you email a YC partner, that email is in the YC database. Every Slack message, every DM, every office hour we've started recording for the last three or four months. every single thing that happens, if it is recorded, it happened to the AI. If it did not get recorded, it is it did not happen to
大家都在探索这里的边界。我不确定有没有谁真的做出了一家在每个职能上都自我改进的公司。我可能错了,你们也许能证明我错了。那我会怎么做?首先,这一点真的非常非常重要:我会让整个组织对 AI 可读(legible)。这是什么意思?意思是你得把一切都记录下来。说得简单点,我们所有的合伙人邮件——现在你给一个 YC 合伙人发邮件,那封邮件就进了 YC 的数据库。每一条 Slack 消息、每一条私信、每一次 office hour,我们过去三四个月都开始录下来了。每一件发生的事,只要被记录下来,对 AI 来说它就发生过。如果没被记录,那它对你的……
[8:42]
your intelligence. You know what I mean? And so, I was talking with some founders over here um just now and we're having like really good conversations about their company, but every conversation I had, I was like, "Fuck, I need to be recording this conversation." Because some guy wanted an introduction to I can't even remember who the introduction was now. Who was that? I was talking to someone about and I promise you an introduction. said yes. And I said, "Email me afterwards cuz I would I'm going to forget this. I'm going to talk to 20 people." Yeah. So, it needs to be on my phone or a clip or or smart glasses or we deck out every room with like microphones. But basically,
……智能体系来说就等于没发生过。你懂我意思吧?所以呢,我刚才还在这边跟几位创始人聊,关于他们公司的对话真的特别好,但每聊完一段我都在想:[ __ ],我得把这段对话录下来啊。因为有个人想让我帮忙引荐一下,我现在都想不起来是要引荐给谁了。那是谁来着?我跟某个人聊过,还答应帮他引荐,对方也答应了。我就跟他说:「回头给我发封邮件,因为我肯定会忘,我要跟二十个人聊呢。」对吧。所以这东西得装在我手机上,或者用个录音夹子、智能眼镜,再不然就把每个房间都布满麦克风。但总之,
[9:15]
everything needs to be recorded so that it can be legible to the AI. And then, as Gary talked about like diorization, you cannot pump in 100,000 hours worth of recordings into a context window. So, you have to diorize it. You have to basically aggregate it down, synthesize it into the important parts, and then give the AI breadcrumbs. It's like, okay, so here's an example. Who's read the user manual? The YC user manual. Hopefully, everyone in this room has at least opened the user manual at one point in time, right? Like, it's fine.
一切都得被记录下来,这样才能对 AI 可读。然后呢,就像 Gary 讲到的 diarization(说话人分离),你没法把十万小时的录音一股脑塞进 context window。所以你得先做 diarization,本质上是把它聚合压缩、提炼出重要的部分,然后给 AI 留下一串「面包屑」线索。举个例子吧:在座有谁读过那本 user manual(用户手册)?YC 那本 user manual。但愿这屋里每个人至少在某个时候翻开过一次,对吧?翻过就行。
[9:41]
It was written 5 to 10 years ago, most of it. It's kind of out of date. So, Haj thought uh last weekend, since now we've got about 2,000 hours of recorded office hours in the last 3 months, why don't we regenerate the user manual? And so you can click like you give it a set of instructions. You basically diorize it down, synthes like categorize it into certain areas like fundraising, hiring, co-founder disputes, whatever. And then write me a new user manual. And by the end of the weekend, he had 150 page user manual, which is dramatically better than the existing user manual. And now we can also update it every single month. So our user manual becomes self-improving. Every new piece of
那本手册大部分内容是五到十年前写的,已经有点过时了。所以上周末 Haj 突发奇想:既然我们现在攒了过去三个月差不多两千小时的 office hour 录音,那干嘛不把 user manual 重新生成一遍呢?你可以这样:给它一组指令,先把录音 diarize 压缩下来,再按几个领域归类,比如融资、招聘、联合创始人纠纷之类的,然后让它给我写一本新的 user manual。到周末结束,他就搞出了一本 150 页的 user manual,比现有那本好太多了。而且现在我们还能每个月都更新它。所以我们的 user manual 就变成了自我改进的——每一条新的
[10:18]
advice we give, it's compared with the existing user manual and either incorporated or thrown away. So the user manual becomes this up-to-date living brain of the advice we give to founders. And obviously it doesn't stop as a user manual. You then pump it in as context to an AI agent and suddenly you can ask a super intelligent AI and get the combined wisdom of 16 YC partners in one, but only if it's legible. So you have to record everything. The second point is kind of the same, right? Like if it creates an artifact that can self-improve, it's legible. If it doesn't, you throw it away. The third point then is that every function can generate this used to say dashboards.
建议给出来后,都会跟现有的 user manual 比对一下,要么被吸收进去,要么被舍弃掉。于是这本 user manual 就成了一个不断更新、活生生的「大脑」,承载着我们给创始人的全部建议。当然它也不止是停在「一本手册」上。你再把它作为 context 喂给一个 AI agent,突然之间你就能去问一个超级聪明的 AI,一下子拿到 16 位 YC 合伙人合在一起的智慧——但前提是它得可读。所以你必须把一切都记录下来。第二点其实和第一点是一回事,对吧?如果它产出的是一个能自我改进的 artifact(产物),那它就是可读的;如果不是,那就扔掉。那么第三点是:每个职能都能生成——我以前会说 dashboard(仪表盘)。
[10:54]
It's not just dashboards. It's on demand software. Codeex 55 is now good enough. You can oneshot most simple inter like most internal software dashboards you can oneshot to a pretty high level of quality. I tried it over the weekend on a bunch of our stuff. It's just unreal. So all of your internal operations teams should be sitting on this layer of like kind of intelligence understanding and then creating their own dashboards and their own workflows. And I would see that those as entirely disposable. I would very preciously store all the data. So as Gary said, he puts it all all of his emails in markdown. Never throw anything away, but then treat the the software as ephemeral. You can you
其实不只是 dashboard,而是按需生成的软件。Codex 现在已经够好用了。大多数简单的内部软件、内部 dashboard,你都能 oneshot(一次生成)出来,而且质量相当高。我上周末拿我们的一堆东西试了试,简直不可思议。所以你们所有的内部运营团队,都应该坐在这层「智能理解」之上,然后自己去生成自己的 dashboard、自己的工作流。而且我会把这些东西都看作完全可丢弃的。我会非常珍惜地把所有数据存好——就像 Gary 说的,他把所有邮件都存成 markdown,什么都别扔——但把软件本身当成转瞬即逝的。你可以
[11:34]
can generate it, you can regenerate it. The valuable part is like the comprehension inside people's heads of like this is how the function works. This is how we run a YC event. Whatever the software to actually run the event, you can generate for the event. You can throw it away. The mo the models get smarter in a month or two. Throw the software away. Give it your original set of instructions and regenerate the software. So I think the business context and and skills are the valuable part. I think the software on top of it is ephemeral. So what what are humans for in this world? I think basically we're talking about a company brain and I know a bunch of people in this room
随时生成它,也可以随时重新生成它。真正有价值的,是人脑子里那份理解:这个职能是怎么运转的,我们是怎么办一场 YC 活动的。至于真正用来办活动的软件,你可以为这场活动现生成一个,活动办完就扔。一两个月后模型变得更聪明了,就把旧软件扔掉,把你最初那组指令再喂给它,重新生成一遍软件。所以我觉得,业务上下文和技能才是有价值的部分,跑在上面的软件是转瞬即逝的。那在这样一个世界里,人是干嘛用的呢?我觉得我们本质上是在讲一个「公司大脑」,我知道这屋里有不少人
[12:08]
are building this but the bit in the middle like all of your data, all of your emails, your DMs, the skills, the knowhow that is like the company brain and I think the humans sit around the edge of this interfacing with the real world. So it's where this intelligence makes contact with reality. Human beings reach into places the models can't go yet. That might be like a conference. It might be a I'm trying to think of examples. I would say a phone call, but I think the AI can reach into phone calls pretty easily now. Um I think it's like novel situations, ethical considerations, high stakes moments, you know, it's like it's where the founder comes to us and is like thinking about
正在做这件事。但处在正中间的那一块——你所有的数据、所有的邮件、私信、技能、专有诀窍——那才是公司大脑。我觉得人是坐在它的边缘,去和真实世界对接。所以那就是这份智能与现实接触的地方。人能伸进模型暂时还到不了的地方。可能是一场会议(conference),也可能是……我想想例子哈,我本来想说打电话,但我觉得现在 AI 已经能挺轻松地介入电话了。嗯,我觉得更像是那些新奇的场景、涉及伦理的考量、高风险的时刻——你懂的,就是那种创始人跑来找我们、正在纠结要不要
[12:47]
breaking up with their co-founder, right? It's like those real high stakes, high emotion moments where you really want a human being. I think that's where the human fits for all of you like sales conversations. I think that's a human being in the room for the next 20 years. So the humans live I think around the edge and I'm over time and cool vision should bullhorn me. I will leave you this one question. If you were building your company today would you start it in this shape for most of you you're small enough to build it right and so I don't think you have any excuse and I know there are a few of you who are in the process of ripping up and rebuilding your company. So with that I will stop
跟联合创始人分道扬镳的时候,对吧?就是那种真正高风险、高情绪的时刻,你特别想要一个活生生的人在场。我觉得那才是人该待的地方——对你们所有人来说,比如销售对话,我觉得未来二十年里那都得是一个真人坐在房间里。所以人是活在边缘的。我时间也差不多了,cool vision(计时的人)该用喇叭轰我下台了,我留给你们最后一个问题:如果今天让你重新打造你的公司,你会一开始就照这个形态去搭吗?你们大多数人规模还够小,完全可以把它搭对,所以我觉得你们没什么借口。我也知道在座有几位正在把自己的公司推倒重建。那么,我就讲到这儿,
[13:25]
um and we'll hand over to Pete. Thank you for listening.
接下来交给 Pete。谢谢大家的聆听。