ESC
↑↓ 选择↵ 打开esc 关闭⌘K 唤起
← 返回速读报告 回声编辑部 · NO.108 · 全文

How Forward Deployed Engineering is done at Factory — Eno Reyes

频道: AI Engineer
视频: https://www.youtube.com/watch?v=wpOA-UXynoM
原文语言: en
统计: 共 22 轮 · 主持人 1 · Eno Reyes 17


[0:01]

[music]

[音乐]


[0:12] 主持人

This is the forward deployed engineering track in case you're in the wrong room. Um as you already know, forward deployed engineering is one of the hottest topics in AI. The most important companies on the planet are building out massive FTE teams. So, think OpenAI, Anthropic, Google DeepMind, you get the idea. Forward deployed engineering was pioneered by Palantir many years ago to embed really strong software engineers directly into their customers' orgs uh to implement customize their platforms around the nuances of the real world. So, today we brought in some amazing speakers from Anthropic, Cursor, Factory, Ramp, Decagon, and many more uh to talk about the current state of forward deployed engineering, how it works at their companies, and where it's going. Our first speaker Our first speaker is Eno Reyes. He's the co-founder and CTO at Factory, which is building autonomous software engineering agents for enterprise teams. Previously, he worked in machine learning and software engineering roles at Hugging Face and Microsoft. Let's give it up for Eno.

这里是 forward deployed engineering 专场,走错房间的朋友请注意一下。大家应该都知道,FDE 是当下 AI 圈最热的话题之一,地球上最重要的那几家公司都在大规模扩建 FDE 团队——OpenAI、Anthropic、Google DeepMind,你懂的。FDE 这个模式最早是 Palantir 在很多年前开创的:把非常强的软件工程师直接嵌进客户的组织内部,围绕现实世界里那些微妙的细节去落地、去定制他们的平台。所以今天我们请来了 Anthropic、Cursor、Factory、Ramp、Decagon 等等公司的一批重量级讲者,来聊聊 FDE 的现状、它在各自公司里到底怎么运转,以及接下来会往哪儿走。第一位讲者是 Eno Reyes,他是 Factory 的联合创始人兼 CTO,Factory 在为企业团队打造自主的软件工程 agent。此前他在 Hugging Face 和 Microsoft 做过机器学习和软件工程。掌声欢迎 Eno。


[1:08]

[applause]

[掌声]


[1:11] Eno Reyes

Yeah, hey everyone. Excited to chat today. Um and you know, basically I I my hope is that at the end of this you guys get a sense of some of the work that we're doing on behalf of our customers and with our customers, and the role of what we call a deployed engineer should hopefully be a little bit clearer since I think that there are honestly tons of different models um for uh for how this should actually operate inside of an org. Um and so I think that when when we start I I I do think that there are some nuances in sort of like the Palantir era playbook. Um and I generally the way that forward deployed goes is I I see that there are lots of different takes on sort of where forward deployed sits within the org, how much it interfaces with the actual product team or the engineering team, how much work is done on behalf of the customers versus with them, and how much work is done on code itself or basically like in the software system versus with the humans and sort of strategizing, right? And so, um generally, I think the there's um in this older model, a lot of the way that software needed to be built was you needed to go and access that codebase.

大家好,很高兴今天来聊这个话题。我的期待是,听完这一段,你们能大致了解我们在为客户、以及和客户一起做的那些事,同时对我们所说的 deployed engineer 这个角色更清楚一点——因为说实话,这个角色在一个组织内部到底该怎么运转,业界有特别多不同的模式。开场我想先说,Palantir 那个时代的打法其实有不少微妙之处。总体上,我看到大家对 forward deployed 的理解差别很大:FDE 在组织里到底坐在什么位置、跟产品团队或工程团队的接口有多深、多少活是替客户做的、多少是跟客户一起做的、又有多少工作落在代码本身、也就是软件系统里,而不是跟人打交道、做策略层面的事。总的来说,在早期那套模式里,很多软件之所以只能那样造出来,是因为你必须真的进到客户的代码库里去。


[2:23] Eno Reyes

You needed to integrate directly in to data streams or software or products that basically you could only access behind the curtain of the customer. And so, if you were building something that was heavily integrated into their environment, yeah, you kind of had the need to send and sort of parachute in individuals into the org. Um but really that has transformed over time into a role that sort of forks out, and you see a lot of people who are sort of quote-unquote forward deployed engineers or deployed engineers or applied AI engineers, and it's always it's always a little bit unclear. Are they doing maybe professional services work on behalf of their customer? Are they transforming like the product around an individual customer? Are they just building entirely net new things in the customer's environment, maybe on top of your product? Uh and I think that the the at least at Factory, we definitely do not want to be doing professional services work on behalf of a customer.

你得直接对接那些数据流、软件或产品,而这些东西基本上只有掀开客户那层帘子才碰得到。所以如果你要做的东西跟他们的环境深度集成,那确实需要把人派过去、空降进对方的组织。但随着时间推移,这个角色慢慢分叉了。你现在看到很多人自称所谓的 forward deployed engineer、deployed engineer,或者 applied AI engineer,界限一直有点模糊:他们到底是在替客户做专业服务(professional services)?是在围绕某一个客户改造产品?还是纯粹在客户的环境里、也许基于你的产品之上,从零造全新的东西?至少在 Factory,我们非常明确地不想替客户做专业服务那类活。


[3:23] Eno Reyes

So, if a customer says, "I want to do a uh modernization of a codebase, uh and it's you know, I just got quoted from all of the big consulting firms, it's going to cost this much. Could you do this consulting work for us?" Uh our goal is not to go and actually do that migration on their behalf, even if we happen to be using our product, right? Um and that is because we don't think that that actually makes our product that much better. Uh and ultimately, that is a great way way get I'd say a decent amount of revenue, but I don't think that that's the way that you can scale a business out enormously, right? And so, what we've done is we've instead said, we need deployed engineers to be the tip of the spear of the product. And I'm going to do this and then go back. But but really when we say the tip of the spear, what we mean is that deployed engineers are basically the stream of information from our largest and most critical customers of the engineering leadership in that org. The on-the-ground tactical engineers, their thought process about how software development and AI is actually happening at that org, and then flowing all of that information back into our product to then rapidly adjust our product in order to then fit into the customer's environment better, right? And so, factory really should be when it gets deployed, and I'll talk about what factory is in a second, but we want that to be effectively self-assembled inside of our customer's environment, right?

比如客户说:我想做一次代码库的现代化改造,几家大咨询公司都给我报过价了,得花这么多钱,你们能不能接下这个咨询的活?我们的目标不是真的替他们把这次迁移做完——哪怕我们用的正好是自己的产品。原因是,我们不认为这么做能让产品变好多少。而且说实话,这确实是一条能赚到不少收入的路子,但我不觉得靠它能把生意做到极大的规模。所以我们的做法反过来:deployed engineer 要成为产品的「矛尖」。我先把这一点讲完再往回说——所谓矛尖,意思是 deployed engineer 本质上是一条信息流:把我们最大、最关键的那些客户那边的信息带回来,包括那个组织里工程领导层的想法、一线战术型工程师的想法、他们对「软件开发和 AI 在我们这儿到底是怎么发生的」的判断,然后把这些信息全部回流到我们的产品里,让我们快速调整产品,更好地嵌进客户的环境。所以 Factory 被部署的时候——我等一下会讲 Factory 到底是什么——我们希望它能在客户的环境里自动装配起来(self-assemble)。


[4:47] Eno Reyes

And then there's a lot of work that goes into understanding that customer's environment, what the flows that happen, and ultimately the ROI story. And what factory really is to our customers is a set of building blocks for building a software software factory, right? And so, when we say software factory, what we mean is there's this implicit process that every organization in the world sits on top of, where signals from the outside world flow in on one side, and those signals could be a lot of different things. It could be customer conversations, it could be bug reports, it could be internal Slack or Teams conversations, it could be an executive saying, "We're going to build this thing," right? All of these are signals. Some of them have higher weight than others, and those signals flow in, and humans implicitly or explicitly then choose to then prioritize, triage, and build plans around those signals. Those plans are then converted, typically by software developers, into changes into some source of truth, a code base, an engineering system. Um and as those changes are actually executed on, they flow through a validation stage where people maybe review the code, they QA, they assess the security implications, they uh pass it through automated validation like SAST tools, linters, type checkers, and ultimately when everything passes, they then ship and deploy. And what do you do with deployed monitored software? Well, it it generates more signals, right? So, this implicit feedback loop is instrumented very poorly, to be honest, at most organizations. And if you're able to take AI and actually transform each of

然后还有大量工作是去理解客户的环境、里面到底在跑哪些流程,以及最终的 ROI 故事。Factory 对客户来说,本质上是一套用来搭建「软件工厂(software factory)」的积木。我们说的软件工厂是指:世界上每个组织其实都隐含地跑在这么一条流程上——外部世界的信号从一端流进来,这些信号可以是很多种东西:客户对话、bug 报告、内部 Slack 或 Teams 里的讨论、某个高管说「我们要做这个东西」,这些都是信号,只是权重有高有低。信号流进来之后,人会或明或暗地去做优先级排序、分诊,然后围绕这些信号做出计划。这些计划再由软件开发者转成对某个事实源(source of truth)的改动——代码库、工程系统。改动被执行出来之后,会走进验证环节:有人 review 代码、做 QA、评估安全影响,再过一遍自动化校验,比如 SAST 工具、linter、类型检查。全部通过之后,就发布上线。那么,已经上线并被监控的软件又会产生什么呢?它会产生新的信号。所以这是一个隐含的反馈闭环——老实说,在大多数组织里,这个闭环的可观测性做得非常差。而如果你能用 AI 真正改造这条流水线的每一个环节,


[6:26] Eno Reyes

these stages of the pipeline and build an understanding of what the workflow looks like at your org from each stage to each stage, then you actually can get to the point where you have a a flow through from signal to deploy that has no human intervention. Now, importantly, that does not mean that humans are not a part of engineering this system, right? But it is that the flow of signal to deploy is uninterrupted by a human. Um and that software factory concept is obviously not something that can just snap your fingers and it appears, right? Instead, it requires an investment from the organization. We we like to say this is built, not bought, right? But what the platform that we've built basically provides to people are the canonical one model independent agent harness that you need to do this, because if you want to build a software factory, if you choose to build that software factory in a vendor locked solution that has like one model available to it, uh that is going to not only be expensive, but two, uh there's open questions about model independence and like what is the role of the model provider in dictating what you can or cannot build with your software factory, right? Um and if you also don't own the traces, the data, everything that flows through your software factory, um then you're probably going to be in trouble as you start to want to evolve your software factory, right? And so with Droid, the hardness that we build, you not only have model independence, but you also have access to every piece of data that flows in in and out of Droid, alongside centralized governance and control at

并且把你们组织里从每一环到下一环的工作流都梳理清楚,那你其实可以做到:从信号到部署整条链路,全程不需要人工介入。这里要强调的是,这不代表人不参与这套系统的搭建——而是说,信号到部署这条流不再被人打断。软件工厂这个概念显然不是打个响指就冒出来的,它需要组织真金白银的投入。我们喜欢说:这东西是「建」出来的,不是「买」来的。而我们做的平台提供给大家的,是搭这件事所必需的那个标准的、与模型无关的 agent harness(agent 的运行支架)。因为如果你要造一座软件工厂,却把它建在一个被厂商锁死、只能用一个模型的方案上,第一,它会很贵;第二,模型独立性会打个问号——模型提供方到底在多大程度上能决定你的软件工厂能造什么、不能造什么?而且如果流经这座工厂的 trace、数据这些东西你都不拥有,那等你想要演进这座工厂的时候,多半会很麻烦。所以我们做的这个 harness 叫 Droid:你不仅拥有模型独立性,还能拿到进出 Droid 的每一份数据,同时在企业层有集中的治理和管控,


[7:56] Eno Reyes

the enterprise layer to be able to dictate where what information flows where. Um you can air gap Droid if you want. Some of our partners um in, you know, the most secure uh environments, uh think finance, health care, uh gov, uh they air gap Droid and they run their software factories entirely contained. Um One of our deployed engineers jokes that you could run Droid in a submarine if you wanted to. And that's that's honestly true. And so when we think about what the role of this deployed engineer is in that context, which I probably should have started with, um you know, you really need somebody who can go in and say, I understand this new model of building software and I understand the building blocks and the pieces. I can help enable building and constructing these software factories with your team, but I ultimately would like to one, make it so that our product effectively, you know, one click self-assembled into your environment, which is needed when you have 45,000 people, maybe hundreds of thousands of uh of of engineers, maybe you have tens of thousands of code bases. Uh you you've got to self-assemble, right? You just can't manually install this level of of complexity. Um and and also on the sort of like end loop, why do all of this, right? I I would argue that there needs to be an ROI or an outcome story that is extremely clear from the beginning so that you can say, well, we know every code change that flows through that gets AI code review, AI QA, AI security analysis is maybe 87% less likely to hit a bug. And what that means is that we can reduce our our bug rate by X, that increases our customer

可以规定什么信息流向哪里。你想的话,还可以把 Droid 做成 air-gapped(物理隔离)。我们有些合作伙伴身处安全要求最高的环境——金融、医疗、政府——他们就把 Droid 隔离部署,整座软件工厂完全跑在封闭环境里。我们有位 deployed engineer 开玩笑说,你要愿意的话,在潜艇里都能跑 Droid。这话说实话是真的。那么在这个语境下,deployed engineer 的角色是什么——这一点我本该一开始就讲——你需要的是这样一个人:他能进去说,我懂这套新的软件构建范式,我懂这些积木和零件,我能帮你的团队把这些软件工厂搭起来。但归根结底我想做到两件事:第一,让我们的产品真正做到一键就在你的环境里自动装配。当你有四万五千人、也许几十万工程师、也许上万个代码库的时候,这是必须的——这种复杂度你根本不可能靠手工去安装。第二,是闭环的另一头:为什么要做这一切?我认为从第一天起就必须有一个极其清晰的 ROI、或者说结果故事,比如你能说:我们知道,凡是经过 AI code review、AI QA、AI 安全分析的代码改动,出 bug 的概率大概会低 87%。这意味着我们能把 bug 率降低 X,从而把客户


[9:35] Eno Reyes

satisfaction by Y, and that leads to revenue or growth or new business, right? Something needs to flow from this software factory process to core business goals. And that often is a complex story that requires engineering knowledge, it requires business knowledge. And so, if those are the types of things that you think are interesting, that is what deployed engineers today are doing for us. Um I've sort of outlined it a little bit here, but that teach the model step is super important because most organizations do not have an autonomy maturity model. They do not have a road map, they don't have a conception of what it means to truly build an autonomous software organization, right? I think a lot of people ask the question, what do the humans do in this world, right? For us, we see an extremely clear role for humans in evolving, refining, and scaling software factories, right? So, you basically the engineers at a company go from directly manipulating software to directly maintaining and managing a system that builds software. And that sort of like upgrade in the level of abstraction that you operate at is actually very difficult. And a lot of people find it extremely challenging. I would argue that in fact most people, even very thoughtful software engineers, will have a learning curve in trying to shift. The people who I think are are well suited for this are DevEx people who have already been thinking about enablement of other developers. I think product managers who want to become very technical very quick can become really great at doing this. And I think that generally like people who are used to

满意度提升 Y,最终带来收入、增长或者新业务。必须有一条链路,能从软件工厂这套流程一路连到核心的业务目标。而这往往是个复杂的故事,既要懂工程,也要懂业务。如果你觉得这类事情有意思,那这就是今天 deployed engineer 在我们这儿做的事。我这里稍微列了一下,其中「教会他们这套模式」这一步特别关键,因为绝大多数组织并没有所谓的自主化成熟度模型(autonomy maturity model):他们没有路线图,也没有概念去想「真正建成一个自主的软件组织」到底意味着什么。很多人会问:那在这个世界里,人还干什么?在我们看来,人的角色极其清晰——去演进、打磨、扩展软件工厂。也就是说,一家公司的工程师,从直接操作软件,变成直接维护和管理一套「造软件的系统」。而这种抽象层级上的跃迁其实非常难,很多人会觉得极具挑战。我甚至认为,绝大多数人,哪怕是非常有思考力的软件工程师,在做这个转变时都会经历一段学习曲线。我觉得比较适合这个角色的,是做 DevEx 的人——他们本来就在琢磨怎么赋能其他开发者;还有那些想快速变得非常技术的产品经理,也能把这件事做得很好。另外我觉得,总体上,那些习惯于


[11:13] Eno Reyes

to working on teams where

在这样的团队里工作的人——


[11:15] Eno Reyes

[clears throat]

[清嗓]


[11:16] Eno Reyes

high quality dev environments were a priority, you will you will get some of the canonical things necessary to enable these agents to succeed. Um I haven't really talked about this last one, which is design the workflows. And I will get to that in a sec, but I I think that when I say tip of the spear of the product, like keep in mind I really do mean everything that is happening inside of factory. So, our product encompasses enterprise controls, the droid harness, the workflows that run on top of it, the observability tools, the cost controls, the auto model routing, the quality of the harness. Like, all of these are pro- potential opportunities of improvement that you will discover when you work very closely in these varied or diverse orgs like how to solve. Um, so making a code base agent ready, right? This is a very challenging thing to do. Uh, most organizations have some degree of consistency in how they've chosen to build deterministic validation loops inside of their company, right? So, your code base runs linters, type checkers, uh, it might run some security scans, and it's like check mark. Like, it passes or it doesn't. The end end test, they pass or it doesn't, right? Or they don't. Um, what agent readiness really is is it's a measure of how many of these deterministic validation loops are present inside of your code base. Uh, when you have a huge volume of these feedback loops, uh, agents are able to operate for greater periods of time on more complex tasks without human intervention. So, we have like a product that we call missions, which I'll also touch on in a sec. But, missions is

只要把高质量的开发环境当成优先事项,你就能拿到那些让 agent 跑得起来的基本条件。最后这一条我还没怎么展开讲——设计工作流。我马上会说到。不过我想先说明一下:我说「产品的最前沿」,指的真的是 Factory 内部正在发生的一切。我们的产品包含企业级管控、Droid harness、跑在它之上的工作流、可观测性工具、成本管控、模型自动路由、harness 本身的质量……这些全都是潜在的改进机会——你只有深入到这些千差万别的组织内部去解决问题,才会发现它们。

那么,怎么把一个代码库变得「agent-ready」?这件事非常难。大多数组织在如何构建确定性的验证回路上,多少已经有了一些一致的做法:代码库会跑 linter、类型检查,可能还会跑一些安全扫描,结果就是一个对勾——过或不过。端到端测试也是,过或不过。

所谓 agent readiness(agent 就绪度),本质上就是在衡量:你的代码库里到底存在多少这样的确定性验证回路。当你拥有大量这类反馈回路时,agent 就能在更复杂的任务上、在没有人工介入的情况下,持续运转更长的时间。我们有个产品叫 missions,我待会儿也会讲到。missions 说白了就是


[12:50] Eno Reyes

basically an extremely elaborate harness built around the concept of working on extremely difficult knowledge work problems that are validatable, right? And so, the quality of the output of these very long-running harnesses of advanced agents is directly proportional to the degree to which you can validate their work. And so, if you introduce the ability to validate at scale, then you introduce increasing autonomy to the org. So, what we'll look at is we have tools that help scan all of these things, but often times, uh, the change is not so simple. Uh, for I'd say maybe 30 to 40% of the low-hanging fruit, you click droid, please fix all of this and it'll go in and it'll fix it, right? But for the other 60% some of them involve workflow changes. Sometimes humans are not used to the degree of I would say like nitpickiness of these automated systems. And so you have to sort of be aware of the concerns, the the humans, you have to think about like the way that people are currently developing systems and say, "How do we introduce some of these more extreme validation strategies without interrupting the dev flow of the humans who are involved in the work?"

一套极其精细的 harness,专门围绕「处理那些极难、但可被验证的知识工作」这个概念搭起来的。这类长时间运行的高级 agent harness,输出质量跟你能验证它工作的程度是直接成正比的。所以,一旦你把「大规模验证」的能力引进来,你就把「越来越高的自主性」引进了这个组织。

我们会怎么做?我们有工具能扫描所有这些东西,但很多时候改造并没那么简单。我估计有 30% 到 40% 属于唾手可得的部分——你点一下 Droid,说「把这些都修了」,它进去就修好了。但剩下那 60% 里,有些涉及工作流的改变。人有时候受不了这些自动化系统那种……我该说是「吹毛求疵」的程度。所以你得留意人的顾虑,得考虑大家现在是怎么开发系统的,然后问:怎么才能在不打断这些人日常开发节奏的前提下,引入这些更严苛的验证策略?


[13:57] Eno Reyes

Um and and I mentioned missions because really I think this is one of the more end game of the agent era at least, pre-software factory era. But the more end game of the agent era style harnesses where it's simply a long running harness that has almost no human intervention except for the planning stage, right? Where you go in and you say, "I would like to have this very bounded task. I know that I want to solve this task and here is what solving this task means. I will now basically push a lever of inference until the task is complete, right? And so that is actually unbelievably competent at solving problems where like is complete is verifiable. So if you can frame any problem as the set of verification uh systems that need to validate it, then you can solve that problem with AI today. Uh and we've seen this work on some pretty insane problem spaces like migrating, you know, 30, 40, 50 million plus line code bases uh fully autonomously, um working on advanced uh like deep learning strategies around biomed, health care uh sort of problems, uh financial institutions that optimize equity research where you can actually build models of different equities and sort of analyze and compare and and build sort of a system that can then back prop and or trade on top of the those equities. Um like it it's mind-blowing to me every day what I what I hear people are using with these tools, but it is not something that you can just download, install, and hit play, right? It does require agent readiness. So, if your code base isn't agent ready, you won't see any of the success of the most capable AI systems in the world today, right? So, this is

我之所以提到 missions,是因为我觉得这算是 agent 时代——至少是「软件工厂时代」之前的 agent 时代——比较接近终局形态的 harness 之一:一个长时间运行的 harness,除了规划阶段之外几乎不需要人介入。你进去说:「我有一个边界很清楚的任务,我知道我要解决它,而且『解决』的标准是什么我也说得明白。」然后我就一路推动推理这个杠杆,直到任务完成。

这套东西在「完成与否可被验证」的问题上,能力强得不可思议。所以说,只要你能把任何问题重新表述成「一组用来验证它的系统」,那这个问题今天就能用 AI 解决。我们已经见过它在一些相当疯狂的问题域里奏效:比如全自主地迁移三千万、四千万、五千万行以上的代码库;比如在生物医学、医疗健康领域做前沿的深度学习工作;比如金融机构用它优化股票研究——你可以给不同的股票建模、分析、比较,搭出一套能反向传播、甚至能据此交易的系统。

我每天听到大家拿这些工具在做什么,都觉得不可思议。但它不是那种下载、安装、按个播放键就能用的东西——它确实需要 agent readiness。如果你的代码库不是 agent-ready 的,那今天这世界上最强的 AI 系统,你一样什么成果都看不到。所以说,这就是


[15:41] Eno Reyes

why we want people to go in and help our customers and say, "Hey, look, you can solve this actually very difficult problem, but it is going to require a different form of investment than you were thinking. Less so solving the problem, more so preparing the environment for verification of the problem." And by the way, if you're familiar with how these models are actually trained, like this makes total sense, right? They they get dense reward when they get post trained on all these complex tasks. Models need dense reward. These verification signals form the basis of that reward that they use to keep them on track over a long-term goal-directed problem. Um So, I sort of mentioned this earlier, but but I think that the the core goal for us really is to say, if we can hand over a model to you of how this should this transformation should go, then we should theoretically be able to say, "Let's do this in a couple of different places, and then let your team actually scale this out across the company."

为什么我们希望有人进到客户那里去,跟他们说:「你看,这个问题其实是能解的,但它需要的投入方式跟你原本想的不一样——重点不在于去解那个问题,而在于把环境准备好,让这个问题变得可验证。」

顺带一提,如果你熟悉这些模型实际是怎么训练出来的,这个道理就再自然不过了:它们在这些复杂任务上做后训练时,拿到的是稠密奖励。模型需要稠密奖励。而这些验证信号,正是那种奖励的来源——是它们在长周期、目标导向的问题上不跑偏的依据。

我前面也提到过,对我们来说核心目标其实是:如果我们能交给你一个「这场转型该怎么走」的样板,那理论上我们就可以说——「我们先在几个地方做出来,然后由你的团队把它推广到整个公司」。


[16:38] Eno Reyes

I always use the analogy of if you're familiar with Walt Disney's Epcot, the the theme park. Like basically that theme park was created originally Disney wanted to create like a master planned exemplary city. He said, "Look, if I can create a city that is the future city, then I can use that as a model to the rest of the world cities, and they can develop entirely new forms of transportation and flourishing." And it became a theme park. But, what's interesting is that in that small example, a lot of other cities actually did cite some of the ideas that he was writing down and sharing about what like centralized urban transit should look like. And now you have like some more contemporary cities built in the last 50 years that basically modeled after that toy example. Um what we want to do is we want to make sure that we get some of that lesson that if you have a working example of a city of the future, of a code base of the future, um people are smart. They're clever. Humans will look at that and they'll say, "Man, that's really cool. Let's bring that to my part of the code base, right?"

我总爱打个比方,不知道你们熟不熟悉 Walt Disney 的 Epcot,那个主题公园。它最初其实不是要做主题公园——Disney 想造一座规划完备的典范城市。他说:「如果我能造出一座『未来之城』,那它就能成为全世界其他城市的样板,让人们发展出全新的交通方式和全新的繁荣形态。」最后它变成了主题公园。但有意思的是,就在这么个小小的样例里,后来确实有不少城市借鉴了他当年写下、分享的那些关于「集中式城市交通该长什么样」的想法。今天你能看到近五十年建起来的一些较新的城市,基本上就是照着那个玩具级样例做出来的。

我们想做的,就是把这个教训用上:只要有一个真正跑得通的样例——一座未来之城、一个未来的代码库——人是聪明的,是机灵的,他们看到之后会说:「哇,这也太酷了,把它搬到我负责的那块代码里去。」


[17:42] Eno Reyes

But if you build too much of an advanced example, then people will say, "That's a theme park. That is not at all how the rest of the world works. I just can't see how that would apply to the way that we currently work today, right?" So it's kind of a delicate balance that you have to walk of building something that demonstrates the future is achievable enough, but ultimately does not scare away uh an org who is thinking, "Man, what is going to be the cost of transforming at this pace, right?" Um I always think about that quote, you know, the the future is here, it's just not evenly distributed. Um there are some code bases, and I I say code bases, not even companies, that are truly remarkable. They are effectively uh beginning to run on autopilot. Uh we ourselves have roughly 15 to 20% of what we call like autonomy, and our autonomy ratio is like in the upper 80%, which means the ratio of actions done by humans to AI systems before interruption, right? So our own code base is fairly agent-ready, pretty autonomous, um but uh the code bases of some of our customers are actually even more autonomous because they operate in more constrained uh ways, right? So it's it's sort of like a uh it is not obvious like who gets 100% autonomy first. I would argue it's probably very contained internal tools. Like we have something we call like legal droid, which is our legal workflow. That is effectively 100% autonomously maintained, but our like core harness, uh we do not yet have validators that can validate some of the hard visual problems of a like terminal based harness. Uh Things like flickering are really hard to catch in a verif- in

但如果你的样例做得太超前,人们又会说:「那是个主题公园,跟真实世界完全不是一回事,我实在看不出它跟我们今天的干活方式有什么关系。」所以这里有个很微妙的平衡要拿捏:你造出来的东西,既要足以证明「未来是够得着的」,又不能把一个正在心里盘算「按这个速度转型,代价得有多大」的组织给吓跑。

我常想起那句话:未来已来,只是分布不均。有一些代码库——我说的是代码库,甚至不是公司——真的非常了不起,基本上开始进入自动驾驶状态了。我们自己大概有 15% 到 20% 属于所谓的自主运行,而我们的自主率在 80% 出头,也就是「在被人打断之前,AI 系统完成的动作与人完成的动作之比」。所以我们自己的代码库算是比较 agent-ready,也相当自主了。但我们有些客户的代码库其实比我们还自主,因为他们的运行方式受到的约束更多。

所以说,谁会第一个做到 100% 自主,其实并不明显。我倾向于认为,最可能的是那些边界极其收敛的内部工具。比如我们有个叫 legal droid 的东西,是我们的法务工作流,那个基本上已经是 100% 自主维护了。但我们的核心 harness 就不行——像终端型 harness 里那些很难搞的视觉问题,我们还没有能验证它们的 validator。比如闪烁(flickering)这种东西,就很难以一种可验证的方式抓出来。


[19:18] Eno Reyes

a verifiable way. So, we're unable to close the loop on some of those challenges. It's an engineering task to build the system that can verify some of those very hard problems. And that might give you a picture into sort of like the weird world of the future where humans are sort of visually our advantages in being visual, our advantages in having context of the outside world provide us a lot of work to do in order to build these systems. So, who's great at this? I If you are a former founder, for sure you should do this. I think it's like the quickest way to basically build out I mean each like stage of the SDLC that Droid has, we think is a billion-dollar business. Like just code review, just incident response, just QA, just testing. Like each of these you will help define basically the nature of these products. If you are someone who is used to tech communication, right? If you are fluent in AI, you understand how to speak to every level, you have business acumen, you have executive presence, that is another great example of someone who should do this. And if you are a systems thinker, if you love designing systems, if you love closing loops, modeling data, and understanding how the flow through a potentially extremely complex org should look, then you are also someone who would thrive at doing this.

所以在这些挑战上我们还闭不了环。要造出能验证这类硬骨头问题的系统,本身就是一项工程任务。这大概也让你窥见了未来那个略显古怪的世界:人类在视觉上的优势、在掌握外部世界上下文上的优势,恰恰给我们留下了大量工作——就是去把这些系统搭出来。

那么,什么样的人适合干这活?如果你当过创始人,那你绝对该来干这个。我觉得这基本上是最快的路径——Droid 覆盖的 SDLC 每一个环节,我们都认为是一门十亿美元级的生意:光是 code review,光是故障响应,光是 QA,光是测试。这里每一块,你都能亲手去定义这些产品的形态。

如果你擅长技术沟通——你对 AI 很熟,知道怎么跟各个层级的人对话,有商业判断力,有面对高管的气场——那也是非常合适的人选。还有,如果你是个系统思考者,喜欢设计系统、喜欢把回路闭上、喜欢建数据模型、喜欢搞清楚一个可能极其复杂的组织里信息该怎么流动,那你在这个岗位上一样会如鱼得水。


[20:35] Eno Reyes

So, if all of this seems interesting, hopefully it does. Please do reach out. And you can reach out to me directly. I'm just Yeah, I'll say it. It's on the slide. I'm eno@factory.ai. And so, you can just email me directly or you can apply on our careers page. It's called engineer, {comma} deployed. So, that's the role. Hopefully this is interesting and gives you a taste of what we're doing at Factory.

所以,如果这些听起来有点意思——但愿是有的——欢迎来找我,可以直接联系我。我就……行吧,我直接说了,反正也写在幻灯片上了:eno@factory.ai。你可以直接给我发邮件,也可以到我们的招聘页面投递。这个岗位叫「engineer, deployed」(工程师,已部署)。希望这些内容对你们有点意思,也让你们大致尝到我们在 Factory 做的事情是什么味道。


[21:19]

[music]

[音乐]


[21:20]

Woo!

哇——!