ESC
↑↓ 选择↵ 打开esc 关闭⌘K 唤起
← 返回速读报告 回声编辑部 · NO.105 · 全文

"The biggest challenge in your stack? Evals, Evals, Evals" - 2026 State of AI Engineering results

频道: AI Engineer
视频: https://www.youtube.com/watch?v=RGe6EjucbzI
原文语言: en
统计: 共 43 轮


[0:01]

[music] Now joining us on stage is the partner at Amplify Barren. [music] Fantastic. You did a great job practicing. I feel very very loved. Um, let's get started. So, like you just heard, my name is Bar.

[音乐] 现在有请 Amplify 的合伙人 Barr Yaron 上台。[音乐] 太棒了。你们排练得真不错。我感觉自己被爱包围了。我们开始吧。就像刚才听到的,我叫 Barr。


[0:48]

I run a survey every year on the state of AI engineering. And the funny thing about running a survey on the state of AI engineering is that the field changes as you make the slides. Just in the past week, we've had Frontier releases treated like national security events. Meta reportedly exploring selling AI compute. By the time I get off stage, maybe something else will happen. So, if I miss a major announcement while I'm up here, please come find me after.

我每年都会做一次 AI 工程现状调研。做这个调研有意思的地方在于,就在你做幻灯片的过程中,这个领域本身就已经变了。就在过去这一周,前沿模型的发布被当成国家安全事件来对待;据报道 Meta 在考虑出售 AI 算力。等我下台的时候,说不定又发生了别的事。所以如果我在台上的这段时间漏掉了什么重大发布,散场之后一定来找我。


[1:18]

But that's exactly why we run the survey every year to cut through the noise, take a moment, step back and understand what AI engineers are actually doing. Uh for the first time this year, we were thrilled to partner with Notion and Verscell to run this survey.

但这恰恰就是我们每年做这份调研的原因——穿透噪音,停下来,退一步看看 AI 工程师到底在做什么。今年是头一回,我们非常高兴能和 Notion 以及 Vercel 一起来做这份调研。


[1:36]

Very quickly on me, uh this is the least interesting slide. I'm an investment partner at Amplify. Very lucky to invest in companies built by and for AI engineers. And I'll make the same promise that I make every single year, which is short time on bar, long time on bar charts. So, let's get right into it with lots of bar charts.

快速介绍一下我自己,这是最没意思的一页。我是 Amplify 的投资合伙人,很幸运能投资那些由 AI 工程师打造、也服务于 AI 工程师的公司。我还是要做和每年一样的承诺:Barr 本人少讲一点,bar chart(柱状图)多放一点。那我们就直接进正题,上一大堆柱状图。


[1:59]

First, let's talk about well, maybe raise your hand. Did you fill out the survey? This is a very large group. Okay. Yes, I see you in the front. Um, if the answer is you, thank you so much. If the answer is not you, I will find you in 2027. But genuinely, this only exists because a thousand of you gave your time. So, thank you. We had 1,048 respondents this year, which is a lot of AI engineers.

首先我们聊聊——要不举个手吧,谁填了这份问卷?在座人真多。好,我看到前排有人举手了。如果你填了,非常感谢。如果你没填,2027 年我会找到你的。但说真的,这份报告能存在,全靠你们一千多人贡献了自己的时间,所以谢谢大家。今年我们收到了 1,048 份回答,这是相当大的一批 AI 工程师。


[2:26]

And to be precise, this is not just AI engineers, as I'm sure you see at the conference. Every year, we see that AI engineering is more of a discipline than a job title. It touches founders, CTO's, engineers, product people, folks across company sizes and experience levels.

准确地说,填问卷的不只是 AI 工程师,就像你们在这个大会上看到的一样。每一年我们都发现,AI 工程更像是一门学科,而不是一个职位头衔。它覆盖创始人、CTO、工程师、产品人,横跨各种公司规模和经验层级。


[2:43]

And that range shows up in experience too. Um for the third year running we see the same pattern which is skew towards senior engineers but newer to AI. Of those with over 10 years of software experience over half have three years or less of AI experience which tracks uh these are very experienced engineers learning a new paradigm in real time. And the newest cohort, the ones who just started uh engineering, the median new engineer has nearly as much AI experience as the median 10-year software veteran. Uh so the newest engineers have never known software without this.

这种跨度在从业经验上也体现出来。连续第三年,我们看到同样的规律:受访者偏资深工程师,但接触 AI 的时间比较短。在软件从业超过 10 年的人里,超过一半的人 AI 经验在三年以内,这也说得通——这些非常有经验的工程师正在实时学习一种新范式。而最新的这一批、刚开始做工程的人,中位数的新工程师所拥有的 AI 经验,几乎和中位数的十年软件老兵一样多。所以对最年轻的这批工程师来说,他们从来没见过没有 AI 的软件世界。


[3:21]

But doing AI doesn't mean one thing. We talked about all these different titles, all these different roles. Before we get into models and agents, I have a more basic question, which is when people say they're doing AI at work, what are they actually doing?

但「做 AI」并不是一件单一的事。我们刚才说了这么多头衔、这么多角色。在进入模型和 agent 的话题之前,我有一个更基础的问题:当人们说自己在工作中做 AI 的时候,他们实际上在做什么?


[3:36]

So, first up, like to start with the modalities. We asked, which modalities are you actively building with at work? Can anyone take a guess? Text dominates. I know. Hold your applause. Um, but one piece of this chart that I always find very interesting and I always look at is the ratio of nope, I'm not using this modality to I'm not using it, but I do plan to.

先从模态说起。我们问:你在工作中正在用哪些模态做开发?有人猜猜看吗?文本一骑绝尘。我知道,先别鼓掌。但这张图里有一部分我一直觉得特别有意思、每年都会看,就是「不用,我不用这个模态」和「我现在不用,但我打算用」这两者的比例。


[4:00]

I call this the intent to adopt ratio. Of the people who are not building with a modality today, how many say they plan to use it? And audio has the strongest intent to adopt this year. Among AI engineers who are not building with audio today, a whopping 56% say they plan to adopt it in the AI applications they build. And this is not a brand new signal. Last year audio also had the highest intent to adopt across modalities but 37%. So audio continues to take the lead and have high interest but that interest is accelerating.

我把它叫做「采用意愿比」。在今天还没有用某个模态做开发的人里,有多少人说自己打算用?今年采用意愿最强的是音频。在今天还没有用音频做开发的 AI 工程师里,有高达 56% 的人说,他们打算在自己做的 AI 应用里用上音频。而且这不是一个全新的信号——去年音频在所有模态里的采用意愿也是最高的,但当时是 37%。所以音频一直领跑、关注度一直很高,而且这种关注还在加速。


[4:38]

Now there has been an audio swing but if we look at what changed most from the last year in the survey the biggest jump is actually in people using image generation. The share of respondents using generative AI for images and feeling really good about it doubled from 18% last year to 36% this year. Makes sense if you look at what we launched in the same window. Over the past year plus survey time, uh we've had models Nano Banana, Nano Banana 2, Chat GPT images 2.0. The products have gotten much better. What used to feel like an efficient way to generate cursed hands is just increasingly becoming a part of real work. Audio may have the strongest intent to adopt, but image generation shows us what happens when a modality crosses that threshold. So, I'm excited to continue watching these adoption curves every single year. I think we're going to see a lot this year.

音频这边确实有一波变化,但如果看这份调研里相比去年变化最大的一项,最大的跳升其实是用图像生成的人。用生成式 AI 做图像、而且感觉相当不错的受访者占比,从去年的 18% 翻倍到今年的 36%。看看同一时间窗口里发布了什么,就不奇怪了。过去这一年多,再加上做调研的这段时间,我们有了 Nano Banana、Nano Banana 2、ChatGPT Images 2.0 这些模型。产品变好太多了。以前它只是个高效生成「诡异手指」的玩意儿,现在正越来越成为真实工作的一部分。音频可能采用意愿最强,但图像生成让我们看到,一个模态跨过那个临界点之后会发生什么。所以我很期待每年继续追踪这些采用曲线,我觉得今年会看到很多东西。


[5:34]

Uh, now models. If you've Who here spends time on Twitter? All right. Yes. I imagine this is a very Twitter pilled uh crowd. If you spend any time on Twitter in this uh in this circle, you've seen a lot written about openweight models these past few months and I think we'll see it even more in the next year. Um so we asked what models are you actually using in production. 94% use closed models. 45% are using openweight models. But here's the thing, you know, openweight models are not replacing closed models for the most part. at least not yet. The respondents using openweight models, over 90% of them are also using closed models. So they're looking like an augmentation. Teams are mixing and matching.

接下来说模型。有没有人常泡在 Twitter 上?好的,是的。我猜这是一群相当「Twitter 中毒」的人。如果你在这个圈子里刷 Twitter,过去几个月你肯定看到了大量关于 open-weight 模型的讨论,我觉得明年会看到更多。所以我们问:你们在生产环境里实际用的是什么模型?94% 在用闭源模型,45% 在用 open-weight 模型。但关键在于,open-weight 模型在大多数情况下并没有在取代闭源模型,至少现在还没有。在用 open-weight 模型的受访者里,超过 90% 同时也在用闭源模型。所以它们看起来更像是一种补充,团队是在混搭着用。


[6:23]

We also asked just to double click on this for the top three considerations when choosing a model. If you're choosing a model, what is important to you? Um and despite the airtime of the open versus closed, it's not what drives model choice. It was a top three consideration for only 5% of the respondents.

为了在这一点上再深入一层,我们还问了选模型时最看重的前三个因素。如果你要选一个模型,你在意什么?尽管开源与闭源之争占了那么多话题量,它其实并不是驱动模型选择的因素——只有 5% 的受访者把它列进了前三。


[6:43]

What matters is actually more straightforward. It's quality. Quality dominates. Followed by agentic capabilities like tool calling and cost tied right with it. We'll money money. We'll get back to that. Um, one thing that I found very interesting is that reliability is not near the top. Only one in five named reliability. That doesn't mean teams stopped caring about reliability. Uh there are different ways to interpret this data. My guess is that it's more likely to become a threshold requirement and the models they're choosing are reliable enough so the decision moves up the stack outside of certain circumstances to quality, capability, cost, but we could talk after. All right, so here's where the model story all comes together.

真正重要的东西其实更直白:是质量。质量遥遥领先。紧随其后的是 agentic 能力,比如 tool calling,而成本和它并列。钱的事儿,我们一会儿再说。有一点我觉得特别有意思,就是可靠性并没有排在前面,只有五分之一的人提到了可靠性。这不代表团队不在乎可靠性了。这个数据有几种解读方式。我的猜测是,可靠性更可能已经变成一个门槛型要求——他们选的模型可靠性已经够用了,所以除了某些特定场景之外,决策就往上走一层,变成看质量、能力和成本。不过这个我们可以会后再聊。好,接下来模型这条线要收拢到一起了。


[7:28]

Like I said, teams are not choosing one model and calling it a day. Earlier I showed that 87% of teams are using more than one model. Uh the model that's the opposite of standardization. Uh and the way that they choose models for given tasks varies. Most popular is routing by task type. Some run multiple models compare outputs. Some route based on cost. Uh but models are good at different things. What was interesting was that more than half of respondents said that their organizations starting to standardize on fewer AI tools.

就像我说的,团队不会选定一个模型就完事儿。前面我说过,87% 的团队在用不止一个模型,这恰恰是标准化的反面。而且他们给具体任务选模型的方式也各不相同。最常见的是按任务类型做路由;有些人同时跑多个模型、对比输出;有些人按成本做路由。因为模型各有所长。有意思的是,超过一半的受访者说,他们的组织开始在更少的 AI 工具上做标准化了。


[8:02]

They're trading standardiz flexibility for standardization. A share of those are mixed. They say they're standardizing on some layers while staying flexible on others. But the headline here is that there's we're in the early great standardization of the platform and tools, not the models.

他们在拿灵活性去换标准化。其中一部分人是混合状态,说自己在某些层做标准化、在另一些层保持灵活。但这里的核心结论是:我们正处在平台和工具大标准化的早期,而不是模型的标准化。


[8:20]

All right, this is the slide where anyone who's opened an AI bill in the last year starts nodding. So it turns out that infinite intelligence still comes with a usagebased bill. Once teams are managing many models and AI workflows, the next question becomes cost.

好,这一页是给过去一年打开过 AI 账单的人看的,看到这儿你就开始点头了。事实证明,无限的智能背后,还是有一张按用量计费的账单。当团队开始管理很多模型和 AI 工作流之后,下一个问题就是成本。


[8:38]

Cost is now a first class engineering constraint. We see this in the data. 40% of respondents say that cost regularly shapes how ambitiously they use AI and another 36% say that it sometimes does. Well, this is pretty straightforward. So all in about uh three out of four respondents are adjusting their AI usage based on cost and maybe the fourth has a company card.

成本现在已经是一等公民级别的工程约束了。数据里能看到这一点:40% 的受访者说,成本经常会影响他们用 AI 时能有多激进,另有 36% 说有时候会影响。这个挺直白的。加起来,大约四分之三的受访者会根据成本来调整自己的 AI 用量,剩下那四分之一可能是有公司报销卡。


[9:06]

That might be surprising or maybe it's obvious but 12 months ago it was not. Token maxing is cool. Being able to find real use cases is amazing but cost is becoming a real big part of the product decision today.

这也许让人意外,也许显而易见,但 12 个月前并不是这样。疯狂堆 token 很酷,能找到真实的落地场景也很棒,但今天成本正在成为产品决策里非常重要的一部分。


[9:21]

And it shows up in monitoring too. What folks are monitoring in production includes cost and token usage as the number two thing they watch for. It's being monitored like an SLA right under quality itself.

这在监控上也有体现。大家在生产环境里监控的东西中,成本和 token 用量排在第二位。它已经像 SLA 一样被盯着了,就排在质量本身的下面。


[9:36]

Which brings us to the biggest line item of them all. Agents. We've been talking about agents for a while. Uh this year, as you've seen, as you'll see today, as you've seen in previous days, you're going to talk a lot about harness engineering. They're escaping demo world.

这就引出了所有开销里最大的那一项:agent。我们聊 agent 已经聊了一阵子了。今年,正如你们已经看到的、今天将会看到的、以及前几天看到的,大家会大量讨论 harness engineering。agent 正在走出 demo 世界。


[9:53]

So, we asked respondents what level of tool permissions their agents typically have. And this is where agents start to look more real. There are two things happening at once. First, and I don't think this is surprising, relative to last year, there are far more teams using agents. This year, 95%, this seems high to me, 95% say they're using agents, roughly double last year.

所以我们问了受访者,他们的 agent 通常拥有什么级别的工具权限。这里 agent 就开始显得更真实了。有两件事在同时发生。第一,我想这并不意外,相比去年,用 agent 的团队多了非常多。今年是 95%——这个数字我自己都觉得高——95% 的人说他们在用 agent,差不多是去年的两倍。


[10:21]

Second, amongst the teams that are using agents, those agents are much more likely to have write access. Last year, 52% of folks building with agents said their agents could actually write data. This year, that number is 89%.

第二,在这些用 agent 的团队里,他们的 agent 拥有写权限的可能性大大提高了。去年,用 agent 做开发的人里有 52% 说自己的 agent 能真的写数据,今年这个数字是 89%。


[10:38]

So when you combine these two shifts, more teams using agents and more of those agents having write permissions, the share of all the respondents and again it's a survey using write enabled agents is up more than three times relative to last year. So this is really the big shift. Agents are no longer reading, summarizing, drafting. They're taking actions inside of systems. And that raises the obvious question, how are we controlling all of this?

把这两个变化叠加起来看——更多团队在用 agent,而且这些 agent 里更多的拥有写权限——在全部受访者中,使用带写权限 agent 的比例相比去年涨了三倍多。再说一句,这毕竟是问卷数据。所以这才是真正的大变化。agent 不再只是读取、总结、起草了,它们在系统内部真正采取行动。这就带出了一个显而易见的问题:我们要怎么控制这一切?


[11:06]

um with pretty blunt instruments. Uh very there are many ways that folks are controlling agents today. The top two are human in the loop approvals and gating permissions which are the right instincts but kind of the same toolkit you'd use to manage an intern.

答案是:用相当粗糙的手段。今天大家控制 agent 的方式有很多种,排前两位的是人在回路的审批,以及权限门禁。这些直觉是对的,但基本上就是你拿来管一个实习生的那套工具。


[11:23]

Below that the results scatter. Task decomposition, retrieval, memory, sandboxing. People are trying everything. Nobody has settled the control layer for agents. uh memory and persistent context is one that I'm watching very carefully right now. I think it's going to evolve a lot in the next year. And when agents fail or when people complain about agents failing to be more precise, it's usually the thinking, not the plumbing. So, uh you know, like twothirds say that hallucination or losing context mid task is what frustrates them the most.

再往下,结果就散开了:任务拆解、检索、记忆、沙箱。大家什么都在试。还没有人把 agent 的控制层给定下来。记忆和持久化 context 是我现在盯得特别紧的一项,我觉得未来一年它会有很大的演进。而当 agent 出问题的时候——或者更准确地说,当人们抱怨 agent 出问题的时候——问题通常出在「思考」上,而不是「管道」上。比如说,大约三分之二的人说,最让他们抓狂的是幻觉,或者是任务做到一半就把 context 丢了。


[11:57]

All right, so agents are out in the wild, which makes it a good time to look at what everyone's actually running underneath. So let's take a peek at the stack. Um, we asked, what is the biggest challenge in your stack? Every single year that I ask this, the answer, the number one answer is eval. Um, so Eval's lead here is same as always, but by a very thin margin, like that margin is getting smaller. And I'll say the quiet part here, which is that 96% of the people in the survey in this room have a problem with the stack. You just can't agree on which one. Um, so if you're deciding what to build next, if you're interested in infrastructure, that scatter is the map.

好,agent 已经在真实世界里跑起来了,那正好可以看看大家底下到底跑的是什么。我们来看一眼这个技术栈。我们问:你的技术栈里最大的挑战是什么?我每一年问这个问题,排第一的答案都是评估(evals)。所以评估依然领先,和往年一样,但优势非常小,这个差距在不断缩小。我把大家心照不宣的话说出来:这份问卷里、这个房间里,有 96% 的人的技术栈是有问题的,你们只是在「到底是哪个问题」上达不成一致而已。所以,如果你在琢磨接下来做什么产品,如果你对基础设施感兴趣,那张散开的图就是你的地图。


[12:37]

And the leading challenge, how to evaluate your AI outputs requires many different methods, but as always, the vibe review is number one. So there are some consistent things that we'll see if they change over the time, but they they have not changed.

而排在首位的这个挑战——怎么评估你的 AI 输出——需要很多种不同的方法,但一如既往,凭感觉人工过一遍(vibe review)排在第一。所以有些东西是很稳定的,我们会继续观察它们会不会随时间改变,但到目前为止,它们没变。


[12:53]

Okay, this is interesting. So across eight layers of the stack, we asked what do people build versus buy. Um again, maybe the corporate card is is going to play a part in this, but there is a wide range and mix for every layer of the stack and a few clear takeaways.

好,这个挺有意思。我们针对技术栈的八个层级,问了大家是自建还是外购。同样,公司报销卡可能在这里也起了点作用。结果是,技术栈的每一层都有很宽的分布和混合状态,但也有几个很清晰的结论。


[13:11]

So the first is that inference and model serving is the layer that people buy the most. Many people don't want to build inference infrastructure and fair enough. Uh prompt management is the opposite. 61% build it themselves. Um apparently everyone's prompts are special. And this is true of a lot of the product logic prompts rag eval. They tend to stay inhouse on a relative basis. fine-tuning is the clearest not yet. Like most people don't have it at all and uh folks are pretty locked in.

第一,推理(inference)和模型服务是大家买得最多的一层。很多人不想自己搭推理基础设施,这也很合理。prompt 管理正好相反,61% 的人自己做。看来每个人的 prompt 都很特别。产品逻辑相关的很多东西也是如此——prompt、RAG、评估,相对来说都更倾向于留在内部自建。fine-tuning(微调)是最明确的「还没到时候」,大多数人压根就没做,而且大家的态度挺固定的。


[13:46]

So those who bought aren't looking as much to build. Those who built aren't looking as much to buy. Uh but those are those are the core takeaways from the usage in our stack. So many of you work on teams and like we said at the start these range from solo founders to large enterprises.

所以买了现成方案的人,就不太想自己造;自己造过的人,也不太想去买。以上就是我们关于技术栈使用情况的核心结论。你们很多人都在团队里工作,就像开头说的,这些团队从单打独斗的创始人到大型企业都有。


[14:07]

What is this doing to teams? And remember this is a builderheavy sample. But among builders the vibes are good which you know I'm sure if you look to your left and your right you're feeling that the vibes are pretty good.

那这一切对团队产生了什么影响?记住,我们这份样本里做工程的人占比很高。而在这些造东西的人当中,气氛是相当不错的——我想你左右看一看,应该也能感觉到大家的状态挺好。


[14:20]

97% report a net positive effect on their organization. The top effect isn't really just speed. It's cheaper failure, more experimentation, more prototypes, more bets. It didn't just make engineers faster, but it made trying things nearly free. And so there's some happy campers as a result of that.

97% 的人表示 AI 对他们的组织有净正面的影响。而最主要的影响其实不只是速度,而是试错变便宜了:更多的实验、更多的原型、更多的下注。它不只是让工程师变快了,而是让「试一试」这件事几乎变成零成本。所以有一批人因此过得挺开心。


[14:44]

But it's not free free. You know, there's no free lunch as nothing is. So the same tool that increases experimentation also increases review burden. Both can be true. And um you know o over nine and 10 respondents are feeling negative downstream effects in some way. The most common ones being wid you know widely discussed at this conference uh online and anywhere that you see AI engineers erosion of deep technical skills and understanding of the codebase.

但它也不是完全免费的。天下没有免费的午餐,凡事都是有代价的。同一个工具在增加实验量的同时,也增加了代码审查的负担。这两件事可以同时成立。超过十分之九的受访者都在某种程度上感受到了负面的下游影响。其中最常见的几项,在这场大会上、在网上、在任何有 AI 工程师出没的地方都被反复讨论过:深层技术能力的退化,以及对代码库理解的退化。


[15:16]

And these are consequences of cheap code generation. And the org chart is really feeling it. So many folks, 81% are saying that AI is blurring the line between their role as engineers and product design and marketing. These stats shocked me. Um, where you feel it the most is shipping software once exclusively the engineers domain. I know folks talk about vibe coding and how that's accessible to more folks than ever before in different roles, but today over a third of teams have non-developers shipping features, which was pretty wild to me. Mostly smaller, mostly internal, but 17% say that non-developers are regularly shipping customerf facing features across the stack.

这些都是代码生成变便宜之后的后果。而组织架构对此感受尤其明显。有相当多的人——81%——表示 AI 正在模糊他们作为工程师,与产品、设计、市场这些角色之间的界线。这些数据让我很震惊。感受最强烈的地方是「发布软件」这件事,它以前是工程师的专属领地。我知道大家都在聊 vibe coding,聊它让比以往任何时候都多的、不同岗位的人都能上手。但今天已经有超过三分之一的团队,有非开发人员在发布功能,这一点让我觉得挺离谱的。这些大多是较小的、内部的东西,但仍有 17% 的人表示,非开发人员会经常在整个技术栈上发布面向客户的功能。


[16:09]

And even when non-developers aren't shipping, a third of teams see them building really useful things, prototypes, front-end mocks, and more. So, shipping software is not gated on being an engineer. We knew this, but uh the extent to which it's being pushed is is higher than I expected.

而即便非开发人员不做发布,也有三分之一的团队看到他们在做出真正有用的东西:原型、前端 mock,等等。所以,发布软件已经不再以「你得是个工程师」为门槛了。这一点我们本来就知道,但它被推进的程度比我预想的要高。


[16:27]

All right, so where does all of this go? We always ask people to place bets rapid fire. So, let's talk about those results. Um, so present tense first. 76% say AI boosted their job satisfaction. So that's good for most of this crowd. I hope you're uh as uh Alphaba and Glenda say, I hope you're happy now. Um, that's great. But 59% fear today's AI code creates long-term liabilities.

好,那这一切会走向哪里?我们每年都会让大家快问快答地押注一下。我们来看看结果。先说现在时。76% 的人说 AI 提升了他们的工作满意度。对在座大多数人来说这是好事。用 Elphaba 和 Glinda 的话说,希望你们现在开心了。这挺好的。但也有 59% 的人担心,今天的 AI 写出来的代码会留下长期的隐患。


[17:00]

Only a third call software engineering a solved problem. Although uh when I have conversations with folks sometimes the way in which they define software engineering is different. So you can read into that stat as you will. Um happier faster but embracing the maintenance build is the TLDDR and people are unsure what's going to happen with hiring.

只有三分之一的人认为软件工程已经是一个被解决的问题。不过我跟大家聊的时候发现,有些人对「软件工程」的定义并不一样。所以这个数据你可以自己去解读。总结一下就是:更开心、更快,但要接受维护成本的上升;另外大家对招聘会怎么变很不确定。


[17:22]

And for the five-year bets we have 67% expect a leading lab will declare AGI in the next five years. Note the wording. We said will, we asked about the press release, not the achievement. So, will they declare it?

至于五年期的押注,67% 的人预计会有一家头部实验室在未来五年内宣布实现 AGI。注意我们的措辞。我们问的是「会不会宣布」,问的是那份新闻稿,而不是这件事真的被做到。所以,他们会宣布吗?


[17:37]

Yes. What does that mean? Not sure. Uh, only 9% bet on Transformers being state-of-the-art in 5 years. Most are unsure. Uh, but that was interesting. And then my favorite, will there be more AI compute in space or on land? 36 yes.

会。那这意味着什么?不好说。只有 9% 的人押注 Transformer 在五年后仍然是最先进的架构。大多数人表示不确定。这一点挺有意思的。然后是我最喜欢的一题:五年后 AI 算力是在太空里更多,还是在地面上更多?36% 的人说是太空。


[17:55]

38 no. The most divisive question in the survey is about outer space. I promised you a lot of bar charts and that was a lot of information. So a review or our 2026 wrapped impact is overwhelmingly positive. Image genen doubled or happy image gen doubled while audio has the highest adoption intent the same as last year. cost really became a first class constraint and we see that everywhere in monitoring in how ambitious folks that are going out and building AI products are behaving.

38% 的人说不是。这份调研里最有争议的问题,居然是关于外太空的。我答应过给你们看很多柱状图,刚才信息量确实很大。我们回顾一下我们的「2026 年度总结」:AI 的影响压倒性地是正面的。图像生成的采用翻了一倍,而音频的采用意愿最高——这跟去年一样。成本真正成了一个头等约束,这一点我们到处都能看到:在监控上,在那些雄心勃勃出去做 AI 产品的人的行为方式上。


[18:35]

Open weights augment but they don't replace. So we're seeing a multimodel future with a consolidation of the stack. Agents got right access more than ever before, tripling relative to last year. While the guardrails stayed pretty primitive, and inference is the buy market, everything closer to product logic tends to relatively stay more in-house. It is a very exciting time to be an AI engineer. I cannot wait to see how the next year unfolds.

开放权重模型是补充,但并没有取代闭源模型。所以我们看到的是一个多模型并存的未来,同时技术栈在收敛。agent 拿到写权限的比例前所未有地高,相比去年翻了三倍。而与此同时 guardrails 还相当原始。推理是一个「买」的市场,而越靠近产品逻辑的部分,越倾向于留在自己手里做。现在做一个 AI 工程师真的是非常令人兴奋的时代。我已经迫不及待想看看明年会怎么展开了。


[19:04]

So, you can find the full report in the link up here. Every chart plus some cuts that we didn't have time for today. Um, I won't ask you to fill out a survey about the survey, but if there's something that you want on the books for 2027, something you're curious about, you can come find me here on the internet. I'm easy to spot. Thank you so much. Uh, we will see you next year or per 36% of you, maybe in orbit. Thank you.

完整报告可以通过上面这个链接找到。里面有全部的图表,还有一些我们今天没时间讲的细分数据。我不会再让你们填一份「关于这份调研的调研」,但如果你有什么想放进 2027 年问卷里的东西,有什么好奇的问题,可以在现场找我,也可以在网上找我,我挺好认的。非常感谢大家。我们明年见——或者按照你们中 36% 的人的说法,也许是在轨道上见。谢谢。