← 返回列表
🔗 原文
即将到来的人工智能利润空间崩塌中的赢家与输家(下)
Winners and losers in the coming AI margin collapse (part 2)

This is the second article in a two part series focusing on what I believe is perhaps the least understood upcoming shift in AI economics. If you haven't read it yet, I'd recommend starting with part one. As always, if you enjoy my writing I'd love it if you subscribe to my newsletter or RSS feed.

As always a week is a long time in AI. In the previous article I discussed the impact of "good enough" models for many agentic workflows - focusing on GLM5.2. In the brief spell of time since I wrote that, Grok 4.5 was released with similar capabilities and is also aggressively priced, which strongly hints at a glut of similar quality models coming out.

Your margin is my opportunity

This is one of Bezos's most famous quotes, and for good reason. It illustrates the dynamic in highly competitive markets - any margin becomes a weakness that others can exploit.

I think Grok 4.5's aggressive pricing - at $6/MTok output, it's being offered at a similar cost to hosted GLM5.2 - shows this up. While xAI is unlikely to beat OpenAI or Anthropic at the very frontier of intelligence, it shows exactly where they can get some traction - price.

It's going to be telling to see what happens to pricing over the next few months. It feels like the market is bifurcating into two - expensive very high end models (Fable, and perhaps GPT5.6 Sol) - and then a broad swath of good (~Opus), cheap models. While this has always been the case - a lag between the frontier and everyone else - I strongly believe the dynamic has switched a bit with these models now becoming good enough for many agentic tasks.

The winners

There is no doubt to me the real winners in all this are semiconductor companies and the entire downstream supply chain to LLM inference. Memory, GPUs, datacentres and the power and cooling needed continue to be severely supply constrained. And with cheaper models, microeconomics tells us that demand increases. But my guess is that the value increasingly accrues to the hardware layer of the value chain, not the software layer.

This isn't what we've typically seen in tech, and is a significant adjustment for many working or analysing the space. Typically hardware was viewed as the ugly duckling value wise - with punishingly low margins and poor ability for suppliers to differentiate products in the market for the most part. Software would sit on top and take all the margin.

Now, by no means am I suggesting that hardware will take all the value. But compared to previous technology waves, perhaps with the limited exception of Apple and the iPhone's cash generation ability[1], this has very much been the exception to the rule.

Apart from the hardware supply chain itself, there are opportunities for the hyperscalers/neoclouds and hosted inference providers to take some value from serving these lower cost models. Serving models like this at serious scale is still difficult, and proprietary efficiency improvements these companies can come up with will give them a competitive advantage. And the quality of these companies relationships with the underlying hardware providers do give them an edge, at least until/if supply starts to catch up with demand.

The most interesting case here is the coding agents - the Cursors of the world. For a long time they faced a brutal path forward: they were reselling frontier inference bought at close to retail API prices, which left them with wafer-thin (or outright negative) margins on their heaviest users. "Good enough" cheap models flip that overnight. A coding agent can now offer something 90% of the way to Opus on a model costing a fraction of the price - and actually make money doing it. But the bigger prize is the data they sit on top of: a firehose of real-world agentic usage - which prompts work, which edits developers accept or throw away, and exactly where the model gets stuck. That is the kind of signal a model provider would kill for to train the next generation. It's no wonder xAI bought Cursor - not for the IDE, but for the cheap-model economics and the analytics flywheel underneath it.

But, I think the real winners out of all this are the users and consumers of LLM inference. Being able to access such high quality intelligence for such a relatively low price is hugely exciting. Back when inference APIs were in their infancy, it looked quite possible that there might be just OpenAI providing inference to any reasonable quality, now we have a multitude of models with substantially better intelligence than GPT4 available for 5-10% the price of that model.

The losers

This is where it gets tricky. You're probably expecting for me to say the frontier AI labs, but I'm extremely torn on this.

Predicting AI market dynamics is tough. On one hand, I do believe that there are suddenly a large chunk of AI use cases that can be moved over to open/cheaper models with little to no loss of quality. This is no doubt a real problem - Anthropic reportedly earns around 80% of its revenue from API usage - which does leave them exposed to people switching out their models for cheaper ones.

On the other hand, I think there are two wildcards at play.

Firstly, I strongly suspect that the frontier labs will increasingly move to not releasing their most powerful models to "everyone". And I don't mean for security/safety reasons, though no doubt that is one factor - and may be one explanation for this move.

I can definitely see a world in the near future where you can only access these frontier models through a higher level of abstraction with no API and no direct coding agent use. Instead, to use these models you have to use their managed agent platforms, which makes it much harder to swap out other models - you don't have true control of the harness. And it substantially reduces the risk of model distillation, which makes it a lot harder for the Chinese providers to keep up.

Secondly, this also assumes that "good enough" is good enough for long. It's highly likely that in a few months we'll get a new set of frontier models which are another leap forward that makes the current set of models look like an antique - either in intelligence terms, or perhaps with some other breakthroughs (speed, context length, continual retraining, who knows).

And if we got these leaps forward then suddenly the way we do and think about agentic workflows totally changes again, much like the leap from chat UIs to coding agents.

In essence, I think it comes down to if the frontier labs can keep innovating, and arguably if they can increase their lead over the open weights models. It feels right now to me that isn't happening and the lead is shrinking, but I don't think it's a good bet to make - we've seen so many shifts and changes.

The B2C wildcard

The other wildcard I've been thinking about is the now very overlooked B2C market. With the explosive growth of coding agents and related use cases, the AI market in general has massively shifted to B2B (and really, enterprise) over the past 12 months.

I think the one thing to keep an eye on is if anyone cracks LLM adjacent advertising. OpenAI have rolled this out, Anthropic have ruled out (as far as I know) running ads, and (surprisingly) I'm barely seeing any ads in my Gemini chat sessions from Google.

There are at least 1 billion MAUs of ChatGPT alone - and while this growth has plateaued - this alone represents an enormous amount of consumer engagement which still hasn't been monetised past subscriptions.

If someone does crack this, expect the hype pendulum to start swinging back to B2C again - changing habits of B2C users is often very difficult, as Google's complete and persistent dominance of web search for so long proved.

Conclusion

So where does this leave us? If I had to bet, the margin in pure model inference is heading towards zero. "Good enough" open models, plus a brutally competitive hosting market, see to that. Bezos's line still holds - but the opportunity is being captured either side of the model layer, not by it. Underneath, by the hardware and power supply chain everyone is scrambling over. And on top, by the users and consumers now getting what would have been unthinkable intelligence a couple of years ago for pennies.

The frontier labs are the one group I'd be wary of betting against. They have two ways out of the commodity trap: keep the intelligence lead wide enough that people happily pay a premium for it, or wall the best models off behind managed platforms where you can't simply swap in something cheaper. My hunch is they'll try both. Whether it works comes down to one thing - whether that lead stops shrinking. Right now, it doesn't look like it is.


  1. Though I'd argue that Apple itself and not the downstream supply chain absorbed much of the margin there. ↩︎

🤖 AI 总结
AI推理利润崩溃:硬件供应链和消费者受益,模型推理商品化,前沿实验室的出路是管理代理和保持领先。

This is the second article in a two part series focusing on what I believe is perhaps the least understood upcoming shift in AI economics. If you haven't read it yet, I'd recommend starting with part one. As always, if you enjoy my writing I'd love it if you subscribe to my newsletter or RSS feed.

As always a week is a long time in AI. In the previous article I discussed the impact of "good enough" models for many agentic workflows - focusing on GLM5.2. In the brief spell of time since I wrote that, Grok 4.5 was released with similar capabilities and is also aggressively priced, which strongly hints at a glut of similar quality models coming out.

Your margin is my opportunity

This is one of Bezos's most famous quotes, and for good reason. It illustrates the dynamic in highly competitive markets - any margin becomes a weakness that others can exploit.

I think Grok 4.5's aggressive pricing - at $6/MTok output, it's being offered at a similar cost to hosted GLM5.2 - shows this up. While xAI is unlikely to beat OpenAI or Anthropic at the very frontier of intelligence, it shows exactly where they can get some traction - price.

It's going to be telling to see what happens to pricing over the next few months. It feels like the market is bifurcating into two - expensive very high end models (Fable, and perhaps GPT5.6 Sol) - and then a broad swath of good (~Opus), cheap models. While this has always been the case - a lag between the frontier and everyone else - I strongly believe the dynamic has switched a bit with these models now becoming good enough for many agentic tasks.

The winners

There is no doubt to me the real winners in all this are semiconductor companies and the entire downstream supply chain to LLM inference. Memory, GPUs, datacentres and the power and cooling needed continue to be severely supply constrained. And with cheaper models, microeconomics tells us that demand increases. But my guess is that the value increasingly accrues to the hardware layer of the value chain, not the software layer.

This isn't what we've typically seen in tech, and is a significant adjustment for many working or analysing the space. Typically hardware was viewed as the ugly duckling value wise - with punishingly low margins and poor ability for suppliers to differentiate products in the market for the most part. Software would sit on top and take all the margin.

Now, by no means am I suggesting that hardware will take all the value. But compared to previous technology waves, perhaps with the limited exception of Apple and the iPhone's cash generation ability[1], this has very much been the exception to the rule.

Apart from the hardware supply chain itself, there are opportunities for the hyperscalers/neoclouds and hosted inference providers to take some value from serving these lower cost models. Serving models like this at serious scale is still difficult, and proprietary efficiency improvements these companies can come up with will give them a competitive advantage. And the quality of these companies relationships with the underlying hardware providers do give them an edge, at least until/if supply starts to catch up with demand.

The most interesting case here is the coding agents - the Cursors of the world. For a long time they faced a brutal path forward: they were reselling frontier inference bought at close to retail API prices, which left them with wafer-thin (or outright negative) margins on their heaviest users. "Good enough" cheap models flip that overnight. A coding agent can now offer something 90% of the way to Opus on a model costing a fraction of the price - and actually make money doing it. But the bigger prize is the data they sit on top of: a firehose of real-world agentic usage - which prompts work, which edits developers accept or throw away, and exactly where the model gets stuck. That is the kind of signal a model provider would kill for to train the next generation. It's no wonder xAI bought Cursor - not for the IDE, but for the cheap-model economics and the analytics flywheel underneath it.

But, I think the real winners out of all this are the users and consumers of LLM inference. Being able to access such high quality intelligence for such a relatively low price is hugely exciting. Back when inference APIs were in their infancy, it looked quite possible that there might be just OpenAI providing inference to any reasonable quality, now we have a multitude of models with substantially better intelligence than GPT4 available for 5-10% the price of that model.

The losers

This is where it gets tricky. You're probably expecting for me to say the frontier AI labs, but I'm extremely torn on this.

Predicting AI market dynamics is tough. On one hand, I do believe that there are suddenly a large chunk of AI use cases that can be moved over to open/cheaper models with little to no loss of quality. This is no doubt a real problem - Anthropic reportedly earns around 80% of its revenue from API usage - which does leave them exposed to people switching out their models for cheaper ones.

On the other hand, I think there are two wildcards at play.

Firstly, I strongly suspect that the frontier labs will increasingly move to not releasing their most powerful models to "everyone". And I don't mean for security/safety reasons, though no doubt that is one factor - and may be one explanation for this move.

I can definitely see a world in the near future where you can only access these frontier models through a higher level of abstraction with no API and no direct coding agent use. Instead, to use these models you have to use their managed agent platforms, which makes it much harder to swap out other models - you don't have true control of the harness. And it substantially reduces the risk of model distillation, which makes it a lot harder for the Chinese providers to keep up.

Secondly, this also assumes that "good enough" is good enough for long. It's highly likely that in a few months we'll get a new set of frontier models which are another leap forward that makes the current set of models look like an antique - either in intelligence terms, or perhaps with some other breakthroughs (speed, context length, continual retraining, who knows).

And if we got these leaps forward then suddenly the way we do and think about agentic workflows totally changes again, much like the leap from chat UIs to coding agents.

In essence, I think it comes down to if the frontier labs can keep innovating, and arguably if they can increase their lead over the open weights models. It feels right now to me that isn't happening and the lead is shrinking, but I don't think it's a good bet to make - we've seen so many shifts and changes.

The B2C wildcard

The other wildcard I've been thinking about is the now very overlooked B2C market. With the explosive growth of coding agents and related use cases, the AI market in general has massively shifted to B2B (and really, enterprise) over the past 12 months.

I think the one thing to keep an eye on is if anyone cracks LLM adjacent advertising. OpenAI have rolled this out, Anthropic have ruled out (as far as I know) running ads, and (surprisingly) I'm barely seeing any ads in my Gemini chat sessions from Google.

There are at least 1 billion MAUs of ChatGPT alone - and while this growth has plateaued - this alone represents an enormous amount of consumer engagement which still hasn't been monetised past subscriptions.

If someone does crack this, expect the hype pendulum to start swinging back to B2C again - changing habits of B2C users is often very difficult, as Google's complete and persistent dominance of web search for so long proved.

Conclusion

So where does this leave us? If I had to bet, the margin in pure model inference is heading towards zero. "Good enough" open models, plus a brutally competitive hosting market, see to that. Bezos's line still holds - but the opportunity is being captured either side of the model layer, not by it. Underneath, by the hardware and power supply chain everyone is scrambling over. And on top, by the users and consumers now getting what would have been unthinkable intelligence a couple of years ago for pennies.

The frontier labs are the one group I'd be wary of betting against. They have two ways out of the commodity trap: keep the intelligence lead wide enough that people happily pay a premium for it, or wall the best models off behind managed platforms where you can't simply swap in something cheaper. My hunch is they'll try both. Whether it works comes down to one thing - whether that lead stops shrinking. Right now, it doesn't look like it is.


  1. Though I'd argue that Apple itself and not the downstream supply chain absorbed much of the margin there. ↩︎

原文
Winners and losers in the coming AI margin collapse (part 2)

This is the second article in a two part series focusing on what I believe is perhaps the least understood upcoming shift in AI economics. If you haven't read it yet, I'd recommend starting with part one. As always, if you enjoy my writing I'd love it if you subscribe to my newsletter or RSS feed.

As always a week is a long time in AI. In the previous article I discussed the impact of "good enough" models for many agentic workflows - focusing on GLM5.2. In the brief spell of time since I wrote that, Grok 4.5 was released with similar capabilities and is also aggressively priced, which strongly hints at a glut of similar quality models coming out.

Your margin is my opportunity

This is one of Bezos's most famous quotes, and for good reason. It illustrates the dynamic in highly competitive markets - any margin becomes a weakness that others can exploit.

I think Grok 4.5's aggressive pricing - at $6/MTok output, it's being offered at a similar cost to hosted GLM5.2 - shows this up. While xAI is unlikely to beat OpenAI or Anthropic at the very frontier of intelligence, it shows exactly where they can get some traction - price.

It's going to be telling to see what happens to pricing over the next few months. It feels like the market is bifurcating into two - expensive very high end models (Fable, and perhaps GPT5.6 Sol) - and then a broad swath of good (~Opus), cheap models. While this has always been the case - a lag between the frontier and everyone else - I strongly believe the dynamic has switched a bit with these models now becoming good enough for many agentic tasks.

The winners

There is no doubt to me the real winners in all this are semiconductor companies and the entire downstream supply chain to LLM inference. Memory, GPUs, datacentres and the power and cooling needed continue to be severely supply constrained. And with cheaper models, microeconomics tells us that demand increases. But my guess is that the value increasingly accrues to the hardware layer of the value chain, not the software layer.

This isn't what we've typically seen in tech, and is a significant adjustment for many working or analysing the space. Typically hardware was viewed as the ugly duckling value wise - with punishingly low margins and poor ability for suppliers to differentiate products in the market for the most part. Software would sit on top and take all the margin.

Now, by no means am I suggesting that hardware will take all the value. But compared to previous technology waves, perhaps with the limited exception of Apple and the iPhone's cash generation ability[1], this has very much been the exception to the rule.

Apart from the hardware supply chain itself, there are opportunities for the hyperscalers/neoclouds and hosted inference providers to take some value from serving these lower cost models. Serving models like this at serious scale is still difficult, and proprietary efficiency improvements these companies can come up with will give them a competitive advantage. And the quality of these companies relationships with the underlying hardware providers do give them an edge, at least until/if supply starts to catch up with demand.

The most interesting case here is the coding agents - the Cursors of the world. For a long time they faced a brutal path forward: they were reselling frontier inference bought at close to retail API prices, which left them with wafer-thin (or outright negative) margins on their heaviest users. "Good enough" cheap models flip that overnight. A coding agent can now offer something 90% of the way to Opus on a model costing a fraction of the price - and actually make money doing it. But the bigger prize is the data they sit on top of: a firehose of real-world agentic usage - which prompts work, which edits developers accept or throw away, and exactly where the model gets stuck. That is the kind of signal a model provider would kill for to train the next generation. It's no wonder xAI bought Cursor - not for the IDE, but for the cheap-model economics and the analytics flywheel underneath it.

But, I think the real winners out of all this are the users and consumers of LLM inference. Being able to access such high quality intelligence for such a relatively low price is hugely exciting. Back when inference APIs were in their infancy, it looked quite possible that there might be just OpenAI providing inference to any reasonable quality, now we have a multitude of models with substantially better intelligence than GPT4 available for 5-10% the price of that model.

The losers

This is where it gets tricky. You're probably expecting for me to say the frontier AI labs, but I'm extremely torn on this.

Predicting AI market dynamics is tough. On one hand, I do believe that there are suddenly a large chunk of AI use cases that can be moved over to open/cheaper models with little to no loss of quality. This is no doubt a real problem - Anthropic reportedly earns around 80% of its revenue from API usage - which does leave them exposed to people switching out their models for cheaper ones.

On the other hand, I think there are two wildcards at play.

Firstly, I strongly suspect that the frontier labs will increasingly move to not releasing their most powerful models to "everyone". And I don't mean for security/safety reasons, though no doubt that is one factor - and may be one explanation for this move.

I can definitely see a world in the near future where you can only access these frontier models through a higher level of abstraction with no API and no direct coding agent use. Instead, to use these models you have to use their managed agent platforms, which makes it much harder to swap out other models - you don't have true control of the harness. And it substantially reduces the risk of model distillation, which makes it a lot harder for the Chinese providers to keep up.

Secondly, this also assumes that "good enough" is good enough for long. It's highly likely that in a few months we'll get a new set of frontier models which are another leap forward that makes the current set of models look like an antique - either in intelligence terms, or perhaps with some other breakthroughs (speed, context length, continual retraining, who knows).

And if we got these leaps forward then suddenly the way we do and think about agentic workflows totally changes again, much like the leap from chat UIs to coding agents.

In essence, I think it comes down to if the frontier labs can keep innovating, and arguably if they can increase their lead over the open weights models. It feels right now to me that isn't happening and the lead is shrinking, but I don't think it's a good bet to make - we've seen so many shifts and changes.

The B2C wildcard

The other wildcard I've been thinking about is the now very overlooked B2C market. With the explosive growth of coding agents and related use cases, the AI market in general has massively shifted to B2B (and really, enterprise) over the past 12 months.

I think the one thing to keep an eye on is if anyone cracks LLM adjacent advertising. OpenAI have rolled this out, Anthropic have ruled out (as far as I know) running ads, and (surprisingly) I'm barely seeing any ads in my Gemini chat sessions from Google.

There are at least 1 billion MAUs of ChatGPT alone - and while this growth has plateaued - this alone represents an enormous amount of consumer engagement which still hasn't been monetised past subscriptions.

If someone does crack this, expect the hype pendulum to start swinging back to B2C again - changing habits of B2C users is often very difficult, as Google's complete and persistent dominance of web search for so long proved.

Conclusion

So where does this leave us? If I had to bet, the margin in pure model inference is heading towards zero. "Good enough" open models, plus a brutally competitive hosting market, see to that. Bezos's line still holds - but the opportunity is being captured either side of the model layer, not by it. Underneath, by the hardware and power supply chain everyone is scrambling over. And on top, by the users and consumers now getting what would have been unthinkable intelligence a couple of years ago for pennies.

The frontier labs are the one group I'd be wary of betting against. They have two ways out of the commodity trap: keep the intelligence lead wide enough that people happily pay a premium for it, or wall the best models off behind managed platforms where you can't simply swap in something cheaper. My hunch is they'll try both. Whether it works comes down to one thing - whether that lead stops shrinking. Right now, it doesn't look like it is.


  1. Though I'd argue that Apple itself and not the downstream supply chain absorbed much of the margin there. ↩︎

中文翻译
即将到来的人工智能利润空间崩塌中的赢家与输家(下)

这是两篇文章系列的第二篇,聚焦于我认为可能是人工智能经济学中最不为人知的即将到来的转变。如果你还没读过,我建议从第一篇开始。一如既往,如果你喜欢我的文章,我很希望你能订阅我的新闻通讯RSS 源

一如既往,一周在人工智能领域是漫长的时间。在上一篇文章中,我讨论了“足够好”的模型对许多代理工作流的影响——重点介绍了 GLM5.2。在我写完那篇文章后的短暂时间里,Grok 4.5 发布了,具有类似的能力,并且定价也很激进,这强烈暗示着一波类似质量的模型即将涌现。

你的利润就是我的机会

这是贝佐斯最著名的名言之一,而且理由充分。它说明了高度竞争市场中的动态——任何利润率都会成为他人可以利用的弱点

我认为 Grok 4.5 的激进定价——输出每百万 token 6 美元,与托管的 GLM5.2 成本相似——就说明了这一点。虽然 xAI 不太可能在智能的最前沿击败 OpenAI 或 Anthropic,但这恰恰显示了他们能在哪里获得一些 traction——价格。

未来几个月定价会发生什么,这将具有启示意义。感觉市场正在分化为两个极端——昂贵的非常高端的模型(Fable,也许还有 GPT5.6 Sol)——以及一大片“足够好”(约等于 Opus)、廉价的模型。虽然情况一直如此——前沿与其他模型之间存在滞后——但我坚信,随着这些模型现在对许多代理任务变得“足够好”,这种动态已经发生了一些转变。

赢家

毫无疑问,在我看来,真正的赢家是半导体公司以及 LLM 推理的整个下游供应链。内存、GPU、数据中心以及所需的电力和冷却仍然严重供应受限。而随着模型变得更便宜,微观经济学告诉我们需求会增加。但我猜测,价值将越来越多地流向价值链的硬件,而不是软件层。

这不是我们在科技领域通常看到的情况,对于许多在该领域工作或分析的人来说,这是一个重大调整。通常硬件在价值上被视为丑小鸭——利润率极低,而且供应商在很大程度上很难在市场上实现产品差异化。软件则坐享其成,攫取所有利润。

现在,我绝不是在暗示硬件会占据全部价值。但与以往的技术浪潮相比——也许苹果和 iPhone 的现金生成能力是一个有限的例外[1]——这一直是非常例外的情况。

除了硬件供应链本身,超大规模云服务商/新云服务商以及托管推理提供商也有机会从提供这些低成本模型中获取一些价值。大规模提供这样的模型仍然困难重重,这些公司能想出的专有效率改进将赋予它们竞争优势。而且这些公司与底层硬件供应商的关系质量确实给了它们优势,至少直到(或除非)供应开始赶上需求。

这里最有趣的案例是编码代理——世界上的 Cursor 们。很长一段时间里,它们面临一条残酷的道路:它们以接近零售 API 价格转售前沿推理,这导致它们在最重度用户身上只有微薄(甚至完全为负)的利润率。而“足够好”的廉价模型在一夜之间扭转了局面。编码代理现在可以提供在能力上达到 Opus 90% 的模型,而成本只是其一小部分——并且实际上可以从中盈利。但更大的奖励是它们所掌握的数据:一个真实世界代理使用情况的洪流——哪些提示有效,哪些编辑被开发者接受或抛弃,以及模型究竟在哪里卡壳。这正是模型提供商梦寐以求用来训练下一代模型的信号。难怪 xAI 收购了 Cursor——不是为了 IDE,而是为了廉价模型的经济效益以及其下的分析飞轮。

但是,我认为所有这一切中真正的赢家是 LLM 推理的用户和消费者。能够以如此相对低廉的价格获得如此高质量的智能,这非常令人兴奋。回想推理 API 还处于萌芽阶段时,似乎很有可能只有 OpenAI 能提供任何合理质量的推理,现在我们有了众多模型,其智能水平远超 GPT4,而价格只有该模型的 5-10%。

输家

这就是棘手的地方。你可能以为我会说是前沿 AI 实验室,但我对此非常纠结。

预测 AI 市场动态很难。一方面,我确实相信突然之间有很大一部分 AI 用例可以转移到开源/更便宜的模型上,而几乎没有质量损失。这无疑是一个真正的问题——据报道,Anthropic 约 80% 的收入来自 API 使用——这确实让它们面临用户转向更便宜模型的风险。

另一方面,我认为有两个变数在起作用。

首先,我强烈怀疑前沿实验室将越来越多地倾向于不向“所有人”发布其最强大的模型。而我的意思不是出于安全原因,尽管这无疑是一个因素——也可能是这种举措的一个解释。

我绝对可以预见到在不久的将来,你只能通过更高层次的抽象来访问这些前沿模型,没有 API,没有直接的编码代理使用。相反,要使用这些模型,你必须使用它们的托管代理平台,这会使切换其他模型变得困难得多——你无法真正控制 harness。而且这大大降低了模型蒸馏的风险,使中国提供商更难跟上。

其次,这也假设“足够好”在长期内仍然是足够好的。很可能在几个月后,我们会得到一组新的前沿模型,它们将是又一次飞跃,使当前一代模型看起来像是古董——无论是在智能方面,还是可能在其他一些突破上(速度、上下文长度、持续再训练,谁知道呢)。

如果我们迎来这些飞跃,那么我们思考和执行代理工作流的方式将再次完全改变,就像从聊天 UI 到编码代理的飞跃一样。

本质上,我认为这归结于前沿实验室是否能够持续创新,并且可以说是否能够扩大它们相对于开源权重模型的领先优势。在我现在看来,这种情况并没有发生,领先优势正在缩小,但我认为这不是一个稳妥的赌注——我们已经看到了太多的转变和变化。

B2C 变数

我一直在思考的另一个变数是现在非常被忽视的 B2C 市场。随着编码代理及相关用例的爆炸式增长,总的来说,AI 市场在过去 12 个月里已大幅转向 B2B(而且实际上是企业级)。

我认为需要关注的一点是,是否有破解了 LLM 相关的广告。OpenAI 已经推出,Anthropic 已经排除了(据我所知)投放广告的可能性,而且(令人惊讶的是)我在 Google 的 Gemini 聊天会话中几乎看不到任何广告。

仅 ChatGPT 就有至少 10 亿月活跃用户——虽然增长已经趋于平稳——但这本身就代表了巨大的消费者参与度,而且仍然没有在订阅之外实现货币化。

如果有人确实破解了这一点,那么热度钟摆预计将开始摆回 B2C——改变 B2C 用户的使用习惯通常非常困难,正如谷歌在网络搜索领域长期以来的完全且持续的主导地位所证明的那样。

结论

那么,这给我们带来了什么?如果我必须下注,纯粹模型推理的利润率正趋于零。“足够好”的开源模型,加上残酷竞争的托管市场,确保了这一点。贝佐斯的那句话仍然成立——但机会正在被模型层的两侧捕获,而不是模型层本身。在底层,是被所有人争抢的硬件和电力供应链。在上层,是用户和消费者,他们现在只需花上几个钱就能获得几年前还难以想象的智能。

前沿实验室是我不敢轻易押注失败的一个群体。它们有两条路走出商品化的陷阱:保持智能领先优势足够大,让人们乐于为此支付溢价;或者将最好的模型封闭在托管平台之后,让你无法简单地用更便宜的东西替换。我的直觉是它们会双管齐下。是否有效取决于一件事——那个领先优势是否停止缩小。目前看来,它并没有停止。


  1. 尽管我认为苹果公司本身,而不是下游供应链,吸收了那里的大部分利润。↩︎