The easiest way to understand Nebius Token Factory is to ask what Nebius does after it has built the GPU cluster.
If the answer is only:
“rent the GPU”
then the business can eventually become a commodity.
Nebius is trying to move further up the software stack.
Token Factory is its managed inference platform for running AI models in production. It provides capabilities including serverless and dedicated endpoints, autoscaling, model serving and production inference.
During 2026 Nebius accelerated that strategy through:
This is the software side of the NBIS thesis.
Training creates or improves a model.
Inference uses that model to answer a request.
A frontier training run can consume massive GPU resources for weeks.
Inference happens repeatedly after deployment:
every question,
every generated line of code,
every AI agent action,
every API request.
As AI moves from experimentation to production, inference can become the larger recurring workload.
If several cloud providers all rent access to the same NVIDIA hardware, the customer can compare:
That invites price competition.
A software layer changes the comparison.
If one provider makes the model:
the customer may care less about the raw hourly GPU rate.
Nebius completed the acquisition of Eigen AI in June 2026.
Eigen specializes in inference and model optimization, including post-training techniques designed to improve production performance.
The strategic logic is:
same underlying model
better optimization
=
more useful output from the same infrastructure
Nebius also brought in Clarifai's core engineering and research team and licensed its inference and compute-orchestration technology.
Nebius described the combination this way:
Eigen focuses on model-level optimization.
Clarifai brings system-level optimization.
Together they support a more complete inference stack.
Suppose an unoptimized model needs twice as much GPU time per million tokens.
The customer pays more.
Nebius uses more infrastructure for the same output.
If software can reduce that requirement, one physical cluster can support more customer workload.
That can improve:
revenue capacity
and
margin
without building the same proportion of additional data centers.
Nebius's Q2 shareholder letter said Token Factory production inference workloads increased more than threefold during Q2.
That does not yet make Token Factory a separately disclosed multibillion-dollar business.
But it provides evidence that the software layer is being used rather than existing only as a product roadmap.
On August 24, Nebius said Token Factory would become the first AI cloud to adopt NVIDIA Groq 3 LPX for generation-focused inference alongside Vera Rubin.
Nebius cited third-party benchmark results of roughly 3,400 output tokens per second on a particular model configuration. These are workload-specific performance figures rather than a universal guarantee for every model.
The larger strategic point is more important than the benchmark:
Nebius wants developers to consume different generations of specialized AI hardware through the same Token Factory interface.
An AI agent can make dozens of sequential model calls.
Latency compounds.
If each inference step takes too long, the entire workflow becomes slow.
That creates demand for both:
high throughput
and
low latency.
Token Factory is positioned around making those infrastructure choices less visible to the developer.
Sarah Chen, MEXC senior crypto industry analyst, believes Token Factory matters because the long-term AI cloud winner may not be the company that owns the most GPUs. Hardware supply should eventually become less scarce. When that happens, margins may depend more heavily on software, utilization and developer lock-in. Sarah's MEXC research is available through her author profile.
Chen therefore sees the Eigen and Clarifai transactions as more than small technology acquisitions. They are a test of whether Nebius can sell a higher-value service on top of expensive infrastructure. If Token Factory improves customer economics and keeps workloads on Nebius after GPU rental prices normalize, the software layer could make the company's future margins more durable. If customers can easily move workloads to whichever provider offers the cheapest hardware, Nebius remains much more exposed to commodity cloud pricing.
Nebius has also integrated agentic search through Tavily.
The idea is to give AI agents access to current web information rather than only static model knowledge.
That pushes the platform another step away from:
rent compute
toward:
build and operate production AI systems.
Token Factory is not a crypto token.
The word Token refers to AI model tokens—the units generated and processed by language models.
It has nothing to do with NBISON being a blockchain token.
The names are similar, but the products belong to completely different layers.
The long-term thesis is straightforward:
If Nebius can earn more revenue and margin from:
software + optimized inference + managed services
than from:
raw GPU capacity alone
then the business may deserve a different economic profile.
That has to be proven over time.
For the broader company structure, see What Is Nebius Group?.
A managed AI inference platform for deploying and operating models in production.
No.
Inference and model-optimization capabilities. The acquisition closed in June 2026.
Its core engineering/research team joined Nebius, and Nebius licensed Clarifai inference and orchestration technology.
Nebius said production inference workloads increased more than threefold.
A stronger software layer could increase customer stickiness and value per unit of infrastructure.
Nebius's software products compete in a fast-changing AI market. Company benchmarks, workload growth and technical capabilities do not guarantee durable pricing power, customer retention or future profitability.

比特币硬件钱包制造商Coinkite已警告用户,Coldcard设备存在种子生成问题,影响从4.0.1版本起的所有Mk3固件版本。该警告是在安全研究人员调查一起涉及594.48 BTC(约合3,800万美元)的协同转移事件时发出的。然而,目前尚无公开的技术证据证实Coldcard的问题导致了这些转账。

辉达的毛利率刚刚连续第三季度维持在接近 75% 的水准。 而在同一份新闻稿里,公司下修了这个数字的指引。 营收仍在加速——截至 2026 年 7 月 26 日的当季达 962 亿美元,比一年前的两倍还多,而下一季的指引则是 1,080 亿美元。 也就是说,公司一边加速增长,一边在每一美元营收上赚得更少;而这一组张力解释了大部分的原因,这组张力也是为什么覆盖同一家公司的分析师,连一年后的目标价都无法

AI 推理需要的不只是算力,大规模部署同样需要不断增长、能够快速访问且具备成本效率的存储容量。 NAND 正逐步成为与 HBM、DRAM 并存的 AI 容量层,而不再只是传统商品型存储产品。 Sandisk 预计,到 2030 年企业数据中心闪存需求将达到 1.2 ZB,并正在开发专门面向 AI 推理的 High Bandwidth Flash。 Sandisk、Samsung 等存储厂商开始采用

Key Takeaways Microsoft、Amazon、Alphabet 和 Meta 在 2026 年仍在大规模投入 AI 基础设施和数据中心。 Nvidia 仍是最直接的 AI 资本支出受益者之一,但 AI 云、服务器、网络和光通信也开始呈现明确增长。 CoreWeave、Nebius、Dell 和 Broadcom 提供了 AI 支出转化为营收、订单和待履约收入最清晰的证据之一。 随着