

Nebius is a Nasdaq-listed AI cloud offering large-scale NVIDIA GPU compute plus the Token Factory inference API for open models. EU data residency, enterprise governance, and an NVIDIA investment make it a serious option for training and production inference.
Nebius is a full-stack AI cloud company, listed on Nasdaq under the ticker NBIS and headquartered in Amsterdam, that rents large-scale NVIDIA GPU compute and runs a managed inference API on top of it. It is not a startup in the usual sense. It emerged from the 2024 restructuring of Yandex N.V., kept roughly 1,300 engineers, a Finnish data center, and the AI intellectual property, divested all Russian assets, and resumed Nasdaq trading as Nebius Group in October 2024. The problem it solves is blunt: AI labs and enterprises need NVIDIA-grade clusters for training and inference without building their own data centers, and Nebius provides that as a managed cloud, from raw GPUs billed hourly to a per-token API endpoint, without making customers babysit Kubernetes or node health.
The offering splits into three layers. Nebius AI Cloud is the raw GPU compute, with H100, H200, B200, and B300 instances on InfiniBand-connected chassis, plus free managed Kubernetes and Slurm. Token Factory, launched in November 2025, is the OpenAI-compatible inference API over 60-plus open models including DeepSeek, Llama, and Qwen, with dedicated endpoints, a 99.9% SLA, and one-click fine-tuning. Aether is the enterprise governance layer adding SOC 2 Type II, ISO 27001, and EU data residency. NVIDIA committed a $2 billion strategic investment in March 2026, and Nebius holds multi-billion-dollar capacity deals with Meta and Microsoft.
What Nebius actually sells in June 2026
Think of Nebius as two front doors into the same infrastructure. The first is raw compute: non-virtualized GPU instances connected over InfiniBand, on-demand or preemptible, with managed Kubernetes, Slurm, and storage included at no extra charge. This is where pre-training, large-scale fine-tuning, and custom distributed workloads run. The second is Token Factory, the managed inference API, which gets a developer from a model name to a production endpoint in minutes rather than the afternoon of setup that raw compute requires. It bills per token with no idle GPU cost, autoscales to enormous request volumes, and offers zero-retention inference in EU and US regions. The corporate backdrop matters here because infrastructure is a long-term bet: NVIDIA's $2 billion investment in March 2026, a $27 billion Meta agreement, and a Microsoft deal worth up to $19.4 billion give Nebius a roughly $45 billion contracted backlog and a credible runway, which is reassuring for anyone committing a training pipeline to it.
Nebius versus CoreWeave and Together AI
CoreWeave is the closest GPU-cloud peer, and the differences are structural. Both offer H100, H200, and B200, but Nebius emphasizes custom ODM chassis for lower total cost of ownership and, crucially, bundles a managed inference API (Token Factory) that CoreWeave does not match. CoreWeave runs a larger North American and European footprint and leans on multi-year hyperscaler contracts, while Nebius offers fully self-service hourly GPUs with no minimum commitment and was profitable in 2025 where CoreWeave ran a large loss. Against Together AI, the contrast is infrastructure depth: Together is primarily an inference API with optional dedicated endpoints, whereas Nebius owns physical data centers and can serve both training-grade multi-node clusters and inference from one account, with explicit EU data residency that the US-centric Together does not emphasize. For a European team under GDPR, that residency is often the deciding factor, and lighter GPU needs can also be met by live alternatives like Crusoe or RunPod.
What the GPU cloud and Token Factory cost
GPU pricing is transparent and competitive: an H100 is $3.85 per hour on demand or $2.15 preemptible, an H200 is $4.50 or $2.45, and a B200 is $7.15 or $3.95, with reserved clusters discounted up to 35% for multi-month commitments. Token Factory inference runs from about $0.08 per million tokens for small models up to roughly $1.93 for premium ones, with batch discounts and large cache savings on repeated context. There is no free GPU trial, and a $25 minimum deposit activates billing.
"By leveraging dedicated endpoints, we secured guaranteed performance. Autoscaling was the game-changer, allowing us to handle 200 billion tokens per day without manual intervention." Zulkuf Genc, Director of AI at Prosus, November 2025.
Credits are available, but they are gated rather than open: NVIDIA Inception startups can apply for up to $150,000 through the AI Lift program, and academic researchers can apply for grants. For most teams, the entry point is simply a deposit and the hourly or per-token meter.
The friction of running on Nebius
The most reported pain is billing surprise. There is no single account-pause button or automatic idle shutdown, so resources left running quietly accrue charges, and an evening experiment can become a months-long bill if you forget to delete every VM, disk, and bucket. The platform's own tone invites this.
"The interface feels built for deploy first, optimize later." Automateed review, 2026.
Beyond billing, Nebius acknowledged that its newest B300 hardware was less stable through mid-2025 while it shipped autohealing improvements, self-service support response times of 24 to 48 hours lag the hyperscalers, documentation trails AWS and GCP for edge features, and there is no Asia-Pacific region as of mid-2026, which is a hard blocker for low-latency APAC workloads rather than a preference.
Who Nebius fits, and who it does not
Nebius is a strong fit for EU-based teams that need GDPR-compliant inference with real data residency, for AI labs running multi-node training or large-scale fine-tuning on InfiniBand clusters, for teams that want one vendor for both training compute and a production inference API, and for enterprises needing SOC 2 Type II, ISO 27001, or HIPAA. NVIDIA Inception startups who qualify for AI Lift credits get a generous on-ramp.
It is not the right choice for teams needing APAC region coverage, for beginners who may not actively monitor and shut down resources, for workloads that require proprietary foundation models like GPT or Claude (Token Factory is open-model only), or for anyone wanting the deep community, tutorials, and mature documentation of a hyperscaler. The product is powerful but still maturing around the edges.
User Reviews
No reviews yet. Be the first to share your experience!
Sign in to write a review.
Related articles
Guides and articles related to Nebius.

Mistral Medium 3.5 Self Host: 77.6% SWE-Bench on 4 GPUs (2026)

Grok 4.3 API for Agents (May 2026): Pricing, Benchmarks, Migration

Google Vision AI Explained (2026): Pricing Per 1,000 Units, Free Tier, and Alternatives

Nous Hermes 4: The Self-Hosted Open-Weight Agent Brain (2026)

Local Agentic Coding May 2026: Qwen 3.6 + BeeLlama.cpp + Star Elastic
