Over the past year, $NBIS has quietly become one of the most powerful players in the AI infrastructure space, not just by rapidly expanding its data center footprint, but by dramatically leveling up its software stack.
Yet despite this momentum, many people still don’t fully get what Nebius actually does, or why its customers are so excited and pleased about it.
To answer that, I’ve put together a breakdown of some of the most impressive ways real teams and companies are using Nebius today.
I won’t waste anymore of your time, so let’s dive in:
1) SGLang
SGLang is an open-source framework designed to serve large language and vision-language models quickly and efficiently. It aims to make working with LLMs faster and easier to control by combining a high-performance backend with a flexible front-end interface for building AI applications. It also includes smart features that help models respond faster, handle multiple requests at once, and run more efficiently on limited hardware.
With the help of Nebius, SGLang was able to:
• Double throughput for DeepSeek R1 whilesignificantly reducing delays, even when handling many users at once or managing long, complex input data.
• Integrate advanced optimization methods such as FlashAttention-3 (which improves memory and speed), FP8 inference with DeepGEMM (which makes calculations more efficient), and partial cache reuse to improve overall performance.
• Took advantage of Nebius AI Cloud’s flexible infrastructure, allowing rapid testing of ideas and faster development through collaborative engineering.
• Deploy at scale, delivering consistently fast and accurate responses across dozens of concurrent users, all without sacrificing output quality.
In essence, SGLang demonstrates that when cutting-edge optimizations are combined with strong infrastructure, even the most demanding AI models can run efficiently in real-world scenarios.
2) vLLM
vLLM is an open-source project backed by theLinux Foundation, built to make running LLMs faster and more cost-efficient at scale. It helps organizations serve LLMs with better performance and lower infrastructure costs. The project is actively developed by contributors from major tech leaders like UC Berkeley, Meta, Hugging Face, NVIDIA, Google, AWS, Intel, and more, making it a strong, community-driven effort.
With the help of Nebius, vLLM was able to:
• Accelerate performance for large-scale AI tasks, experimenting with cutting-edge methods such as multi-token prediction and multi-latent attention, which improve how quickly and accurately models respond.
• Run performance tests on a wide variety of hardware setups using Nebius’ cloud platform, helping them fine-tune for maximum throughput (speed), low latency (delay), and high concurrency (ability to handle many users).
• Streamline model loading and data access with fast local storage, reducing delays and increasing efficiency.
• Use on-demand computing resources to support continuous testing and development, without interruptions or infrastructure bottlenecks.
• Focus their engineering efforts on innovation, leaving the infrastructure and operations to Nebius.
In short, vLLM turned state-of-the-art AI research into fast, practical tools ready for production, all while keeping costs under control.
3) SieveStack
SieveStack is building the world’s largest dataset of molecular simulations to train foundational models for AI-powered drug discovery, tackling one of the most complex and critical challenges in medicine.
With support from Nebius + TractoAI, SieveStack was able to:
• Maintain over 90% GPU efficiency during simulations and training by using a strategy that smartly spreads out tasks, adjusts job sizes on the fly, and ensures smooth data flow using high-performance storage.
• Speed up training by up to 4×, taking full advantage of Nebius’ infrastructure, including powerful H100 GPUs and precision-focused optimization techniques.
• Reduce the time from testing to real-world use by 30–50%, thanks to TractoAI’s tools that allow direct deployment from research notebooks, no environment switching or outdated systems required.
• Prototype, debug, and scale foundational models faster than before, enabling a tightly integrated workflow from molecular simulation to training and evaluation, significantly shortening the time from experiment to insight.
All of this gives them a huge edge in discovering new medicines for hard-to-treat conditions, far beyond what traditional labs can achieve.
4) SynthLabs
SynthLabs is a startup improving large language models using open science. They combine human feedback with AI-generated feedback to create scalable, interpretable methods that make models easier to control, reduce the need for human labeling, and help models adapt more efficiently to different tasks.
With the help of Nebius, SynthLabs was able to:
• Rapidly develop and scale Big Math, the largest open-source dataset of challenging math problems specifically designed for training AI using reinforcement learning.
• Evaluate over 650,000 problems on a serverless infrastructure without managing hardware directly, using TractoAI’s distributed platform.
• Simplify complex data processing by using a system that automatically distributes tasks across hundreds of GPUs, allowing researchers to focus on model development rather than infrastructure details.
• Experiment with new model alignment methods like RLAIF (Reinforcement Learning from AI Feedback), which guide models to reason better using structured feedback loops.
In effect, SynthLabs turned advanced AI research into usable tools that help bridge the gap between model training and real-world reasoning capabilities.
5) YerevaNN
YerevaNN is a non-profit AI research lab based in Yerevan, Armenia. Since 2016, it has focused on advancing machine learning by developing scalable models for use in biotechnology, including areas like molecular data, multispectral imagery, and radar signals.
With the help of Nebius, YerevaNN was able to:
• Train LLMs for AI-assisted drug designusing a unique dataset of 110 million molecules. This required highly efficient use of memory and speed, made possible by technologies from Nebius.
• Achieve extremely high inference speeds of up to 180,000 words per second, thanks to techniques such as optimized batching, token processing in parallel, and multi-threaded data loading.
• Accelerate molecular simulations by 12×using powerful CPUs (without interfering with ongoing GPU training) allowing them to simulate real-life drug interactions more efficiently.
• Validate massive datasets with ease, by using H100 GPUs and TractoAI’s distributed systems to run hundreds of thousands of AI tasks simultaneously.
Thanks to Nebius, YerevaNN is scaling open-source research in drug discovery, showing how lean, high-impact AI teams can tackle complex problems with the right infrastructure, from molecule to model, and from data to discovery.
6) Chatfuel
Chatfuel is a no-code platform that helps small and medium businesses automate customer communication on WhatsApp, Facebook Messenger, Instagram, and websites. It streamlines tasks like appointment booking, lead qualification, personalized sales, and targeted re-engagement campaigns, no coding needed.
With the support of Nebius (via TractoAI and Nebius Studio), Chatfuel was able to:
• Replace expensive models like GPT-4 with a cascade of smaller Llama-based models. This switch improved chatbot accuracy by 24%, reduced waiting time, and significantly lowered operational costs.
• Deploy production-ready AI systems in under 30 days, thanks to close support from Nebius engineers and streamlined workflows in TractoAI Studio.
• Thoroughly evaluate chatbot performance through a three-stage testing process: technical benchmarks, internal product team reviews, and direct customer feedback.
• Reduce data labeling efforts by training high-quality models on carefully curated small datasets, which Nebius’ tools handled with high efficiency.
• Integrate new models into their platform in just three days, using a custom software development kit (SDK) provided by Nebius.
Thanks to Nebius, Chatfuel now delivers faster, smarter, and more affordable AI-powered chat experiences, helping thousands of businesses turn conversations into conversions.
7) Positronic Robotics
Positronic Robotics is a startup building AI-powered control systems for cleaning robots. Their goal is to make robots smart enough to eventually outperform humans at cleaning. For now, the robots assist human cleaners, but over time they’re designed to become fully autonomous.
With Nebius, they accomplished:
• Training and scaling their third-generation (Gen 3) robot control models usingNVIDIA H100 GPUs in containerized environments for consistency and portability.
• Reducing training time throughautomated, distributed computing pipelines, with each training cycle completing in just 12–16 hours.
• Processing over 1 terabyte of sensory and virtual-reality-captured data directly on Nebius' infrastructure, enabling growth as models and use cases expand.
• Deploying their unique Action Chunking Transformers (ACT) models, which allow robots to turn tasks like scrubbing into smooth, lifelike sequences.
• Minimizing manual setup with scripts that automate the full training lifecycle, including smart shutdowns to save resources when not in use.
By leveraging Nebius’ cloud infrastructure, Positronic Robotics is steadily bringing fully autonomous cleaning robots closer to reality, starting with hotel bathrooms and expanding toward broader commercial and home environments.
8) Recraft
Recraft is the first generative AI model built specifically for designers. Recently backed by Khosla Ventures and former GitHub CEO Nat Friedman, it was trained from scratch on Nebius AI and features 20 billion parameters, nearly 8× more than leading open-source alternatives like Stable Diffusion XL.
With the support of Nebius, Recraft was able to:
• Access to a fully managed Kubernetes environment, which simplified the process of scaling and orchestrating GPU clusters needed for large-scale model training.
• Leveraged NVIDIA’s high-performance networking tools to dramatically improve how fast GPUs communicate, speeding up training cycles.
• Resolved complex hardware/software issues (such as network slowdowns and rare bugs) through direct collaboration with Nebius solution architects.
• Implemented a robust system of alerts and performance monitoring that reduced downtime and ensured high resource efficiency.
As a result, Recraft successfully trained a massive, state-of-the-art design model with stable, secure infrastructure, delivering top-tier generative results for professional design use.
9) TheStage AI
TheStage AI helps teams improve the speed and cost-efficiency of their deep learning models, especially during inference. Their tools identify performance bottlenecks and apply mathematical methods to make models run faster, even on lower-cost hardware.
With Nebius’ support, they were able to:
• Cut GPU costs by up to 3× in practical cases by applying advanced methods such as quantization (reducing model size without hurting quality) and sparsification (skipping unnecessary calculations).
• Ensure constant access to H100 GPUs, letting researchers instantly run large experiments without waiting for resources.
• Base their product, ANNA, on a rigorous mathematical framework recognized at top conferences like CVPR 2023, ensuring scientific soundness.
• Work directly with clients on Nebius virtual machines, fine-tuning models to run faster with little or no loss in accuracy, and often without the need for full retraining.
In conclusion, TheStage AI is helping companies deploy smarter, faster AI with less infrastructure hassle and lower operating costs.
10) Krisp
Krisp is an AI company known for its real-time noise cancellation technology that removes background sounds and voices during calls, improving audio quality for professionals and businesses. Founded in 2017, the company has since expanded into on-device Speech-to-Text and Accent Localization, a technology that adjusts call center agents’ speech to sound more native to their audience, improving clarity and comprehension.
With Nebius, Krisp was able to:
• Train large voice transformation models up to 80% faster using high-performance GPUs.
• Efficiently manage over 2 terabytes of audio data with fast, reliable SSD storage.
• Deploy lightweight, real-time models onlow-powered devices without sacrificing voice quality, ideal for global call centers.
• Measure model performance using a mix of automated tests and custom metrics for speech clarity and naturalness.
• Meet demanding R&D timelines thanks tostable cloud systems and rapid technical support.
As a result, Krisp is now delivering high-quality, real-time voice enhancements, even on modest hardware, enabling better communication around the world.
11) Dubformer
Dubformer is a secure, AI-powered dubbing and end-to-end localization platform that delivers studio-quality audio in over 70 languages. Designed for the global media industry, Dubformer helps content creators, broadcasters, and streaming services break through language barriers, faster, more affordably, and at scale. Its pipeline includes intelligent transcription, context-aware translation, neural speech synthesis, and audio mastering, all integrated with human review for maximum quality.
With Nebius, Dubformer was able to:
• Train advanced speech synthesis and recognition models using powerful A100 and H100 GPUs, which enabled the production of realistic, expressive AI voiceovers in multiple languages.
• Seamlessly stream and process hundreds of terabytes of audio data using Object Storage, allowing for distributed training at the same scale as large language models.
• Run training jobs continuously, 24/7, using an automated queue of experimental tasks, ensuring constant improvement of their voice models.
• Achieve results that now meet or exceedbroadcast standards, with Dubformer’s AI-generated content airing on major international networks and platforms.
Thanks to Nebius, Dubformer continues to redefine what’s possible in content localization, blending cutting-edge AI with human precision to bring high-fidelity voiceovers to a truly global audience.
12) Unum
Unum is an independent AI lab that builds compact yet powerful models capable of understanding and processing text, images, and video together, a field known as multimodal AI. Unlike larger models that require enormous resources, Unum focuses on makingsmaller, smarter systems that are easy to deploy on phones or edge devices.
With Nebius, Unum accomplished:
• Open-sourced four high-performing multimodal models, that perform as well or better than models 10 to 20 times larger, according to standard benchmarks like MM-Vet and ScienceQA.
• Efficiently trained these models on H100 GPUs, while keeping workflows lightweight enough for use in mobile or embedded devices.
• Solved one of their biggest technical challenges (slow data loading) by designing a fast and tailored infrastructure to move visual data to GPUs quickly and reliably.
• Improved the quality of AI conversations across languages by fine-tuning smaller models with Direct Preference Optimization (DPO), enabling real-time interaction without compromising privacy.
Thanks to Nebius, Unum is setting new standards for efficient, high-performance multimodal AI, showing that with the right data, design, and compute, small can be powerful.
13) The London Institute for Mathematical Sciences (LIMS)
LIMS is Britain’s only independent, non-profit research institute in physics and mathematics. Based at the historic Royal Institution in central London, it supports curiosity-driven research and empowers scientists to pursue fundamental discoveries, from theoretical physics to the foundations of artificial intelligence.
With Nebius’ support, LIMS was able to:
• Investigate how well LLMs can understand and apply logical rules, through experiments involving tasks like cellular automata and algorithmic reasoning.
• Train models on billions of rule-based configurations, made possible by Nebius’ cloud infrastructure, enabling simulations that far exceed the complexity of typical AI benchmarks.
• Achieve over 93% accuracy in certain tasks, demonstrating that LLMs can indeed grasp abstract rules when trained properly, challenging the belief that their architecture inherently limits their reasoning ability.
• Use mathematical and complexity-science tools to explore how models generalize, providing a deeper theoretical understanding of AI cognition and its limits.
As a result, LIMS is pushing the boundaries of how we understand intelligence, both natural and artificial, using rigorous scientific methods supported by scalable infrastructure.
14) Simulacra AI
Simulacra AI is a deep-tech startup developing next-generation AI4Science models that simulate molecules at the quantum level. By combining deep learning with real quantum chemistry, they aim to unlock faster and more accurate ways to discover new materials and drugs.
With Nebius, Simulacra AI was able to:
• Train powerful models that predict molecular behavior using the laws of quantum physics, without needing huge, expensive simulations.
• Scale these models beyond 100 million parameters by using Nebius’ distributed GPU infrastructure, with advanced data and weight partitioning strategies.
• Reduce model compilation time fromhours to minutes using ahead-of-time (AOT) compilation, dramatically increasing development speed.
• Shift their business from model development to offering commercial datasets of pre-computed quantum-accurate molecular properties, saving time for pharmaceutical and material science companies.
In short, Simulacra AI is making it faster and cheaper to discover new molecules by replacing traditional simulations with AI, all powered by Nebius’ flexible cloud infrastructure.
15) Quantori
Quantori is an end-to-end digital solutions provider transforming life sciences and healthcare with advanced AI, cloud engineering, and DevOps. Their latest innovation focuses on generating 3D molecular structures with precise shapes, a critical advancement for modern drug discovery and material design.
With Nebius, Quantori succeeded in:
• Train AI models that can generate new 3D molecules that look and behave like real ones, based on a database of 1.6 million known compounds.
• Achieving over 98% uniqueness in their generated molecules after 1,500 training epochs, and generating new molecular structures in about 2.1 seconds each.
• Moving from older string-based molecule representations to more advanced graph-based 3D models, which better reflect real molecular geometry and chemical properties.
• Trained and scaled their models using 8× H200 GPUs on Nebius infrastructure, enabling high-throughput development with minimal delays.
In conclusion, Quantori is using AI to generate new drug-like molecules quickly and reliably, helping researchers invent better medicines, faster.
Thanks for reading!