NVIDIA DGX Cloud Lepton
Global GPU compute marketplace and AI model deployment by NVIDIA
About NVIDIA DGX Cloud Lepton
NVIDIA DGX Cloud Lepton (formerly Lepton AI, acquired by NVIDIA) is an AI-centric compute marketplace and model serving platform that connects developers to tens of thousands of GPUs across a global network of NVIDIA Cloud Partners (including CoreWeave, Lambda, and tier-1 clouds). Founded by Yangqing Jia (creator of Caffe) and acquired by NVIDIA, DGX Cloud Lepton functions like a high-performance compute marketplace for AI engineering teams. It allows developers to discover available GPU compute across regions and seamlessly deploy, fine-tune, and scale AI workloads with zero Kubernetes overhead.
DGX Cloud Lepton integrates directly with the full NVIDIA enterprise software stack, including NVIDIA NIM (Inference Microservices), NeMo, and NVIDIA Cloud Functions. Developers use Python Photons and simple CLI commands to turn arbitrary PyTorch scripts into auto-scaling microservices running on NVIDIA H100, H200, and Blackwell B200 clusters. The platform provides heterogeneous multi-cloud abstraction, automatic load balancing, scale-to-zero serverless runtimes, and distributed key-value storage, giving enterprise teams instant access to reserved and spot GPU capacity with guaranteed NVIDIA driver and CUDA acceleration.
Install the Lepton CLI with pip install leptonai and authenticate your workspace.
Wrap your model logic in a standard Python Photon class with simple decorators.
Test your Photon locally using the built-in local development server.
Deploy to the cloud in one command: lep photon run -n my-model --resource-shape gpu.a10g.
Interact with your live auto-scaling endpoint via standard REST APIs or Python SDK.
Capabilities & Features
Common Use Cases
python-model-deployment
serverless-photon-containers
private-llm-hosting
batch-ml-pipelines
auto-scaling-gpu-microservices
Frequently Asked Questions
What is a Lepton Photon?
A Photon is a lightweight Python package that encapsulates your AI model, dependencies, and API handlers into an executable, portable container.
Can Lepton scale GPU instances down to zero when idle?
Yes, Lepton supports scale-to-zero autoscaling, automatically pausing containers during inactivity to save cloud budget.
Can I use custom PyTorch models and proprietary weights?
Yes, Lepton allows you to package and deploy any custom PyTorch, Hugging Face, or ONNX model with full security and data isolation.
Free Plan
$10 free monthly cloud credits with full access to standard serverless photon runtimes.
Paid Plan
Pay-as-you-go GPU compute starting at $0.40/hr for T4/A10G up to $2.80/hr for H100 SXM5 instances.
Direct link · Verified & reader-supported
Pros & Cons
Pure Python developer experience with zero Docker or Kubernetes complexity required
Single command deployment from local script to auto-scaling cloud microservice
Extensive library of pre-built Photons for popular open-source models
Instant zero-scaling to eliminate idle GPU compute waste and cut cloud costs
Multi-cloud GPU availability ensuring dependable capacity and zero provisioning delays
Tailored primarily for Python and PyTorch ML developers
Complex multi-cloud networking configurations require enterprise tier
Looking for the best alternatives to NVIDIA DGX Cloud Lepton?
Side-by-side feature matrix, pricing models, and decision frameworks.
Alternatives to NVIDIA DGX Cloud Lepton
Deep comparison hubTogether AI
The fastest cloud for open-source AI
A cloud platform for fine-tuning and running the world's leading open-source AI models at scale.
Baseten
Transform text prompts into stunning images.
Baseten converts textual prompts into high‑quality images using diffusion models. Users can generate visuals for marketing, design, and content.
Modal Labs
Serverless cloud for AI models, batch jobs, and GPU workloads in Python
Modal Labs is a high-performance serverless cloud platform that enables AI engineers and developers to run Python code in the cloud with instant access to thousands of CPUs, GPUs, and persistent network volumes. Founded by former Spotify CTO Erik Bernhardsson, Modal reimagines cloud computing with sub-second cold starts and zero infrastructure configuration. With Modal, you define your container image, dependencies, and GPU hardware directly inside standard Python code using simple decorators (e.g. `@app.function(gpu="H100")`). Modal handles container building, volume mounting, GPU scheduling, and automatic scaling down to zero in milliseconds, making it the premier choice for running generative AI models, ComfyUI video pipelines, and massive parallel batch jobs.
Fireworks AI
Production-grade serverless inference platform for open AI models
Fireworks AI is an enterprise AI inference and model serving platform built to run open-weights LLMs, vision models, and multimodal architectures with lightning-fast speeds and lowest cost. Created by former Meta AI and PyTorch infrastructure engineers, Fireworks powers millions of daily AI requests with sub-100ms time-to-first-token (TTFT) and high token throughput. Fireworks allows developers to seamlessly deploy, fine-tune, and serve models like Llama 3.1/3.3, DeepSeek-R1/V3, Mixtral, Qwen 2.5, and Flux.1 with zero cold starts. It uniquely supports instant LoRA fine-tuning switching on shared GPU infrastructure, allowing thousands of custom fine-tuned adapters to run without paying for dedicated hardware.
LlamaIndex
Leading data framework for connecting custom data sources to LLMs and Agentic RAG workflows.
LlamaIndex is the premier open-source data framework designed to bridge private, enterprise, and unstructured data with large language models. By providing sophisticated data connectors, automated parser modules, semantic chunking algorithms, and multi-document index structures, LlamaIndex enables developers to build context-augmented LLM applications and autonomous knowledge retrieval engines with minimal boilerplate. From parsing complex multi-page PDF documents and financial spreadsheets to orchestrating complex Agentic RAG workflows that query multiple disparate databases, LlamaIndex handles the complete data ingestion, indexing, and query evaluation lifecycle.
Kestra
Declarative event-driven workflow orchestrator for microservices, AI agents, and data pipelines
Kestra is an open-source, event-driven orchestration platform built to automate and coordinate complex data pipelines, microservices, and multi-agent AI systems. With a modern declarative YAML-first architecture, Kestra enables engineering teams to manage scheduled tasks, webhook triggers, distributed compute jobs, and LLM agent pipelines through code or a rich interactive UI. The platform provides over 600+ pre-built plugins spanning major cloud providers (AWS, GCP, Azure), databases (Postgres, Snowflake, BigQuery), and modern AI ecosystems (OpenAI, LangChain, Hugging Face, Vector DBs). Workflows can execute parallel compute tasks, branch conditionally, manage secrets securely, and handle automated retries with exponential backoff. Kestra eliminates the operational overhead of legacy orchestrators by running statelessly on top of modern container runtimes and Kubernetes, providing real-time workflow visualizers, sub-millisecond execution triggers, and enterprise-grade role-based access control.
Compare NVIDIA DGX Cloud Lepton with Alternatives
Side-by-side feature, pricing, and pros & cons breakdowns
NVIDIA DGX Cloud Lepton vs Together AI
Side-by-side comparison
NVIDIA DGX Cloud Lepton vs Baseten
Side-by-side comparison
NVIDIA DGX Cloud Lepton vs Modal Labs
Side-by-side comparison
NVIDIA DGX Cloud Lepton vs Fireworks AI
Side-by-side comparison
NVIDIA DGX Cloud Lepton vs LlamaIndex
Side-by-side comparison
NVIDIA DGX Cloud Lepton vs Kestra
Side-by-side comparison
