Inference Endpoints
Catalog
Deploy
F
Log In
Model Catalog
Collection
NVIDIA
NVIDIA's open and NVFP4-optimized models for reasoning, agentic, and long-context workloads. Fast and memory-efficient with native Blackwell acceleration. A strong fit for serving frontier models at scale, cost-effectively.
Inference Task
All Tasks
Price
$ 0 - 40 / hour
0
0.1
0.5
1
5
40
Deploy On
ALL
CPU
GPU
INF2
Inference Engine
All
Llama.cpp
TEI
vLLM
SGLang
License
All Licenses
Hub Models
Browse All Models
2 items
Order by:
Most Recent
nvidia
NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
Text Generation
vLLM Engine
OpenMDW-1.1 License
GPU
8x Nvidia RTX PRO 6000 Blackwell
$
22
nvidia
GLM-5.2-NVFP4
Text Generation
Default Engine
MIT License
GPU
8x Nvidia RTX PRO 6000 Blackwell
$
22