NVIDIA NIM Fleet Overview
Real-time status and reliability for every free NVIDIA NIM endpoint, probed continuously.
- 22
- Endpoints
- 2
- Healthy
- 7
- Incidents
- —
- Last probe
Partial degradation
2 of 22 endpoints healthy
Avg latency
1333ms
Avg throughput
19.2tok/s
Incidents 24h
7critical
Updated —
Loading…
Recommended Now
nemotron-3.5-content-safety
NVIDIA
Why this model
TTFT
161ms
Throughput
24.7 tok/s
Uptime
100%
Volatility
highly unstable
Queue
low
Reliability
100
Active Incidents
20nemotron-3.5-content-safety recovered to healthy
9h ago
INFOnemotron-3.5-content-safety congestion rising
9h ago
WARNllama-3.2-90b-vision-instruct partially recovered
9h ago
INFOllama-3.2-90b-vision-instruct degraded (timeout)
9h ago
CRITICALnemotron-3.5-content-safety recovered to healthy
8h ago
INFOnemotron-3.5-content-safety congestion rising
8h ago
WARNModel Fleet
Search, filter, pin & export · 22 endpoints
Reliability & SLA
Uptime history · time-of-day latency · error budgets
Loading reliability history…
Loading latency history…
About this dashboard
Measured, not reported
What you are looking at
NIM Stats is an independent operational dashboard for the free NVIDIA NIM API endpoints. A collector sends a real chat-completion request to every tracked model on a fixed cadence — roughly every ten minutes — and records what actually happened: time to first token, end-to-end latency, sustained tokens per second, whether the call succeeded or timed out, and how congested the endpoint appeared. Every figure above is one of those measurements, taken from outside NVIDIA's network. None of it is a vendor-published status claim.
Healthy means the endpoint is serving normally. Busy means it is serving but with elevated latency or congestion. Jammed means it is failing or timing out on probe. TTFT is milliseconds until the first token arrives, which is what an interactive chat feels; throughput is sustained tokens per second, which is what a batch job feels. The two rarely rank endpoints the same way, so the fleet table reports both.
How to use it
Pick the highest-reliability endpoint with a healthy status, or read the recommendation at the top of the page. If a call you were already making starts failing, check whether that endpoint is jammed here before you go debugging your own client. Measurements come from a single vantage point on a fixed cadence, so treat them as a strong prior for which endpoint to try first rather than a service-level guarantee — your own latency will vary with geography, network path, and prompt size.
Agents and scripts can read every page on this site as clean Markdown at the same URL by sending Accept: text/markdown, or by appending .md to the path. Start at /llms.txt for what this site covers and when to reach for it. For a time series or per-endpoint history rather than a summary, the NIM Stats API reference documents the public read-only JSON API — no key, no rate limit — with an OpenAPI 3.1 spec at /openapi.json. NIM Stats is not affiliated with, endorsed by, or operated by NVIDIA Corporation.