Fleet Operations
Discover
Live operational intelligence across every free NVIDIA NIM endpoint — ranked by throughput, reliability, and congestion, refreshed continuously.
- 22
- Endpoints
- 5
- Providers
- 20
- Degraded
- —
- Last probe
Fleet Degraded2 of 22 endpoints operational · 20 degraded
2713
LIVEThroughput▲34%
23.6tok/s
P95 Latency▼14%
5,764ms
Reliability▲0.1pt
58.6%
Congestion▼0.1pt
48%
Model Registry
22 endpoints
Fleet Performance · 12h
1926ms
▼ 48.5% · 12havg1951min635max3737
Loading…
Provider Intelligence
5 providers · 22 models · 20 degraded
ProviderMCongestionUptimeDegRouteRecovery
DOther5
69%
39.33%
50/5
DNVIDIA10
33%
78.00%
80/10
DMeta4
43%
56.66%
40/4
DGoogle2
60%
43.34%
20/2
DDeepSeek1
100%
0.00%
10/1
Overall Rankings
22 ranked · throughput · reliability · congestion
Runner Up
gpt-oss-20b
Other
510 pts
Best Overall
nemotron-3-super-120b-a12b
NVIDIA
572 pts
Third
llama-3.2-11b-vision-instruct
Meta
510 pts
1572A
nemotron-3-super-120b-a12b
NVIDIA
2510A
gpt-oss-20b
Other
3510A
llama-3.2-11b-vision-instruct
Meta
4508A
nemotron-3-nano-omni-30b-a3b-reasoning
NVIDIA
5493A
nemotron-3.5-content-safety
NVIDIA
6471A
nemotron-3-ultra-550b-a55b
NVIDIA
7463A
diffusiongemma-26b-a4b-it
Google
8439A
nemotron-3.5-lightning-30b-a3b
NVIDIA
9436A
riva-translate-4b-instruct-v2
NVIDIA
10423A
muse-glimmer-30b
Meta
11408A
llama-3.1-nemoguard-8b-topic-control
NVIDIA
12289A
laguna-xs-2.1
Other
13261B
ising-calibration-1.5-31b
NVIDIA
14224C
llama-3.2-90b-vision-instruct
Meta
15129D
llama-3.1-nemotron-safety-guard-8b-v3
NVIDIA
1660D
kimi-k3
Other
1730D
glm-5.3
Other
189D
glm-5.3-flash
Other
190D
llama-3.1-nemoguard-8b-content-safety
NVIDIA
200D
llama-guard-4-12b
Meta
210D
gemma-4-31b-it
Google
220D
deepseek-v4.1-flash
DeepSeek
Trend Analysis
22 models tracked
Rising1
riva-translate-4b-instruct-v2100%
Improving7
laguna-xs-2.163%
llama-3.1-nemotron-safety-guard-8b-v317%
kimi-k320%
muse-glimmer-30b90%
llama-3.2-90b-vision-instruct43%
Stable14
glm-5.3-flash3%
glm-5.310%
gpt-oss-20b100%
nemotron-3.5-lightning-30b-a3b83%
nemotron-3.5-content-safety100%
Declining0
All clear
Reliability Matrix
recent · by model
gpt-oss-20b
89
89
89
90
90
90
90
90
92
92
92
93
riva-translate-4b-instruct-v2
89
90
90
91
91
91
90
89
89
90
90
90
nemotron-3.5-content-safety
98
98
98
98
98
98
98
98
98
98
98
98
nemotron-3-ultra-550b-a55b
93
93
93
93
93
96
96
96
96
96
97
97
nemotron-3-super-120b-a12b
97
96
97
96
96
97
97
97
97
97
96
96
nemotron-3-nano-omni-30b-a3b-reasoning
95
94
93
94
94
93
93
95
95
94
94
94
llama-3.1-nemoguard-8b-topic-control
78
79
79
79
79
76
76
77
77
76
76
77
llama-3.2-11b-vision-instruct
92
92
92
92
92
92
93
0
92
93
92
93
muse-glimmer-30b
89
85
0
82
82
82
0
78
77
78
78
79
ising-calibration-1.5-31b
81
81
81
81
81
81
81
82
82
83
87
0
diffusiongemma-26b-a4b-it
87
87
0
83
83
87
84
84
85
0
80
80
nemotron-3.5-lightning-30b-a3b
0
73
73
73
73
73
73
73
73
74
76
77
laguna-xs-2.1
0
54
56
60
58
58
58
61
62
65
0
61
llama-3.2-90b-vision-instruct
56
55
56
58
0
0
0
0
0
0
54
57
kimi-k3
42
0
0
41
0
0
43
45
0
0
46
0
llama-3.1-nemotron-safety-guard-8b-v3
0
0
0
0
41
0
42
0
43
0
44
46
glm-5.3
0
0
0
0
0
0
0
0
0
0
0
0
glm-5.3-flash
0
0
0
0
0
0
0
0
0
0
0
0
llama-3.1-nemoguard-8b-content-safety
0
0
0
0
0
0
0
0
0
0
0
0
llama-guard-4-12b
0
0
0
0
0
0
0
0
0
0
0
0
gemma-4-31b-it
0
0
0
0
0
0
0
0
0
0
0
0
deepseek-v4.1-flash
0
0
0
0
0
0
0
0
0
0
0
0
Incident Timeline
20 active
Congestion spike detected — consider failover
glm-5.3-flash · Other
Congestion spike detected — consider failover
glm-5.3 · Other
Congestion spike detected — consider failover
laguna-xs-2.1 · Other
Congestion spike detected — consider failover
nemotron-3.5-lightning-30b-a3b · NVIDIA
Congestion spike detected — consider failover
llama-3.1-nemotron-safety-guard-8b-v3 · NVIDIA
Congestion spike detected — consider failover
llama-3.1-nemoguard-8b-content-safety · NVIDIA
Congestion spike detected — consider failover
ising-calibration-1.5-31b · NVIDIA
Congestion spike detected — consider failover
kimi-k3 · Other
Congestion spike detected — consider failover
llama-guard-4-12b · Meta
Congestion spike detected — consider failover
llama-3.2-90b-vision-instruct · Meta
Congestion spike detected — consider failover
gemma-4-31b-it · Google
Congestion spike detected — consider failover
diffusiongemma-26b-a4b-it · Google
Congestion spike detected — consider failover
deepseek-v4.1-flash · DeepSeek
Throughput degradation on primary endpoint
gpt-oss-20b · Other
Throughput degradation on primary endpoint
riva-translate-4b-instruct-v2 · NVIDIA
Throughput degradation on primary endpoint
nemotron-3-ultra-550b-a55b · NVIDIA
Throughput degradation on primary endpoint
nemotron-3-super-120b-a12b · NVIDIA
Throughput degradation on primary endpoint
llama-3.1-nemoguard-8b-topic-control · NVIDIA
Throughput degradation on primary endpoint
muse-glimmer-30b · Meta
Throughput degradation on primary endpoint
llama-3.2-11b-vision-instruct · Meta