Fleet Operations

Discover

Live operational intelligence across every free NVIDIA NIM endpoint — ranked by throughput, reliability, and congestion, refreshed continuously.

22
Endpoints
5
Providers
20
Degraded
—
Last probe
Fleet Degraded2 of 22 endpoints operational · 20 degraded
Throughput▲34%
23.6tok/s
P95 Latency▼14%
5,764ms
Reliability▲0.1pt
58.6%
Congestion▼0.1pt
48%
22 endpoints
1926ms
▼ 48.5% · 12h
avg1951min635max3737

Loading…

5 providers · 22 models · 20 degraded
ProviderMCongestionUptimeDegRouteRecovery
DOther5
69%
39.33%
5
0/5
DNVIDIA10
33%
78.00%
8
0/10
DMeta4
43%
56.66%
4
0/4
DGoogle2
60%
43.34%
2
0/2
DDeepSeek1
100%
0.00%
1
0/1
22 ranked · throughput · reliability · congestion
Runner Up
gpt-oss-20b
Other
510 pts
Best Overall
nemotron-3-super-120b-a12b
NVIDIA
572 pts
Third
llama-3.2-11b-vision-instruct
Meta
510 pts
1
nemotron-3-super-120b-a12b
NVIDIA
572A
2
gpt-oss-20b
Other
510A
3
llama-3.2-11b-vision-instruct
Meta
510A
4
nemotron-3-nano-omni-30b-a3b-reasoning
NVIDIA
508A
5
nemotron-3.5-content-safety
NVIDIA
493A
6
nemotron-3-ultra-550b-a55b
NVIDIA
471A
7
diffusiongemma-26b-a4b-it
Google
463A
8
nemotron-3.5-lightning-30b-a3b
NVIDIA
439A
9
riva-translate-4b-instruct-v2
NVIDIA
436A
10
muse-glimmer-30b
Meta
423A
11
llama-3.1-nemoguard-8b-topic-control
NVIDIA
408A
12
laguna-xs-2.1
Other
289A
13
ising-calibration-1.5-31b
NVIDIA
261B
14
llama-3.2-90b-vision-instruct
Meta
224C
15
llama-3.1-nemotron-safety-guard-8b-v3
NVIDIA
129D
16
kimi-k3
Other
60D
17
glm-5.3
Other
30D
18
glm-5.3-flash
Other
9D
19
llama-3.1-nemoguard-8b-content-safety
NVIDIA
0D
20
llama-guard-4-12b
Meta
0D
21
gemma-4-31b-it
Google
0D
22
deepseek-v4.1-flash
DeepSeek
0D
22 models tracked
Rising1
riva-translate-4b-instruct-v2100%
Improving7
laguna-xs-2.163%
llama-3.1-nemotron-safety-guard-8b-v317%
kimi-k320%
muse-glimmer-30b90%
llama-3.2-90b-vision-instruct43%
Stable14
glm-5.3-flash3%
glm-5.310%
gpt-oss-20b100%
nemotron-3.5-lightning-30b-a3b83%
nemotron-3.5-content-safety100%
Declining0
All clear
recent · by model
gpt-oss-20b
89
89
89
90
90
90
90
90
92
92
92
93
riva-translate-4b-instruct-v2
89
90
90
91
91
91
90
89
89
90
90
90
nemotron-3.5-content-safety
98
98
98
98
98
98
98
98
98
98
98
98
nemotron-3-ultra-550b-a55b
93
93
93
93
93
96
96
96
96
96
97
97
nemotron-3-super-120b-a12b
97
96
97
96
96
97
97
97
97
97
96
96
nemotron-3-nano-omni-30b-a3b-reasoning
95
94
93
94
94
93
93
95
95
94
94
94
llama-3.1-nemoguard-8b-topic-control
78
79
79
79
79
76
76
77
77
76
76
77
llama-3.2-11b-vision-instruct
92
92
92
92
92
92
93
0
92
93
92
93
muse-glimmer-30b
89
85
0
82
82
82
0
78
77
78
78
79
ising-calibration-1.5-31b
81
81
81
81
81
81
81
82
82
83
87
0
diffusiongemma-26b-a4b-it
87
87
0
83
83
87
84
84
85
0
80
80
nemotron-3.5-lightning-30b-a3b
0
73
73
73
73
73
73
73
73
74
76
77
laguna-xs-2.1
0
54
56
60
58
58
58
61
62
65
0
61
llama-3.2-90b-vision-instruct
56
55
56
58
0
0
0
0
0
0
54
57
kimi-k3
42
0
0
41
0
0
43
45
0
0
46
0
llama-3.1-nemotron-safety-guard-8b-v3
0
0
0
0
41
0
42
0
43
0
44
46
glm-5.3
0
0
0
0
0
0
0
0
0
0
0
0
glm-5.3-flash
0
0
0
0
0
0
0
0
0
0
0
0
llama-3.1-nemoguard-8b-content-safety
0
0
0
0
0
0
0
0
0
0
0
0
llama-guard-4-12b
0
0
0
0
0
0
0
0
0
0
0
0
gemma-4-31b-it
0
0
0
0
0
0
0
0
0
0
0
0
deepseek-v4.1-flash
0
0
0
0
0
0
0
0
0
0
0
0
20 active
Congestion spike detected — consider failover
glm-5.3-flash · Other
2m ago
Congestion spike detected — consider failover
glm-5.3 · Other
2m ago
Congestion spike detected — consider failover
laguna-xs-2.1 · Other
3m ago
Congestion spike detected — consider failover
nemotron-3.5-lightning-30b-a3b · NVIDIA
3m ago
Congestion spike detected — consider failover
llama-3.1-nemotron-safety-guard-8b-v3 · NVIDIA
2m ago
Congestion spike detected — consider failover
llama-3.1-nemoguard-8b-content-safety · NVIDIA
2m ago
Congestion spike detected — consider failover
ising-calibration-1.5-31b · NVIDIA
2m ago
Congestion spike detected — consider failover
kimi-k3 · Other
2m ago
Congestion spike detected — consider failover
llama-guard-4-12b · Meta
2m ago
Congestion spike detected — consider failover
llama-3.2-90b-vision-instruct · Meta
2m ago
Congestion spike detected — consider failover
gemma-4-31b-it · Google
2m ago
Congestion spike detected — consider failover
diffusiongemma-26b-a4b-it · Google
2m ago
Congestion spike detected — consider failover
deepseek-v4.1-flash · DeepSeek
2m ago
Throughput degradation on primary endpoint
gpt-oss-20b · Other
3m ago
Throughput degradation on primary endpoint
riva-translate-4b-instruct-v2 · NVIDIA
2m ago
Throughput degradation on primary endpoint
nemotron-3-ultra-550b-a55b · NVIDIA
2m ago
Throughput degradation on primary endpoint
nemotron-3-super-120b-a12b · NVIDIA
2m ago
Throughput degradation on primary endpoint
llama-3.1-nemoguard-8b-topic-control · NVIDIA
2m ago
Throughput degradation on primary endpoint
muse-glimmer-30b · Meta
2m ago
Throughput degradation on primary endpoint
llama-3.2-11b-vision-instruct · Meta
2m ago