Exposed Nvidia GPU monitors can reveal AI infrastructure secrets

Tags:

A component of Nvidia’s GPU monitoring software that enterprises use to keep tabs on their AI training and inference infrastructure has been vulnerable to denial-of-service (DoS) and information disclosure attacks.

The Nvidia DCGM Exporter contains an unauthenticated resource exhaustion vulnerability that could allow remote attackers to crash the monitoring service and potentially disrupt AI workloads running on the same host.

Nvidia released a fix for the flaw in May when researchers at Lava Security found thousands of DCGM Exporter instances reachable from the internet, with about a quarter of them exposing the profiling endpoints associated with the vulnerability.

“During our research, we found more than 2,000 GPU servers exposing Nvidia DCGM Exporter directly to the public internet without authentication,” Lava researcher Michael Katchinskiy said in a blog post. “Across four scans, those servers reported more than 12,000 unique GPUs, representing an estimated $100 million in hardware.”

DCGM Exporter is a telemetry agent Nvidia uses to collect metrics from Nvidia GPUs and make them available to monitoring systems such as Prometheus. Lava had reported the flaw to Nvidia, which had assigned it CVE-2026-47483 with a high severity rating of CVSS 8.2.

Exposed profiling endpoints led to monitoring crashes

The problem lies in the way an unpatched DCGM Exporter handles concurrent, unauthenticated requests to its “/debug/pprof/” profiling endpoints. A large number of concurrent profiling requests can drive memory consumption high enough to exhaust resources and crash the exporter.

The immediate consequence is a loss of visibility into GPU health and activity, Lava said. But the potential impact is not necessarily limited to monitoring because the exporter shares a host with GPU workloads, Katchinskiy noted in the blog post. The CPU and memory pressure could also affect AI training or inference processes running alongside it.

Nvidia’s security bulletin classifies CVE-2026-47483 as an uncontrolled resource-consumption vulnerability, identifying DCGM Exporter version 4.8.2 as the updated version addressing the issue.

Telemetry could be used for reconnaissance

Nvidia DCGM Exporter collects GPU telemetry, including utilization, memory usage, power consumption, and error events, and exposes the metrics over HTTP for monitoring platforms such as Prometheus to scrape.

When publicly accessible without authentication, these endpoints could reveal a lot about an organization’s computing infrastructure, Lava says. The systems that were found exposed included Nvidia Blackwell Ultra B300 GPUs, H200s and H100s used for large-scale AI workloads, as well as consumer RTX 5090 and 4090 systems.

“Anyone who could reach these endpoints could see what hardware organizations were running, how heavily it was being used, and details about the AI infrastructure around it,” Katchinskiy warned. While such information does not provide access to model weights, training data, or other protected assets, it can help an attacker profile a target and identify weaker components or software versions.

Lava has recommended restricting public access to DCGM Exporter and Prometheus unless external access is necessary, binding exporters to loopback or private interfaces, and enforcing access controls through firewalls, security groups, or other network measures.

These steps should follow upgrading to version 4.8.2 or later and verifying that the “–enable–pprof” option is disabled. Katchinskiy noted that this profiling feature may be enabled by default in unpatched software.

Categories

No Responses

Leave a Reply

Your email address will not be published. Required fields are marked *