Post

NVIDIA AI Factory(1) - Software Components

NVIDIA AI Factory(1) - Software Components

πŸ” How Do the Main NVIDIA AI Factory Software Components Relate to Each Other?

NVIDIA AI Factory management software is split across multiple control and observability domains. NICo, Config Manager, NetQ, UFM, and NVSentinel are not equivalent components: they are deployed in different places, manage different infrastructure layers, and interact with different hardware or workload domains.


β‘  High-Level Architecture

NVIDIA AI Factory Management Software Architecture Figure 1. NVIDIA AI Factory Software Components.

The main relationship can be summarized as follows:

1
2
3
4
5
6
7
8
9
10
11
12
13
AI Factory Management Services
β”œβ”€β”€ NICo
β”‚   └── Managed Hosts / DPU / SuperNIC / provisioning / network lifecycle
β”œβ”€β”€ NVIDIA Config Manager
β”‚   └── Spectrum / Spectrum-X switch configuration lifecycle
β”œβ”€β”€ NetQ
β”‚   └── Ethernet fabric telemetry / validation / operational state
└── UFM
    └── InfiniBand fabric management / monitoring

Customer / Tenant Kubernetes Cluster
└── NVSentinel
    └── GPU / node health monitoring and Kubernetes remediation

The important distinction is that these components do not form one flat management stack. Each component owns a different management domain.


β‘‘ NICo

NICo (NVIDIA Infra Controller) is the infrastructure control plane for managed-host lifecycle, DPU/SuperNIC-related networking, provisioning, IPAM, and tenant infrastructure resources.

1
2
3
4
5
6
7
8
9
NICo
 β”œβ”€β”€ Core / API
 β”œβ”€β”€ provisioning workflows
 β”œβ”€β”€ DHCP / IPAM
 β”œβ”€β”€ host and DPU lifecycle
 └── network / partition resources
        β”‚
        β–Ό
Managed Hosts / DPU / SuperNIC

NICo should therefore be viewed primarily as an infrastructure orchestration and lifecycle controller, not as a general-purpose switch configuration manager.


β‘’ NVIDIA Config Manager

NVIDIA Config Manager manages the configuration lifecycle of NVIDIA network devices. Its main flow is based on source-of-truth data, rendering, configuration storage, and deployment workflows.

1
2
3
4
5
6
7
8
9
10
11
12
13
Nautobot / DCIM
      β”‚
      β–Ό
Render Service
      β”‚
      β–Ό
Config Store
      β”‚
      β–Ό
Deployment Workflow
      β”‚
      β–Ό
Spectrum / Spectrum-X Switch

Its role is closest to desired-state network configuration management.


β‘£ NetQ

NetQ provides Ethernet fabric observability, telemetry, validation, and operational-state visibility.

1
2
3
4
5
Spectrum / Spectrum-X Fabric
          β”‚
          β”‚ telemetry / state
          β–Ό
         NetQ

Config Manager and NetQ therefore represent different views of the same Ethernet infrastructure:

  • Config Manager: intended configuration
  • NetQ: observed operational state

β‘€ UFM

UFM (Unified Fabric Manager) manages and monitors the InfiniBand fabric.

1
2
3
4
5
6
7
8
9
10
UFM
 β”‚
 β”œβ”€β”€ topology
 β”œβ”€β”€ routing
 β”œβ”€β”€ partitioning
 β”œβ”€β”€ telemetry
 └── fabric operations
        β”‚
        β–Ό
InfiniBand Switches / HCAs

UFM is the primary management domain for InfiniBand scale-out networking.


β‘₯ NVSentinel

NVSentinel is associated with the customer or tenant Kubernetes cluster rather than the central AI Factory management services.

1
2
3
4
5
6
7
8
9
10
11
Customer / Tenant Kubernetes Cluster
β”‚
β”œβ”€β”€ NVSentinel
β”‚   β”œβ”€β”€ GPU Health Monitor
β”‚   β”œβ”€β”€ Platform Connector
β”‚   β”œβ”€β”€ Fault Quarantine
β”‚   β”œβ”€β”€ Node Drainer
β”‚   └── Fault Remediation
β”‚
β”œβ”€β”€ Kubernetes API
└── GPU Nodes / Workloads

Its primary role is GPU and node health monitoring, fault handling, quarantine, drain, and remediation through Kubernetes-native workflows.


⑦ Responsibility Boundaries

1
2
3
4
5
6
7
8
9
10
11
12
13
14
NICo
β†’ infrastructure lifecycle / managed hosts / DPU / SuperNIC

Config Manager
β†’ Spectrum / Spectrum-X switch configuration

NetQ
β†’ Ethernet fabric observability

UFM
β†’ InfiniBand fabric management

NVSentinel
β†’ Kubernetes GPU-node health and remediation

These boundaries are important because the products are related operationally but are not interchangeable.

For example, NICo and Config Manager should be treated as separate management domains, while NetQ complements Config Manager by observing the resulting network state. NVSentinel belongs closer to the workload cluster because its actions are Kubernetes- and node-health-oriented.


Key Takeaway

The simplest way to understand the NVIDIA AI Factory software stack is to classify each component by deployment location, control target, and responsibility.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
Infrastructure orchestration
β†’ NICo

Ethernet configuration
β†’ NVIDIA Config Manager

Ethernet observability
β†’ NetQ

InfiniBand management
β†’ UFM

GPU / Kubernetes remediation
β†’ NVSentinel

The architecture is therefore a set of domain-specific management planes rather than a single hierarchical controller.


TL;DR

Q. What does NICo manage?

Managed-host lifecycle, DPU/SuperNIC-related infrastructure, provisioning, and tenant network resources.

Q. What does NVIDIA Config Manager manage?

Desired configuration and deployment workflows for NVIDIA Ethernet switches.

Q. What does NetQ manage?

Ethernet telemetry, validation, and operational-state visibility.

Q. What does UFM manage?

InfiniBand topology, routing, partitioning, monitoring, and fabric operations.

Q. Where does NVSentinel belong?

Inside or alongside the customer Kubernetes workload cluster, where it monitors GPU/node health and drives remediation.


References

This post is licensed under CC BY 4.0 by the author.