TLDR: A blueprint-anchored study plan for NCP-AIO. Two findings drive it: NVIDIA names the three hands-on labs outright in the study guide, and Base Command Manager has a free-to-use license — so the labs everyone assumes need a DGX can be built on a workstation with no GPU.


WIP — I have not sat this exam yet. Everything numbered comes from the official NVIDIA-Certified Professional: AI Operations Exam Study Guide, doc 4929800, FEB26. Download it from the certification page; it is free and seven pages. Read it before this.


1. The Exam

Format 30 multiple choice + 3 performance-based labs
Weighting Each section is 50% of the score
Duration 120 minutes total, labs included
Price $500
Delivery Online, remotely proctored (Certiverse)
Result Pass/Fail, with a per-topic score report
Validity 2 years, recertify by retaking

The three labs are published

The performance-based section puts you in a live VM with a scenario panel. The study guide then says this, which I have not seen quoted anywhere:

"To complete the performance-based test you will need to be comfortable using Slurm in BCM to troubleshoot a cluster, using Kubernetes to run a workload, and using BCM to add a node to a cluster."

Three clauses, three labs, published in plain English.

Lab Task GPU needed
A Troubleshoot a Slurm cluster, inside BCM No
B Run a workload on Kubernetes No
C Add a node to a cluster using BCM No

Time budget: ~35–40 min for the MCQ, 80+ for the labs. Running out of lab time is the failure mode, not running out of knowledge.


2. The Blueprint

# Domain Weight Covers
1 Administration 28% Fleet Command, Slurm, DC architecture, Run:ai, MIG
2 Workload Management 20% Kubernetes, DCGM/NVSM/nvidia-smi, BCM provisioning
3 Install and Deploy 32% BCM install, K8s via BCM, NGC, cloud VMI, storage, DOCA
4 Troubleshooting and Optimization 20% Docker, Fabric Manager, BCM, Magnum IO, storage perf

BCM appears in all four domains — items 2.3, 3.1, 3.2, 4.3, plus two of three labs. Budget by frequency of appearance, not by domain weight.

Two document defects worth knowing, because third-party guides have hallucinated scope from them:

  • The contents page lists a fifth domain, "Cognition, Planning, and Memory, 10%". No such section exists in the body, and the four real domains already sum to 100%. It is agentic-AI vocabulary left over from the NCP-AAI guide (NVIDIA's own filename is ...copy-updates-v2, and the body still contains a MModule typo). Ignore it.
  • The trademark footer names NeMo, NIM, Nsight, TensorRT, Triton — none appear in the body. This is not an inference-serving exam.

Weights of 31/23/23/23 circulate on third-party sites. The study guide says 32/28/20/20.

What to skip — and why a list of absences earns its place

A negative list looks like padding until you notice what it costs to be wrong. My own first outline for this article opened with NCCL deadlock debugging and a full Xid table, because that is what an operations background assumes a GPU exam tests. Neither appears in the blueprint. That assumption was worth about two wasted weeks.

So: the things you will be tempted to study, and what the blueprint actually says about them.

Topic In blueprint? Do
NCCL tuning & debugging Not mentioned at all Know it only as a Magnum IO component
Xid taxonomy Not named Common ones via DCGM. Don't memorise the table
Nsight Systems Footer only Skip
Pyxis / Enroot Not mentioned Skip
Triton / TensorRT / NeMo Footer only Skip
Fleet Command Item 1.1 Budget real time — nobody studies it
DOCA on DPU-Arm Item 3.6 Same
Cloud VMI Item 3.4 Narrow, cheap to learn
Magnum IO / GPUDirect Item 4.4 Named in Troubleshooting

The pattern, stated concretely. GPU internals — why a GPU crashed, how collectives move data between GPUs, how a CUDA kernel spends its time — barely appear. NVIDIA's management products — BCM, Fleet Command, Run:ai, NGC, DOCA — are most of the blueprint.

That is the difference between "Xid 79 appeared, what died" and "add this node to the cluster and get Slurm scheduling on it." The first is hardware forensics. The second is administering NVIDIA's tooling, and it is what gets examined.


3. Domain by Domain

Markers: [LAB] reproducible free, no GPU · [HW] needs real hardware · [DOC] read-only.

3.1 Administration — 28%

1.1 — Administer Fleet Command [DOC] A cloud-hosted control plane for edge AI fleets: enrolled edge systems, containerised applications pushed to them over the air. Know locations, deployments, update rollout, and how to open an interactive remote console — the study guide cites that doc page by name, which is usually a tell.

Being [DOC]-only makes it pure MCQ material, and almost nobody preparing for this exam studies it. Cheap points.

1.2 — Administer Slurm clusters [LAB] The most practisable item on the blueprint, and half of Lab A. Know cold: sinfo -R, the squeue REASON codes, scontrol show node|job, scontrol update state=drain|resume, sacct for finished jobs, sacctmgr for associations, and DOWN vs DRAINED.

1.3 — Design data center architecture for AI workloads [DOC] Power and cooling envelopes, leaf-spine and rail-optimised fabric, storage tiering, scalable-unit building blocks. Mostly revision from article 51.

1.4 — Administer Run:ai [LAB] partial NVIDIA-owned since January 2025 and being open-sourced. Know projects and departments, fair-share quota with over-quota borrowing, node pools, GPU fractions (not MIG, not time-slicing), and the workload types — training, interactive, inference — and why each is scheduled differently.

The scheduler half is open source as KAI Scheduler, installable on kind, so this item is partly practisable for free.

1.5 — Configure MIG for AI and HPC [HW] Hardware partitioning on A100/H100-class silicon. Know GPU instance vs compute instance, the profile naming (1g.10gb, 3g.40gb), that enabling MIG mode needs no running processes, and — on the Kubernetes side — that the single vs mixed strategy changes which resource name pods request (nvidia.com/gpu vs nvidia.com/mig-1g.10gb).

MIG is the wrong answer when one large job needs the whole GPU: it partitions memory and SMs statically, so it suits many small inference tenants and hurts a single big training run.

3.2 Workload Management — 20%

2.1 — Administer Kubernetes [LAB] Scheduling, resource requests, node readiness, pod-failure triage. Note the lab says run a workload, not CNI internals.

2.2 — DCGM / NVSM / nvidia-smi [HW] for output, [DOC] for the map The scope split below, plus these by shape:

dcgmi discovery -l          # what GPUs does DCGM see
dcgmi diag -r 1             # -r 1/2/3 = escalating diagnostic depth
dcgmi dmon                  # live telemetry
nvidia-smi -q -d ECC        # ECC counters and remapping state
nvidia-smi topo -m          # GPU/NIC affinity matrix
nvsm show health            # chassis health, not GPU compute health

2.3 — Administer BCM and cluster provisioning [LAB] The commit model and the three config layers, below.

Tool scope — the testable part:

Tool Scope Not for
nvidia-smi One host's GPUs, right now Fleet health, history
DCGM GPU compute health, fleet-wide, over time Chassis hardware
NVSM System health of a DGX chassis — PSU, fans, storage, BMC GPU compute health

Evidence bundles: nvidia-bug-report.sh for driver/GPU/Xid → RMA. nvsm dump health for chassis → platform support case.

GPU Operator chain, in dependency order — which is also the debug order:

NFD → GPU Feature Discovery → driver → container toolkit
    → device plugin → DCGM + exporter → MIG manager → validators

Validators are symptoms, never causes. Read down the chain, not at the loudest pod. Use driver.enabled=false when the host already has a driver (DGX OS) — running both is a classic self-inflicted outage.

BCM's mental model — cmsh is modal and staged. Enter a mode (device, category, softwareimage, network, partition, monitoring, user, jobs), use an object, set fields — and nothing happens until commit. refresh discards. It is kubectl apply semantics with the editor built in.

BCM layer What it is K8s analogy
Software image OS filesystem nodes boot Container image
Category Nodes sharing an image + settings Node pool
Configuration overlay Targeted config on top Kustomize patch

Provisioning flow — and Lab C is this flow: PXE/DHCP → node-installer → image sync → disk layout → finalize → hand off to the workload manager.

cm-wlm-setup is how BCM stands up Slurm or Kubernetes, which also means hand-editing slurm.conf afterwards fights the tool. imageupdate syncs an image to a live node — it has a dry-run, and you never run it against a node executing a job.

3.3 Install and Deploy — 32%

3.1 — Install and configure BCM [LAB] Head node from the ISO, licensing, the split between provisioning and external networks, disk layout. Free license — see §4.

3.2 — Install and initialize Kubernetes on NVIDIA hosts using BCM [LAB] Note the using BCM: this is cm-wlm-setup, not kubeadm and not a distro. Genuinely different from how most of us install clusters, which is why it gets asked. Do it once for real.

3.3 — Deploy containers from NGC [LAB] NVIDIA's container and model registry. Know the nvcr.io path, how API keys work with docker login (the username is the literal string $oauthtoken), the split between the public catalogue and entitled NVIDIA AI Enterprise content, tag conventions, and the ngc CLI. The registry mechanics are practisable with no GPU — you just cannot usefully run the images.

3.4 — Deploy cloud VMI containers [DOC] NVIDIA's Virtual Machine Images: preconfigured cloud VM images carrying the GPU driver and container stack. Narrow and obscure, therefore cheap to learn. Know what a VMI is and why it exists — driver and toolkit version-matching pain.

3.5 — Understand storage requirements for AI data centers [DOC] Throughput vs IOPS vs capacity per tier, bursty checkpoint writes, random-read dataloader patterns, parallel filesystems. Revision from article 51.

3.6 — Deploy DOCA services on DPU-Arm [DOC] DOCA is the SDK and runtime for BlueField DPUs. The phrase DPU-Arm is the point: a BlueField DPU has its own Arm cores running their own Linux, and DOCA services are containers that run on the DPU, not on the host. Know that distinction, and that BlueMan is the DPU's web management service.

Almost nobody preparing for this exam has a BlueField card. Everyone gets asked. Read the DOCA installation guide once.

3.4 Troubleshooting and Optimization — 20%

4.1 — Troubleshoot Docker [LAB] The GPU-specific failure is container cannot see the GPU. Check in this order:

  1. Container toolkit not installed
  2. Runtime not registered in the Docker / containerd config
  3. --gpus flag or RuntimeClass missing
  4. Device cgroup restrictions
  5. Driver / toolkit version mismatch

4.2 — Troubleshoot Fabric Manager [HW] nvidia-fabricmanager is the service people forget. Without it, NVSwitch-class GPUs form no NVLink, and it surfaces as a confusing peer-to-peer error rather than anything mentioning fabric. Its version must match the driver exactly.

4.3 — Troubleshoot BCM [LAB] CMDaemon logs, node stuck in the installer, image sync failures, PXE failures, and split-brain recovery in an HA head-node pair — the study guide cites that doc page by name, which usually means a question exists.

4.4 — Troubleshoot Magnum IO [HW] Magnum IO is the umbrella for NVIDIA's data-movement stack: GPUDirect RDMA (NIC straight to GPU memory), GPUDirect Storage (NVMe straight to GPU memory), and NCCL. Know what each one bypasses, so you recognise the symptom when it silently falls back to a host-memory bounce path — it works, but at a fraction of the expected bandwidth.

4.5 — Troubleshoot storage performance [LAB] partial Sawtooth GPU utilisation means the GPU is repeatedly waiting for data. Utilisation is not SM occupancy. Dataloader worker count and prefetch depth are the cheapest fixes.

Firmware update order — out of order surfaces as Fabric Manager refusing to start or version-mismatch errors:

BMC/BIOS → NIC firmware → GPU VBIOS → driver → CUDA → container toolkit

4. Building the Three Labs

NVIDIA Base Command Manager has a free-to-use license. Register at the NVIDIA Enterprise Product Registration portal. No Enterprise Support, valid up to eight accelerators per node at any cluster size, one-year renewable term. NVIDIA's FAQ lists evaluation, education, demos and production as covered. A corporate email address is required — personal domains are rejected.

Eight accelerators per node is the ceiling. Zero is inside it. BCM provisions nodes that have no GPU at all, because provisioning is a PXE-and-image problem, not a CUDA problem.

Lab Build with GPU
A BCM free license on VMs + Slurm via cm-wlm-setup; or slurm-101 for pure CLI reps No
B kind or k3s + KWOK + fake-gpu-operator No
C BCM head node VM + one blank VM to PXE-boot into it No

The real barrier is RAM: 16 GB host is the floor for a head plus one compute node, 32 GB to be comfortable.

Honesty marker: I have not yet stood up BCM on VMs. The free license and its terms are documented by NVIDIA; the nested-VM path is my plan, not my experience. PXE inside a hypervisor needs the head node owning DHCP on its own network, and I expect that to be the fight.

Lab A drills — the scenario shape is jobs pending / node not accepting work / job fails immediately:

  • sinfo -R — why nodes are down or drained, with the reason string. Highest-value single command on the exam
  • squeue --start, and the REASON codes: Resources, Priority, Dependency, QOSMaxJobsPerUserLimit, ReqNodeNotAvail
  • scontrol show node <n> — State, Reason, CfgTRES vs AllocTRES
  • scontrol update nodename=<n> state=resume — the recovery verb
  • sacct -j <id> --format=JobID,State,ExitCode,DerivedExitCode — post-mortem
  • sacctmgr show assoc where user=<u> — when the answer is "no association"

Lab B: rehearse the full loop — write or fix a manifest, apply, watch it fail to schedule, diagnose, fix, confirm Running. KWOK fabricates nodes the scheduler treats as real; fake-gpu-operator makes them advertise nvidia.com/gpu.

Lab C — the end state, rehearsed until boring:

  1. Head node up and licensed
  2. cmsh → device mode → add physicalnode, set MAC, set category, commit
  3. Boot the blank VM: PXE → node-installer → image sync
  4. Node UP; confirm from both layers — cmsh -c "device; list" and sinfo
  5. Break it deliberately: wrong MAC, wrong category, wrong network. Read the failure. Fix it.

Step 5 is the exam. A live lab hands you a broken state, not a blank one.


5. How to Actually Practise This

Rule 1: do, then read. Never open a document as the day's task. Run the command, get confused, then look up one section. A study guide used as a lookup table sticks; the same guide used as homework does not.

Rule 2: every session produces text. One append-only log, three lines minimum:

## S07 — 2026-10-06 — scontrol update state=drain
$ scontrol update nodename=cpu-0 state=drain reason="test"
$ sinfo -R
> Surprise: pending jobs stayed PENDING with no new reason string.

That file is exam practice and proof of progress at the same time.

The labs themselves live in a repo, not in this article. Twenty-seven exercises in order, one file each, covering Slurm reflexes → BCM install and node provisioning → Kubernetes with a simulated GPU layer → a timed full rehearsal: slurm-101/lab/. They are task sheets — what to run and what to look for — not transcripts, because I have not run all of them yet and inventing output would defeat the purpose.

The gate: do not pay the $500 until you have rehearsed all three labs end to end, under 80 minutes. Three labs at 50% of the score is not a section you can hope your way through.


6. Quick Recall

Tool → job

Tool Use for Not for
nvidia-smi One host, right now Fleet health, history
dcgmi GPU compute health, diagnostics, telemetry Chassis hardware
nvsm DGX chassis: PSU, fans, storage, BMC GPU compute health
cmsh Everything BCM: nodes, images, categories, networks Scheduling decisions
sinfo/squeue/scontrol Slurm state and triage Anything K8s
kubectl K8s workloads and nodes Bare-metal provisioning
ngc NVIDIA registry: containers, models Running them

cmsh → kubectl

Intent cmsh kubectl
List device; list kubectl get nodes
Describe use node001; show kubectl describe node node001
Edit set <field> <value> kubectl edit
Apply commit kubectl apply
Discard refresh quit the editor
Non-interactive cmsh -c "device; list" kubectl get -o json

Slurm — symptom → first command

Symptom First command
Nodes missing from the pool sinfo -R
Job stuck PENDING squeue --start, read REASON
Node refusing work scontrol show node <n>
Node drained, needs to return scontrol update nodename=<n> state=resume
Job died, cause unknown sacct -j <id> --format=JobID,State,ExitCode
User cannot submit at all sacctmgr show assoc where user=<u>

Appendix: The Zero-Hardware Lab Stack

Covers 10 of the blueprint's 19 numbered items and all three labs, for the price of RAM:

Piece Gets you Cost
slurm-101 Real Slurm on kind, two operators, CPU-only Free
BCM free license Head node, provisioning, cmsh, cm-wlm-setup Free, corporate email
kind / k3s Kubernetes Free
KWOK Fabricated nodes the scheduler believes Free
fake-gpu-operator Nodes advertising nvidia.com/gpu Free
KAI Scheduler Open-source Run:ai scheduling Free
ai-factory-ops-lab Six packaged lessons, 0–5 GPU-free Free
One rented GPU VM, once Confirm the real driver stack behaves ~$5–10

The honest ceiling: MIG, NVLink/NVSwitch and Fabric Manager, InfiniBand, real ECC/Xid failures, BlueField DPUs. Those stay [DOC] — and per §2, the blueprint leans lighter on that column than instinct expects.


Sources. Exam structure, weights and all numbered items: NVIDIA-Certified Professional: AI Operations Exam Study Guide, doc 4929800, FEB26, from the certification page. BCM free-license terms: NVIDIA's free license FAQ and BCM licensing docs. Verify before booking — NVIDIA revises these and the weights have moved once already.


Published

Category

Knowledge Base

Tags

Contact