git clone --depth 1 https://github.com/NVIDIA/skills /tmp/doca-bare-metal-deployment && cp -r /tmp/doca-bare-metal-deployment/skills/doca-bare-metal-deployment ~/.claude/skills/doca-bare-metal-deploymentSKILL.md
# DOCA bare-metal deployment **Where to start:** This skill is the bundle's home for *operating* a DOCA-linked application binary **directly on hardware** — no container, no kubelet, no static-pod manifest. It is the parallel of [`doca-container-deployment`](../doca-container-deployment/SKILL.md) for the non-container path. If the user has a DOCA-linked binary they built (per the canonical workflow in [`doca-programming-guide`](../doca-programming-guide/SKILL.md)) and they want to know *how to actually run it on the host or on the BlueField Arm cores correctly*, open [`TASKS.md`](TASKS.md) and start at [`## configure`](TASKS.md#configure). If the question is *what shape does the bare-metal runtime even have and what is the deployment contract*, start at [`CAPABILITIES.md`](CAPABILITIES.md). If the user is not yet sure whether their target system shape is the container path or the bare-metal path, route the recognition step to [`doca-setup`](../doca-setup/SKILL.md) first; only return here once *bare-metal* is the confirmed shape. ## Audience This skill serves **external DOCA developers and operators who have a DOCA-linked application binary they built and want to run it directly on hardware** — i.e., people who already have: - a DOCA-linked application binary they built per [`doca-programming-guide ## build`](../doca-programming-guide/TASKS.md#build), - a real BlueField NIC and a host that talks to it (the **host x86** path — DOCA host install on the host talks to the BlueField NIC over PCIe), OR a BlueField with a console or SSH to the Arm side (the **BlueField Arm bare-metal** path — DOCA installed on the DPU Arm cores; the binary runs there directly), and - a desire to RUN that binary directly on the hardware, not inside a kubelet-standalone-managed container. It is **not** for: - kernel-driver developers contributing to `mlx5_*` or the BlueField OS, - DOCA library contributors (those changes go to the internal DOCA tree, not to a bare-metal deployment), - full-Kubernetes-cluster operators managing a fleet of BlueFields (the bundle covers [`doca-container-deployment`](../doca-container-deployment/SKILL.md) for the single-host kubelet-standalone shape; **fleet/production-scale deployment is fleet-orchestration scope** — route to the orchestration entry-point in [`doca-public-knowledge-map ## Deploying DOCA services at scale`](../doca-public-knowledge-map/references/map.md#deploying-doca-services-at-scale--orchestration-entry-point-personascale-routing) (DPF / Network Operator / Launch Kit), not hand-rolled static-pod loops), - fresh-laptop-no-hardware users with no DOCA install yet — those belong on [`doca-setup ## no-install`](../doca-setup/TASKS.md#no-install). The skill teaches the agent the bare-metal-deployment *procedure* and the rules for quoting documented commands from the public DOCA Programming Guide and the public BlueField / DPU User Manual via [`doca-public-knowledge-map`](../doca-public-knowledge-map/SKILL.md); it does not invent flag names, PCI BDFs, NUMA numbers, devlink paths, representor strings, or systemd `Restart=` mode names from memory. ## When to load this skill Load this skill when the user is doing **hands-on bare-metal deployment of a DOCA-linked application binary** on either of the two supported host modes (host x86 or BlueField Arm), or asking a cross-cutting bare-metal question that is not specific to one library's API. Concretely: - Launching a DOCA-linked binary for the first time on a host with a BlueField NIC in a PCIe slot, with DOCA installed on the host. - Launching a DOCA-linked binary on the BlueField Arm cores directly (BlueField Arm bare-metal mode), with DOCA installed on the Arm side per the BlueField OS image. - Deciding which launch mode to use (direct foreground for interactive debug; tmux/screen for long-running with manual reattach; systemd-supervised for restart-after-reboot, journald-integrated logs, and Restart= policy). - Binding the DOCA process to the right PCIe function, the right representor, the right NUMA node, and the right CPU set — and pinning IRQs to match — without inventing the addresses or the flag names. - Setting up per-tenant isolation (cgroup-v2 cpu / memory / io controllers, network namespaces for multi-tenant deployments, `numactl` / `taskset` for CPU + NUMA binding) so multiple DOCA processes co-tenant on the same BlueField without crushing each other. - Diagnosing a bare-metal launch that is misbehaving — won't start, starts and exits immediately, runs but can't find the device, attaches to the device but the workload errors, OOMs or is signal-killed, is in a restart loop under a supervisor, or is being interfered with by a co-tenant. - Cross-cutting questions: *"should I run this in tmux or as a systemd unit"*, *"what is the smoke-before-bulk loop for a binary on bare metal"*, *"my binary works in a container on the BlueField but not when I run it directly on the Arm — what changed"*. Do **not** load this skill for the container-path equivalent (those questions go to [`doca-container-deployment`](../doca-container-deployment/SKILL.md)); for full-Kubernetes-cluster operations (out of scope per the bundle's non-goals); for library-API questions (route to the matching `libs/<library>` skill); for env-preparation questions including hugepages, IOMMU, pkg-config, and devlink mode flips (use [`doca-setup`](../doca-setup/SKILL.md)); for any hardware-state-changing operation including `mlxconfig` writes and BFB reflashes (route to [`doca-hardware-safety`](../doca-hardware-safety/SKILL.md) for the cross-cutting meta-policy); or for cross-library programming questions (use [`doca-programming-guide`](../doca-programming-guide/SKILL.md)). ## What this skill provides This is a **thin loader**. Substantive material lives in two companion files: - `CAPABILITIES.md` — the bare-metal deployment runtime contract for a DOCA-linked bina
>-
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
|
|
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrate a new dataset from pre-recorded video files via the AutoMagicCalib REST API. Use when user has local MP4s and says 'calibrate my videos', 'run AMC on these videos', or similar. For RTSP/live streams, use amc-run-rtsp-calibration instead.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.