git clone --depth 1 https://github.com/NVIDIA/skills /tmp/doca-bench-extension && cp -r /tmp/doca-bench-extension/skills/doca-bench-extension ~/.claude/skills/doca-bench-extensionSKILL.md
# DOCA Bench Extension **Where to start:** This is a tool skill for the **extension / plug-in framework** that augments [`doca-bench`](../doca-bench/SKILL.md) — NOT a workload-shape skill on its own. Open [`TASKS.md`](TASKS.md) and start at [`## configure`](TASKS.md#configure) to commit to the three-axis decision (workload class is genuinely outside doca-bench's built-in modes × extension API surface fits × parent-tool co-load is acceptable), then [`## build`](TASKS.md#build) for how a custom extension is compiled and laid out, then [`## run`](TASKS.md#run) for how `doca-bench` discovers and invokes the extension, then [`## test`](TASKS.md#test) for the smoke-before-bulk loop the agent applies to every new extension. Open [`CAPABILITIES.md`](CAPABILITIES.md) when the question is *what an extension can do that built-in `doca-bench` modes cannot*, *what the extension API surface looks like in broad strokes (the `DOCA_EXPERIMENTAL` C entry points the shipped reference exposes)*, *how the build / registration / discovery flow works*, or *how the extension's lifetime is bounded by the parent `doca-bench` invocation*. If `doca-bench` itself is the question, route to [`doca-bench`](../doca-bench/SKILL.md). If the question is "which built-in `doca-bench` mode do I pick?", that is also [`doca-bench`](../doca-bench/SKILL.md) — extensions are the *exit ramp* for workloads built-in modes do not cover. ## Example questions this skill answers well - *"My workload class is `<X>` — does `doca-bench` measure it natively, or do I need an extension?"* — the extension-vs-built-in decision question. The agent walks the user back to [`doca-bench`](../doca-bench/SKILL.md)'s built-in mode inventory FIRST and only routes to the extension framework when no built-in mode applies. - *"I want to benchmark a CUDA / GPU-side workload that drives DOCA GPUNetIO RX and TX queues. Where do I start? Is there a reference extension I can copy?"* — the agent surfaces the shipped `doca_bench_cuda` extension under `/opt/mellanox/doca/tools/bench_extension/doca_bench_cuda/` as the reference exemplar and walks the operator through its API surface and build shape. - *"How does `doca-bench` actually discover and load my custom extension at runtime? Is it a versioned shared library? What does my entry-point need to look like?"* — the build / registration / discovery flow question. The agent walks the Meson-built shared library shape, the versioning, and the parent-tool's runtime discovery path (which the agent does NOT invent from memory — the shipped extension's `meson.build` and the public DOCA Bench documentation on `docs.nvidia.com` are the source of truth). - *"The API headers I have are marked `DOCA_EXPERIMENTAL`. What does that mean for my extension's stability across DOCA releases? Am I going to have to rebuild it every release?"* — the experimental-surface and version compatibility question. - *"Once I build my extension, what is the cheapest possible smoke I can run before pointing my real workload at it? How do I know `doca-bench` actually loaded it, called into it, and that the call returned the data the parent tool expected?"* — the smoke-before-bulk question. - *"My custom extension builds, but `doca-bench` says it cannot find / load / call it. Where do I look first?"* — the layered-debug question that distinguishes build-failures, load-failures, registration-mismatches, and runtime-call-failures. ## Audience Experienced AI agents and platform / performance engineers who already use [`doca-bench`](../doca-bench/SKILL.md) for the built-in workload modes and now have a workload class that the built-in modes do not cover. Readers are expected to be comfortable with native build systems (Meson, in this codebase), shared-library packaging on Linux, and the `DOCA_EXPERIMENTAL` API stability contract. If the user asks about GPU-side benchmarking via the shipped `doca_bench_cuda` reference extension, the reader is also expected to be familiar with DOCA GPUNetIO and CUDA toolchain basics — those domains live in their own skills, not here. This skill is NOT for: - operators who can express their workload with one of `doca-bench`'s built-in modes — that is [`doca-bench`](../doca-bench/SKILL.md); - operators who want to benchmark a different DOCA primitive (Flow, Comch, RMAX) via that primitive's own measurement tool — route to that tool; - contributors authoring or modifying the in-tree extensions themselves (this skill is for external operators consuming the framework, not for internal DOCA contributors). ## Language scope A `doca-bench` extension surfaces as: 1. A **versioned shared library** on Linux (`.so` with `soversion` matching the DOCA release), built via the `doca-bench-extension` Meson rules in the shipped `/opt/mellanox/doca/tools/bench_extension/meson.build` and the per-extension subdirectory (the reference exemplar is `doca_bench_cuda/`). 2. A small set of **`DOCA_EXPERIMENTAL`-marked C entry points** that the parent `doca-bench` invokes — i.e. the API surface declared in the extension's header file. The shipped `doca_bench_cuda/doca_bench_cuda.h` is the reference for what that surface shape looks like in practice (`*_init`, `*_device_query`, `*_device_synchronize`, and per-workload kernel-start entry points such as `*_start_nop_kernel`, `*_start_eth_recv_kernel`, `*_start_eth_send_kernel`, `*_start_eth_bidir_kernel`). 3. A set of **per-workload settings structs** that the parent passes through (e.g. the reference exemplar's `doca_bench_cuda_kernel_settings`, `doca_bench_cuda_eth_rx_kernel_settings`, `doca_bench_cuda_eth_tx_kernel_settings`, `doca_bench_cuda_eth_bidir_kernel_settings` carry block counts, threads-per-block, RX / TX queues, buffer address / mkey / size, a stop flag, and a stats pointer). The skill itself is Markdown. The user's extension source is wh
>-
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
|
|
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras; VIOS records clips, AMC ingests them, then runs calibration.
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Calibrate a new dataset from pre-recorded video files via the AutoMagicCalib REST API. Use when user has local MP4s and says 'calibrate my videos', 'run AMC on these videos', or similar. For RTSP/live streams, use amc-run-rtsp-calibration instead.
Launch AutoMagicCalib microservice and web UI from NGC release images via Docker Compose. Use when user says 'deploy auto calibration', 'launch auto calibration', 'launch AMC', 'start MS+UI', or 'set up auto-magic-calib'. Requires NGC API key.