Skip to main content
ClaudeWave
Skill3.2k repo starsupdated 3d ago

doca-bench-extension

>

Install in Claude Code
Copy
git clone --depth 1 https://github.com/NVIDIA/skills /tmp/doca-bench-extension && cp -r /tmp/doca-bench-extension/skills/doca-bench-extension ~/.claude/skills/doca-bench-extension
Then start a new Claude Code session; the skill loads automatically.

SKILL.md

# DOCA Bench Extension

**Where to start:** This is a tool skill for the **extension /
plug-in framework** that augments
[`doca-bench`](../doca-bench/SKILL.md) — NOT a workload-shape
skill on its own. Open [`TASKS.md`](TASKS.md) and start at
[`## configure`](TASKS.md#configure) to commit to the three-axis
decision (workload class is genuinely outside doca-bench's
built-in modes × extension API surface fits × parent-tool
co-load is acceptable), then [`## build`](TASKS.md#build) for
how a custom extension is compiled and laid out, then
[`## run`](TASKS.md#run) for how `doca-bench` discovers and
invokes the extension, then [`## test`](TASKS.md#test) for the
smoke-before-bulk loop the agent applies to every new
extension. Open [`CAPABILITIES.md`](CAPABILITIES.md) when the
question is *what an extension can do that built-in
`doca-bench` modes cannot*, *what the extension API surface
looks like in broad strokes (the `DOCA_EXPERIMENTAL` C entry
points the shipped reference exposes)*, *how the
build / registration / discovery flow works*, or *how the
extension's lifetime is bounded by the parent `doca-bench`
invocation*. If `doca-bench` itself is the question, route to
[`doca-bench`](../doca-bench/SKILL.md). If the question is
"which built-in `doca-bench` mode do I pick?", that is also
[`doca-bench`](../doca-bench/SKILL.md) — extensions are the
*exit ramp* for workloads built-in modes do not cover.

## Example questions this skill answers well

- *"My workload class is `<X>` — does `doca-bench` measure it
  natively, or do I need an extension?"* — the
  extension-vs-built-in decision question. The agent walks
  the user back to [`doca-bench`](../doca-bench/SKILL.md)'s
  built-in mode inventory FIRST and only routes to the
  extension framework when no built-in mode applies.
- *"I want to benchmark a CUDA / GPU-side workload that
  drives DOCA GPUNetIO RX and TX queues. Where do I start?
  Is there a reference extension I can copy?"* — the agent
  surfaces the shipped `doca_bench_cuda` extension under
  `/opt/mellanox/doca/tools/bench_extension/doca_bench_cuda/` as the
  reference exemplar and walks the operator through its
  API surface and build shape.
- *"How does `doca-bench` actually discover and load my
  custom extension at runtime? Is it a versioned shared
  library? What does my entry-point need to look like?"* —
  the build / registration / discovery flow question. The
  agent walks the Meson-built shared library shape, the
  versioning, and the parent-tool's runtime discovery path
  (which the agent does NOT invent from memory — the
  shipped extension's `meson.build` and the public DOCA Bench
  documentation on `docs.nvidia.com` are the source of
  truth).
- *"The API headers I have are marked `DOCA_EXPERIMENTAL`.
  What does that mean for my extension's stability across
  DOCA releases? Am I going to have to rebuild it every
  release?"* — the experimental-surface and version
  compatibility question.
- *"Once I build my extension, what is the cheapest possible
  smoke I can run before pointing my real workload at it?
  How do I know `doca-bench` actually loaded it, called
  into it, and that the call returned the data the parent
  tool expected?"* — the smoke-before-bulk question.
- *"My custom extension builds, but `doca-bench` says it
  cannot find / load / call it. Where do I look first?"* —
  the layered-debug question that distinguishes
  build-failures, load-failures, registration-mismatches,
  and runtime-call-failures.

## Audience

Experienced AI agents and platform / performance engineers
who already use [`doca-bench`](../doca-bench/SKILL.md) for
the built-in workload modes and now have a workload class
that the built-in modes do not cover. Readers are expected
to be comfortable with native build systems (Meson, in this
codebase), shared-library packaging on Linux, and the
`DOCA_EXPERIMENTAL` API stability contract. If the user
asks about GPU-side benchmarking via the shipped
`doca_bench_cuda` reference extension, the reader is also
expected to be familiar with DOCA GPUNetIO and CUDA toolchain
basics — those domains live in their own skills, not here.

This skill is NOT for:

- operators who can express their workload with one of
  `doca-bench`'s built-in modes — that is
  [`doca-bench`](../doca-bench/SKILL.md);
- operators who want to benchmark a different DOCA primitive
  (Flow, Comch, RMAX) via that primitive's own
  measurement tool — route to that tool;
- contributors authoring or modifying the in-tree extensions
  themselves (this skill is for external operators consuming
  the framework, not for internal DOCA contributors).

## Language scope

A `doca-bench` extension surfaces as:

1. A **versioned shared library** on Linux (`.so` with
   `soversion` matching the DOCA release), built via the
   `doca-bench-extension` Meson rules in the shipped
   `/opt/mellanox/doca/tools/bench_extension/meson.build` and the
   per-extension subdirectory (the reference exemplar is
   `doca_bench_cuda/`).
2. A small set of **`DOCA_EXPERIMENTAL`-marked C entry
   points** that the parent `doca-bench` invokes — i.e. the
   API surface declared in the extension's header file.
   The shipped `doca_bench_cuda/doca_bench_cuda.h` is the
   reference for what that surface shape looks like in
   practice (`*_init`, `*_device_query`,
   `*_device_synchronize`, and per-workload kernel-start
   entry points such as `*_start_nop_kernel`,
   `*_start_eth_recv_kernel`, `*_start_eth_send_kernel`,
   `*_start_eth_bidir_kernel`).
3. A set of **per-workload settings structs** that the
   parent passes through (e.g. the reference exemplar's
   `doca_bench_cuda_kernel_settings`,
   `doca_bench_cuda_eth_rx_kernel_settings`,
   `doca_bench_cuda_eth_tx_kernel_settings`,
   `doca_bench_cuda_eth_bidir_kernel_settings` carry block
   counts, threads-per-block, RX / TX queues, buffer
   address / mkey / size, a stop flag, and a stats
   pointer).

The skill itself is Markdown. The user's extension source
is wh