Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 9 additions & 9 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,11 +48,11 @@ python3 -m venv .venv
.venv/bin/python -m pytest
```

Validate checked-in configurations through the production parser and common semantic checks
without initializing hardware:
The standard pre-PR check runs the portable tests, validates both retained and
generated configurations without initializing hardware, and checks the documentation:

```bash
python3 scripts/check_daqiri_configs.py --validator build/tools/daqiri_config_validate
scripts/check_pr.sh
```

The default pytest suite collects only `tests/portable/`. Build-backed C++ tests live under
Expand All @@ -67,19 +67,19 @@ Integration and performance verification is done via the benchmark executables i

| Executable | Source | Typical config |
|---|---|---|
| `daqiri_bench_raw_gpudirect` | `raw_gpudirect_bench.cpp` | `daqiri_bench_raw_tx_rx.yaml`, `daqiri_bench_raw_tx_rx_4q.yaml`, `daqiri_bench_raw_tx_rx_spark.yaml`, `daqiri_bench_raw_{tx,rx}_spark_xhost.yaml`, `daqiri_bench_raw_sw_loopback.yaml`, `daqiri_bench_raw_hw_loopback_ibverbs.yaml`, `daqiri_bench_raw_rx_multi_q.yaml`, `daqiri_bench_raw_tx_rx_vxlan.yaml`, `daqiri_bench_raw_tx_rx_vlan.yaml`, `daqiri_bench_raw_tx_rx_gre.yaml`, `daqiri_bench_raw_tx_rx_nvgre.yaml`, `daqiri_bench_raw_tx_rx_spark_mq.yaml` (mq base; `run_spark_mq_bench.sh` derives the 4 cells via `scripts/gen_spark_mq_config.py`), `daqiri_bench_raw_tx_rx_pacing.yaml` (per-queue `pacing_mbps`; DPDK engine only) |
| `daqiri_bench_raw_gpudirect` | `raw_gpudirect_bench.cpp` | Canonical examples: `daqiri_bench_raw_tx_rx.yaml`, `daqiri_bench_raw_tx_rx_4q.yaml`, `daqiri_bench_raw_sw_loopback.yaml`, `daqiri_bench_raw_hw_loopback_ibverbs.yaml`, `daqiri_bench_raw_rx_multi_q.yaml`, `daqiri_bench_raw_tx_rx_pacing.yaml` (per-queue `pacing_mbps`; DPDK engine only). Generate Spark, cross-host, multi-queue matrix, and VLAN/VXLAN/GRE/NVGRE variants with `scripts/gen_daqiri_config.py`; the Spark harnesses invoke it directly. |
| `daqiri_bench_raw_latency` | `raw_latency_bench.cpp` | `daqiri_bench_raw_latency_ibverbs.yaml` — caller-driven direct TX/RX, RX hardware timestamps, 64–8192-byte power-of-two latency sweep |
| `daqiri_example_dynamic_rx_flow` | `dynamic_rx_flow_example.cpp` | `daqiri_example_dynamic_rx_flow.yaml` — `flow_isolation: true` startup followed by runtime scalar queue steering, multi-queue RSS, and raw-engine decap/pop flow add/delete |
| `daqiri_example_dynamic_resource` | `dynamic_resource_example.cpp` | Any ibverbs config with at least one RX queue and one single-region TX queue (for example `daqiri_bench_raw_hw_loopback_ibverbs.yaml`) — initializes without RX queues, repeatedly adds/removes the first RX queue's steering flow, and exercises runtime MR and RX/TX queue add/delete |
| `daqiri_example_named_endpoints` | `named_endpoints_example.cpp` | `daqiri_example_named_endpoints_tx_rx.yaml` — raw ibverbs runtime named endpoints with inline Ethernet/IPv4/UDP headers and one gathered payload segment |
| `daqiri_bench_raw_hds` | `raw_hds_bench.cpp` | `daqiri_bench_raw_tx_rx_hds.yaml` |
| `daqiri_bench_raw_reorder_seq` | `raw_reorder_seq_bench.cpp` | `daqiri_bench_raw_tx_rx_reorder_seq_1024*.yaml`, `daqiri_bench_raw_rx_reorder_seq_*.yaml` |
| `daqiri_bench_raw_reorder_quantize` | `raw_reorder_quantize_bench.cpp` | `daqiri_bench_raw_tx_rx_reorder_quantize_seq_batch.yaml` |
| `daqiri_bench_rdma` | `rdma_bench.cpp` | `daqiri_bench_rdma_tx_rx.yaml`, `daqiri_bench_rdma_tx_rx_spark.yaml`, `daqiri_bench_rdma_tx_rx_spark_xhost.yaml`, `daqiri_bench_rdma_tx_rx_spark_netns.yaml` (combined-role netns base; `run_spark_bench.sh` splits per role via `scripts/gen_spark_netns_config.py`) |
| `daqiri_bench_socket` | `socket_bench.cpp` | `daqiri_bench_socket_{udp,tcp}_tx_rx.yaml`, `daqiri_bench_socket_{udp,tcp}_tx_rx_spark_netns.yaml` (combined-role netns bases), `daqiri_bench_socket_{udp,tcp}_{client,server}_spark_xhost.yaml` (cross-host role configs) |
| `daqiri_bench_rdma` | `rdma_bench.cpp` | Canonical example: `daqiri_bench_rdma_tx_rx.yaml`. Generate Spark netns/cross-host client and server roles with `scripts/gen_daqiri_config.py socket-pair --transport roce`. |
| `daqiri_bench_socket` | `socket_bench.cpp` | Canonical examples: `daqiri_bench_socket_{udp,tcp}_tx_rx.yaml`. Generate Spark netns/cross-host client and server roles with `scripts/gen_daqiri_config.py socket-pair`. |
| `daqiri_pool_ring_bench` | `pool_ring_bench.cpp` | none — microbenchmark comparing `daqiri::Ring`/`daqiri::ObjectPool` vs DPDK `rte_ring`/`rte_mempool` (SPSC/MPMC, single/bulk, thread sweep). The `rte_*` comparison arm compiles only in a DPDK-enabled build; takes no YAML/CLI args |

The four `raw_*` benches share `raw_bench_common.{cpp,h}` and accept `--seconds N`. `daqiri_bench_rdma` and `daqiri_bench_socket` also take `--mode {tx,rx,both}`. `daqiri_bench_raw_gpudirect`, `daqiri_bench_raw_hds`, `daqiri_bench_rdma`, and `daqiri_bench_socket` additionally accept `--workload none|fft|gemm|gemm_fp16` — a reusable representative GPU workload (`examples/bench_workload.{h,cu}`, cuFFT/cuBLAS) run once per received reorder window on the **actual received payload**. Each backend first assembles the burst's payloads into one contiguous GPU buffer via `examples/bench_pipeline.{h,cu}` (`ReorderPipeline`): a sequence-number reorder kernel for the out-of-order transports (DPDK raw, UDP) and an arrival-order gather for the in-order ones (RoCE RC, TCP); sockets stage host→device first since their payloads land in pageable host memory. The reorder/gather kernels (`packet_reorder_copy_payload_by_sequence`, `packet_gather_copy_payload`) live in `src/kernels.cu`. `gemm` is FP32 `cublasSgemm`; `gemm_fp16` is the same-size mixed-precision FP16/tensor-core `cublasGemmEx` (inference-style); the contiguous buffer supplies the FFT input / GEMM A operand. `--workload-gemm-dim N` pins the square GEMM side length (default 1024), so the FLOP count per call (2·n³) is FIXED and the compute working set is exactly n·n·elem_size, read from the front of each received I/O unit (the unit must be at least that large). `--workload-fft-len N` pins the 1-D C2C transform length for `fft` (default 1024; independent of the GEMM dimension); the working set is fanned out across as many batched length-N transforms as fit. Used by `run_spark_bench.sh`'s `WORKLOAD` / `GEMM_DIM` / `FFT_LEN` env (all backends) to fill the CSV `post_process` / `post_process_gemm_dim` columns (issue #15).
The four `raw_*` benches share `raw_bench_common.{cpp,h}` and accept `--seconds N`. `daqiri_bench_rdma` and `daqiri_bench_socket` also take `--mode {server,client,both}`. `daqiri_bench_raw_gpudirect`, `daqiri_bench_raw_hds`, `daqiri_bench_rdma`, and `daqiri_bench_socket` additionally accept `--workload none|fft|gemm|gemm_fp16` — a reusable representative GPU workload (`examples/bench_workload.{h,cu}`, cuFFT/cuBLAS) run once per received reorder window on the **actual received payload**. Each backend first assembles the burst's payloads into one contiguous GPU buffer via `examples/bench_pipeline.{h,cu}` (`ReorderPipeline`): a sequence-number reorder kernel for the out-of-order transports (DPDK raw, UDP) and an arrival-order gather for the in-order ones (RoCE RC, TCP); sockets stage host→device first since their payloads land in pageable host memory. The reorder/gather kernels (`packet_reorder_copy_payload_by_sequence`, `packet_gather_copy_payload`) live in `src/kernels.cu`. `gemm` is FP32 `cublasSgemm`; `gemm_fp16` is the same-size mixed-precision FP16/tensor-core `cublasGemmEx` (inference-style); the contiguous buffer supplies the FFT input / GEMM A operand. `--workload-gemm-dim N` pins the square GEMM side length (default 1024), so the FLOP count per call (2·n³) is FIXED and the compute working set is exactly n·n·elem_size, read from the front of each received I/O unit (the unit must be at least that large). `--workload-fft-len N` pins the 1-D C2C transform length for `fft` (default 1024; independent of the GEMM dimension); the working set is fanned out across as many batched length-N transforms as fit. Used by `run_spark_bench.sh`'s `WORKLOAD` / `GEMM_DIM` / `FFT_LEN` env (all backends) to fill the CSV `post_process` / `post_process_gemm_dim` columns (issue #15).

```bash
./build/examples/daqiri_bench_raw_gpudirect ./build/examples/daqiri_bench_raw_tx_rx.yaml --seconds 10
Expand Down Expand Up @@ -174,7 +174,7 @@ The web docs live in `docs/` and are built with [MkDocs Material](https://squidf
- `docs/getting-started.md` — system requirements, build instructions, and first benchmark smoke-test guidance. Only add information to Getting Started when it directly affects requirements, library build steps, or benchmark smoke-test instructions.
- `docs/concepts.md` — terminology glossary (stream types and endpoint URI schemes, GPUDirect, packet/burst/segment, flow/queue, memory region, zero-copy ownership, RX reorder). Meant to be opened in parallel with the rest of the docs.
- `docs/api-reference/index.md` — API guide (6-step application lifecycle, configuration-first model)
- `docs/api-reference/configuration.md`, `docs/api-reference/cpp.md`, `docs/api-reference/python.md` — YAML schema, C++ API, and Python bindings docs
- `docs/api-reference/configuration.md`, `docs/api-reference/cpp.md`, `docs/api-reference/python.md` — YAML reference, C++ API, and Python bindings docs
- `docs/tutorials/` — tutorial walkthroughs (system config, config-file walkthrough, Holoscan integration, ResNet inference)
- `docs/benchmarks/` — benchmark guide pages, surfaced as a top-level "Benchmarking" nav section in `mkdocs.yml` and the landing page (`docs/index.md`):
- `docs/benchmarks/index.md` — overview and engine-selection decision tree
Expand All @@ -183,7 +183,7 @@ The web docs live in `docs/` and are built with [MkDocs Material](https://squidf
- `docs/benchmarks/performance-dgx-spark.md` — per-platform performance report for DGX Spark stream/protocol combinations (the long internal report lives outside the repo in `projects/daqiri-notes/`)
- `docs/stylesheets/extra.css` — custom theme overrides

**User-facing vocabulary:** the YAML schema uses `stream_type` (`raw`, `socket`, future `pcie`); for socket streams the transport is encoded in the endpoint URI scheme (`udp://`, `tcp://`, `roce://`) in `socket_config.local_addr`/`remote_addr`, **not** a separate `protocol` field. (`SocketProtocol` still exists internally, derived from the scheme.) **"Engine"** is the standard term for the specific library backing an implementation; it replaced the former "manager" and "backend" terms and is now used consistently across code (`src/engines/<name>/`, the `Engine` ABC, CMake `DAQIRI_ENGINE`), the API reference, tutorials, the landing page, and concept pages. The mapping: `stream_type: "raw"` is implemented by the `dpdk` engine; `stream_type: "socket"` with `udp://`/`tcp://` endpoints by the always-built `socket` engine; `stream_type: "socket"` with `roce://` endpoints by the `ibverbs` engine.
**User-facing vocabulary:** the YAML format uses `stream_type` (`raw`, `socket`, future `pcie`); for socket streams the transport is encoded in the endpoint URI scheme (`udp://`, `tcp://`, `roce://`) in `socket_config.local_addr`/`remote_addr`, **not** a separate `protocol` field. (`SocketProtocol` still exists internally, derived from the scheme.) **"Engine"** is the standard term for the specific library backing an implementation; it replaced the former "manager" and "backend" terms and is now used consistently across code (`src/engines/<name>/`, the `Engine` ABC, CMake `DAQIRI_ENGINE`), the API reference, tutorials, the landing page, and concept pages. The mapping: `stream_type: "raw"` is implemented by the `dpdk` engine; `stream_type: "socket"` with `udp://`/`tcp://` endpoints by the always-built `socket` engine; `stream_type: "socket"` with `roce://` endpoints by the `ibverbs` engine.

**Keeping docs in sync with code:** before committing changes, scan for the recurring drift hotspots:
- **Stream-type list** (`src/engines/*/`) — README Engines table, `docs/getting-started.md`, `docs/concepts.md` (Stream Types section + Support and testing admonition), `docs/api-reference/configuration.md`
Expand Down
21 changes: 21 additions & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -82,6 +82,27 @@ if(DAQIRI_BUILD_APPLICATIONS)
add_subdirectory(applications)
endif()

set(DAQIRI_CONFIG_GENERATOR_DATA_SUBDIR "daqiri/config-generator")
if(IS_ABSOLUTE "${CMAKE_INSTALL_DATADIR}")
set(DAQIRI_CONFIG_GENERATOR_MODULE_HINT
"${CMAKE_INSTALL_DATADIR}/${DAQIRI_CONFIG_GENERATOR_DATA_SUBDIR}")
elseif(IS_ABSOLUTE "${CMAKE_INSTALL_BINDIR}")
set(DAQIRI_CONFIG_GENERATOR_MODULE_HINT
"${CMAKE_INSTALL_PREFIX}/${CMAKE_INSTALL_DATADIR}/${DAQIRI_CONFIG_GENERATOR_DATA_SUBDIR}")
else()
file(RELATIVE_PATH DAQIRI_CONFIG_GENERATOR_MODULE_HINT
"/${CMAKE_INSTALL_BINDIR}"
"/${CMAKE_INSTALL_DATADIR}/${DAQIRI_CONFIG_GENERATOR_DATA_SUBDIR}")
endif()
configure_file(scripts/gen_daqiri_config.py
${CMAKE_CURRENT_BINARY_DIR}/gen_daqiri_config.py
@ONLY)
install(PROGRAMS ${CMAKE_CURRENT_BINARY_DIR}/gen_daqiri_config.py
DESTINATION ${CMAKE_INSTALL_BINDIR})
install(DIRECTORY scripts/daqiri_config
DESTINATION ${CMAKE_INSTALL_DATADIR}/${DAQIRI_CONFIG_GENERATOR_DATA_SUBDIR}
PATTERN "__pycache__" EXCLUDE)

install(
EXPORT daqiriTargets
NAMESPACE daqiri::
Expand Down
4 changes: 2 additions & 2 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -89,8 +89,8 @@ runners. See `tests/README.md` for the container dependency command, supported
invocations, and marker policy.

Build `daqiri_config_validate` in the required project container before running
`scripts/check_pr.sh`. The check script validates representative checked-in configurations
through the production C++ parser and hardware-independent semantic checks. Set
`scripts/check_pr.sh`. The check script validates the retained and generated configuration
matrices through the production C++ parser and hardware-independent semantic checks. Set
`DAQIRI_CONFIG_VALIDATOR` when the executable is not at `build/tools/daqiri_config_validate`.

When a new example config exercises a configuration form these cases do not cover, add a
Expand Down
2 changes: 2 additions & 0 deletions Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -323,6 +323,8 @@ RUN cmake -S . -B build \
&& cmake --build build -j "$(nproc)" \
&& python3 scripts/check_daqiri_configs.py \
--validator build/tools/daqiri_config_validate \
&& python3 scripts/check_generated_configs.py \
--validator build/tools/daqiri_config_validate \
&& cmake --install build

# ==============================
Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -93,6 +93,7 @@ Reference material for the DAQIRI codebase:
- [Concepts](https://nvidia.github.io/daqiri/concepts/) — Glossary of DAQIRI terminology (kernel bypass, GPUDirect, packet/burst/segment, flow/queue, memory region, zero-copy ownership, RX reorder). Meant to be opened in parallel with the rest of the docs.
- [API Guide](https://nvidia.github.io/daqiri/api-reference/) — Six-step DAQIRI application lifecycle and configuration-first model
- [Configuration YAML Reference](https://nvidia.github.io/daqiri/api-reference/configuration/) — Full YAML config reference for all engines
- [Configuration Generation](https://nvidia.github.io/daqiri/config-generation/) — Deterministic production, benchmark, multi-queue, and cross-host configurations
- [C++ API Usage](https://nvidia.github.io/daqiri/api-reference/cpp/) — C++ RX/TX workflows, buffer lifecycle, file writing, utilities, and status codes
- [Python API Usage](https://nvidia.github.io/daqiri/api-reference/python/) — Python bindings, workflow examples, enums, config classes, and helper functions
- [Performance: DGX Spark](https://nvidia.github.io/daqiri/benchmarks/performance-dgx-spark/) — Per-platform throughput, drop, and utilization numbers for stream/protocol combinations on DGX Spark
Expand Down
16 changes: 11 additions & 5 deletions docs/api-reference/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,9 +10,13 @@ Either form defines memory regions, NIC interfaces, TX/RX queues, and flow rules
is passed to `daqiri_init()` at startup. The struct form is useful for customers who
want to interoperate with existing configuration code.

See `examples/daqiri_bench_*.yaml` for complete working examples.
Start with the commented configurations under `examples/` to see complete,
readable configurations and understand how the fields fit together. When you
need repeatable production, benchmark, multi-queue, or cross-host variants, use
[Configuration Generation](../config-generation.md) to apply system and topology
parameters consistently.

## Validate without hardware initialization
## Optional validation without hardware initialization

Use `daqiri_config_validate` to check YAML files before running an application, such as in CI or
on a machine without the target NIC. `daqiri_init()` performs these checks during startup. The
Expand All @@ -24,8 +28,9 @@ daqiri_config_validate config.yaml another-config.yaml
```

The command exits with status `0` when every file is valid, `1` when any file is invalid, and
`2` when no file was provided. It is built and installed even when
`DAQIRI_BUILD_EXAMPLES=OFF`.
`2` when no file was provided. Use `daqiri_config_validate --list-engines` to print the
engines compiled into the validator and exit successfully without checking files.
It is built and installed even when `DAQIRI_BUILD_EXAMPLES=OFF`.

OpenTelemetry metrics do not add YAML fields. Metrics-enabled builds use the
same interface, queue, and flow names from the active configuration as metric
Expand Down Expand Up @@ -55,6 +60,7 @@ These settings apply globally to both TX and RX:
- **`log_level`**: Engine log level.
- type: `string`
- values: `trace`, `debug`, `info`, `warn` (default), `error`, `critical`, `off`
- any other value is rejected during configuration parsing
- **`loopback`**: Select a loopback mode for local testing.
- type: `string`
- values: `""` (disabled, default), `"sw"` (DPDK software loopback, no NIC),
Expand Down Expand Up @@ -406,7 +412,7 @@ weights.
RSS is flow-affine: every packet with an unchanged five tuple stays on one
queue. Roughly even packet counts require enough distinct tuples with reasonably
balanced traffic; this is not packet striping or exact round-robin delivery.
There is no queue-action mode field in schema v1; a future stripe mode can be
There is no queue-action mode field in configuration version 1; a future stripe mode can be
added without changing the multi-ID RSS default. If
the NIC rejects an RSS action, static initialization or the dynamic flow
completion fails rather than falling back to one queue.
Expand Down
2 changes: 1 addition & 1 deletion docs/api-reference/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ endpoints after initialization, resolve their opaque `EndpointId` handles, and
select a configured TX queue for each submission. These endpoints provide
payload-only TX buffers and are separate from socket endpoint URIs.

The configuration schema lives in the
The configuration format is documented in the
[Configuration YAML Reference](configuration.md). For an annotated
end-to-end example, see the
[configuration walkthrough tutorial](../tutorials/configuration-walkthrough.md).
Expand Down
Loading
Loading