Exact local/private reconstruction

Recreate the 13 omitted tasks exactly as used in the benchmark

This workflow exists so a researcher can clone this release repo and still rebuild the full 30-task benchmark surface locally without us directly redistributing the higher-risk corpora.

What you need

  • This release repo clone.
  • A separate local clone of dctabgan-benchmark-expansion.
  • Either the already-materialized omitted processed CSVs inside that benchmark clone, or the upstream inputs needed for the benchmark prep scripts.
Important: the omitted-task scripts copy and verify local files. They do not download upstream corpora, do not bypass source access requirements, and do not grant redistribution rights for omitted data.

Fast path: copy exact processed CSVs from the benchmark clone

python3 scripts/recreate_omitted_datasets.py \
  --benchmark-repo /path/to/dctabgan-benchmark-expansion

python3 scripts/verify_omitted_datasets.py \
  --benchmark-repo /path/to/dctabgan-benchmark-expansion

This path is preferred because it recreates the omitted tasks by copying the exact processed benchmark CSVs and verifying them against expected checksums, row counts, label counts, and column counts.

If the processed omitted CSVs are not present yet

python3 scripts/recreate_omitted_datasets.py \
  --benchmark-repo /path/to/dctabgan-benchmark-expansion \
  --run-benchmark-prep

This wrapper asks the benchmark repo to run its own materialization scripts first and then performs the same copy-and-verify process. It is still only a local workflow; it does not fetch upstream corpora for you.

The underlying benchmark scripts are:

python3 repo/scripts/prepare_preserved_ratio_train500_test500_single_label_benchmark.py
python3 repo/scripts/prepare_modern_ids_preserved_ratio_train500_test500_benchmark.py
python3 repo/scripts/prepare_additional_modern_ids_preserved_ratio_train500_test500_benchmark.py

If you need to go this route, the benchmark worktree must already contain the upstream raw or staged inputs expected by those scripts. The most important cases are:

Omitted corpus What the benchmark prep expects
CIC-UNSW-NB15The benchmark repo expects local raw inputs such as Data.csv and Label.csv in its UNSW raw-data location.
CICIoMT2024Small mirrorThe benchmark repo expects the modern IDS staging layout and manifests used by the modern7 preparation script, including staged train.csv, test.csv, and split metadata.
Edge-IIoTsetThe benchmark repo expects the modern IDS staging layout for both the modern7 and additional10 preparation surfaces, depending on task.

For the exact manifest pointers, read metadata/omitted-datasets-reproduction.md and the protocol notes in the execution runbook.

Exact manifests behind the workflow

The omitted-task script reads the machine-readable manifest:

metadata/omitted-datasets.json

That file records, for each omitted task:

  • exact processed source path inside the benchmark repo
  • local target path in this release repo
  • expected byte size and SHA256
  • expected row count, column count, and label counts
  • the benchmark rebuild script to use if needed

Read the full omitted-task reproduction note

After local recreation: export ML splits

Once an omitted CSV exists locally under the ignored data/cic-unsw-nb15/, data/ciciomt2024small-mirror/, or data/edge-iiotset/ path, the same split helper can export benchmark-consistent train/test files. This remains a local/private workflow and does not make omitted CSVs part of the public bundle.

python3 scripts/export_benchmark_splits.py \
  --task cic_unsw_nb15_exploits \
  --output-dir ml_exports/cic_unsw_nb15_exploits