Recreate the 13 omitted tasks exactly as used in the benchmark
This workflow exists so a researcher can clone this release repo and still rebuild the full 30-task benchmark surface locally without us directly redistributing the higher-risk corpora.
What you need
- This release repo clone.
- A separate local clone of
dctabgan-benchmark-expansion. - Either the already-materialized omitted processed CSVs inside that benchmark clone, or the upstream inputs needed for the benchmark prep scripts.
Fast path: copy exact processed CSVs from the benchmark clone
python3 scripts/recreate_omitted_datasets.py \
--benchmark-repo /path/to/dctabgan-benchmark-expansion
python3 scripts/verify_omitted_datasets.py \
--benchmark-repo /path/to/dctabgan-benchmark-expansion
This path is preferred because it recreates the omitted tasks by copying the exact processed benchmark CSVs and verifying them against expected checksums, row counts, label counts, and column counts.
If the processed omitted CSVs are not present yet
python3 scripts/recreate_omitted_datasets.py \
--benchmark-repo /path/to/dctabgan-benchmark-expansion \
--run-benchmark-prep
This wrapper asks the benchmark repo to run its own materialization scripts first and then performs the same copy-and-verify process. It is still only a local workflow; it does not fetch upstream corpora for you.
The underlying benchmark scripts are:
python3 repo/scripts/prepare_preserved_ratio_train500_test500_single_label_benchmark.py
python3 repo/scripts/prepare_modern_ids_preserved_ratio_train500_test500_benchmark.py
python3 repo/scripts/prepare_additional_modern_ids_preserved_ratio_train500_test500_benchmark.py
If you need to go this route, the benchmark worktree must already contain the upstream raw or staged inputs expected by those scripts. The most important cases are:
| Omitted corpus | What the benchmark prep expects |
|---|---|
| CIC-UNSW-NB15 | The benchmark repo expects local raw inputs such as Data.csv and Label.csv in its UNSW raw-data location. |
| CICIoMT2024Small mirror | The benchmark repo expects the modern IDS staging layout and manifests used by the modern7 preparation script, including staged train.csv, test.csv, and split metadata. |
| Edge-IIoTset | The benchmark repo expects the modern IDS staging layout for both the modern7 and additional10 preparation surfaces, depending on task. |
For the exact manifest pointers, read metadata/omitted-datasets-reproduction.md and the protocol notes in the execution runbook.
Exact manifests behind the workflow
The omitted-task script reads the machine-readable manifest:
metadata/omitted-datasets.json
That file records, for each omitted task:
- exact processed source path inside the benchmark repo
- local target path in this release repo
- expected byte size and SHA256
- expected row count, column count, and label counts
- the benchmark rebuild script to use if needed
After local recreation: export ML splits
Once an omitted CSV exists locally under the ignored data/cic-unsw-nb15/, data/ciciomt2024small-mirror/, or data/edge-iiotset/ path, the same split helper can export benchmark-consistent train/test files. This remains a local/private workflow and does not make omitted CSVs part of the public bundle.
python3 scripts/export_benchmark_splits.py \
--task cic_unsw_nb15_exploits \
--output-dir ml_exports/cic_unsw_nb15_exploits