Clone the public bundle
Start here to get the 17 directly shipped CSVs, the full 30-task metadata, benchmark definitions, and omitted-task recreation tooling.
This page is for researchers who want to reproduce the same benchmark inputs, not just reuse a few cleaned CSVs.
Start here to get the 17 directly shipped CSVs, the full 30-task metadata, benchmark definitions, and omitted-task recreation tooling.
You need a separate local clone of dctabgan-benchmark-expansion if you want to recreate the omitted tasks exactly.
Use the included scripts to copy and verify the 13 omitted processed CSVs into your local release clone.
The full task surface is inherited from the May 11 no-DoS 30D benchmark definition, and the active operational equal-dose definition inherits that same dataset surface:
metadata/source_manifests/benchmark_definitions/
PRESERVED_RATIO_TRAIN500_TEST500_SINGLE_LABEL_NODOS_30D_BENCHMARK_DEFINITION_2026-05-11.json
metadata/source_manifests/benchmark_definitions/
PRESERVED_RATIO_TRAIN500_TEST500_SINGLE_LABEL_NODOS_30D_CLEAN_DCT_AND_GAN_DOSE_MATCHED_OPERATIONAL_BENCHMARK_DEFINITION_2026-05-20.json
Those files define the exact task list, processed dataset names, data folders, and split semantics used to construct the release.
For immediate ML use, this repository includes a metadata-driven splitter that uses the mirrored benchmark definition and task counts, preserves the benchmark classwise_temporal row-order semantics, and strips obvious audit columns such as source_row_index by default.
python3 scripts/list_tasks.py --status all
python3 scripts/export_benchmark_splits.py \
--task friday_bot \
--output-dir ml_exports/friday_bot
python3 scripts/export_benchmark_splits.py \
--task friday_bot \
--output-dir ml_exports/friday_bot_encoded \
--encoded
The optional --encoded path writes basic X_train.csv, y_train.csv, X_test.csv, and y_test.csv files. Numeric fills and categorical one-hot levels are fit on train rows only, then applied to test rows. This helper is a convenience for scikit-learn-style experiments, not a replacement for the benchmark repo's full treatment/evaluation pipeline.
git clone https://github.com/rayborg/dctabgan-ids-benchmark-datasets.git
# Obtain a separate local checkout of dctabgan-benchmark-expansion
# by whatever route is available to you. This site does not assume
# that benchmark worktree is publicly clonable.
cd dctabgan-ids-benchmark-datasets
shasum -a 256 -c SHA256SUMS.txt
python3 scripts/recreate_omitted_datasets.py \
--benchmark-repo ../dctabgan-benchmark-expansion
python3 scripts/verify_omitted_datasets.py \
--benchmark-repo ../dctabgan-benchmark-expansion
python3 scripts/export_benchmark_splits.py \
--task friday_bot \
--output-dir ml_exports/friday_bot
After that, your local clone contains the same 30-task cleaned dataset surface used by the benchmark definitions.
If your goal is exact replication rather than just dataset reuse, read these mirrored benchmark notes in addition to the dataset manifests:
Those files help explain how the benchmark worktree uses the cleaned CSVs, manifests, and method surface when running the actual modeling and evaluation scripts.
| This dataset repo | The benchmark repo |
|---|---|
| Cleaned benchmark input CSVs | Training/evaluation code and manuscript-facing benchmark logic |
| Public-safe 17-task downloadable subset | Full materialized benchmark worktree |
| Machine-readable task metadata and provenance | Experiment scripts, modeling pipelines, and output tables |
| Omitted-task copy/verify scripts | Optional materialization/rebuild scripts for upstream-derived processed inputs |