Synthetic-data method evaluation
Use the benchmark as a fair IDS comparison surface
If you are testing a new generator or augmentation method, this page highlights the choices that matter most for a fair comparison against the DCTABGAN benchmark setup.
Fair-comparison checklist
| Benchmark choice | What to do |
|---|---|
| Task surface | Use the same 30 no-DoS/no-DDoS/no-flood binary IDS tasks, or explicitly state if you use only the 17 public tasks. |
| Class semantics | Keep label = 0 as majority benign/normal and label = 1 as minority attack. |
| Split-first protocol | Do not fit preprocessing or generate synthetic rows using the test rows. |
| Dose matching | If comparing against the equal-dose headline result, align to the 1x minority-generation target. |
| Evaluation rows | Evaluate on untouched real test rows only. |
| Metrics | Report recall, FPR, precision, MCC, F1, balanced accuracy, AUROC, and AP/PR-AUC. |
| Dataset-level reporting | Prefer dataset-level paired comparisons over pooled-row claims. |
Minimum practical workflow
# 1. Clone the dataset release
git clone https://github.com/rayborg/dctabgan-ids-benchmark-datasets.git
# 2. Optionally recreate the omitted 13 tasks locally
cd dctabgan-ids-benchmark-datasets
python3 scripts/recreate_omitted_datasets.py \
--benchmark-repo /path/to/dctabgan-benchmark-expansion
# 3. Enumerate task keys and availability
python3 scripts/list_tasks.py --status all
# 4. Export benchmark-consistent train/test files for a task
python3 scripts/export_benchmark_splits.py \
--task friday_bot \
--output-dir ml_exports/friday_bot
# 5. Run your own split-first training/evaluation pipeline
If you want to align with the original benchmark structure as closely as possible, use the benchmark definitions and mirrored protocol docs here rather than inventing your own task surface or split semantics:
Metadata fields worth using programmatically
downloadable: whether the task is directly bundledpublic_bundle_status: downloadable vs omittedtrain_countsandtest_counts: exact majority/minority supportssource_benchmark_csv: exact benchmark-processed source pathlocal_reproduction: how to rebuild or copy omitted tasks privatelyredistribution_caveat: per-task/corpus licensing and redistribution note
Suggested reporting checklist for a new generator
- State whether you used the 17 public tasks only or the full 30-task surface.
- State whether your augmentation is dose-matched to the equal-dose 1x framing.
- Keep preprocessing train-only and evaluate on untouched real test rows.
- Report recall, FPR, precision, MCC, F1, balanced accuracy, AUROC, and AP/PR-AUC.
- Prefer dataset-level paired comparisons over pooled-row claims.