Synthetic-data method evaluation

Use the benchmark as a fair IDS comparison surface

If you are testing a new generator or augmentation method, this page highlights the choices that matter most for a fair comparison against the DCTABGAN benchmark setup.

Fair-comparison checklist

Benchmark choice What to do
Task surfaceUse the same 30 no-DoS/no-DDoS/no-flood binary IDS tasks, or explicitly state if you use only the 17 public tasks.
Class semanticsKeep label = 0 as majority benign/normal and label = 1 as minority attack.
Split-first protocolDo not fit preprocessing or generate synthetic rows using the test rows.
Dose matchingIf comparing against the equal-dose headline result, align to the 1x minority-generation target.
Evaluation rowsEvaluate on untouched real test rows only.
MetricsReport recall, FPR, precision, MCC, F1, balanced accuracy, AUROC, and AP/PR-AUC.
Dataset-level reportingPrefer dataset-level paired comparisons over pooled-row claims.

Minimum practical workflow

# 1. Clone the dataset release
git clone https://github.com/rayborg/dctabgan-ids-benchmark-datasets.git

# 2. Optionally recreate the omitted 13 tasks locally
cd dctabgan-ids-benchmark-datasets
python3 scripts/recreate_omitted_datasets.py \
  --benchmark-repo /path/to/dctabgan-benchmark-expansion

# 3. Enumerate task keys and availability
python3 scripts/list_tasks.py --status all

# 4. Export benchmark-consistent train/test files for a task
python3 scripts/export_benchmark_splits.py \
  --task friday_bot \
  --output-dir ml_exports/friday_bot

# 5. Run your own split-first training/evaluation pipeline

If you want to align with the original benchmark structure as closely as possible, use the benchmark definitions and mirrored protocol docs here rather than inventing your own task surface or split semantics:

Metadata fields worth using programmatically

  • downloadable: whether the task is directly bundled
  • public_bundle_status: downloadable vs omitted
  • train_counts and test_counts: exact majority/minority supports
  • source_benchmark_csv: exact benchmark-processed source path
  • local_reproduction: how to rebuild or copy omitted tasks privately
  • redistribution_caveat: per-task/corpus licensing and redistribution note

Open the full JSON manifest

Suggested reporting checklist for a new generator

  • State whether you used the 17 public tasks only or the full 30-task surface.
  • State whether your augmentation is dose-matched to the equal-dose 1x framing.
  • Keep preprocessing train-only and evaluate on untouched real test rows.
  • Report recall, FPR, precision, MCC, F1, balanced accuracy, AUROC, and AP/PR-AUC.
  • Prefer dataset-level paired comparisons over pooled-row claims.