Method stress-testing, not just demo data
You can evaluate synthetic-data generators, oversamplers, undersamplers, and cost-sensitive approaches on 30 controlled IDS tasks instead of a single convenience dataset.
This release is built for academics who want something stronger than a single toy dataset. It gives you a controlled IDS benchmark surface for testing resampling, cost-sensitive learning, synthetic tabular generation, and downstream classifier robustness under explicit majority/minority imbalance.
The full benchmark contains 30 binary attack-vs-benign/normal tasks across eight source corpora. This public release directly ships 17 lower-risk cleaned CSVs and gives you exact scripts, manifests, and copy-and-verify workflows to recreate the remaining 13 omitted tasks locally.
The CSVs are benchmark-ready inputs, not pre-split or model-encoded matrices: labels, selected rows, no-DoS scope, task counts, and row order are already set, while train/test export, scaling, encoding, and estimator-specific preprocessing remain up to you.
label, with 0 as majority benign/normal and 1 as minority attack.classwise_temporal split reconstruction.Offset, SrcTCPBase, and DstTCPBase.You can evaluate synthetic-data generators, oversamplers, undersamplers, and cost-sensitive approaches on 30 controlled IDS tasks instead of a single convenience dataset.
Every task has explicit majority/minority train and test counts. That makes it easier to reason about dose-matching, augmentation volume, and how severe the imbalance really is.
Even where we do not directly redistribute a corpus, the repository still documents the exact processed source paths and provides scripts to recreate the omitted benchmark tasks locally.
Clone the repo, verify the checksums, inspect the task metadata, and use the 17 directly downloadable cleaned datasets immediately.
Use the public bundle for the 17 shipped tasks, then run the included scripts to copy and verify the 13 omitted tasks from a separately obtained benchmark clone.
Follow the benchmark scope, split semantics, and evaluation conventions so your generator can be compared against the same IDS task surface and treatment framing.
The table below mirrors the benchmark-facing dataset composition view used in the paper, with one added column for the source publication year. Publication year refers to the source record or citation year used to describe the corpus, while the row counts are the exact majority/minority supports used in the benchmark surface. For per-task classes, availability, and provenance notes, open the 30-task catalog.
| Year | Source corpus | Benchmark dataset | Attack/minority label | Train maj./min. | Test maj./min. | Redistribution status |
|---|---|---|---|---|---|---|
| 2018 | CIC-IDS-2017 | Friday Bot vs BENIGN | Friday Bot | 48,000/500 | 48,000/500 | Public CSV; CIC provenance/citation caveat |
| 2018 | CIC-IDS-2017 | Thursday Web Attack - Brute Force vs BENIGN | Thursday Web Attack - Brute Force | 55,500/500 | 55,500/500 | Public CSV; CIC provenance/citation caveat |
| 2018 | CIC-IDS-2017 | Tuesday FTP-Patator vs BENIGN | Tuesday FTP-Patator | 27,000/500 | 27,000/500 | Public CSV; CIC provenance/citation caveat |
| 2018 | CIC-IDS-2017 | Tuesday SSH-Patator vs BENIGN | Tuesday SSH-Patator | 36,500/500 | 36,500/500 | Public CSV; CIC provenance/citation caveat |
| 2018 | CSE-CIC-IDS2018 | Bot vs BENIGN | Bot | 23,500/500 | 23,500/500 | Public CSV; CIC/UNB provenance/citation caveat |
| 2018 | CSE-CIC-IDS2018 | FTP-BruteForce vs BENIGN | FTP-BruteForce | 34,500/500 | 34,500/500 | Public CSV; CIC/UNB provenance/citation caveat |
| 2018 | CSE-CIC-IDS2018 | Infilteration vs BENIGN | Infilteration | 41,500/500 | 41,500/500 | Public CSV; CIC/UNB provenance/citation caveat |
| 2018 | CSE-CIC-IDS2018 | SSH-Bruteforce vs BENIGN | SSH-Bruteforce | 35,500/500 | 35,500/500 | Public CSV; CIC/UNB provenance/citation caveat |
| 2015 | CIC-UNSW-NB15 | Exploits vs BENIGN | Exploits | 5,500/500 | 5,500/500 | Omitted; not redistributed; local reproduction only |
| 2015 | CIC-UNSW-NB15 | Fuzzers vs BENIGN | Fuzzers | 6,000/500 | 6,000/500 | Omitted; not redistributed; local reproduction only |
| 2015 | CIC-UNSW-NB15 | Generic vs BENIGN | Generic | 38,500/500 | 38,500/500 | Omitted; not redistributed; local reproduction only |
| 2015 | CIC-UNSW-NB15 | Reconnaissance vs BENIGN | Reconnaissance | 10,500/500 | 10,500/500 | Omitted; not redistributed; local reproduction only |
| 2015 | CIC-UNSW-NB15 | Shellcode vs BENIGN | Shellcode | 85,000/500 | 85,000/500 | Omitted; not redistributed; local reproduction only |
| 2021 | HIKARI-2021 | Bruteforce vs BENIGN | Bruteforce | 43,500/500 | 43,500/500 | Public CSV; CC BY 4.0 attribution route |
| 2021 | HIKARI-2021 | Probing vs BENIGN | Probing | 11,000/500 | 11,000/500 | Public CSV; CC BY 4.0 attribution route |
| 2024 | CICIoMT2024Small mirror | ARP Spoofing vs BENIGN | ARP Spoofing | 12,500/500 | 12,500/500 | Omitted; not redistributed; local reproduction only |
| 2024 | CICIoMT2024Small mirror | MQTT Malformed Data vs BENIGN | MQTT Malformed Data | 8,500/500 | 8,500/500 | Omitted; not redistributed; local reproduction only |
| 2022 | Edge-IIoTset | SQL injection vs NORMAL | SQL_injection_attack | 15,500/500 | 15,500/500 | Omitted; not redistributed; local reproduction only |
| 2022 | Edge-IIoTset | Password vs NORMAL | Password_attack | 16,000/500 | 16,000/500 | Omitted; not redistributed; local reproduction only |
| 2022 | Edge-IIoTset | Uploading vs NORMAL | Uploading_attack | 21,000/500 | 21,000/500 | Omitted; not redistributed; local reproduction only |
| 2022 | 5G-NIDD | TCPConnectScan vs BENIGN | TCPConnectScan | 11,500/500 | 11,500/500 | Public CSV; Fairdata CC BY 4.0 route; IEEE gated |
| 2022 | 5G-NIDD | SYNScan vs BENIGN | SYNScan | 11,500/500 | 11,500/500 | Public CSV; Fairdata CC BY 4.0 route; IEEE gated |
| 2022 | 5G-NIDD | UDPScan vs BENIGN | UDPScan | 15,000/500 | 15,000/500 | Public CSV; Fairdata CC BY 4.0 route; IEEE gated |
| 2022 | Edge-IIoTset | Vulnerability scanner vs NORMAL | Vulnerability_scanner_attack | 16,000/500 | 16,000/500 | Omitted; not redistributed; local reproduction only |
| 2022 | Edge-IIoTset | Backdoor vs NORMAL | Backdoor_attack | 32,000/500 | 32,000/500 | Omitted; not redistributed; local reproduction only |
| 2022 | Edge-IIoTset | Port Scanning vs NORMAL | Port_Scanning_attack | 35,500/500 | 35,500/500 | Omitted; not redistributed; local reproduction only |
| 2023 | RT-IoT2022 | NMAP UDP SCAN vs BENIGN | NMAP_UDP_SCAN | 2,000/500 | 2,000/500 | Public CSV; UCI CC BY 4.0 attribution route |
| 2023 | RT-IoT2022 | NMAP XMAS TREE SCAN vs BENIGN | NMAP_XMAS_TREE_SCAN | 3,000/500 | 3,000/500 | Public CSV; UCI CC BY 4.0 attribution route |
| 2023 | RT-IoT2022 | NMAP OS DETECTION vs BENIGN | NMAP_OS_DETECTION | 3,000/500 | 3,000/500 | Public CSV; UCI CC BY 4.0 attribution route |
| 2023 | RT-IoT2022 | NMAP TCP scan vs BENIGN | NMAP_TCP_scan | 6,000/500 | 6,000/500 | Public CSV; UCI CC BY 4.0 attribution route |
Each task is a cleaned binary IDS CSV with label as the target, where 0 is the majority benign/normal class and 1 is the minority attack class.
The release includes CSV and JSON manifests that record corpus, task key, exact processed source path, counts, availability status, and redistribution caveats.
The release includes scripts that copy the omitted processed CSVs from a separate benchmark checkout and verify the recreated files against expected checksums, counts, and column shapes.
The downloadable CSVs are not raw packet captures, and they are not synthetic outputs. They are the exact cleaned benchmark input tables used to define the task surface, so several benchmark-preparation steps have already been done before any model-specific preprocessing begins.
label target where 0 is majority benign/normal and 1 is minority attack.source_row_index in the full CSVs for auditability.What still remains for you is model-facing preprocessing, not benchmark-input cleaning. Use scripts/export_benchmark_splits.py to export train/test files immediately, and add --encoded if you want dependency-free train-only one-hot encoded matrices for scikit-learn-style experiments.
This release follows the active no-DoS equal-dose IDS benchmark surface used in the DCTABGAN study. The benchmark intentionally excludes DoS, DDoS, and flood-like tasks as a scope-control decision. That is a benchmark design choice, not a claim that those attacks are unimportant.
| Corpus | Full tasks | Downloadable now | Omitted | Redistribution status |
|---|---|---|---|---|
| CIC-IDS-2017 | 4 | 4 | 0 | Downloadable with CIC provenance/citation caveat; no project relicensing |
| CSE-CIC-IDS2018 | 4 | 4 | 0 | Downloadable with CIC/UNB provenance/citation caveat; no project relicensing |
| HIKARI-2021 | 2 | 2 | 0 | Downloadable under documented CC BY 4.0 route with attribution |
| 5G-NIDD | 3 | 3 | 0 | Downloadable under Fairdata CC BY 4.0 route; IEEE route remains gated |
| RT-IoT2022 | 4 | 4 | 0 | Downloadable under UCI CC BY 4.0 route with attribution |
| Edge-IIoTset | 6 | 0 | 6 | Omitted; not redistributed; local/private reproduction only |
| CIC-UNSW-NB15 | 5 | 0 | 5 | Omitted; not redistributed; local/private reproduction only |
| CICIoMT2024Small mirror | 2 | 0 | 2 | Omitted; not redistributed; local/private reproduction only |