IDS benchmark release for reproducible ML research

Stress-test imbalanced-data methods on a real IDS benchmark with exact task counts and reproducible splits.

This release is built for academics who want something stronger than a single toy dataset. It gives you a controlled IDS benchmark surface for testing resampling, cost-sensitive learning, synthetic tabular generation, and downstream classifier robustness under explicit majority/minority imbalance.

The full benchmark contains 30 binary attack-vs-benign/normal tasks across eight source corpora. This public release directly ships 17 lower-risk cleaned CSVs and gives you exact scripts, manifests, and copy-and-verify workflows to recreate the remaining 13 omitted tasks locally.

The CSVs are benchmark-ready inputs, not pre-split or model-encoded matrices: labels, selected rows, no-DoS scope, task counts, and row order are already set, while train/test export, scaling, encoding, and estimator-specific preprocessing remain up to you.

What has already been done to the CSVs

  • Binary attack-vs-benign/normal task selection is complete under the no-DoS/no-DDoS/no-flood benchmark scope.
  • The target is normalized to label, with 0 as majority benign/normal and 1 as minority attack.
  • Preserved-ratio selected rows are materialized for 500 minority train rows and 500 minority test rows per task.
  • CSV row order is retained for benchmark classwise_temporal split reconstruction.
  • Recorded materialization drops are already applied, such as 5G-NIDD Offset, SrcTCPBase, and DstTCPBase.

Why this benchmark is useful for imbalanced-learning research

Method stress-testing, not just demo data

You can evaluate synthetic-data generators, oversamplers, undersamplers, and cost-sensitive approaches on 30 controlled IDS tasks instead of a single convenience dataset.

Exact task composition is visible

Every task has explicit majority/minority train and test counts. That makes it easier to reason about dose-matching, augmentation volume, and how severe the imbalance really is.

Full-surface replication is still possible

Even where we do not directly redistribute a corpus, the repository still documents the exact processed source paths and provides scripts to recreate the omitted benchmark tasks locally.

Choose your path

I want benchmark-ready IDS CSVs for my own ML experiments

Clone the repo, verify the checksums, inspect the task metadata, and use the 17 directly downloadable cleaned datasets immediately.

Fastest path 17 public CSVs ML-ready

Go to the download guide

I want the full 30-task dataset surface we used

Use the public bundle for the 17 shipped tasks, then run the included scripts to copy and verify the 13 omitted tasks from a separately obtained benchmark clone.

Exact task recreation Checksum verified No relicensing claim

Go to the recreation walkthrough

I want to compare a new synthetic-data method fairly

Follow the benchmark scope, split semantics, and evaluation conventions so your generator can be compared against the same IDS task surface and treatment framing.

Benchmark-focused Generator comparison Reproducible

Go to the comparison guide

Full task composition

The table below mirrors the benchmark-facing dataset composition view used in the paper, with one added column for the source publication year. Publication year refers to the source record or citation year used to describe the corpus, while the row counts are the exact majority/minority supports used in the benchmark surface. For per-task classes, availability, and provenance notes, open the 30-task catalog.

Year Source corpus Benchmark dataset Attack/minority label Train maj./min. Test maj./min. Redistribution status
2018CIC-IDS-2017Friday Bot vs BENIGNFriday Bot48,000/50048,000/500Public CSV; CIC provenance/citation caveat
2018CIC-IDS-2017Thursday Web Attack - Brute Force vs BENIGNThursday Web Attack - Brute Force55,500/50055,500/500Public CSV; CIC provenance/citation caveat
2018CIC-IDS-2017Tuesday FTP-Patator vs BENIGNTuesday FTP-Patator27,000/50027,000/500Public CSV; CIC provenance/citation caveat
2018CIC-IDS-2017Tuesday SSH-Patator vs BENIGNTuesday SSH-Patator36,500/50036,500/500Public CSV; CIC provenance/citation caveat
2018CSE-CIC-IDS2018Bot vs BENIGNBot23,500/50023,500/500Public CSV; CIC/UNB provenance/citation caveat
2018CSE-CIC-IDS2018FTP-BruteForce vs BENIGNFTP-BruteForce34,500/50034,500/500Public CSV; CIC/UNB provenance/citation caveat
2018CSE-CIC-IDS2018Infilteration vs BENIGNInfilteration41,500/50041,500/500Public CSV; CIC/UNB provenance/citation caveat
2018CSE-CIC-IDS2018SSH-Bruteforce vs BENIGNSSH-Bruteforce35,500/50035,500/500Public CSV; CIC/UNB provenance/citation caveat
2015CIC-UNSW-NB15Exploits vs BENIGNExploits5,500/5005,500/500Omitted; not redistributed; local reproduction only
2015CIC-UNSW-NB15Fuzzers vs BENIGNFuzzers6,000/5006,000/500Omitted; not redistributed; local reproduction only
2015CIC-UNSW-NB15Generic vs BENIGNGeneric38,500/50038,500/500Omitted; not redistributed; local reproduction only
2015CIC-UNSW-NB15Reconnaissance vs BENIGNReconnaissance10,500/50010,500/500Omitted; not redistributed; local reproduction only
2015CIC-UNSW-NB15Shellcode vs BENIGNShellcode85,000/50085,000/500Omitted; not redistributed; local reproduction only
2021HIKARI-2021Bruteforce vs BENIGNBruteforce43,500/50043,500/500Public CSV; CC BY 4.0 attribution route
2021HIKARI-2021Probing vs BENIGNProbing11,000/50011,000/500Public CSV; CC BY 4.0 attribution route
2024CICIoMT2024Small mirrorARP Spoofing vs BENIGNARP Spoofing12,500/50012,500/500Omitted; not redistributed; local reproduction only
2024CICIoMT2024Small mirrorMQTT Malformed Data vs BENIGNMQTT Malformed Data8,500/5008,500/500Omitted; not redistributed; local reproduction only
2022Edge-IIoTsetSQL injection vs NORMALSQL_injection_attack15,500/50015,500/500Omitted; not redistributed; local reproduction only
2022Edge-IIoTsetPassword vs NORMALPassword_attack16,000/50016,000/500Omitted; not redistributed; local reproduction only
2022Edge-IIoTsetUploading vs NORMALUploading_attack21,000/50021,000/500Omitted; not redistributed; local reproduction only
20225G-NIDDTCPConnectScan vs BENIGNTCPConnectScan11,500/50011,500/500Public CSV; Fairdata CC BY 4.0 route; IEEE gated
20225G-NIDDSYNScan vs BENIGNSYNScan11,500/50011,500/500Public CSV; Fairdata CC BY 4.0 route; IEEE gated
20225G-NIDDUDPScan vs BENIGNUDPScan15,000/50015,000/500Public CSV; Fairdata CC BY 4.0 route; IEEE gated
2022Edge-IIoTsetVulnerability scanner vs NORMALVulnerability_scanner_attack16,000/50016,000/500Omitted; not redistributed; local reproduction only
2022Edge-IIoTsetBackdoor vs NORMALBackdoor_attack32,000/50032,000/500Omitted; not redistributed; local reproduction only
2022Edge-IIoTsetPort Scanning vs NORMALPort_Scanning_attack35,500/50035,500/500Omitted; not redistributed; local reproduction only
2023RT-IoT2022NMAP UDP SCAN vs BENIGNNMAP_UDP_SCAN2,000/5002,000/500Public CSV; UCI CC BY 4.0 attribution route
2023RT-IoT2022NMAP XMAS TREE SCAN vs BENIGNNMAP_XMAS_TREE_SCAN3,000/5003,000/500Public CSV; UCI CC BY 4.0 attribution route
2023RT-IoT2022NMAP OS DETECTION vs BENIGNNMAP_OS_DETECTION3,000/5003,000/500Public CSV; UCI CC BY 4.0 attribution route
2023RT-IoT2022NMAP TCP scan vs BENIGNNMAP_TCP_scan6,000/5006,000/500Public CSV; UCI CC BY 4.0 attribution route

What this repository gives you

1

Cleaned benchmark inputs

Each task is a cleaned binary IDS CSV with label as the target, where 0 is the majority benign/normal class and 1 is the minority attack class.

2

Machine-readable metadata

The release includes CSV and JSON manifests that record corpus, task key, exact processed source path, counts, availability status, and redistribution caveats.

3

Exact omitted-task reconstruction

The release includes scripts that copy the omitted processed CSVs from a separate benchmark checkout and verify the recreated files against expected checksums, counts, and column shapes.

What we did to make these files benchmark-ready

The downloadable CSVs are not raw packet captures, and they are not synthetic outputs. They are the exact cleaned benchmark input tables used to define the task surface, so several benchmark-preparation steps have already been done before any model-specific preprocessing begins.

  • Selected the exact 30 binary IDS tasks used in the benchmark surface.
  • Applied the benchmark no-DoS, no-DDoS, and no-flood scope rule.
  • Normalized every task to a binary label target where 0 is majority benign/normal and 1 is minority attack.
  • Materialized preserved-ratio benchmark subsets with exact majority/minority row counts.
  • Preserved benchmark row order so train/test reconstruction stays consistent with the benchmark definitions.
  • Applied task-specific materialization drops already recorded in the manifests, such as the 5G-NIDD capture-order artifact columns and the high-cardinality Edge-IIoTset artifact columns.
  • Kept provenance columns such as source_row_index in the full CSVs for auditability.

What still remains for you is model-facing preprocessing, not benchmark-input cleaning. Use scripts/export_benchmark_splits.py to export train/test files immediately, and add --encoded if you want dependency-free train-only one-hot encoded matrices for scikit-learn-style experiments.

Benchmark scope

This release follows the active no-DoS equal-dose IDS benchmark surface used in the DCTABGAN study. The benchmark intentionally excludes DoS, DDoS, and flood-like tasks as a scope-control decision. That is a benchmark design choice, not a claim that those attacks are unimportant.

Corpus Full tasks Downloadable now Omitted Redistribution status
CIC-IDS-2017440Downloadable with CIC provenance/citation caveat; no project relicensing
CSE-CIC-IDS2018440Downloadable with CIC/UNB provenance/citation caveat; no project relicensing
HIKARI-2021220Downloadable under documented CC BY 4.0 route with attribution
5G-NIDD330Downloadable under Fairdata CC BY 4.0 route; IEEE route remains gated
RT-IoT2022440Downloadable under UCI CC BY 4.0 route with attribution
Edge-IIoTset606Omitted; not redistributed; local/private reproduction only
CIC-UNSW-NB15505Omitted; not redistributed; local/private reproduction only
CICIoMT2024Small mirror202Omitted; not redistributed; local/private reproduction only