Open Access

On the Real-World Applicability of Automotive CAN Intrusion Detection Systems - Dataset

Abstract

Description

Machine Learning (ML)-based CAN Intrusion Detection Systems (IDSs) are widely studied and used to detect and mitigate attacks on in-vehicle communication. Despite frequent reports on high detection performance for well-established datasets, the practical applicability and robustness against real-world attackers of the trained solutions is rarely discussed. This work presents a preliminary analysis of the practical robustness of ML-based CAN IDS solutions, with a focus on per-frame, binary intrusion detection approaches. Our approach consists of three steps: (1) We analyze common datasets used for training and evaluating ML-based CAN IDS, (2) train human-interpretable decision trees on the datasets and investigate their learned classification rules, and (3) critically evaluate the models against a practical attacker model. We observe that many commonly used CAN IDS datasets only include simple and easily distinguishable attacks. Consequently, ML-based models trained on these datasets tend to learn dataset-specific statistical artifacts or shortcuts rather than robust or semantically meaningful detection logic, a finding we demonstrate specifically within the scope of per-frame, binary detection. While our results are preliminary and limited to this detection paradigm, they show that high detection performance is not inherently indicative of practical robustness. We therefore argue that new detection strategies founded upon vehicle specifications and domain knowledge, in combination with ML tools, may offer more transparent and reliable protection than purely ML-driven approaches, and we call for broader evaluation across more diverse detection settings in future work.

Citation

Endorsement

Project(s)

Faculty

Collections

License

Except where otherwise noted, this license is described as CC BY 4.0 - Attribution 4.0 International