Abstract
Hardware attacks exploit microarchitectural vulnerabilities at the CPU level, and RISC-V processors are no exception. These attacks induce anomalous behavior by stressing specific hardware components beyond the patterns observed in legitimate applications. Such deviations can be captured through internal monitoring units, namely Hardware Performance Counters (HPCs). This work analyzes a diverse set of hardware attacks on the RISC-V XuanTie C910 core and introduces a methodology that integrates HPC logging with Machine Learning (ML) techniques for automated attack detection. We present a reproducible dataset composed of 16 benign and 16 malicious applications, enabling systematic evaluation of classifiers such as Random Forest, Decision Trees, Support Vector Machines, and Naive Bayes. Feature-selection strategies, including mRMR and RFECV, are applied to identify the most informative counters. Experimental results show a detection precision of 99% for both benign and malicious samples, while classification performance for malicious samples remains around 95%. In zero-day scenarios, feature selection proves essential: mRMR achieves the best performance with Random Forest, whereas RFECV is more effective for Decision Trees. Overall, the study demonstrates that HPC-based monitoring combined with machine learning provides a hardware-centric and proactive defense mechanism against microarchitectural threats in RISC-V processors.
Similar content being viewed by others
Introduction
Modern information systems are deeply integrated into daily life, from enterprise servers to personal mobile devices. This pervasive reliance underscores the importance of ensuring system security to protect sensitive data and maintain reliable operation. Although many contemporary threats target software-level vulnerabilities, such as privilege escalation in operating systems, these can often be mitigated through timely updates. In contrast, hardware-level attacks represent a more subtle and challenging class of exploits because hardware cannot be easily modified or patched.
Attacks such as Meltdown (Lipp et al. 2018) and Spectre (Kocher et al. 2020) have shown how microarchitectural vulnerabilities can be exploited to access sensitive information without compromising the system software. These attacks leverage speculative execution to access privileged data, producing behavioral patterns that differ significantly from those of legitimate applications. Although relatively new, the RISC-V instruction set architecture (ISA) is not immune to such hardware-based threats, as demonstrated in recent studies (Gerlach et al. 2023).
This work proposes an approach for early detection of hardware attacks on RISC-V architectures by combining HPC-based activity logging with ML techniques for automated detection.
The main contributions of this work are:
-
Dataset construction: We build a reliable and reproducible dataset containing representative samples of both benign applications and known hardware attacks. All samples are labeled, and the process includes selecting relevant binaries, identifying informative HPCs, and automating data collection under controlled conditions.
-
ML model development: We design and evaluate ML models capable of distinguishing between malicious and benign behaviors based on HPC data. This includes algorithm selection, data preprocessing, hyperparameter optimization, and the deployment of a robust classification pipeline.
The remainder of this paper is organized as follows. Section Hardware attacks and RISC-V describes the hardware attacks and benign programs considered, and reviews related work combining HPCs with machine learning. Section System setup and dataset creation details the data collection process. Section AI models presents the ML models and evaluation metrics. Section Experiments and results discusses the experimental setup and the results obtained. Section Future work outlines future research directions. Finally, Section Conclusions summarizes the main findings and concludes the paper.
Hardware attacks and RISC-V
Threat model and attacks under study
Although hardware attacks are a reality, the RISC-V processor family does not implement mitigation techniques like speculative execution disabling or memory access ordering enforcement. Following (Gerlach et al. 2023), we assume an unprivileged attacker (i.e., without kernel-level access) and the absence of kernel-level mitigations against microarchitectural vulnerabilities such as Kernel Address Space Layout Randomization (KASLR) (e.g., KAISER (Corbet 2025), LAZARUS (Gens et al. 2017) or FLARE (Canella et al. 2020)) (which is currently the case for RISC-V processors to date). Throughout this work, we use the RISC-V attack primitives that have proven effective on RISC-V processors: (Gerlach et al. 2023; Thomas et al. 2025)0 and Cheng et al. (2022). The implementations are publicly available at Gerlach et al. (2023), CISPA (2025), and Researchers of Southeast University (2025). A detailed description of these attacks is provided in Table 1.
Related work
Previous research has shown that Hardware Performance Counters (HPCs) combined with machine learning can effectively detect hardware-level attacks (Li and Gaudiot 2022; Carnà et al. 2023; Chiappetta et al. 2016; Bhattacharya and Mukhopadhyay 2015). However, many of these works lack reproducibility, often omitting access to datasets, models, or clear instructions for local replication. Moreover, most prior studies focus exclusively on x86 architectures, which benefit from mature PMUs capable of concurrent multi-counter sampling and from well-established datasets. In contrast, RISC-V is an emerging architecture whose microarchitectural monitoring capabilities remain limited, making systematic evaluation more challenging.
In the RISC-V domain, existing work has explored only narrow subsets of attacks. For example, (Chenet et al. 2024) use HPCs to detect buffer-overflow attacks, while (Le et al. 2022) focus solely on Spectre variants and Prime+Probe (five attacks in total) on a non-commercial chip. More recently, (Gerlach et al. 2023) introduced a comprehensive set of attack primitives (see Section System setup and dataset creation) and validated their effectiveness across several RISC-V CPU implementations. Building on this foundation, our work is the first to investigate the detection of a broad set of new hardware-level attacks–including the Flush family and the GhostWrite benchmark–directly on a real RISC-V processor.
Beyond these studies, we also include recent HPC-based detection efforts on RISC-V that rely on simulation. Several works use GEM5 to collect HPC traces and typically focus on specific attack families. For instance, (Palumbo 2025) evaluates Spectre-type and Return-Flush attacks, without generalizing the methodology to heterogeneous attack classes. Similarly, Hassan et al. (Hassan et al. 2025) detects Flush+Flush attacks on RISC-V using Association Rule Mining, reporting accuracies between 95 and 99% (see their Table II). While competitive for this specific threat model, these approaches do not address a diverse attack set comparable to ours.
Other studies, such as Iamundo (2022), apply Isolation Forest on GEM5-based RISC-V traces to detect anomalies in cryptographic workloads. Although they achieve high accuracy (\(\approx 99\%\)), these methods are tailored to cryptography-specific leakage patterns and do not consider heterogeneous microarchitectural attacks. Furthermore, Isolation Forest and related outlier-detection techniques are unsupervised, whereas our dataset is fully labeled, making supervised learning more appropriate. Unsupervised methods are typically advantageous when attacks cannot be labeled or when datasets are highly imbalanced–conditions that do not characterize our experimental setting. For completeness, we implemented Isolation Forest on our dataset and observed that it only identifies attacks whose HPC signatures do not overlap with benign behavior, reinforcing the suitability of supervised models for our scenario.
We also considered (Dipta and Gulmezoglu 2023), a recent Intel study that operates on real RISC-V hardware but focuses on power-trace monitoring for navigation and video-related attacks using a CNN-based architecture. Although their accuracy (\(\approx 99\%\)) is competitive, the attack set is limited, the threat model differs substantially from ours, and the proposed neural architecture is significantly more complex to deploy in real systems. Consequently, their approach is not directly applicable to our setting.
In summary, no existing work simultaneously (i) targets a diverse set of hardware-level attacks, (ii) operates directly on real RISC-V hardware, and (iii) evaluates supervised learning models using HPCs. Our study provides the first systematic and reproducible evaluation of a broad and diverse set of microarchitectural attacks on RISC-V, characterizing the standalone behavior of individual counters under realistic hardware constraints and establishing a foundational baseline for future research on more advanced RISC-V platforms.
System setup and dataset creation
The dataset was constructed by collecting HPC data from both benign and malicious applications. Benign programs were selected to emulate representative workloads while stressing processor structures commonly targeted by attacks. This ensures that the model can distinguish malicious code even when execution patterns resemble legitimate processes.
To provide broad coverage, the dataset includes memory-intensive workloads to differentiate harmful memory behaviors and compute-intensive tasks to simulate realistic CPU demands. Tables 1 (Section Threat model and attacks under study) and 2 (below) summarize the attack scenarios and benign applications that constitute the foundation of our dataset and subsequent machine learning training. The final dataset comprises 32 applications:16 benign and 16 malicious.
To further assess the diversity and realism of the selected applications, we applied dimensionality-reduction techniques (PCA and UMAP) to the full HPC dataset. As shown in Figure 1, both projections reveal a heterogeneous distribution of benign and malicious programs. Benign workloads form multiple well-separated clusters, reflecting the variety of computational and memory behaviors present in multimedia processing, cryptographic routines, numerical kernels, and stress-test applications. Similarly, malicious applications appear in distinct regions corresponding to different microarchitectural attack primitives (e.g., cache-based, TLB-based, or speculative-execution attacks). This dispersion confirms that the dataset captures a broad spectrum of execution patterns rather than a narrow or redundant set of benchmarks, thereby supporting the robustness of the subsequent machine-learning evaluation.
HPC validation and logging
The experiments were carried out on a XuanTie C910 RISC-V processor (Chen et al. 2020), running Debian 13 on a SiPEED LicheePi 4 A development board, using the second revision of the official user manual (Xuantie n.d.).
To generate the dataset, the perf tool was used to record multiple Hardware Performance Counters (HPCs) during the execution of the benchmarks. The |perf stat| command outputs HPC counts in CSV format. To ensure consistency, a single CPU core was enforced using the |taskset| Linux tool, preventing performance data from being skewed by multi-core scheduling.
First, we validate the HPCs using the benchmarks presented in Tables 1 and 2 to ensure the reliability and reproducibility of the measurements. Table 3 compares the processor specifications (Xuantie n.d.) with the perf event descriptions from |perf stat|. The final column reflects the actual event considered for each counter. We found that most perf event descriptions are inaccurate and the C910 specification provides the accurate mapping (except for counter 0x7, which behaves according to the C906 specs).
Next, we evaluated |perf| interference in the actual counts by logging a varying number of HPCs per run (from 1 to all). As an example, Figure 2 illustrates the relative increase in the Instruction Retired and the Conditional Branch Misprediction counter values for the Matrix Multiply benchmark. The recorded values of some counters increased tremendously (over 300%) as more counters are tracked simultaneously, indicating a significant interference from the |perf| internal routines. Consequently, due to the reproducible nature of the benchmarks, each counter was collected in isolation (i.e., one |perf| run per benchmark per counter). This ensured reliable data collection across all available HPCs.
After quantifying the interference effects, measuring each event in isolation allows us to obtain clean, attribution-correct microarchitectural counters that would be obscured under concurrent sampling in our platform. This isolated characterization provides two benefits. First, it reveals the intrinsic sensitivity of each counter to the studied attacks, enabling a precise understanding of which microarchitectural components are affected and how. Second, it establishes a hardware-agnostic baseline that can guide systems equipped with richer PMUs–where concurrent sampling is possible–in selecting the most informative subset of events for real-time monitoring. Thus, while single-counter runs are a limitation of the specific chip, the resulting per-counter ground truth is essential for designing, prioritizing, and validating multi-counter detection strategies on more capable platforms.
Although the platform provides reliable HPC measurements, it does not expose power sensors or interfaces that allow accurate energy-consumption monitoring. For this reason, our evaluation focuses on latency and microarchitectural behavior, leaving direct power measurements for future work on RISC-V platforms equipped with on-chip energy instrumentation.
Sampling rate and HPC relevance
Previous studies have used sampling rates ranging from 1 to 100 ms per sample. For example, (Li and Gaudiot 2022) dynamically adjusts the sampling rate to counteract evasive malware. In Gerlach et al. (2023), the authors report successful attack durations as small as 800 ms. Thus, to maximize the number of samples for training and for detection with that time-window in mind, we use a sampling interval of 10 ms. All programs are executed multiple times in a single processor by perf in order to track the values from every counter individually. More than 2000 samples were obtained from each iteration. Some programs were modified to repeat their execution to run long enough to generate the desired amount of data. The samples were then merged to build the dataset that contains the trace produced by all counters. Figure 3 depicts the overall process for dataset generation.
The raw datasets resulting from every program execution, which contain the perf readings of all HPCs of the Xuantie c910 processor sampled by 10 ms time intervals, can be found at \citep{harpy-v}. Table 4 provides an overview of the dataset, including the number of programs, samples per program, and the microarchitectural structures exercised by each workload.
To assess the relevance of each Hardware Performance Counter (HPC) in distinguishing between malicious and benign applications, we applied the feature selection methods–mRMR (Ding and Peng 2003; Mazzanti 2024) and RFECV (Guyon and Elisseeff 2003). These methods classify HPCs according to their predictive importance, producing an ordered list from most to least relevant. The resulting classifications are presented in Tables 5 and 6.
AI models
Machine learning is preferred over deep learning due to the relatively small size of the dataset. Deep learning models typically require hundreds of thousands to millions of samples, whereas our dataset contains only 64, 000 samples from 32 different benchmarks. We focus on supervised learning methods because we have a tagged dataset (benign/malicious). Furthermore, the fixed 10 ms sampling interval introduces implicit temporal structure in the HPC traces, but the present work does not leverage models capable of learning explicit temporal dependencies. Architectures such as RNNs, LSTMs, or transformers could capture fine-grained temporal correlations across consecutive samples, potentially improving detection of short-lived or rapidly evolving attacks. Integrating such temporal models constitutes a natural extension of this work, particularly in scenarios where higher-frequency sampling or larger datasets are available.
ML models
We employ Naive Bayes, Decision Tree, Random Forest, and Support Vector Machine classifiers for both binary and multi-class classification. A brief overview of each follows:
-
Naive bayes (NB) is a simple, efficient probabilistic classifier based on the Bayes’ theorem, assuming feature independence (Manning et al. 2008).
-
Decision trees (DT) are popular supervised classifiers that organize decisions in a hierarchical tree structure, where the internal nodes are split into feature values, branches represent outcomes, and leaves assign class labels (Rokach and Maimon 2005).
-
Random forest (RF) builds an ensemble of decision trees using bootstrap samples of the training data (Breiman 2001). At each split, it selects a random subset of features. The final prediction is based on the majority vote among the trees.
-
Support vector machine (SVM) is a supervised classifier that finds the optimal hyperplane maximizing the margin between classes (Press et al. 2007). It uses support vectors, key data points near the boundary, and supports both linear and non-linear kernels such as polynomial and RBF.
Evaluation metrics
The performance of the models was evaluated using metrics derived from the confusion matrix (Table 7).
These metrics are defined as follows:
-
Accuracy: Reflects the overall proportion of correctly classified instances, calculated by
$$\begin{aligned} Accuracy = \frac{TP + TN}{TP + TN + FP + FN}. \end{aligned}$$(1) -
Recall: Indicates the proportion of actual positive instances that are correctly identified, defined as
$$\begin{aligned} Recall = \frac{TP}{TP + FN}. \end{aligned}$$(2) -
Precision: Quantifies the proportion of predicted positive instances that are actually positive, given by
$$\begin{aligned} Precision = \frac{TP}{TP + FP}. \end{aligned}$$(3) -
F1-Score: Represents the harmonic mean of Precision and Recall, balancing both metrics, calculated as
$$\begin{aligned} F1 = \frac{2 \times Precision \times Recall}{Precision + Recall}. \end{aligned}$$(4) -
Matthews correlation coefficient: Reflects the quality of predictions, calculated by
$$\begin{aligned} MCC = \frac{ (TP \cdot TN) - (FP \cdot FN) }{ \sqrt{ (TP+FP)(TP+FN)(TN+FP)(TN+FN) } }. \end{aligned}$$(5)
Experiments and results
We evaluated the dataset described in Section System setup and dataset creation. This balanced dataset is used for training ML models: Naive Bayes (NB), Decision Tree (DT), Random Forest (RF) and Support Vector Machine (SVM). 80% of the samples from each program were used for training and the remaining 20% were reserved for testing. Figure 4 describes the process of training and testing ML models. We adopt the following terminology: Detection corresponds to the binary task of distinguishing benign from malicious instances, while Classification denotes the multi-class problem of determining the exact attack type.
For each combination of HPCs and ML model, hyperparameter tuning was performed using GridSearchCV (Lerman 2018) and GridSearch Documentation (2025). Table 8 presents the optimal parameter values for each method using the most significant counters 5 according to the rankings.
Figures 5 (Detection) and 6 (Classification) show the precision of ML models as a function of the number of HPCs selected in the order shown in Tables 5 and 6. For each method, the results are separated for both ranking techniques (mRMR and RFECV) as they select different counters for each combination. For completeness in the detection analysis, we also show the performance of RFECV for 1–4 counters even if it reports a minimum group of 5 counters (rank 1). We took the first 4 counters in the order returned by RFECV (same as in Table 5). Similarly, in the multi-class classification scenario, RFECV does not eliminate any Hardware Performance Counter (HPC). As shown in Table 6, “RFECV assigns a rank 1 to all counters”, indicating that the recursive elimination process concludes that every HPC contributes discriminative information to the multi-class decision boundaries. This behavior contrasts with the binary detection task, where RFECV identifies a compact subset of informative counters. For classification, however, the method consistently retains the full HPC set, which explains why its performance coincides with that of mRMR when all counters are used.
Figure 5 illustrates that both Random Forest and Decision Tree classifiers maintain an accuracy greater than 95% when at least 3 counters are used. SVM achieves over 90% accuracy when 5 or more counters are employed, except when they are selected using mRMR combined with Linear or Polynomial kernels, where the accuracy is significantly lower. This suggests that certain counter combinations correspond to a more effective separation of the hyperplane between benign and malicious samples. None of the Naive Bayes classifiers achieves 90% accuracy, indicating that this approach is not optimal for detecting malicious behavior. Naive Bayes performance is highly impacted when the features (in our case HPCs) are strongly interdependent, revealing that some counters influence others. DT, RF, and SVM classifiers, with the exception of those using linear and polynomial kernels trained on features selected by mRMR, achieve very high detection scores starting from five counters. Although the overall improvement gained by adding more counters is modest, in more demanding scenarios–such as zero-day detection–using a larger number of features does lead to higher accuracy. This indicates that additional counters improve class discrimination when the classification task becomes more complex. In the classification analysis, Figure 6, incorporating additional counters as features consistently improves the accuracy in all ML models. Decision Tree and Random Forest classifiers approach near-perfect performance when more than seven counters are included, while Support Vector Machines exhibit similar behavior starting from twelve counters. Naive Bayes classifiers obtain the worst results compared to the other models; however, the Gaussian variant achieves over 95% accuracy when nine or more counters are included. This observation suggests that, despite feature dependencies, each class conditional feature distribution tends to approximate a Gaussian pattern.
Tables 9, 10 and 11 present the accuracy, recall, precision, F1-score and MCC for each method using the parameters from Table 8 and the HPCs ranked first 5 (see Table 5). As shown in Table 9, Random Forest achieves the highest performance in the balanced dataset, followed by Decision Tree and SVM. Naive Bayes, by contrast, performs significantly worse in this context. In comparison, Table 10 illustrates that the counters used positively influence the performance of the models when features are selected using RFECV, generally leading to better results compared to mRMR for the same number of counters. It can also be observed that the Matthews Correlation Coefficient (MCC), a correlation measure between observed and predicted classifications, corroborates the high accuracy obtained for RF, DT, and most SVM models. In contrast, the Linear and Polynomial SVM variants derived from the selection of the mRMR feature exhibit MCC values close to 50%, indicating a moderate correlation. Classification metrics shown in Table 11 indicate that 5 counters are sufficient to achieve a performance above 90% across all reported metrics when RF, DT and SVM models are employed. However, this number of counters is insufficient for NB, whose effectiveness is considerably lower. Figure 7 is included for completeness, presenting the confusion matrices corresponding to the different experimental setups of the Random Forest classifier when trained with the samples of the five most relevant counters.
Bootstrap-based statistical confidence analysis
To assess the statistical reliability of the reported performance metrics, we employed a nonparametric bootstrap procedure applied uniformly across all experimental configurations. For each program, 1000 bootstrap replicates were generated by sampling with replacement from the original dataset. Each replicate was evaluated by training a Random Forest classifier using an 80/20 stratified train–test split, with all hyperparameters and feature subsets fixed to those used in the primary experiments. The resulting estimates are summarized in Table 12, which consolidates the mean and 95% confidence intervals for all metrics and experimental settings. A detailed examination of the table reveals several noteworthy patterns.
First, the RFECV-based detection model exhibits extremely narrow confidence intervals across all metrics, with variability consistently below \(3x10^{-3}\). This indicates that the model’s performance is effectively invariant under resampling perturbations. Notably, Recall shows the smallest dispersion (±0.0015%), suggesting that the detector’s ability to identify malicious executions is particularly stable. The tight clustering of Accuracy, F1-score, and MCC further confirms that the selected HPC subset yields a highly robust decision boundary.
The mRMR-based detection model displays similarly compact intervals, but with slightly higher mean values in Accuracy, Precision, and MCC. The reduction in interval width for Accuracy (±0.0011%) and F1-score (±0.0011%) indicates that mRMR produces a feature subset with marginally lower variance than RFECV. This suggests that the counters selected by mRMR may capture more stable discriminative patterns across bootstrap replicates, despite both methods achieving near-saturation performance.
In contrast, the multi-class classification task shows confidence intervals that are approximately one order of magnitude wider than those of the detection models. This increase in dispersion is consistent across all metrics, with MCC exhibiting the largest interval (±0.0087%). These wider intervals reflect the inherently higher complexity of distinguishing among multiple benign program classes, where inter-class variability introduces additional uncertainty. Nevertheless, the intervals remain sufficiently narrow to indicate that the classifier’s performance is statistically reliable, and the mean values remain stable across replicates.
Overall, the results confirm that the observed performance is not driven by a particular train–test split. The narrow confidence intervals across all metrics demonstrate that the selected Hardware Performance Counters yield consistent and reproducible predictive behavior under repeated resampling.
Simulating zero-day attacks
To assess the generalization ability of the classifiers, which would reflect their capacity to correctly identify previously unseen malicious programs within the same domain and demonstrate the applicability of the system in real-world scenarios, we adopt a leave-one-program-out strategy. Specifically, the two best-performing models from our previous analysis (Decision Tree and Random Forest) are trained on a dataset comprising 2000 samples from every benign and malicious program except one malicious program, which is reserved for testing. The same hyperparameters from the original experiment are reused to evaluate intra-domain generalization, as the previous validation was performed using the entire dataset, ensuring consistency and preventing bias from a new optimization process. For each malicious program, we evaluated 4 different configurations of counters (i.e. top 5, 10, 15 and 20, according to the ranking in Table 5). The high accuracy in this context indicates that these samples are correctly classified as malicious. Figure 8 illustrates the precision of the Decision Tree model using mRMR, while Figure 9 presents the results obtained with RFECV. The selected counters differ notably between the two methods. In general, RFECV exhibits more stable and robust performance than mRMR. Despite the fact that mRMR shows high variability depending on the number of counters–particularly for attacks such as Spectre V1, Spectre V2, Interrupt Timing, Flush+Flush, or TLB Eviction, RFECV maintains consistent detection rates even with smaller subsets, achieving optimal performance around 10 counters and avoiding the abrupt drops seen with mRMR.
On the other hand, Figures 10 and 11 show the precision of the Random Forest model using mRMR and RFECV, respectively, for subsets of hardware counters 5, 10, 15 and 20. As in the previous figures, the subset sizes, feature selection methods, and the horizontal axis, indicating attacks not included in the model’s training, are consistent, allowing a direct comparison between the approaches.
Although several attacks – Flush+Fault-ret, Flush+Fault, Page Walk, Spectre V1, Flush+Reload, IFlush+Reload, TLB Eviction– maintain near-perfect accuracy across most configurations, others exhibit substantial variability with mRMR. For example, Fence+Flush and Ghostwrite show abrupt swings in accuracy depending on the subset size, from near-zero with 20 counters to over 90% with smaller subsets. Attacks such as Spectre RSB, Interrupt Timing, Spectre V2, Evict+Reload, Flush+Flush, and Timer Drift, generally yield low accuracy, with occasional peaks when particularly relevant counters are selected. This behavior reflects mRMR’s sensitivity to redundancy and correlation, which can lead to suboptimal subset selection when discriminative signals are weak or distributed. In particular, increasing the number of counters does not guarantee improved accuracy; in some cases, accuracy decreases when moving from 10 to 20 counters. Overall, these results suggest that mRMR is effective for attacks with highly distinctive patterns but less stable than RFECV for diffuse or complex traces, highlighting the critical importance of selecting the right counters rather than simply increasing their number.
When comparing the Decision Tree method (Figures 8 and 9) to the Random Forest method (Figures 10 and 11), we can link the differences in precision to the interaction between the feature selection method and the classifier. RFECV is particularly effective with Decision Trees (DT), which are sensitive to irrelevant features; removing each feature produces a clear structural change, facilitating the identification of an optimal subset. In contrast, Random Forest (RF) is more resistant to noise and redundancy due to its use of subsets of random characteristics, which mitigates the effect of progressive elimination and limits the impact of RFECV.
In contrast, mRMR performs better with RF, as maximizing relevance while minimizing redundancy enhances forest diversity and improves aggregated voting. Applied to a DT, mRMR can discard features that are individually weak but critical for specific decision paths. These methodological differences explain why RFECV pairs better with DT, while mRMR is more effective with RF, highlighting the importance of aligning feature selection techniques with the classifier to ensure reliable detection of microarchitectural attacks.
The zero-day results reveal that the interaction between feature-selection strategies and model architectures plays a decisive role in generalization to unseen attacks. mRMR achieves the best performance when combined with Random Forest because it provides a compact set of non-redundant counters that aligns with the ensemble’s intrinsic feature-decorrelation mechanism, enhancing robustness under distribution shifts. In contrast, RFECV yields superior results with Decision Trees: by recursively eliminating unstable or noisy counters, RFECV reduces variance and prevents overfitting to attack-specific artifacts present in the training set, producing trees that generalize better to novel attack behaviors. This explains the complementary patterns observed in Figures 8- 11, where ensemble-based models benefit from relevance-driven feature diversity, while single-path models benefit from aggressive pruning. Overall, these results highlight that zero-day resilience depends not only on the selected counters but also on the compatibility between the selection method and the inductive biases of each classifier.
The proposed approach leverages hardware performance counters that capture low-level behavioral characteristics of program execution, which are well aligned with the requirements of zero-day generalization (i.e. no signatures, behavioral abstraction, robustness, anomaly sensitivity, noise tolerance -see Section Robustness under multi-process interference, distribution shift, high-frequency sampling). Unlike signature-based methods, these features are less dependent on specific malware instances and may therefore extend to previously unseen threats. In addition, high-frequency sampling implicitly captures short-term temporal dynamics, enabling the model to reflect evolving execution patterns even without explicit sequence-based modeling.
Computational performance analysis
The computational performance evaluation provides a comprehensive view of the feasibility and operational constraints of deploying HPC-based malware detection systems in production environments. The results extracted from the full benchmark dataset reveal consistent patterns across feature selection methods, model families, and task types, offering clear guidance for real-time deployment.
On deployment, the method proposed relies on two steps. First, reading the HPCs and then the ML inference. The average overhead of running |perf| is of 0.2ms per sample (i.e. every 10ms); this corresponds to a 2% execution time overhead.
The next sections detail the latency of inference, memory footprint and energy-efficiency considerations.
Real-time inference enabled by sub-microsecond latency
A key outcome of the analysis is the exceptionally low per-sample latency achieved by lightweight models. Both tree-based and Naive Bayes classifiers consistently operate in the hundreds of nanoseconds range, even when evaluated across different feature selection strategies. For instance, the attached results report prediction time (per sample) 704 ns for Multinomial Naive Bayes mRMR and 1231 ns for Decision Tree mRMR. Similarly, RFECV-based configurations achieve comparable performance, with Multinomial Naive Bayes reaching 673 ns per sample.
These latencies are orders of magnitude below the sampling periods typically used in HPC monitoring pipelines, confirming that real-time inference is not only feasible but sustainable under high-frequency sampling regimes. This is particularly relevant for cloud workloads, containerized microservices, and HPC nodes, where rapid detection of anomalous behavior is essential to minimize propagation and reduce incident response time.
Minimal memory footprint supports deployment on constrained systems
The memory analysis further reinforces the practicality of the proposed approach. Most detection models exhibit extremely small memory footprints, often in the range of a few KiB to 1 MiB. For example, the Random Forest mRMR detector reports a load model space of 1MiB, while several Naive Bayes configurations require only a few KiB (Bernoulli NB mRMR). Even the heaviest tree-based configuration, Decision Tree mRMR, remains modest at 11MiB, well within the capabilities of typical endpoint security agents.
This compact memory usage makes the models suitable for deployment on resource-constrained platforms such as embedded devices, IoT nodes, and edge computing systems. The prediction-time memory footprint is equally small–often below 1 KiB per sample–ensuring minimal pressure on cache hierarchies and memory bandwidth, which is critical in multi-tenant or NUMA-aware environments.
Model complexity introduces significant latency penalties
The analysis also reveals a clear trade-off between model complexity and inference latency. While lightweight models deliver excellent performance, more complex families–particularly SVMs and One-Class classifiers–incur substantial computational overhead. The attached results show per sample prediction time 4086\(\upmu\)s for SVM Linear mRMR and 13 ms for OneClass Benign mRMR, representing several orders of magnitude slowdown compared to Naive Bayes or Decision Trees.
These models may still be valuable for offline analysis, forensic triage, or periodic scanning, where latency is less critical. However, their computational cost makes them unsuitable for continuous real-time monitoring or inline detection pipelines. This distinction is essential for designing multi-layered detection architectures, where heavier models can complement–but not replace–fast, low-overhead detectors.
Energy efficiency implications for large-scale deployments
Although energy consumption was not directly measured, the combination of nanosecond-scale inference times and small memory footprints strongly suggests low power overhead for the most efficient models. In large-scale deployments–such as data centers, HPC clusters, or distributed edge infrastructures–these characteristics can translate into significant cumulative energy savings. For battery-powered devices, the reduced computational load directly contributes to longer operational lifetimes and lower thermal impact.
Overall, the computational results validate the practicality of the proposed approach and provide clear guidance for selecting models that balance accuracy, latency, and resource usage. Lightweight models such as Naive Bayes and Decision Trees consistently emerge as the best candidates for real-time, low-overhead detection, while more complex models are better suited for offline or high-capacity environments. Table 13 summarizes the latency–size trade-offs for the best-performing configurations under each feature selection method and task.
Robustness under multi-process interference
In realistic environments, hardware performance counter (HPC) measurements are affected by noise due to concurrent execution and shared resource contention. In order to overcome this limitation, we consider two approaches to improve robustness: temporal aggregation and normalization.
Temporal aggregation techniques, such as moving averages over multiple samples, can effectively smooth high-frequency fluctuations in the signal, but introduce a tradeoff by increasing detection latency and reducing temporal resolution. In contrast, normalization of HPC values (e.g., expressing counters as ratios over instruction counts) reduces sensitivity to absolute variations in event counts while preserving the fine-grained (10 ms) temporal characteristics of the signal.
In this scenario, we recomputed our dataset by converting raw counts into normalized values. We divided each raw HPC value by the number of instructions retired in that sample (i.e. r002). We trained the Random Forest method with the full -normalized- dataset. On the evaluation part, for each hardware attack, we evaluate 2 multi-process settings: (1) one attack and one benign application; and (2) one attack and 3 benign applications (for any combination of attack/benign applications).
The results show a high accuracy for the normalized dataset 98.56%. Whereas accuracy drops for the raw count dataset (60% accuracy). These results clearly show that, while raw counter values become significantly noisier, normalized metrics remain stable and preserve the characteristic patterns of the hardware attacks analyzed. This allows reliable detection without the need for temporal aggregation. While aggregation could further enhance robustness in highly noisy conditions, normalization alone proves sufficient in our setting and avoids compromising responsiveness.
Limitations of the work
Despite the positive results obtained, this study presents several limitations that should be considered when interpreting its scope and generalizability.
First, the experiments were conducted on a single RISC-V CPU model. This constraint prevents us from evaluating the robustness of the approach against microarchitectural variations such as differences in memory hierarchies, branch predictors, or functional units. The lack of hardware diversity limits the extent to which the results can be generalized to other RISC-V implementations or heterogeneous architectures.
Second, although we evaluated generalization using a leave-one-attack-out strategy, this method does not cover all possible scenarios and remains a partial approximation. Excluding entire attack classes provides a reasonable approximation of a true zero-day setting, but it does not account for modified, obfuscated, or polymorphic variants of known attacks. The absence of such evaluation leaves open the question of how the system would behave against adaptive variants specifically designed to evade hardware-counter-based detection.
Third, the approach relies exclusively on supervised models, which requires the availability of prior attack labels. In dynamic environments or in the presence of emerging threats, this assumption may not hold. Unsupervised or anomaly-detection methods–which do not require prior knowledge of attack types–represent an important direction for extending the applicability of the system.
Furthermore, although the 10 ms sampling interval implicitly introduces temporal structure into the HPC traces, the present work does not incorporate explicit temporal-sequence models. Architectures such as RNNs, LSTMs, or transformer-based encoders could exploit fine-grained temporal dependencies across consecutive samples, enabling the detection of short-duration or rapidly evolving attacks whose signatures may be diluted under coarse sampling. Exploring temporal deep-learning pipelines–potentially combined with higher-frequency sampling or multi-resolution temporal aggregation–constitutes a technically relevant extension of this work.
In addition, the dataset construction relies on instruction-count normalization, which, while effective for aligning traces across workloads, may not adequately capture heterogeneous behaviors such as memory-bound, branch-intensive, or I/O-dominated execution phases. Alternative normalization strategies–including cycle-based scaling, operation-type normalization, or temporal-window aggregation–could preserve microarchitectural dynamics that are not proportional to retired instructions. Evaluating these normalization schemes is essential for improving feature stability across diverse workload profiles.
Moreover, the current evaluation does not incorporate direct energy-consumption measurements. As discussed in Section HPC validation and logging, the experimental platform lacks on-chip power instrumentation, preventing the extraction of power-related behavioral signatures. Future work should explore energy-aware detection pipelines on RISC-V platforms equipped with integrated power sensors, and investigate whether combining HPC-derived features with energy-consumption metrics can enhance robustness, reduce false positives, or reveal attack-specific power anomalies.
Finally, this work does not address cross-architecture generalization. Hardware performance counters are inherently microarchitecture-dependent: differences in pipeline design, cache hierarchies, branch prediction mechanisms, and event definitions can lead to significant variation in counter semantics and statistical distributions across platforms. Consequently, models trained on one architecture may not directly transfer to another without accuracy degradation. From an implementation perspective, these challenges may be addressed through strategies such as lightweight per-architecture retraining (e.g. transfer learning), and the use of architecture-agnostic metrics (e.g. rank-based features), or domain adaptation techniques (e.g. normalization and distribution alignment).
Future work
As a future line of work, it would be valuable to explore additional strategies to improve the classification of all attacks, particularly those with weaker or highly correlated signals. Such strategies include combining different feature selection methods–for example, hybrid approaches integrating mRMR, RFECV, and model-based selectors–to obtain more robust and stable subsets. It would also be relevant to evaluate deep learning techniques capable of extracting latent representations from execution traces, as well as incorporating temporal analysis of the counters, given the dynamic nature of certain attacks.
Other promising directions include the inclusion of new types of microarchitectural counters, the application of supervised dimensionality reduction methods, the use of combined selection techniques to enhance subset stability, and the integration of multimodal information beyond hardware events. Collectively, these extensions could improve the discriminability of more attacks and contribute to the development of more robust and generalizable classifiers.
Conclusions
This work introduces a reliable and reproducible dataset built from hardware performance counters (HPCs), collected from the execution of 16 benign applications and 16 verified microarchitectural attacks. The dataset enables rigorous evaluation of ML-based detection techniques and can be downloaded at \citep{harpy-v}.
A range of ML models was evaluated on this dataset. Random Forest and Decision Trees consistently achieved the highest accuracy across different numbers of counters, while SVM performed competitively, particularly when seven or more counters were used. Naive Bayes achieves lower but relatively stable accuracy. Even when models are trained exclusively on either benign or malicious samples, they retain high overall accuracy, although detection capability decreases when only benign data is available for training. In particular, the models analyzed maintained an accuracy greater than 99% in detecting both benign and malicious samples, while the overall classification performance remained around 95%. These findings reinforce that combining HPC logs with machine learning is an effective strategy for detecting hardware attacks and that optimal performance is obtained when both benign and malicious samples are included during training. In the zero-day scenario, the interplay between the feature-selection strategy and the classifier becomes especially pronounced. With Random Forest, mRMR produces counter subsets that generalize well, correctly detecting most attacks except those with weak or ambiguous signatures (e.g., Spectre v2). However, using RFECV, the attacks Evict+Reload and Timer Drift are not reliably classified, indicating that progressive feature elimination may discard weak but important signals for this model. In contrast, with Decision Trees, RFECV clearly outperforms mRMR–the opposite of what we observe with Random Forest. This reversal stems from the inherent robustness of the Random Forest to noisy or redundant counters. Its random feature sub-sampling diminishes the relative impact of RFECV. Decision Trees lack this robustness and therefore benefit more from RFECV’s incremental pruning of non-informative counters.
Overall, the findings show that HPC-based detection can achieve high fidelity even with a small number of counters and that careful selection of features is often more critical than the choice of classifier. At the same time, the difficulty in detecting certain attacks suggests that additional or more expressive microarchitectural events–and possibly more advanced or hybrid learning approaches–may be needed to ensure resilience against stealthy or previously unseen threats. Future work should explore richer temporal features, ensemble or deep-learning architectures, and evaluations under realistic workloads and adversarial conditions to further advance the practicality and robustness of hardware-assisted attack detection.
Data availibility
The datasets generated and analyzed during the current study are publicly available in the CORA Research Data Repository. The HARPY-V dataset (version V2), containing RISC-V hardware attack traces collected from on-chip hardware performance counters, can be accessed at the following DOI: 10.34810/DATA2538. Repository URL: https://doi.org/10.34810/DATA2538
References
Gerlach L, Weber D, Zhang R (2023) Schwarz, M.: A security risc: Microarchitectural attacks on hardware risc-v cpus. In: 2023 IEEE Symposium on Security and Privacy (SP), pp. 2321–2338. IEEE Computer Society, Los Alamitos, CA, USA. https://doi.org/10.1109/SP46215.2023.10179399
Cheng X, Jiang F, Sun Y, Zhou Z, Mao Y, Wang H, Tong F (2022) A study of mcu-level attacks and defences on power distribution fusion terminals. In: CIRED 2022 Shanghai Workshop, vol. 2022, pp. 610–615. https://doi.org/10.1049/icp.2022.2220
Aguilar-Hernández AI, Marin-Mendoza CA, Valdez R, Müller MF (2025) Uniform approach to reproduce the Mandelbrot set, some of its inner structure as well as its complement with high sensitivity using Fourier phases. Chaos Solitons Fractals 198:116575. https://doi.org/10.1016/j.chaos.2025.116575
Bhattacharya S, Mukhopadhyay D (2015) Who watches the watchmen?: Utilizing performance monitors for compromising keys of rsa on intel platforms. In: Güneysu T, Handschuh H (eds) Cryptographic Hardware and Embedded Systems - CHES 2015. Springer, Berlin, Heidelberg, pp 248–266
Breiman L (2001) Random forests. Mach Learn 45:5–32
Canella C, Schwarz M, Haubenwallner M, Schwarzl M, Gruss D (2020) Kaslr: break it, fix it, repeat. In: the 15th ACM Asia conference on computer and communications security, pp. 481–493. Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3320269.3384747
Carnà S, Ferracci S, Quaglia F, Pellegrini A (2023) Fight hardware with hardware: systemwide detection and mitigation of side-channel attacks using performance counters. Digit Threats Res Pract 4:1–24. https://doi.org/10.1145/3519601
Chen C, Xiang X, Liu C, Shang Y, Guo R, Liu D, Lu Y, Hao Z, Luo J, Chen Z, Li C, Pu Y, Meng J, Yan X, Xie Y, Qi X (2020) Xuantie-910: A commercial multi-core 12-stage pipeline out-of-order 64-bit high performance risc-v processor with vector extension: industrial product. In: 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), pp. 52–64. https://doi.org/10.1109/ISCA45697.2020.00016
Chiappetta M, Savas E, Yilmaz C (2016) Real time detection of cache-based side-channel attacks using hardware performance counters. Appl Soft Comput 49:1162–1174. https://doi.org/10.1016/j.asoc.2016.09.014
Ding C, Peng H (2003) Minimum redundancy feature selection from microarray gene expression data. In: Computational Systems Bioinformatics. CSB2003. Proceedings of the 2003 IEEE Bioinformatics Conference. CSB2003, pp. 523–528. https://doi.org/10.1109/CSB.2003.1227396
Dipta DR, Gulmezoglu B (2023) Mad-en: microarchitectural attack detection through system-wide energy consumption. IEEE Trans Inf Forensics Secur 18:3006–3017. https://doi.org/10.1109/TIFS.2023.3272748
Gens D, Arias O, Sullivan D, Liebchen C, Jin Y, Sadeghi A-R (2017) Lazarus: Practical side-channel resilient kernel-space randomization. In: Dacier M, Bailey M, Polychronakis M, Antonakakis M (eds) Research in Attacks, Intrusions, and Defenses (RAID), vol vol 10453. Springer, Cham, pp 238–258
Guyon I, Elisseeff A (2003) An introduction to variable and feature selection. J Mach Learn Res 3:1157–1182
Corbet J. KAISER: hiding the kernel from user space. LWN. Accessed: 2025-11-03
Kocher P, Horn J, Fogh A, Genkin D, Gruss D, Haas W, Hamburg M, Lipp M, Mangard S, Prescher T, Schwarz M, Yarom Y (2020) Spectre attacks: exploiting speculative execution. Commun ACM 63(7):93–101. https://doi.org/10.1145/3399742
Le A-T, Hoang T-T, Dao B-A, Tsukamoto A, Suzaki K, Pham C-K (2022) Spectre attack detection with neutral network on risc-v processor. In: 2022 IEEE International Symposium on Circuits and Systems (ISCAS), pp. 2467–2471. https://doi.org/10.1109/ISCAS48785.2022.9937212
Lerman PM (2018) Fitting segmented regression models by grid search. J R Stat Soc: Ser C: Appl Stat 29(1):77–84. https://doi.org/10.2307/2346413
Li C, Gaudiot J-L (2022) Detecting spectre attacks using hardware performance counters. IEEE Trans Comput 71(6):1320–1331. https://doi.org/10.1109/TC.2021.3082471
Lipp M, Schwarz M, Gruss D, Prescher T, Haas W, Fogh A, Horn J, Mangard S, Kocher P, Genkin D, Yarom Y, Hamburg M (2018) Meltdown: Reading kernel memory from user space. In: 27th USENIX Security Symposium (USENIX Security 18), pp. 973–990. USENIX Association, Baltimore, MD
Manning CD, Raghavan P, Schütze H (2008) Introduction to information retrieval
Press WH, Teukolsky SA, Vetterling WT, Flannery BP (2007) Numerical Recipes 3rd edition: The Art of Scientific Computing
Thomas F, Arribas EG, Hetterich L, Weber D, Gerlach L, Zhang R, Schwarz M (2025) RISCover: automatic discovery of user-exploitable architectural security vulnerabilities in closed-source RISC-V CPUs. p CCS
Rokach L, Maimon O (2005) Top-down induction of decision trees classifiers - a survey. IEEE Trans Syst Man Cyber Part C (Appl Rev) 35(4):476–487. https://doi.org/10.1109/TSMCC.2004.843247
Bitcoin.org: How Bitcoin works. https://bitcoin.org/en/how-it-works Accessed: 2025-11-03 (n.d.)
Chenet CP, Zhang Z, Savino A, Carlo SD (2024) Hardware-based stack buffer overflow attack detection on RISC-V architectures. https://arxiv.org/abs/2406.10282
CISPA researchers: Access Retired. https://github.com/cispa/Security-RISC/tree/main/access-retired Accessed: 2025-11-03 (n.d.)
CISPA researchers: Evict+Reload. https://github.com/cispa/Security-RISC/tree/main/evict_reload_histogram Accessed: 2025-11-03 (n.d.)
CISPA researchers: Fence+Flush. https://github.com/cispa/Security-RISC/tree/main/fence-flush Accessed: 2025-11-03 (n.d.)
CISPA researchers: Flush+Fault. https://github.com/cispa/Security-RISC/tree/main/flush-fault Accessed: 2025-11-03 (n.d.)
CISPA researchers: Flush+Fault-ret. https://github.com/cispa/Security-RISC/tree/main/flush-fault Accessed: 2025-11-03 (n.d.)
CISPA researchers: Flush+Flush. https://github.com/cispa/Security-RISC/tree/main/flush_flush_histogram Accessed: 2025-11-03 (n.d.)
CISPA researchers: Flush+Reload. https://github.com/cispa/Security-RISC/tree/main/flush_reload_histogram Accessed: 2025-11-03 (n.d.)
CISPA researchers: GhostWrite. https://github.com/cispa/GhostWrite Accessed: 2025-11-03 (n.d.)
CISPA researchers: iFlush+Reload. https://github.com/cispa/Security-RISC/tree/main/iflush_reload_histogram Accessed: 2025-11-03 (n.d.)
CISPA researchers: Interrupt Timing. https://github.com/cispa/Security-RISC/tree/main/interrupt-timing Accessed: 2025-11-03 (n.d.)
CISPA researchers: Page Walk. https://github.com/cispa/Security-RISC/tree/main/page-walk Accessed: 2025-11-03 (n.d.)
CISPA researchers: Spectre v1. https://github.com/cispa/Security-RISC/tree/main/spectre-v1 Accessed: 2025-11-03 (n.d.)
CISPA researchers: Spectre v2. https://github.com/cispa/Security-RISC/tree/main/spectre Accessed: 2025-11-03 (n.d.)
CISPA researchers: Timer Drift. https://github.com/cispa/Security-RISC/tree/main/timer-drift Accessed: 2025-11-03 (n.d.)
CISPA researchers: TLB Eviction. https://github.com/cispa/Security-RISC/tree/main/tlb_evict_histogram Accessed: 2025-11-03 (n.d.)
EEMBC: CoreMark Benchmark. https://www.eembc.org/coremark/ Accessed: 2025-11-03 (n.d.)
FFmpeg Developers: FFmpeg: A Complete, Cross-Platform Solution to Record, Convert and Stream Audio and Video. https://ffmpeg.org/ Accessed: 2025-11-03 (n.d.)
Gerlach L, Weber D, Zhang R, Schwarz M (2023) Security RISC: microarchitectural attacks on hardware RISC-V CPUs artifact repository. https://github.com/cispa/Security-RISC
GNU Project: GNU Coreutils Manual: SHA2 Utilities. https://www.gnu.org/software/coreutils/manual/html_node/sha2-utilities.html Accessed: 2025-11-03 (n.d.)
GridSearch Documentation. https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GridSearchCV.html#sklearn.model_selection.GridSearchCV Accessed: 13-05-2025
Hassan M, Mushtaq M, Raik J, Ghasempouri T (2025) DRsam: detection of fault-based microarchitectural side-channel attacks in RISC-V using statistical preprocessing and association rule mining. arXiv e-prints, 1–5. https://doi.org/10.48550/arXiv.2510.18612arXiv:2510.18612 [cs.CR]
Iamundo M (2022) A machine learning-based security architecture to detect microarchitectural side-channel attacks in microprocessors. Master’s thesis, Politecnico di Milano, Milano, Italy. https://www.politesi.polimi.it/handle/10589/190001
Mazzanti S (2024) mRMR. GitHub. Accessed: 2025-11-03
Ookla: Speedtest. https://www.speedtest.net/apps/cli/ Accessed: 2025-11-03 (n.d.)
Palumbo A (2025) Machine learning-based detection of microarchitectural attacks on risc-v via gem5, pp. 1–7. https://api.semanticscholar.org/CorpusID:282923491
Paul Black, Dictionary of Algorithms and Data Structures (DADS), National Institute of Standards and Technology: Bubble Sort. https://xlinux.nist.gov/dads/HTML/bubblesort.html Accessed: 2025-11-03 (2017)
Paul Black, Dictionary of Algorithms and Data Structures (DADS), National Institute of Standards and Technology: Matrix Multiplication. https://xlinux.nist.gov/dads/HTML/matrixMultiply.html Accessed: 2025-11-03 (2017)
Paul Black, Dictionary of Algorithms and Data Structures (DADS), National Institute of Standards and Technology: Sieve of Eratosthenes. https://xlinux.nist.gov/dads/HTML/sieve.html Accessed: 2025-11-03 (2017)
Pouchet L-N (2025) PolyBench: The Polyhedral Benchmark Suite. https://sourceforge.net/projects/polybench/ Accessed: 2025-11-03 (n.d.)
Researchers of Southeast University, Nanjing: spectre RSB. https://github.com/zznjupt/Spectre-RISCV Accessed: 2025-11-03 (n.d.)
Resurrecting Open Source Projects, A.W.M.: stress. https://github.com/resurrecting-open-source-projects/stress Accessed: 2025-11-03 (n.d.)
Seward J (2025) bzip2: A high-quality data compressor. https://sourceware.org/bzip2/ Accessed: 2025-11-03 (n.d.)
T-HEAD Xuantie C910 Manual. https://occ-intl-prod.oss-ap-southeast-1.aliyuncs.com/resource/XuanTie-OpenC910-UserManual.pdf Accessed: 13-05-2025
University of Michigan researchers: MyBench v1.0, Automotive, Bitcount. https://vhosts.eecs.umich.edu/mibench/ Accessed: 2025-11-03 (n.d.)
Virginia JDM (2025) STREAM. https://www.cs.virginia.edu/stream/ Accessed: 2025-11-03 (n.d.)
Weicker RP (2025) Dhrystone Benchmark (C version). https://www.netlib.org/benchmark/dhry-c Accessed: 2025-11-03 (n.d.)
Acknowledgements
The authors would like to thank the Spanish Ministry of Science and Innovation for supporting this work under contracts PID2024-156150OB-I00 and EQC2024-008344-P, and the Generalitat de Catalunya through grants 2025-SGR-0040 and 2025-SGR-0651. Additional research contributing to these results received funding from the European Union’s Horizon 2020 research and innovation programme under the VITAMIN-V (101093062), WISE4 (101298767), MARE (101191436) and ACCOMPLISH (101189763) projects.
Funding
This work was partially supported by the Spanish Ministry of Science and Innovation through the projects PID2024-156150OB-I00 (MICIU/AEI/10.13039/501100011033 and by ERDF) and EQC2024-008344-P, as well as by the Generalitat de Catalunya under grants 2025-SGR-0040 and 2025-SGR-0651. Furthermore, research contributed from Horizon 2020 research and innovation programs under the HE VITAMIN-V (101093062), WISE4 (101298767), MARE (101191436) and ACCOMPLISH (101189763) projects.
Author information
Authors and Affiliations
Contributions
All authors contributed to the conception and design of the study. Material preparation, data collection, and analysis were performed by the authors. All authors read and approved the final manuscript.
Corresponding author
Ethics declarations
Ethics approval and consent to participate
This study did not involve human participants, personal data, or any procedures requiring institutional ethics approval. Therefore, ethics approval and consent to participate are not applicable.
Consent for publication
No human participants or identifiable personal data are included in this study. Consent for publication is therefore not applicable.
Competing interests
The authors declare that they have no competing interests.
Additional information
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/
About this article
Cite this article
Pou, A., Otero, B., Robledo, M. et al. RISC-V hardware attack detection using on-chip hardware performance counters. Cybersecurity 9, 222 (2026). https://doi.org/10.1186/s42400-026-00662-8
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1186/s42400-026-00662-8
Facts Only
* The study analyzed hardware attacks on the RISC-V XuanTie C910 core.
* A dataset of 16 benign and 16 malicious applications was constructed for evaluation.
* Hardware Performance Counters (HPCs) were logged during application execution.
* Feature selection methods applied were mRMR and RFECV.
* Machine Learning models evaluated included Naive Bayes, Decision Trees, Random Forest, and Support Vector Machines.
* Detection precision reached 99% for both benign and malicious samples in the ML evaluation.
* Random Forest achieved the highest performance in the balanced dataset among tested models.
* The feature selection method influenced results: mRMR paired best with Random Forest, while RFECV was more effective for Decision Trees.
* Real-time inference latency for lightweight models (NB, DT) was in the hundreds of nanoseconds per sample.
* Normalized HPC values showed 98.56% accuracy when used for detection, compared to 60% for raw counts.
Executive Summary
Full Take
Sentinel — Human
This is a highly technical, empirically-driven research paper detailing an experimental methodology for detecting hardware attacks on RISC-V processors using Machine Learning and Hardware Performance Counters.
