Signal Study Finds Most Explainable-AI Methods for ECG Diagnosis Miss Clinically Relevant Signal Regions
Summary
Researchers evaluated 13 post-hoc explainable-AI (XAI) methods used to interpret ECG classification models, testing them against clinically defined regions of diagnostic interest rather than sample-level plausibility alone. Using four binary classifiers trained on the PTB-XL dataset, they compared explanations across low-amplitude segments and high-amplitude QRS morphology patterns. The results show a systematic failure among methods transferred from computer vision: explanations tend to track signal amplitude rather than clinical relevance, with correlations to amplitude reaching Spearman values up to 0.69. For ischemia detection, one common method (LRP-ε) assigned only 4.6% of relevance to the clinically decisive ST segment, compared with 63.8% for an alternative method (LRP-SIGN). Nine of the 13 evaluated methods performed below chance level for at least one cardiac condition, indicating inconsistent reliability. The authors conclude that global, guideline-grounded evaluation can reveal systematic explanation failures invisible in individual sample-level heatmaps.
Classification
Evidence 1
- arXiv (cs.CY, cs.LG) 2026-07-27 accessed 2026-07-28T15:00:27+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-7422f8ff67cc