Abstract
Class attribution maps (CAMs) are vital for explaining CNN decisions, but they are limited by unreliable
evaluation metrics and low-resolution outputs. To address the evaluation challenge, this paper
introduces a synthetic dataset with ground-truth attributions alongside ARCC, a new composite metric
that more reliably identifies faithful explanations. To fix the resolution issue, the paper proposes
RefineCAM, a method that aggregates CAMs across multiple network layers to produce highly detailed
attribution maps. Ultimately, benchmarking demonstrates that RefineCAM consistently outperforms existing
methods under this new, rigorous evaluation framework.
Evaluate your CAM
Comparing different CAM metrics is challenging due to the lack of ground-truth attributions (we do not know
if
the model is focusing on the region of the image we are expecting or if it is taking a shortcut).
In order to address this, we created a simple synthetic dataset where we know the ground-truth attributions.
We also define a new metric that improves over the existing ADCC by substituting the Average Drop with the ROAD metric, which is more robust to noise and better captures the faithfulness. These new metric is called ARCC and is defined as follows: $$\text{ARCC} = 3 \left( \frac{1}{\text{Coherency}} \right. \left. + \frac{1}{1-\text{Complexity}} + \frac{1}{\text{ROAD}} \right) ^{-1}$$
We also define a new metric that improves over the existing ADCC by substituting the Average Drop with the ROAD metric, which is more robust to noise and better captures the faithfulness. These new metric is called ARCC and is defined as follows: $$\text{ARCC} = 3 \left( \frac{1}{\text{Coherency}} \right. \left. + \frac{1}{1-\text{Complexity}} + \frac{1}{\text{ROAD}} \right) ^{-1}$$
4 examples from our synthetic dataset.
Top row: RGB synthetic images
Bottom row: Ground-truth attributions.
Evaluation results of all the different evaluation metrics on our synthetic dataset.
To compare the
metrics we compute the correlation between each metric
and the golden metric (in this
case
the
cosine similarity between the CAM and the ground-truth attribution).
Refine your CAM
If you want to compute a high resolution CAM, you need to use a shallow layer of the network. However, this
usually leads to low quality and very noisy CAMs. To address this,
we propose RefineCAM, a method that mix the CAMs obtained from multiple layers of the network to produce a
faithful high resolution CAM, defined as:
$$L_{l, \text{Refine-CAM}}^c = \prod_{l' \in \mathcal{L}, l' \geq l} L_{l', \text{Your-CAM}}^c$$
where $\mathcal{L}$ is the set of selected layers of the network, $c$ is the class to explain, $L_{l',
\text{Your-CAM}}^c$ is the CAM
generated by your favorite CAM method and the notation $l' \geq l$ means that $l'$ is deeper than $l$ in the
network.
RefineCAM applied to Grad-CAM++ on VGG11 on ImageNet. Note that our metric, ARCC, is more
faithful
w.r.t. ADCC.
Top row: original CAMs
Bottom
row: RefineCAM outputs.
RefineCAM applied to Grad-CAM++ on VGG11 on our synthetic dataset.
Top row: original CAMs
Bottom row:
RefineCAM
outputs.
RefineCAM applied to LayerCAM on ResNet18 on our synthetic dataset.
Top row: original CAMs
Bottom
row:
RefineCAM
outputs.
Results
The following results shows that RefineCAM consistently improves the quality of the CAMs across different
methods and models, as measured by our ARCC metric.
ARCC performance of our RefineCAM method (OURS) against baseline
attribution methods across 4 models on 1,000 ImageNet images.
Try it out!
The implementation of the metric ARCC and the new meta method RefineCAM are available on the
repository:
Pytorch GradCAM
Pytorch GradCAM
BibTeX
@misc{domeniconi2026evaluaterefinecam,
title={How to Evaluate and Refine your CAM},
author={Luca Domeniconi and Alessandra Stramiglio and Michele Lombardi and Samuele Salti},
year={2026},
eprint={2605.14641},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.14641},
}