
TOTAL VIEWS: 924
As CNNs become increasingly complex, understanding their decision-making processes is crucial for improving model transparency. This study evaluates various visualization tools to enhance interpretability, including saliency maps, Grad-CAM, and feature map visualizations. The goal is to compare these tools based on effectiveness, usability, transparency, and integration with deep learning frameworks. The methodology involved reviewing these tools to assess how well they provide insights into CNNs' learned features and decision pathways. Key findings reveal that NetDissect emerges as the top-performing tool due to its comprehensive capabilities in visualization, user interface, data processing, transparency, and research impact. Other top performers, such as SHAP, AI Explainability 360, and LIME, demonstrate strong performance, particularly in transparency, data processing, and integration with frameworks. Tools such as TensorBoard and Activation Atlas also perform well, excelling in visualization and interactivity. However, tools such as Foolbox and CNN-Fixations have limitations in transparency and interpretability, making them less suitable for in-depth model analysis. The study concludes that NetDissect, SHAP, AI Explainability 360, and LIME are the most effective tools for interpreting CNN models, with NetDissect standing out as the most comprehensive and impactful for research.
Convolutional Neural Networks; Deep Learning; Model Interpretability; Feature Visualization; Explainable Artificial Intelligence
[1] Taye MM. Theoretical understanding of convolutional neural network: concepts, architectures, applications, future directions. Computation. 2023;11(3):52.
[2] Alzubaidi L, Zhang J, Humaidi AJ, Al-Dujaili A, Duan Y, Al-Shamma O, et al. Review of deep learning: concepts, CNN architectures, challenges, applications, future directions. J Big Data. 2021;8(1):53.
[3] Parmar UPS, Surico PL, Singh RB, Goparaju AS. Artificial intelligence (AI) for early diagnosis of retinal diseases. Medicina. 2024;60(4):527.
[4] Ali S, Abuhmed T, El-Sappagh S, Muhammad K, Alonso-Moral JM, Confalonieri R, et al. Explainable artificial intelligence (XAI): what we know and what is left to attain trustworthy artificial intelligence. Inf Fusion. 2023;99:101805.
[5] Kim B, Wattenberg M, Gilmer J, Cai C, Wexler J, Viegas F, et al. Interpretability beyond feature attribution: quantitative testing with concept activation vectors (TCAV). arXiv. 2017. Available from: https://arxiv.org/abs/1711.11279v5
[6] LeCun Y, Bottou L, Bengio Y, Haffner P. Gradient-based learning applied to document recognition. Proc IEEE. 1998;86(11):2278-324.
[7] Krizhevsky A, Sutskever I, Hinton GE. ImageNet classification with deep convolutional neural networks. Adv Neural Inf Process Syst. 2012;25:1097-105.
[8] Simonyan K, Zisserman A. Very deep convolutional networks for large-scale image recognition. arXiv. 2014. Available from: https://arxiv.org/abs/1409.1556
[9] He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. arXiv. 2015. Available from: https://arxiv.org/abs/1512.03385
[10] Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, et al. Going deeper with convolutions. arXiv. 2014. Available from: https://arxiv.org/abs/1409.4842
[11] Jamtsho Y, Yangden P, Wangmo S, Lhaden, Pema, Tshering K, et al. Deep learning-based Dzongkha handwritten digit classification. J Inf Technol Comput Eng. 2024;8(1):1-7.
[12] DEEPLIZARD. Interactive demo - convolution operation [Internet]. Available from: https://deeplizard.com/resource/pavq7noze2
[13] DEEPLIZARD. Interactive demo - max pooling operation [Internet]. Available from: https://deeplizard.com/resource/pavq7noze3
[14] Wang ZJ, Turko R, Shaikh O, Park H, Das N, Hohman F, et al. CNN explainer: learning convolutional neural networks with interactive visualization. IEEE Trans Vis Comput Graph. 2021;27(2):1396-406.
[15] Harley AW. An interactive node-link visualization of convolutional neural networks. In: Bebis G, Boyle R, Parvin B, Koracin D, editors. Advances in visual computing. Cham: Springer International Publishing; 2015. p. 867-77.
[16] Roeder L. Netron: visualizer for neural network, deep learning, and machine learning models [Internet]. Zenodo; 2017. Available from: https://doi.org/10.5281/zenodo.5854962
[17] Abadi M, Barham P, Chen J, Chen Z, Davis A, Dean J, et al. TensorFlow: a system for large-scale machine learning. arXiv. 2016. Available from: https://arxiv.org/abs/1605.08695
[18] Gildenblat J. PyTorch library for CAM methods [Internet]. GitHub; 2021. Available from: https://github.com/jacobgil/pytorch-gradcam
[19] Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-CAM: visual explanations from deep networks via gradient-based localization. Int J Comput Vis. 2020;128(2):336-59.
[20] Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30:4765-74.
[21] Ozbulak U. PyTorch CNN visualizations [Internet]. GitHub; 2019. Available from: https://github.com/utkuozbulak/pytorch-cnn-visualizations
[22] Arya V, Bellamy RKE, Chen PY, Dhurandhar A, Hind M, Hoffman SC, et al. One explanation does not fit all: a toolkit and tax-onomy of AI explainability techniques. arXiv. 2019. Available from: https://arxiv.org/abs/1909.03012
[23] Bojarski M, Choromanska A, Choromanski K, Firner B, Jackel L, Muller U, et al. VisualBackProp: efficient visualization of CNNs. arXiv. 2017. Available from: https://arxiv.org/abs/1611.05418
[24] Yosinski J, Clune J, Nguyen A, Fuchs T, Lipson H. Understanding neural networks through deep visualization. arXiv. 2015. Available from: https://arxiv.org/abs/1506.06579
[25] Ribeiro MT, Singh S, Guestrin C. “Why should I trust you?”: explaining the predictions of any classifier. arXiv. 2016. Available from: https://arxiv.org/abs/1602.04938
[26] Shrikumar A, Greenside P, Kundaje A. Learning important features through propagating activation differences. arXiv. 2019. Available from: https://arxiv.org/abs/1704.02685
[27] Zhang L, Zhang P, Ma X, Zhang S, Tao Y, Li W, et al. A generalized language model in tensor space. arXiv. 2019. Available from: https://arxiv.org/abs/1901.11167
[28] Carter S, Armstrong Z, Schubert L, Johnson I, Olah C. Activation atlas. Distill. 2019;4(3):e15.
[29] Rauber J, Zimmermann R, Bethge M, Brendel W. Foolbox native: fast adversarial attacks to benchmark the robustness of machine learning models in PyTorch, TensorFlow, and JAX. J Open Source Softw. 2020;5(53):2607.
[30] Alber M, Lapuschkin S, Seegerer P, Hägele M, Schütt KT, Montavon G, et al. iNNvestigate neural networks! J Mach Learn Res. 2019;20(93):1-8.
[31] Bau D, Zhou B, Khosla A, Oliva A, Torralba A. Network dissection: quantifying interpretability of deep visual representations. arXiv. 2017. Available from: https://arxiv.org/abs/1704.05796
[32] Mopuri KR, Garg U, Babu RV. CNN fixations: an unraveling approach to visualize the discriminative image regions. arXiv. 2018. Available from: https://arxiv.org/abs/1708.06670
A Comprehensive Survey on Visualization Techniques for Convolutional Neural Networks
How to cite this paper: Yonten Jamtsho. (2025) A Comprehensive Survey on Visualization Techniques for Convolutional Neural Networks. Future Trends in AI Research, 2(1), 26-36.
DOI: http://dx.doi.org/10.26855/ftair.2025.12.006