Leaf Classification for Local Area Network Deployment: Comparison and Pruning of Resnet18 and Attention Modules (SE/CBAM)

Authors

  • Yanze Guo School of Mathematics, Harbin Institute of Technology, Harbin, 150006, China
  • Mingyue Ma Shanghai Jinqiu A-level Center, Shanghai, 200041, China
  • Zhijie Wang Sino-German College, University of Shanghai for Science and Technology, Shanghai, 200093, China

DOI:

https://doi.org/10.54097/ah0ykv71

Keywords:

Leaf Classification; ResNet18; SE; CBAM; Grad-CAM; Model Pruning; Edge Deployment.

Abstract

This paper focuses on resource-constrained edge deployment for the 15-class grayscale image classification task using the Swedish Leaf dataset. It systematically compares the standard ResNet18, the channel attention SE (Squeeze-and-Excitation) model, and the channel and spatial attention CBAM model. Structured pruning and distillation fine-tuning are then applied to the best-performing model. The data is divided into train, verification, and test sets, with class balance, each comprising 55, 10, and 10 images per class (totalling 825/150/150 images, respectively). Under a unified training strategy (frozen backbone, AdamW, label smoothing, no Mixup), SE achieved a test accuracy of 0.9133, outperforming the Baseline (0.8733) and CBAM (0.8667). Paper further performed progressive channel pruning on stages 3 and 4 of the SE models (6% per step, over three steps, resulting in total reduction parameters of approximately 23.3% and 19.1% FLOPs). This was followed by low-learning-rate, long-cycle fine-tuning. The final test accuracy was 0.8933, which remains higher than the Baseline despite significant model compression. Using Grad-CAM, the paper quantified interpretability through three metrics: leaf coverage, leaf IoU, and leaf margin IoU, with the following results: CBAM > SE > Baseline, indicating that attention modules focus more on the leaf regions. Experiments also showed that eight-view TTA significantly reduced accuracy in this setup, suggesting TTA can have side effects when mismatched with data augmentation/ normalisation. This paper presents comprehensive training/validation curves, confusion matrices, interpretability metrics, and the pruning-accuracy trade-off, providing a reproducible reference for the engineering implementation of lightweight plant phenotyping classification.

Downloads

Download data is not yet available.

References

[1] Hu Jie, Shen Li, Sun Gang. Squeeze-and-Excitation Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018: 7132–7141.

[2] Paszke Adam, Gross Sam, Massa Francisco, et al. PyTorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems (NeurIPS), 2019, 32: 8024–8035.

[3] Loshchilov Ilya, Hutter Frank. Decoupled weight decay regularization. International Conference on Learning Representations (ICLR), 2019.

[4] Zhang Hongyi, Cisse Moustapha, Dauphin Yann N., Lopez-Paz David. mixup: Beyond empirical risk minimization. International Conference on Learning Representations (ICLR), 2018.

[5] He Yihui, Zhang Xiangyu, Sun Jian. Channel pruning for accelerating very deep neural networks. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017: 1389–1397.

[6] Molchanov Pavlo, Tyree Stephan, Karras Tero, Aila Timo, Kautz Jan. Pruning CNNs for resource efficient inference. International Conference on Learning Representations (ICLR), 2017.

[7] Hinton Geoffrey, Vinyals Oriol, Dean Jeff. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.

[8] Chattopadhyay Aditya, Sarkar Anirban, Howlader Prantik, Balasubramanian Vineeth N. Grad-CAM++: Improved visual explanations for deep convolutional networks. IEEE Winter Conference on Applications of Computer Vision (WACV), 2018: 839–847.

[9] Söderkvist Olof J. O. Computer vision classification of leaves from Swedish trees. Master’s thesis, Linköping University, 2001.

[10] He Kaiming, Zhang Xiangyu, Ren Shaoqing, Sun Jian. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016: 770–778.

[11] Woo Sanghyun, Park Jongchan, Lee Joon-Young, Kweon in So. CBAM: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), 2018: 3–19.

[12] Li Hao, Kadav Asim, Durdanovic Igor, Samet Hanan, Graf Hans Peter. Pruning filters for efficient ConvNets. International Conference on Learning Representations (ICLR), 2017.

[13] Selvaraju Ramprasaath R., Cogswell Michael, Das Abhishek, et al. Grad-CAM: Visual explanations from deep networks via gradient-based localization. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017: 618–626.

[14] Girshick Ross B. Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2015: 1440–1448.

Downloads

Published

30-12-2025

How to Cite

Guo, Y., Ma, M., & Wang, Z. (2025). Leaf Classification for Local Area Network Deployment: Comparison and Pruning of Resnet18 and Attention Modules (SE/CBAM). Highlights in Science, Engineering and Technology, 160, 107-118. https://doi.org/10.54097/ah0ykv71