Transformer-Based Hybrid Convolution Attention Framework for Plant Species Classification
DOI:
https://doi.org/10.46604/peti.2026.16186Keywords:
fine-grained visual categorization, hybrid inductive bias, plant species identification, leakage-safe evaluation, CoAtNet architectureAbstract
Fine-grained plant species identification remains challenging because of high intra-class variation and strong inter-class similarity among closely related taxa. Existing hybrid convolutional neural network (CNN)-Transformer architectures progressively attenuate discriminative texture cues during deep attention processing. To address this limitation, this study aims to develop the Transformer-based hybrid convolution attention framework (TB-HCAF), a targeted extension of CoAtNet that preserves complementary local morphological and global structural representations through a dual-path architecture, a parameter-efficient spatial alignment block, and a channel-wise hybrid fusion. The framework is evaluated using a leakage-safe, group-stratified k-fold protocol eliminates specimen-level data leakage and validated through comparative experiments, ablation studies, effective receptive field (ERF) analysis, and cross-dataset evaluation. TB-HCAF achieves high Top-1 accuracy on PlantCLEF 2015, surpassing representative baseline models, and further demonstrates strong zero-shot cross-dataset generalization on Oxford 102 Flowers and iNaturalist 2018. ERF analysis confirms improved discrimination of morphologically similar species and robustness across long-tailed distributions.
References
A. Joly, H. Goëau, H. Glotin, C. Spampinato, P. Bonnet, W.-P. Vellinga, et al., “LifeCLEF 2015: Multimedia Life Species Identification Challenges,” Proceedings of Experimental IR Meets Multilinguality, Multimodality, and Interaction, vol. 9283, pp. 462-483, 2015.
G. Van Horn, O. Mac Aodha, Y. Song, Y. Cui, C. Sun, A. Shepard, et al., “The iNaturalist Species Classification and Detection Dataset,” Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8769-8778, 2018.
T.-Y. Lin, A. RoyChowdhury, and S. Maji, “Bilinear CNN Models for Fine-Grained Visual Recognition,” Proceedings of IEEE International Conference on Computer Vision (ICCV), pp. 1449-1457, 2015.
H. Zheng, J. Fu, T. Mei, and J. Luo, “Learning Multi-attention Convolutional Neural Network for Fine-Grained Image Recognition,” Proceedings of IEEE International Conference on Computer Vision (ICCV), pp. 5209-5217, 2017.
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” arXiv preprint arXiv:1704.04861, 2017.
W. Luo, Y. Li, R. Urtasun, and R. Zemel, “Understanding the Effective Receptive Field in Deep Convolutional Neural Networks,” Proceedings of Advances in Neural Information Processing Systems, vol. 29, pp. 4898-4906, 2016.
Y. S. Tong, T. H. Lee, and K. S. Yen, “Recognition of Ginger Seed Growth Stages Using a Two-Stage Deep Learning Approach,” Proceedings of Engineering and Technology Innovation, vol. 26, pp. 1-17, 2024.
V. Choudhary and A. Thakur, “Multiclass Plant Leaf Disease Prediction Using Fuzzy Multimodal Feature Extraction,” Advances in Technology Innovation, vol. 10, no. 4, pp. 370-382, 2025.
F. O. Isinkaye, M. O. Olusanya, and A. A. Akinyelu, “A Multi-Class Hybrid Variational Autoencoder and Vision Transformer Model for Enhanced Plant Disease Identification,” Intelligent Systems with Applications, vol. 26, article no. 200490, 2025.
M. A. K. Sergio, A. B. S. Siallagan, and E. Joelianto, “Potato Leaf Disease Detection by Means of Vision Transformer and CoAtNet Architecture,” Proceedings of IEEE International Conference on Technology, Informatics, Management, Engineering and Environment (TIME-E), pp. 154-159, 2024.
Y. Chen, Y. Bai, W. Zhang, and T. Mei, “Destruction and Construction Learning for Fine-Grained Image Recognition,” Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5157-5166, 2019.
Z. Yang, T. Luo, D. Wang, Z. Hu, J. Gao, and L. Wang, “Learning to Navigate for Fine-Grained Classification,” Proceedings of European Conference on Computer Vision (ECCV), pp. 420-435, 2018.
Y. Cui, M. Jia, T.-Y. Lin, Y. Song, and S. Belongie, “Class-Balanced Loss Based on Effective Number of Samples,” Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9268-9277, 2019.
M.-E. Nilsback and A. Zisserman, “Automated Flower Classification over a Large Number of Classes,” Proceedings of the Sixth Indian Conference on Computer Vision, Graphics and Image Processing (ICVGIP), pp. 722-729, 2008.
Z. Dai, H. Liu, Q. V. Le, and M. Tan, “CoAtNet: Marrying Convolution and Attention for All Data Sizes,” Advances in Neural Information Processing Systems, vol. 34, pp. 3965-3977, 2021.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, et al., “Attention Is All You Need,” Advances in Neural Information Processing Systems, vol. 30, pp. 5998-6008, 2017.
M. Tan and Q. V. Le, “EfficientNetV2: Smaller Models and Faster Training,” Proceedings of the 38th International Conference on Machine Learning (ICML), pp. 10096-10106, 2021.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception Architecture for Computer Vision,” Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2818-2826, 2016.
I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” Proceedings of International Conference on Learning Representations (ICLR), 2019.
S. G. Müller and F. Hutter, “TrivialAugment: Tuning-Free Yet State-of-the-Art Data Augmentation,” Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV), pp. 774-782, 2021.
F. Zhuang, Z. Qi, K. Duan, D. Xi, Y. Zhu, H. Zhu, et al., “A Comprehensive Survey on Transfer Learning,” Proceedings of the IEEE, vol. 109, no. 1, pp. 43-76, 2021.
A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, et al., “Big Transfer (BiT): General Visual Representation Learning,” Proceedings of European Conference on Computer Vision (ECCV), pp. 491-507, 2020.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, et al., “An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale,” Proceedings of International Conference on Learning Representations (ICLR), 2021.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, et al., “Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows,” Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV), pp. 10012-10022, 2021.
H. Zhang and G. Ren, “Intelligent Leaf Disease Diagnosis: Image Algorithms Using Swin Transformer and Federated Learning,” The Visual Computer, vol. 41, pp. 4815-4838, 2025.
S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I. S. Kweon, et al., “ConvNeXt V2: Co-Designing and Scaling ConvNets with Masked Autoencoders,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 16133-16142, 2023.
E. D. Cubuk, B. Zoph, J. Shlens, and Q. V. Le, “RandAugment: Practical Automated Data Augmentation with a Reduced Search Space,” Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 3008-3017, 2020.
The Angiosperm Phylogeny Group, “An Update of the Angiosperm Phylogeny Group Classification for the Orders and Families of Flowering Plants: APG IV,” Botanical Journal of the Linnean Society, vol. 181, no. 1, pp. 1-20, 2016.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Karnan Arunasalam, Ragupathy Rengaswamy

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Submission of a manuscript implies: that the work described has not been published before that it is not under consideration for publication elsewhere; that if and when the manuscript is accepted for publication. Authors can retain copyright of their article with no restrictions. Also, author can post the final, peer-reviewed manuscript version (postprint) to any repository or website.

Since Oct. 01, 2015, PETI will publish new articles with Creative Commons Attribution Non-Commercial License, under The Creative Commons Attribution Non-Commercial 4.0 International (CC BY-NC 4.0) License.
The Creative Commons Attribution Non-Commercial (CC-BY-NC) License permits use, distribution and reproduction in any medium, provided the original work is properly cited and is not used for commercial purposes
