A Vision-Transformer-based Ensemble Model for Multi-label Disease Classification on CT Scans and Chest Radiographs
DOI:
https://doi.org/10.5614/itbj.ict.res.appl.2026.20.2.2Keywords:
C-T scans, classification, deep learning, normalization, vision transformer, stacked ensembleAbstract
The increase in the volume of X-rays and computed tomography (CT) images has drastically increased the workload on radiologists. As a result, a computer-aided solution with the capability to classify CT scans is needed to reduce this workload. In this paper, a vision-transformer (ViT) based model for multi-disease, ensemble-based classification is proposed. ViT, Data2Vec, and SegFormer models were fine-tuned to carry out the classification of selected diseases, namely, effusion, pneumonia, and pneumothorax. Normal cases of the selected diseases were included in the dataset. The datasets were obtained from two sources: the chest X-ray dataset from the Nigerian Institute of Health Chest Clinic and the optical coherence tomography (OCT) images dataset containing 6,621 images from University of California, San Diego. The images were preprocessed using random cropping, horizontal flipping, and normalization. The dataset was partitioned into training, validation, and testing sets. Model training was done in 10 epochs. The evaluation metrics showed a better performance from ensemble learning compared to other individual transformer models. The weighted average performance for all metrics was 90.56% precision, 90.58% recall, and 90.48% F1 score. The model is useful for classification of multiple diseases and can be used by radiologists.
Downloads
References
Hassaan, M., Tayyaba, A., Ahmad, S A., Salman Z.A., Wajeeha, K. & Adnan, A., Deep Learning-based Classification of Chest Diseases using X-rays, CT Scans, and Cough Sound Images, Diagnostics, 13, 2772, 2023. https://doi.org/10.3390/diagnostics13172772
Liu, Z., Tong, T., Chen, L., Jiang, Z., Zhou, F. Zhang, Q., Zhang, X., Jin, Y. & Zhou, H., Deep Learning Based Brain Tumor Segmentation: A Survey, Complex & Intelligent Systems, 9, pp. 1001-1026, 2023. doi: https://doi.org/10.1007/s40747-022-00815-5.
Fagbuagun, O.A., Nwankwo, O., Samson S.A. & Folorunsho, O., Model Development for Pneumonia Detection from Chest Radiographs using Transfer Learning, Telkomnika (Telecommunication Computing Electronics and Control), 20(3), pp. 544-550, 2022. doi: 10.12928/TELKOMNIKA.v20i3.23296.
Huang, L., Ma, J., Yang, H., & Wang, Y., Research and Implementation of Multi-disease Diagnosis on Chest X-ray based on Vision Transformer, Quantitative Imaging in Medicine and Surgery, 14(3), pp. 2539-2555, 2024, https://qims.amegroups.org/article/view/122245.
Jack, V. & Tu, Advantagesa Disadvantages of using Artificial Neural Networks Versus Logistic Regression for Predicting Medical Outcomes, Journal of Clinical Epidemiology, 49(11), pp. 1225-1231, 1996. doi: 10.1016/S0895-4356(96)00002-9.
Teodoro, M.N., Fix P., Mart-Valdivia, M.T. Christine O.M. & Antonio L. Strengths, Weaknesses, Opportunities, and Threats Analysis of Artificial Intelligence and Machine Learning Applications in Radiology, Journal of the American College of Radiology, 16(9), Part B, pp. 1239-1247, 2019. doi: 10.1016/j.jacr.2019.05.047.
Sarker, I.H., Deep Learning: A Comprehensive Overview on Techniques, Taxonomy, Applications and Research Directions, SN Computer Science, 2, 420, 2021. https://doi.org/10.1007/s42979-021-00815-1.
Azad, R., Kazerouni, A, Heidari, M., Aghdam, E.K., Molaei, A., Jia, Y., Jose, A., Roy, R. & Merhof, D., Advances in Medical Image Analysis with Vision Transformers: A Comprehensive Review, Med. Image Anal, 91, 103000, 2024.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, arXiv 2020, arXiv:2010.11929.
Azhar, T., Application of Vision Transformer for Brain Stroke Classification based on CT Images, Eng. Technol. Appl. Sci. Res., 15(6), pp. 29435-29439, Dec. 2025.
Aboghanem, A., Abd Elfattah, M.M., Amer, H. & Tawkol, K.A.A., Hybrid Resnet50-vision Transformer Model with an Attention Mechanism for Aerial Image Classification, Sci Rep, 16, 5940, 2026. doi: 10.1038/s41598-026-36492-4
Garcia, A., Zhou, J., Pinero-Crespo, G., Beachkofsky, T. & Huang, X., Clinical Application of Vision Transformers for Melanoma Classification: A Multi-dataset Evaluation Study, Cancers (Basel), 17(21), 3447, 2025. doi: 10.3390/cancers17213447. PMID: 41228240; PMCID: PMC12607522.
Malik, H., Anees, T., Al-Shamaylehs, A.S., Alharthi, S.Z., Khalil, W. & Akhunzada, A., Deep Learning-based Classification of Chest Diseases using X-rays, CT Scans, and Cough Sound Images, Diagnostics (Basel), 13(17), 2772, 2023. doi: 10.3390/diagnostics13172772. PMID: 37685310; PMCID: PMC10486427.
Islam, S., Elmekki, H., Elsebai, A., Bentahar, J., Drawel, N., Rjoub, G. & Pedrycz, W., A Comprehensive Survey on Applications of Transformers for Deep Learning Tasks, arXiv:2306.07303, 2023.
Pu, Q., Xi, Z., Yin, S., Zhao, Z. & Zhao, L., Advantages of Transformer and its Application for Medical Image Segmentation: A Survey. BioMedical Engineering OnLine, 23, 14, 2024. .
Vaswani, A., Shazeer, N., Parmer, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L. & Polosukhin, I., Attention is All You Need. 31st Conference on Neural Information Processing System, Long Beach, CA, USA, 2023.
Singh, S., Kumar, M., Kumar, A., Verma, B. K., Abhishek, K., & Selvarajan, S., Efficient Pneumonia Detection using Vision Transformers on Chest X-rays, Scientific Reports, 14, 2487, 2024. doi: 10.1038/s41598-024-52703-2.
Chen, C., Gong, D., Wang, H., Li, Z. & Wong, K.Y.K., Learning Spatial Attention for Face Super-resolution, IEEE Trans. Image Process, 30, 1219?1231, 2020.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T. & Houlsby, N., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, arXiv preprint arXiv, 11929. 2010.
Behrendt, F., Bhattacharya, D., Kruger, J. Opfer, R. & Schlaefer, A., Data-efficient Vision Transformers for Multi-label Disease Classification on Chest Radiographs, Current Directions in Biomedical Engineering, 8(1), pp. 34-37, 2022. doi: 10.1515/cdbme-2022-0009
Huang, L., Ma, J., Yang, H. & Wang, Y., Research and Implementation of Multi-disease Diagnosis on Chest X-ray based on Vision Transformer, Quantitative Imaging in Medicine and Surgery, 14(3), pp. 2539-2555, March 15, 2024. https://qims.amegroups.org/article/view/122245.
Ashraf, S.M.N., Hasnat, A.M, & Alam, G.R., SynthEnsemble: A Fusion of CNN, Vision Transformer, and Hybrid Models for Multi-label Chest X-ray Classification, 26th International Conference on Computer and Information Technology (ICCIT), 2023.
Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M. & Summers, R., Hospital-scale Chest X-ray Database and Benchmarks on Weakly-supervised Classification and Localization of Common Thorax Diseases, in IEEE CVPR, 7, pp. 2097-2106, 2017.
Kermany, D.S., Goldbaum, M., Cai, W., Valentim, C.C., Liang, H., Baxter, S.L. & Zhang, K., Identifying Medical Diagnoses and Treatable Diseases by Image-based Deep Learning, Cell, 172(5), pp.1122-1131, 2018.


