Cross-Species Generalization and Comparative Performance Analysis of Deep Neural Network Architectures in Histological Image Classification
Tarih
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
Erişim Hakkı
Özet
Histological image classification plays a critical role in biomedical research and diagnostic processes. Advances in the field of deep learning present significant opportunities for enhancing diagnostic accuracy and developing automated decision support systems. This study aims to comparatively evaluate the out-of-distribution generalization and cross-domain classification performance of different deep neural network encoders. In this study, models were trained on an internal dataset of 4307 hematoxylin and eosin (H&E) stained images of male rat lung, cerebellum, and adipose tissues, obtained under ethical committee approval (Ege University, HADYEK 2026-06). To evaluate genuine generalization, testing was conducted on a separate, diverse external cohort of 600 mixed human-and-animal images. Nine different deep learning architectures-MobileNetV3-S, MobileNetV3-L, DenseNet121, DenseNet201, ConvNeXt-S, ConvNeXt-B, DINOv3 ViT-S, DINOv3 ViT-H+, and the pathology foundation model UNI2-h-were evaluated using frozen feature extraction combined with a linear probe. For performance evaluation, accuracy, sensitivity, specificity, F1 score, Cohen's κ, and ROC-AUC metrics were analyzed along with per-image inference times. While all models achieved near-perfect results during internal cross-validation, their performance diverged significantly on the external dataset, confirming that the observed differences reflect genuine cross-domain generalization capabilities rather than under-fitting. All competitive models demonstrated high performance in classifying adipose and cerebellum tissues, achieving an F1 score of 95% or higher for these tissues. In distinguishing the lung tissue, which has the most diverse structure and is the most difficult to classify, the UNI2-h model emerged as the most successful, achieving an F1 score of 97.4% and a recall of 95.0%. When evaluated in terms of computational efficiency, the MobileNetV3-Small model stood out as having the lowest processing time among all scenarios, demonstrating an inferenc











