Transfer Learning from ImageNet-Pretrained Vision Models for Arabic Text Classification Using Glyph-Based Image Representations
DOI:
https://doi.org/10.54361/ajmas.2584134Keywords:
Arabic Text Classification, Transfer Learning, Glyph-Based Representation, EfficientNet-B3, XceptionAbstract
Arabic text classification is challenging because of rich morphology, orthographic variation, and complex tokenization. This study investigates an alternative approach that represents Arabic documents as rendered images and classifies them using ImageNet-pretrained vision models. We developed a rendering pipeline to preserve contextual letter shaping and right-to-left layout; we also divided long documents into image tiles and classified them by averaging tile-level probabilities. EfficientNet-B3 and Xception were fine-tuned using the same two-stage training protocol without automated hyperparameter optimization. Experiments were conducted on 6,000 balanced Modern Standard Arabic news articles across six categories. EfficientNet-B3 achieved 98.77% test accuracy, compared with 98.34% for Xception, while requiring approximately half the parameters and one-fifth of the floating-point operations. The approach eliminates tokenization, vocabulary construction, and manual feature engineering while achieving competitive classification performance. These results demonstrate the potential of rendered Arabic text as a compact and effective representation for visual transfer learning in Arabic text classification.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Hosam Alzawam, Tarik Idbeaa

This work is licensed under a Creative Commons Attribution 4.0 International License.











