Nawah-VL — وصف الصور بالعربية

SigLIP2 vision tower grafted onto an Arabic language model. Pick a backbone, upload an image, get a Modern Standard Arabic caption.

Model
1 5
10 60

What to expect. It names the main subject and its colour reliably and stops cleanly. It reads text inside images poorly and will occasionally state something confidently wrong — the language model is 25M/50M parameters, so coarse-but-correct is the ceiling.

chrF++ perplexity grounding gap
25M 27.4 12.2 2.09
50M 29.0 8.6 2.36
جرب صورة من دول
الصورة Model