Making eight centuries of Arabic scientific manuscripts computationally
accessible — beginning with al-Bīrūnī's al-Qānūn al-Mas'ūdī.
PhD in AI · IRIT / IMT / CLLEUniversity of Toulouse · 2025–2028
The Arabic scientific tradition carries eight centuries of knowledge in astronomy, algebra,
and medicine. My PhD makes these manuscripts computationally accessible, beginning with
al-Bīrūnī's al-Qānūn al-Mas'ūdī (written c. 1030) as a first corpus.
The problem space
What makes this hard
Calligraphic diversity
Six copies, six hands. Each scribe used a distinct style with elongation strokes and unique ligatures.
Scientific vocabulary
Astronomical and mathematical terms absent from any modern Arabic training corpus.
Eight centuries of copying
The same text survives in copies from the 12th to the 19th century. Models must generalise across all of them.
Document wear
Water stains, faded ink, marginal annotations and physical damage introduce significant noise.
System design
The pipeline
01
Manuscript imageRaw scan input
→
02
HTR modelVision-language
→
03
TranscriptionArabic text output
→
04
Medieval LLMSearch & translation
Methodology
My approach
ATHAAR is the first manually annotated benchmark for medieval Arabic scientific manuscript recognition. It compiles 1,031 lines from six witnesses of al-Bīrūnī's al-Qānūn, spanning the 12th to 19th century, held at institutions including the Bibliothèque nationale de France and the British Library.
Once text is recognised, the goal is language models pre-trained on medieval Arabic scientific prose. Standard Arabic LLMs struggle with archaic vocabulary and domain-specific terminology. A dedicated model enables accurate translation, semantic search, and scholarly indexing.
Beyond models, the goal is a usable tool: historians submit a manuscript image and get a searchable, correctable transcription. Accessibility for non-technical scholars is a first-class requirement.