PhD Research · 2025–2028

Research

Making eight centuries of Arabic scientific manuscripts computationally accessible — beginning with al-Bīrūnī's al-Qānūn al-Mas'ūdī.

PhD in AI · IRIT / IMT / CLLE University of Toulouse · 2025–2028

The Arabic scientific tradition carries eight centuries of knowledge in astronomy, algebra, and medicine. My PhD makes these manuscripts computationally accessible, beginning with al-Bīrūnī's al-Qānūn al-Mas'ūdī (written c. 1030) as a first corpus.

What makes this hard

Calligraphic diversity

Six copies, six hands. Each scribe used a distinct style with elongation strokes and unique ligatures.

Scientific vocabulary

Astronomical and mathematical terms absent from any modern Arabic training corpus.

Eight centuries of copying

The same text survives in copies from the 12th to the 19th century. Models must generalise across all of them.

Document wear

Water stains, faded ink, marginal annotations and physical damage introduce significant noise.

The pipeline

01
Manuscript image Raw scan input
02
HTR model Vision-language
03
Transcription Arabic text output
04
Medieval LLM Search & translation

My approach

ATHAAR is the first manually annotated benchmark for medieval Arabic scientific manuscript recognition. It compiles 1,031 lines from six witnesses of al-Bīrūnī's al-Qānūn, spanning the 12th to 19th century, held at institutions including the Bibliothèque nationale de France and the British Library.
Once text is recognised, the goal is language models pre-trained on medieval Arabic scientific prose. Standard Arabic LLMs struggle with archaic vocabulary and domain-specific terminology. A dedicated model enables accurate translation, semantic search, and scholarly indexing.
Beyond models, the goal is a usable tool: historians submit a manuscript image and get a searchable, correctable transcription. Accessibility for non-technical scholars is a first-class requirement.
IRIT IMT CLLE University of Toulouse · 2025–2028