An Alraqeem AI project

From the image of a page to structured knowledge.

A tool that uses computer vision and AI to understand the structure of Arabic documents—their text, tables, and images—then produce formats that can be reviewed, edited, and processed.

Project statusLimited prototype
The role of AIAssistance, not human replacement

Recognizing characters is not enough for many Arabic books and documents. Meaning depends on title placement, column order, tables, images, and notes. Much of that structure is lost when a page becomes flat text.

Warraq combines page vision, text extraction, and structural analysis, producing formats such as HTML, Word, and JSON. Its pipeline is layered so output can be inspected at each stage rather than hiding errors inside one opaque process.

Vision and language models help classify page regions, read content, and recover relationships between parts. Complex documents still require layered verification and human review; we make no claim of perfect accuracy.

  • Preserve the source image and link it to output for verification.
  • Independent checks for text, tables, images, and reading order.
  • Surface uncertainty instead of presenting it as correct.
  • No publication or attribution to a source before appropriate human review.

Expand the Arabic evaluation set, measure quality by page type and output format, and build a review interface that makes corrections faster and clearer.

Share an idea that can serve Islam and Muslims

Share an idea that can serve Islam and Muslims