مشباح
MYSHBAH  · AI
Bridging Arabic Heritage & Artificial Intelligence

Classical Arabic heritage — sourced, scholar-verified, and delivered AI-ready. Pre-Islamic and early Islamic content, structured and authenticated by named Arab scholars, for the sovereign Arabic AI programs building the next generation.

Request a Scoping Conversation How It Works
~0.5% Arabic on the Web · Fanar 2.0
~18% Native Arabic Tokens · Jais
~41% Arabic Tokens · Fanar 1.0
13% Arabic Books · ALLaM
Undocumented Classical Heritage · Every Published Report
The Problem

Arabic AI has never seen the language at its roots

The pre-Islamic Jahiliyya era produced the foundational texts of the Arabic language — the Mu'allaqat, the tribal histories, the poetry that preserved an entire civilization's memory for over fourteen centuries. Arab scholars have verified, annotated, and taught this heritage for generations.


Arabic AI encounters it only in fragments. Scattered retellings and isolated verses circulate online. What does not exist is the verified layer — complete, attributable, scholar-checked, and structured for machines. MYSHBAH.AI was built to build that bridge, from the shelf to the system.

"The question is not whether AI will mediate the world's relationship with Arabic heritage. It already does. The question is whether that mediation will be grounded in verified scholarly knowledge — or in Wikipedia and the contents of the internet."

Targeted Demonstration — Five AI Systems · Four Questions

Five AI systems tested in Arabic on the War of al-Basus — no prior context — including two purpose-built Arabic LLMs — 2026

Claude Could not identify the camel's owner. Could not reproduce the Verses of Death. Admitted ignorance — the content does not exist in its training data.
Copilot Misattributed the camel's owner. Declined to fabricate the verses — correctly noting the sources preserve only fragments — and quoted one genuine transmitted line.
Gemini Misattributed the camel. Disclosed its source: Wikipedia and Facebook. Illustrated historical figures with actors from a modern television drama.
Fanar 2.0 ★ Misattributed the camel and confused the woman's name with the camel's name. Inverted the war's causation entirely.
Jais v2 ★ Misattributed the camel. Fabricated martial verses in the wrong genre. Contradicted itself on the most basic relationship in the story within a single response.
A conflict taught in Arab school curricula. No system tested identified the camel's owner correctly, and responses ranged from appropriate refusal to confident fabrication.
★ Purpose-built Arabic-centric LLM
Tested via publicly available web interfaces in Arabic, with no prior context. Decoding parameters are not user-exposed in these interfaces; a controlled API-based evaluation with fixed parameters and repeated trials is planned next.

Verified datasets — civilizational memory, machine-readable

A verified dataset is not a technology product. It is a scholarly judgement, made permanent. When a professor of classical Arabic confirms that a root attribution is correct, that confirmation becomes part of what Arabic AI will learn. The scholar's name, institution, and published expertise are embedded in every licensing agreement.


The mishkah (المشكاة) holds the light. The misbah (المصباح) is the light. The scholar is not a service provider in this framework — the scholar is the product.


"The dataset does not exist without the scholar."

Step 01
Verified Arab Academic Sources
Content drawn from Arab academic critical editions, research center publications, and university faculty research and doctoral theses across the Arab world — provided each source carries verifiable academic attribution. No machine-translated content.
Step 02
Automated First-Pass Annotation
Each text is computationally segmented and morphologically tagged. Where classical vocabulary falls outside standard Arabic NLP tools — 86% of tokens in our first corpus — analysis is completed manually against the classical dictionaries and logged as engineering-level work.
Step 03
Scholar Verification
Morphology, cultural context, and literary structure reviewed by named experts with published academic credentials. Only a scholar can mark a record verified — that boundary is enforced in the schema itself.
Step 04
Delivery — Base Format & Pre-embedded
Standard structured data format for any AI pipeline, or pre-embedded for immediate deployment in Azure AI Search and other vector databases.

One text. Three dimensions of verified knowledge.

Full verification covers three distinct dimensions. Each is a separate scholarly competence, and at scale each is reviewed by a named expert with published credentials in that specific area.

01
Morphology & Diacritics
Classical Arabic morphology and pre-Islamic prosody. Root attribution, diacritical rendering, grammatical tagging, and disambiguation of ambiguous classical forms.
02
Cultural & Tribal Context
Pre-Islamic Arabian history and Jahiliyya social structures. Tribal references, honor codes, historical allusions, and the social meaning of specific lexical choices.
03
Literary Structure
Classical Arabic poetry and the Mu'allaqat canon. Meter, rhyme, intertextual references, thematic classification, and verse-level literary significance.
Research Agenda

The gap is not just in the data — it is in how we measure what is missing

01
The Benchmark Problem
Existing Arabic LLM benchmarks — including ArabicMMLU and ACVA — measure factual recall, mathematical reasoning, and dialectal comprehension. No shared standard exists for civilizational awareness: the capacity to represent the pre-Islamic and classical foundations upon which the Arabic language was built. These benchmarks represent real progress. They do not measure the foundational layer. Recent benchmarks have begun measuring poetry comprehension. None measures whether a model can reproduce a canonical text accurately, or distinguish an authentic verse from a fabricated one.
02
The Digitization Imperative
Decades of verified classical Arabic scholarship sit in university archives across the Arab world, largely unstructured for AI training pipelines. Arabic represents approximately 0.5% of indexed web content despite 400 million native speakers. The material exists. The pipeline does not. Nationally sponsored digitization programs, developed in partnership between Arab universities, language research centers, and national archives, represent the most direct path forward.
03
The Partnership Call
A properly scaled benchmark for classical Arabic heritage knowledge requires hundreds of questions spanning multiple episodes, poetic genres, and historical periods — designed through coordinated collaboration between classical Arabic scholars, Arabic NLP researchers, and AI evaluation teams. MYSHBAH.AI is building the data infrastructure that makes this possible. We call on Arab universities, language research centers, and Arabic LLM developers to build it together.
Built For

Three pathways to Arabic heritage intelligence

نموذج
Sovereign Arabic LLM Programs
Training and deployment rights for national Arabic AI models. Verified classical Arabic heritage content that measurably improves cultural and historical knowledge performance.
بحث
Retrieval & Enterprise Deployment
Annual deployment licenses for universities, ministries, and organizations that need a verified Arabic heritage knowledge base without training their own model. LLM-agnostic — indexed search or RAG.
تراث
Heritage Institutions & Archives
Archives, libraries, and universities holding undigitized or uncommercialized heritage. We source, verify, annotate, and structure it. Scholar attribution throughout, and a revenue share if the institution chooses to license onward.

'Antara ibn Shaddad's Mu'allaqa — the opening commission

The pilot corpus is 'Antara ibn Shaddad's Mu'allaqa — one of the seven canonical pre-Islamic odes, among the most studied texts in Arabic literary scholarship, and the centerpiece of Arabic literature curricula across the Arab world.


This is not a test of whether the content matters. Every Arabic AI that encounters a question about Antara, pre-Islamic Arabian tribal culture, or the Mu'allaqat tradition will demonstrate the gap immediately. The pilot demonstrates that MYSHBAH.AI can close it.


Corpus v1 is built. 1,293 verified tokens — 698 across the 75 verses of the Mu'allaqa itself, plus 595 tokens of contextual verses drawn from the classical narrative sources. Root, part of speech, gender, number, diacritics, and classical context on every token. Zero empty required fields, zero QA errors on structural audit, produced under a documented six-phase workflow.

"The difference between an AI that has read the poem and an AI that has studied under its greatest living interpreter."

What the Pilot Delivers

  • Verified structured corpus of the full Mu'allaqa — morphologically annotated, culturally contextualized, scholar-authenticated
  • Pre-embedded version compatible with Azure AI Search and major vector databases for immediate deployment
  • Academic verification certificate with named scholars, institutional affiliations, and methodology documentation
  • Live before/after demonstration: model responses with and without the verified corpus
  • Full schema documentation, ready to scale to the Seven Mu'allaqat and subject-specific datasets
  • Corrections register capturing every manual analysis — a compounding asset across corpora
  • A clear pathway to the Seven Mu'allaqat corpus and extended subject datasets — available to scope upon completion of the pilot
Scope
Scoped to your program's requirements — corpus size, annotation depth, and deployment format.
Request a Scoping Conversation
Academic Partnership · Currently in Formation

For scholars of Arabic heritage — an invitation

There was a time when Arabic was to the world what English is today. From the 8th to the 13th century, to be a scholar anywhere between the Atlantic and the Indian Ocean was to read and write in Arabic. Persian scientists, Turkish administrators, and Jewish and Christian scholars in Andalusia all conducted their intellectual lives in Arabic — because Arabic was where human knowledge lived.


That civilization did not begin with Islam. It began before it — in the poetry and oral histories of the Jahiliyya era. The scholars who have spent careers in this heritage are the only ones who can ensure that Arabic AI learns it correctly.

We invite scholars to:

  • Verify a defined classical Arabic heritage corpus — with permanent attribution on every licensing agreement
  • Contribute as a named co-author on the benchmark evaluation study — the next stage of MYSHBAH.AI's research program on classical Arabic heritage knowledge in AI
  • Open archive materials — manuscripts, unpublished annotated editions, rare texts — for structured digital preservation

The Arabic Heritage Data Report

MYSHBAH.AI has produced a cross-model evaluation of the pre-Islamic Arabic heritage data gap in large language models — five AI systems including two purpose-built Arabic LLMs, tested on foundational pre-Islamic knowledge. The evaluation was submitted to ArabicNLP 2026 (ACL) and examined in detail by three independent reviewers. Academic partnerships with Arab universities, language research centers, and NLP institutions are currently in formation. Scholars joining now will be primary named authors on the benchmark study that follows.

Academic Partnership Enquiry

Request a Pilot, Demo, or Partnership

مشباح
MYSHBAH  · AI
The mishkah holds the light. We are building the mishkah.
Website
Email
LinkedIn
Response Time
We respond to all inquiries within 48 hours.