The Middle East and Europe, 1097–1517 · Waseda University, Tokyo · 9–10 December 2023

Ṣundūq al-ġarāʾib fī-l-ʿaṣr al-raqāʾim

Reading Medieval Texts in the Digital Age

Maxim Romanov

December 9, 2023 · The Evolution of Islamic Societies (c. 600–1600 CE), Universität Hamburg

Mountain of research…

Mountain of research

Mountain of research

Mountain of research

  • Ignaz Goldziher (d. 1921): only about 1%
  • Vasily/Wilhelm Bartold (d. 1930): about 2%
  • Ignatii Krachkovskii (d. 1951): about 4%
  • Joseph Schacht (d. 1969): about 8,5%
  • [in 1973: 10% threshold]
  • Stanislav M. Prozorov (defended: 1967): about 8%
  • Alexander D. Knysh (defended: 1986): about 20% (2,5 times more)
  • My cohort (defended: ±2013): about 80% (10 times more and 4 times more, respectively)

Mountain of research

  • since 2001 (less than 20 years) the number of publications doubled
  • Even if we consider the most conservative growth rate, there may be twice as many publications about the Islamic world—1,2 billion!—by the year 2040.

Mountain of primary sources as well…

Open Islamicate Texts Initiative [OpenITI] as a proxy to the Arabic Written Tradition

OpenITI as a proxy to AWT

OpenITI Corpus (EIS1600 Subcorpus)

OpenITI as a proxy to AWT

OpenITI as a proxy to AWT

OpenITI as a proxy to AWT

  • * 889 mln tokens: is roughly comparable to an English corpus of between 0.99 and 1.07 billion tokens.

OpenITI as a proxy to AWT

OpenITI as a proxy to AWT

OpenITI as a proxy to AWT

OpenITI as a proxy to AWT

OpenITI as a proxy to AWT

OpenITI as a proxy to AWT

Solution(s)

Memex, Thinking Machine, Zettelkasten

Zettelkasten: Niklas Luhmann (1927-1998)

Networks of Knowledge

Mem[ex|ex]periment, MasterChronicle, EIS1600

Memex: a quick demo

( http://localhost:1506/_web_pages/_html/details.html?pub_id=MallettMedieval2015?page=0096 )

MasterChronicle Model

Mem[ex|ex]periment, MasterChronicle, EIS1600

EIS1600 Status

Mem[ex|ex]periment, MasterChronicle, EIS1600

EIS1600 :: DFG Emmy Noether Project (2021-2027)

  • The Emmy-Noether Project (#445975300) undertakes an innovative study of “The Evolution of Islamic Societies (c. 600-1600 CE)” [EIS1600] through the computational analysis of a large number of historical and biographical texts, which are treated holistically as a unified corpus of historical information. EIS1600 employs a series of advanced computational methods of text analysis and data modeling, which are the key to discovering, evaluating, and modeling all relevant textual evidence at an unprecedented scale.
  • The EIS1600 team works on identifying and analyzing long-term historical trends through three closely connected research areas. The first area focuses on social factors—major ethnic, religious, and professional groups and how they shaped the development of local communities and fused them into what we call the Islamic world. The second one focuses on various economic factors that had an effect on local communities. The third one traces patterns of environmental factors—plagues, famines, droughts, pest infestations, earthquakes, and climate change—and their effect on the life of local communities. Complementing and informing each other, these case studies will be the foundation for the PI’s robust synthesis of the evolution of the Islamic world over the period under study.

EIS1600 :: DFG Emmy Noether Project (2021-2027)

  • Team:
  • Hamid Reza Hakimi, PhD Researcher (Arabic Studies)
  • Data annotation
  • Economic Data Analysis
  • Lisa Mischer, PhD Researcher (CS / Arabic Studies)
  • Advanced Coding (Annotation Routines; Data Analysis)
  • Environmental Data Analysis
  • Tariq Yousef, PhD Researcher (CS, until Sept 2023)
  • Deep Learning Tasks
  • Alicia González Martínez, PostDoc Researcher (CLinguistics / CS, from Oct 2023)
  • Machine Learning; Development;
  • Maxim Romanov, Research Group Leader
  • Data Annotation; Advanced Coding; Data Analysis;
  • Social History Data Analysis

EIS1600: Current Progress

  • OpenITI Corpus and Texts Preparation
  • OpenITI Corpus > EIS1600 Subcorpus;
  • MIUs, and EIS1600 mARkdown Tagging Scheme;
  • Automated Processing: Active Learning Cycle
  • Rule-based annotation; Deep-Learning / Machine-Learning - driven annotation;
  • Data for research experiments;
  • MasterChronicle Concept;
  • Biographical Data:
  • Extraction of biographical information;
  • Identification of duplicate biographies;
  • Geographical Data:
  • NER Model to identify all toponyms;
  • Operational Gazetteer;
  • DL-Based Identification of Places via Proxies;
  • Historical Data:
  • Punctuation and segmentation Model;

OpenITI Corpus and EIS1600 Subcorpus

  • EIS1600 subcorpus of OpenITI includes historical, biographical, and geographical texts (around 1,500 texts of potential relevance).
  • *Date statements are a proxy to the number/volume of biographies.
biographical records date statements* tokens
top 100 texts c. 477 thousand (477,880) 331 thousand (331,370) 71 million (70,972,858)
top 343 texts c. 611 thousand (610,966) 416 thousand (416,597) 112 million (111,537,708)
… … … …
top 1,500 texts (?) c. 707 thousand (707,369) 426 thousand (426,614) 250 million (250,595,377)

Approach to Textual Sources: Minimal Information Units

  • the main focus is on Minimal Information Units, e.g., descriptions of events from chronicles, and biographical records from biographical collections:
  • premise: narratives are consciously constructed, while myriads of details scattered across vast texts simply cannot be subjected to the same level of agenda-driven editing; in large quantities, such agenda-resistant data is likely to provide more reliable historical evidence
  • MUI will be aggregated into clusters of related historical information, which will be consequently studied
  • Each primary source is broken into such units and they all to be re-aggregated into MasterChronicle, which will serve as a unified research ecosystem for the project and its collaborators

Conceptual Data Processing Pipeline

Conceptual Research Pipeline

Example of Pattern Tracing

Linked Open Data > Linked Local Data

  • URIs: 0748Dhahabi.TarikhIslam.MGR20180917-ara1.785357193186.EIS1600
  • Author URI: 0748Dhahabi
  • Book URI: 0748Dhahabi.TarikhIslam
  • Edition URI: 0748Dhahabi.TarikhIslam.MGR20180917-ara1
  • MIU URI: 0748Dhahabi.TarikhIslam.MGR20180917-ara1.785357193186

Linked Open Data > Linked Local Data

  • Biographical Entities:
  • Onomastics + an Onomastic Gazetteer
  • Toponyms + Geographical Gazetteer (https://althurayya.github.io/
  • Dates, Thematic Keywords, Misc Entities
  • Persons + linked back to Biographies (future)
  • Historical Events Entities:
  • Subject - Predicate - Object
  • Toponyms + Geographical Gazetteer (https://althurayya.github.io/
  • Dates, Thematic Keywords, Misc Entities
  • Persons + linked back to Biographies (future)

EIS1600 mARkdown (OpenITI mARkdown Flavor) / dynamic /

EIS1600 mARkdown (OpenITI mARkdown Flavor) / static /

Annotation Cycle: Active Learning

Data Annotation

  • Persons, toponyms, miscellaneous (NER)
  • Part of Speech (POS) and Lemmas (Lemmatization)
  • Dates (RBC)
  • String expression is parsed to numerical value
  • Onomastic section (TC)
  • Onomastic information (TC)
  • Relationship classification between biographee and toponyms (TC), persons (TC), and dates (RBC) — work in progress

  • RBC: Rule-Based Classification

  • TC: Token Classification — Fine-Tuned CamelBERT Models
  • NER: Named Entity Recognition — Fine-Tuned CamelBERT Models

Pre-trained Arabic NER Models Benchmarking

  • Data: 1,125 texts (~8,000 entities)
  • Initial annotation was done with pre-trained CamelBERT ca
  • 3 Labels
  • Train / Test split: 80 / 20 %
CamelBert msa CamelBert ca CamelBert mix
LOC Precision 94.37% 97.48% 95.05%
Recall 96.06% 97.13% 96.42%
F1 95.20% 97.31% 95.73%
MISC Precision 76.40% 79.78% 76.40%
Recall 83.95% 87.65% 83.95%
F1 80.00% 83.53% 80.00%
PERS Precision 94.39% 95.92% 94.84%
Recall 95.58% 96.61% 95.82%
F1 94.98% 96.26% 95.33%
OVERALL Precision 93.42% 95.31% 93.89%
Recall 95.08% 96.25% 95.33%
F1 94.24% 95.78% 94.60%
Accuracy 98.62% 98.96% 98.65%

Onomastic Section Detection

  • Detect the boundaries of the onomastic section in the biography
  • Data: 2,242 biographies
  • 1,193 manually annotated
  • 1,059 automatically annotated and manually corrected (ONOM section only)
  • Model’s accuracy: 88.32%
  • Most inaccuracies, however, are extremely minor (a single token difference or a punctuation sign), which makes the usability of this model significantly higher than the given number of 88.32%; in fact, all results are perfectly usable, even with an extra token or two that should not be included.

Onomastic Information Classification

  • Data: 1,047 texts
  • Rule-based annotated and manually corrected
  • 6 Labels
  • Train / Test split: 80 / 20 %
ISM KUN NSB
Precision Recall F1 Precision Recall F1 Precision Recall F1
97.59% 96.81% 97.20% 96.88% 94.90% 95.88% 95.78% 96.99% 96.38%
97.98% 96.81% 97.39% 97.12% 94.39% 95.73% 95.74% 96.22% 95.98%
NAS LQB SHR
Precision Recall F1 Precision Recall F1 Precision Recall F1
97.88% 99.22% 98.55% 91.30% 77.78% 84.00% 87.50% 77.78% 82.35%
97.24% 97.42% 97.33% 78.26% 75.00% 76.60% 68.57% 85.71% 76.19%
Overall
Precision Recall F1
96.65% 96.90% 96.77%
95.85% 96.09% 95.97%

Annotated Text

Data Extraction

  • Based on added tags from the annotation pipeline
  • Extracted entities are added as metadata
  • Toponyms:
  • Identification (Regex matching)
  • Classification into settlement or province
  • For settlements: adding corresponding province
  • Calculating edges between settlements and between provinces
  • Dates:
  • Birth and death dates as boundaries for active period
  • If no birth or death date: approximation for active period

Ongoing Work

Biographical Data

  • Conflating biographies into individuals
  • Creating training data (pre-generated, using different text comparison algorithms);
  • Revising results: same, not-same;
  • Using resulting training data to train a model for classifications
  • Splitting model (nearest future):
  • Model to split untagged texts into biographies automatically

Biographical Data: Same-Person Annotation Tool

Biographical Data: Same-Person Identification Model

  • Data: 3.305 biography pairs
  • 2.035 Not the same Person, 1.270 The same person
  • 8 Features: tfidf, tfidf_bigrams, dice3gram, dice4gram, dice5gram, length_rate, min_length, ent_overlap
  • * Dice: Sørensen-Dice Coefficient; ent_overlap: Entity Overlap; RBF: Radial Basis Function; SVM: Support Vector Machine; NN: Nearest Neighbour
0 (348) 1 (234) Accuracy
Precision Recall F1-score Precision Recall F1-score
Logistic Regression 0.91 0.97 0.94 0.95 0.86 0.90 0.9261
SVM Linear 0.85 0.97 0.9 0.94 0.75 0.83 0.8780
SVM RBF 0.91 0.93 0.92 0.89 0.85 0.87 0.9003
Random Forest 0.96 0.98 0.97 0.97 0.94 0.96 0.9673
NN 0.93 0.97 0.95 0.95 0.89 0.92 0.9347
Gradient Boosting 0.94 0.97 0.96 0.96 0.91 0.93 0.9485

Biographical Data: Onomastics Model

  • All onomastic elements, mainly nisbaŧs are to be categorized
  • This will allow to identify larger social, religious, tribal, etc. groups;
NISBAŦ CATEGORY METACATEGORY
mālikī Fiqh Religious
ḥanafī Fiqh Religious
šāfiʿī Fiqh Religious
ḥanbalī Fiqh Religious
muqriʾ Qurʾān Religious
qāriʾ Qurʾān Religious
mufassir Qurʾān Religious
rāwī Ḥadīṯ Religious
muḥaddiṯ Ḥadīṯ Religious
ḥāfiẓ Ḥadīṯ Religious

Geographical Data

  • 101 sources:
  • 477,880 thousand biographical records;
  • 854,394 toponyms were found with our NER model;
  • 235,466 MIUs contain toponyms (~49%; geographical nisbaŧs were not checked yet)
  • 45,386 Unique Toponyms (included prefixed variants)
  • 20,195 Unique Toponyms with MIU occurrences > 1
  • 13,084 Unique Toponyms with MIU occurrences > 2
  • 10,144 Unique Toponyms with MIU occurrences > 3
  • 8,397 Unique Toponyms with MIU occurrences > 4
  • 7,308 Unique Toponyms with MIU occurrences > 5

Geographical Data: Linking Toponyms

  • Gazetteer to identify toponyms
  • al-Thurayya as a starting point
  • Georgette Cornu’s Atlas du monde arabo-islamique à l'époque classique: IXe-Xe siècles
  • https://althurayya.github.io/

Geographical Data: Current Results

  • Althurayya is a starting point, but only about half of it is “operational”;
  • Together with newly identified toponyms, only about 2,100 occur in at least 10 biographical records;
  • This is still a lot for expanding the gazetteer manually…
  • Identifying the location and its coordinates is one of the most time-consuming tasks;

Geographical Data: Operational Gazetteer

  • Premise: rare toponyms are problematic to medieval authors as well; they often define them through some proxy:
  • For example, “Bashan is a village near Herat”;
  • for most cases, the mention of Herat will be most sufficient;
  • Plan:
  • We can train a model to recognize such descriptions in texts;
  • We can collect them from the entire corpus (or, at least, EIS1600 subcorpus);
  • Extracted information can be processed automatically; ChatGPT can convert the above example (in Arabic) into a structured data like:
  • Name of place: Bashan;
  • Type of place: village;
  • Location (Graviton): Herat;
  • Primary experiments are very promising
  • Other information, like distances can also be extracted in this manner;
  • Nisbaŧs can be automatically generated from toponyms (with some minor post-correction)

Geographical Data: Operational Gazetteer

  • Since collecting coordinates is an extremely time consuming task, we can limit it only for the most frequent places;
  • Other locations can be added as proxies to larger locations:
  • Village > its city;
  • Quarter > it city;
  • Market > its city;
  • etc.
  • The gazetteer will thus allow to preserve the original geographical information, yet allowing to effectively group it into meaningful regions;

Geographical Data: Modeling Regions - Voronoi Magnets

  • Althurayya offers divisions into provinces
  • however, they are relevant to 10-11th century;
  • Alternative:
  • Use most prominent urban centers as magnets and automatically associate other locations with them using Voronoi diagram:
    • Select magnets (highest frequencies);
    • Generate distances to those magnets > the closest one is the main magnet;
    • Approach parametrical:
    • Potentially, we can factor in other parameters:
      • Large bodies of water between place x and a magnet
      • Mountain ranges;
      • Deserts;
      • etc.

Geographical Data: Modeling Regions - Voronoi Magnets

Data: Prosopografía de los ulemas de al-Andalus. Disponible en [https://www.eea.csic/pua/](https://www.eea.csic/pua/)

Geographical Data: Modeling Regions - Voronoi Magnets

Data: Prosopografía de los ulemas de al-Andalus. Disponible en [https://www.eea.csic/pua/](https://www.eea.csic/pua/)

Geographical Data: Modeling Regions - Voronoi Magnets (380AH)

Data: Prosopografía de los ulemas de al-Andalus. Disponible en [https://www.eea.csic/pua/](https://www.eea.csic/pua/)

Geographical Data: Modeling Regions - Voronoi Magnets (570 AH)

Data: Prosopografía de los ulemas de al-Andalus. Disponible en [https://www.eea.csic/pua/](https://www.eea.csic/pua/)

Historical Data: Punctuation / Segmentation

  • Blobs of events in chronicles:
  • Punctuation model (BAM):
    • Baššār ʿAwād Maʿrūf’s editions as training data
    • Rather decent results;
  • Segmentation model (SEGM0):
    • some training data > first SEGM0 model to simplify things;
    • training data: blobs of text are automatically punctuated and split with BAM model into paragraph-like sections;
    • Human annotator rechecks all generated results;
  • Now: the second iteration of active learning cycle;
  • More training data has been recently prepared for further training;

Appendix Algorithmic Analysis Examples

[https://eis1600.github.io/presentations/TB_algorithmicAnalysis_HAexample.html#1](https://eis1600.github.io/presentations/TB_algorithmicAnalysis_HAexample.html#1)

Cultural Geography

Taʾrīḫ al-islām of al-Ḏahabī

Patterns: Regional Representations / Floruits (?)

Patterns: Regional Representations / Floruits (?)

Connections

Taʾrīḫ al-islām of al-Ḏahabī

Patterns: Overall Transregional Mobility

Patterns: Connections

Patterns: Connections

Patterns: Connections

Patterns: Connections

Patterns: Connections

Patterns: Connections

Patterns: Connections

Patterns: Connections in the 6th century AH

Patterns: Connections in the 6th century AH

Geography

Hadiyaŧ al-ʿārifīn of Ismāʿīl Bāšā al-Baġdādī

Patterns: Regions

Patterns: Regions

Patterns: Regions

Patterns: Regions

Patterns: Connections (700–1200 CE)

Patterns: Connections

Patterns: Connections

Patterns: Connections

Patterns: Connections (1700–1900)

Great Cities of Islam

Taʾrīḫ al-islām of al-Ḏahabī

Patterns: Great Cities of Islam - Early

Patterns: Great Cities of Islam - Early

Patterns: Great Cities of Islam - Early

Patterns: Great Cities of Islam - Early

Patterns: Great Cities of Islam - Middle

Mongols sack Baghdad in 656/1258

Patterns: Great Cities of Islam - Middle

Mongols sack Baghdad in 656/1258

Patterns: Great Cities of Islam - Middle

Mongols sack Baghdad in 656/1258

Patterns: Great Cities of Islam - Middle

Mongols sack Baghdad in 656/1258

Patterns: Great Cities of Islam - Late

Mongols sack Baghdad in 656/1258

Patterns: Great Cities of Islam - Late

Mongols sack Baghdad in 656/1258

[Miṣr / “Egypt” — also a reference to al-Qāhiraŧ / Cairo]

Patterns: Great Cities of Islam - Late

Mongols sack Baghdad in 656/1258

Societal Transformations

Taʾrīḫ al-islām of al-Ḏahabī

Patterns:

Patterns:

Patterns:

Patterns:

Patterns:

Patterns:

Patterns:

Patterns:

Patterns:

Patterns:

Patterns: Islamic Sciences

Patterns: Geography of Legal Schools

Patterns: Life Spans of Biographees

Patterns: People Vs Floruits - Traces of Demographics

  • According to archeological evidence (ceramics and coin finds) as well as evidence from narrative sources, al-Raqqa gains prominence during the early ‘Abbasid period, then goes into decline in the 4th Islamic century, then recovers during the Saljuq and later Zengid periods (Heidemann 2006). Based on al-Dhahabi’s “History”, the prosopographical graph reflects these developments.
EIS1600

Thank you!

1 / 138

Keys

→ space PgDnnext slide
← PgUpprevious slide
Home / Endfirst / last
Ffullscreen
Ooverview of all slides
Sspeaker notes in a second window
Ttype tester: fonts, weights, size
+ / − / 0text larger / smaller / reset
?this help

Print (⌘P) gives one slide per page. Add #12 to the address to open a slide.