THE SINDHI TOKEN TAX: ORTHOGRAPHIC PENALTY AND TOKENIZATION INEQUITY BETWEEN ENGLISH AND SINDHI IN LARGE LANGUAGE MODELS

Authors

  • Dr. Ali Siddiqui English, Hyderabad Institute for Technology and Management Sciences (HITMS), Hyderabad, Sindh, Pakistan.
  • Dr. Shauban Ali Solangi Department of Computer Science & Related Studies, Hyderabad Institute for Technology and Management Sciences (HITMS), Hyderabad, Sindh, Pakistan.
  • Uzma Ehsan Senior Lecturer, NUST-Military College of Signals.

DOI:

https://doi.org/10.59075/jsrd.v7i8.547

Keywords:

computational linguistics; tokenization; large language models; Sindhi; English; token premium; Orthographic Penalty Index; language equity; Pakistan

Abstract

Large language models (LLMs) do not read words; they read tokens, and the tokenizer decides how many tokens a language needs to say the same thing. This paper introduces the notion of a Sindhi token tax and tests it empirically for Pakistan’s English–Sindhi context. Using 1,004 parallel sentences from the FLORES-200 benchmark in English, Sindhi and Urdu, the study measured five widely deployed tokenizers (GPT-2, Llama 2, Llama 3, Qwen2 and Command-R) with a reproducible Python pipeline. A convergent mixed-methods design combined quantitative indices with a rule-assisted qualitative analysis of how frequent Sindhi words are segmented. Quantitatively, Sindhi required 2.87 to 4.98 times as many tokens as English for the same content in total (Wilcoxon signed-rank, r = .87, p < .001), and a 4,096-token context window held about 770 to 1,335 Sindhi words against 2,890 to 3,330 English words. A new Orthographic Penalty Index showed that the 18 letters unique to the Sindhi alphabet, such as ڪ, ٿ, ڻ and ڳ, were split into byte fragments in 100% of their occurrences in four of the five tokenizers, while letters shared with Arabic and Urdu were almost never split. Qualitatively, five patterns emerged, including the fragmentation of core grammatical words, the “Urdu shortcut” by which Urdu code points are cheaper than Sindhi ones, and the invisibility of Sindhi morphology to subword boundaries. The study also found Unicode inconsistency in digital Sindhi, with 238 of 1,004 benchmark sentences mixing forms of the letter heh. The paper proposes the Script-Aware Tokenization Equity Framework (SATEF) and argues that script-level design, not only data scarcity, makes AI costlier and weaker for Sindhi speakers, with implications for language policy, education and technology in Pakistan.

Downloads

Details

    Abstract Views: 42
    PDF Downloads: 13

Published

01-10-2026

How to Cite

Dr. Ali Siddiqui, Dr. Shauban Ali Solangi, & Uzma Ehsan. (2026). THE SINDHI TOKEN TAX: ORTHOGRAPHIC PENALTY AND TOKENIZATION INEQUITY BETWEEN ENGLISH AND SINDHI IN LARGE LANGUAGE MODELS. JOURNAL OF SOCIAL RESEARCH DEVELOPMENT, 7(8), 01–26. https://doi.org/10.59075/jsrd.v7i8.547