THE SINDHI TOKEN TAX: ORTHOGRAPHIC PENALTY AND TOKENIZATION INEQUITY BETWEEN ENGLISH AND SINDHI IN LARGE LANGUAGE MODELS
DOI:
https://doi.org/10.59075/jsrd.v7i8.547Keywords:
computational linguistics; tokenization; large language models; Sindhi; English; token premium; Orthographic Penalty Index; language equity; PakistanAbstract
Large language models (LLMs) do not read words; they read tokens, and the tokenizer decides how many tokens a language needs to say the same thing. This paper introduces the notion of a Sindhi token tax and tests it empirically for Pakistan’s English–Sindhi context. Using 1,004 parallel sentences from the FLORES-200 benchmark in English, Sindhi and Urdu, the study measured five widely deployed tokenizers (GPT-2, Llama 2, Llama 3, Qwen2 and Command-R) with a reproducible Python pipeline. A convergent mixed-methods design combined quantitative indices with a rule-assisted qualitative analysis of how frequent Sindhi words are segmented. Quantitatively, Sindhi required 2.87 to 4.98 times as many tokens as English for the same content in total (Wilcoxon signed-rank, r = .87, p < .001), and a 4,096-token context window held about 770 to 1,335 Sindhi words against 2,890 to 3,330 English words. A new Orthographic Penalty Index showed that the 18 letters unique to the Sindhi alphabet, such as ڪ, ٿ, ڻ and ڳ, were split into byte fragments in 100% of their occurrences in four of the five tokenizers, while letters shared with Arabic and Urdu were almost never split. Qualitatively, five patterns emerged, including the fragmentation of core grammatical words, the “Urdu shortcut” by which Urdu code points are cheaper than Sindhi ones, and the invisibility of Sindhi morphology to subword boundaries. The study also found Unicode inconsistency in digital Sindhi, with 238 of 1,004 benchmark sentences mixing forms of the letter heh. The paper proposes the Script-Aware Tokenization Equity Framework (SATEF) and argues that script-level design, not only data scarcity, makes AI costlier and weaker for Sindhi speakers, with implications for language policy, education and technology in Pakistan.
Downloads
Details
-
Abstract Views: 42
PDF Downloads: 13
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 JOURNAL OF SOCIAL RESEARCH DEVELOPMENT

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

