Lost but Potentially Recoverable Sources for the Xiongnu/Hunnic Language Problem: A Survey of Recovery Programs

Abstract

The Xiongnu confederation of the eastern steppe (roughly third century BCE through second century CE) and the European Huns of the fourth and fifth centuries CE remain linguistically opaque despite occupying central positions in Eurasian political history. The total surviving direct linguistic corpus for both groups, taken together, fits comfortably on a single page: a few hundred personal names, a small set of titles, perhaps a dozen common nouns, and one fragmentary verse. This poverty has not prevented vigorous classification proposals — Turkic, Mongolic, Yeniseian, Iranian, “para-Mongolic,” composite, isolate — but it has prevented their adjudication. This paper surveys categories of evidence that are known or strongly suspected to exist, that are not currently in the active comparative dossier, and that could in principle be recovered through deliberate research programs rather than fortuitous discovery. The argument is not that decisive material is around the corner. It is that the gap between current evidence and adequate evidence is smaller than the field’s pessimism suggests, and that the bottleneck is more institutional than evidentiary.

I. The Current Evidentiary Floor

Any honest accounting begins with what we actually have. For the Xiongnu, the corpus consists almost entirely of Chinese transcriptions: tribal and clan names, the title chanyu (or shanyu) and a stable inventory of subordinate titles (tuqi, guli, danghu, gudu, etc.), personal names of rulers and nobles, a handful of glossed common nouns (“heaven,” “wife,” “milk”), and the famous Jie couplet preserved in the Jin Shu — ten Chinese characters allegedly transcribing a four-line oracular verse by the monk Fotudeng. The Jie are usually considered a Xiongnu-affiliated population, though the equation is not airtight, which is itself part of the problem: scholars cannot agree whether the only surviving running text in the language is in fact in the language.

For the European Huns, we have Greek and Latin transcriptions of personal names from Priscus, Jordanes, Ammianus Marcellinus, and a thin scatter of later authors. We have strava (funeral feast), medos (a fermented drink), kamos (a grain-based drink), and titles like logades (which is Greek but applied to a Hunnic institution). We have nothing approaching a sentence.

This corpus is not merely small. It is structurally deficient in ways that ordinary smallness would not produce. We have almost no verb morphology, almost no inflectional evidence, no clearly attested syntactic clauses, and no minimal pairs that would let us isolate phonemic contrasts. The names and titles can be — and have been — etymologized into half a dozen language families with comparable plausibility, because the constraint set is too thin to falsify any of them.

II. Subsurface Archaeological Corpora

The most underexploited category of potential evidence is material that already exists, in the ground, in conditions that should preserve writing surfaces and inscribed objects, and that has either not been excavated or has been excavated without linguistic-recovery as a research priority.

Elite Xiongnu kurgans in northern Mongolia and Transbaikalia. The Noin-Ula necropolis, Gol Mod 1 and 2, Tsaraam, and Duurlig Nars have yielded textiles, lacquers, chariot fittings, and bronzes preserved by groundwater seal and sometimes by partial permafrost. Every well-excavated elite Xiongnu tomb has produced inscribed objects — Chinese seals, mirror inscriptions, lacquer cartouches — but also non-Chinese marks: tamgas, isolated graphs on bronze plaques, scratched signs on bone. These have been catalogued primarily as ownership marks or apotropaic symbols. They have not been systematically analyzed as candidate writing. Several hundred elite tombs in the region remain unexcavated, and a comparable number have been disturbed but not exhausted. The probability that some carry inscribed wood, leather, or birch — the writing supports actually used in inner Asia — is not negligible. The probability that any single tomb yields such material is low; the probability across the unexcavated population is materially higher.

Frozen tombs of the Altai-Sayan complex. The Pazyryk culture predates the Xiongnu proper but the same mortuary technology — log chambers under stone cairns, sealed by ice lenses — was used into the Xiongnu period at sites like Olon-Kurin-Gol and Ak-Alakha. Permafrost preserves leather, felt, wood, and even tattooed skin. Climate change is now degrading these deposits at an accelerating rate, which makes salvage urgent rather than optional. A single inscribed birch fragment from such a context would more than double the running-text corpus.

The Xiongnu use of writing. Chinese sources repeatedly mention that the Xiongnu kept records, exchanged written messages with the Han, and used some form of notation for census and military purposes. The standard scholarly default has been that they used Chinese-literate scribes (often defectors) for diplomatic correspondence and tallies for everything else. This default is convenient but underdetermined. The discovery of even a small Xiongnu-internal notation system — whether a true script, a syllabary, or a sophisticated tally — would constitute evidence about phonological structure if it can be paired with Chinese-transcribed glosses. Tamga repertoires from Xiongnu sites have not been comprehensively published, much less subjected to the kind of distributional analysis that would distinguish writing from heraldry.

III. Han Chinese Administrative Slips

The single most productive category over the last forty years has been the wooden and bamboo slips excavated from Han frontier garrisons: Juyan, Dunhuang, Xuanquan, Suoyang, and the more recent finds at sites along the limes from the Etsin-gol to the Hexi corridor. The Xuanquanzhi corpus alone runs to tens of thousands of slips, of which only a fraction are fully published and a much smaller fraction translated.

These slips contain operational, not literary, material: dispatches, ration accounts, prisoner logs, intelligence reports, lists of envoys and gifts, and routine bureaucratic correspondence with Xiongnu chiefdoms during periods of nominal submission. They preserve Xiongnu personal names, titles, place names, and occasionally glossed terms in transcriptions made by clerks who were hearing the language at the frontier rather than copying canonical texts. This is precisely the kind of corpus where phonological detail survives — because the clerks were not standardizing against literary models — and where the same name recurs across multiple hands, allowing reconstruction of underlying forms by triangulation.

The recovery program here is not archaeological; it is editorial. The slips exist. They are in institutional storage. What is missing is a sustained, well-funded effort to extract every Xiongnu-language item from the unpublished and underpublished portions, to assemble them in a single searchable corpus with reconstructed Old Chinese phonological values attached, and to make this corpus available to comparative linguists rather than only to Han historians. A serious scholar with five years and a small team could plausibly triple the Xiongnu name and title corpus from this source alone.

IV. Lost and Partially Lost Greek and Latin Histories

The European Hunnic situation is shaped by the catastrophic loss of Priscus of Panium’s history, which survives only in excerpts preserved by later Byzantine compilers — primarily the Excerpta de Legationibus commissioned by Constantine VII in the tenth century, and scattered citations in Jordanes, the Suda, and Theophanes. Priscus visited Attila’s court in 449. He recorded Hunnic words and described Hunnic institutions with the eye of a participant. The portions we have are tantalizing and clearly represent a small fraction of the original.

Three recovery vectors merit attention. First, Byzantine palimpsests in the libraries of Mount Athos, Sinai, the Vatican, the Patriarchate of Jerusalem, and various Italian and Spanish monastic collections have not been comprehensively imaged. Multispectral and hyperspectral imaging has, in the last twenty years, recovered substantial portions of Galen, Archimedes, and various scriptural and patristic texts from undertexts that scholars had given up on. The Constantinian Excerpta tradition, which is where Priscus survives, is exactly the kind of material that was being recopied in the period when palimpsesting was common, and the loss patterns suggest that Priscus’s full text was extant in some monastic libraries as late as the eleventh or twelfth century. The probability that some substantial fragment survives as undertext in an unexamined manuscript is non-trivial.

Second, the Suda and other Byzantine lexicographic compilations contain alphabetized fragments whose source attributions have not been fully reconciled. Some entries that look like glosses on common Greek words are actually preserving foreign vocabulary with a Greek paraphrase. A systematic comparison of unattributed or thinly attributed Suda entries against the known Hunnic lexical inventory could identify additional items.

Third, the Latin tradition preserves Hunnic material in places it has not been systematically catalogued: Cassiodorus’s lost Gothic History (which Jordanes summarized) had Priscus as a source; saints’ lives from the Pannonian and Rhaetian regions occasionally preserve Hunnic personal names and place names in forms that differ from the Greek transmission and may be closer to the spoken substrate. A targeted philological program — Hunnic prosopography across Latin hagiography — has not been done.

V. Central Asian Textual Corpora

The textual cultures of the Tarim Basin and Sogdiana operated in a world that included Xiongnu successor states, “Iranian Huns” (Kidarites, Hephthalites, Alchon, Nezak), and ongoing political relationships with steppe powers. Several corpora are still being processed and contain potentially relevant material.

Sogdian. The Sogdian Ancient Letters from the early fourth century mention xwn (Huns) and discuss events around the sack of Luoyang. The Turfan Sogdian Manichaean and Buddhist corpora contain references to steppe peoples. Sogdian was the Silk Road lingua franca for much of the relevant period, and Sogdian transcriptions of Hunnic names and titles would be independently valuable because they preserve different phonological information than Greek or Chinese transcriptions. The unpublished portion of the Berlin Turfan collection alone is substantial.

Bactrian. The corpus of Bactrian documents from northern Afghanistan, mostly published by Sims-Williams over the last three decades, includes legal and administrative texts from the period of Hunnic rule. They contain Hunnic personal names, titles, and occasionally glossed terms in the local language — i.e., a contemporary, bilingual interface that is exactly what we lack for the European Huns. New Bactrian material continues to surface from the antiquities market, much of it unpublished and some of it in private hands.

Khotanese. Khotanese Saka texts mention hūṇa in various contexts and preserve some titles and ethnic terms. The Khotanese corpus is small but well-edited.

Tocharian. The Tocharian A and B corpora include occasional references to steppe peoples. Tocharian transcription practices have not been systematically mined for steppe lexical items.

Tibetan. The Old Tibetan Annals and the Dunhuang Tibetan manuscripts preserve names and titles from Inner Asian peoples. The relevant sections have been studied for Tibetan history but not systematically for what they preserve about steppe languages.

VI. South Asian and Iranian Sources

The “Iranian Huns” — Kidarites, Hephthalites, Alchon, Nezak — produced their own coins, seals, and a small number of inscriptions. Whether they are linguistically continuous with the European Huns and Xiongnu is precisely the question we want to answer, but their material counts as evidence under any of the live hypotheses.

Coin legends survive in Bactrian, Brahmi, Pahlavi, and occasionally in scripts that have been partially deciphered (the so-called “Hephthalite script” or “Bactrian-derived cursive”). Personal names of rulers, titles, mint marks, and tribal names are preserved across thousands of coins. Comprehensive corpora exist (Vondrovec, Alram, Pfisterer) but have not been fully integrated into linguistic comparison. The Tochi Valley bilingual inscriptions, the Jaghatu inscription, and various rock inscriptions from Afghanistan and Pakistan have been individually studied but not synthesized.

Sanskrit sources — the Mahabharata, the Puranas, Gupta-era inscriptions, the Mudrarakshasa — preserve forms of Hūṇa and associated names that diverge from the Greek-Latin and Chinese-Iranian transmissions. These divergences are themselves data: the differences between Mihirakula in Sanskrit, the same name on coins in Bactrian script, and similar-looking names in Chinese sources triangulate the underlying form better than any single transmission.

VII. Caucasian and Armenian Sources

Armenian historiography from Movses Khorenatsi, Agathangelos, Elishe, and Lazar P’arpec’i preserves Hunnic personal names and ethnonyms in transcriptions that follow Armenian phonological conventions. These are systematically different from Greek and Latin transcriptions of the same names and provide independent triangulation. Georgian chronicles add a further layer. Caucasian Albanian — recovered from the Sinai palimpsests in the last twenty-five years — opens a fresh transmission channel that is just beginning to be exploited.

The “Caucasian Huns” of the fifth and sixth centuries (Sabirs, Onogurs, and others) appear in Armenian, Georgian, and Syriac sources. The Syriac material in particular — Pseudo-Zacharias, John of Ephesus — has not been mined for ethnolinguistic content with the same intensity as the Greek tradition.

VIII. Indirect Attestation: Loanwords and Substrate

Even if no new direct attestation surfaces, the language can in principle be partially reconstructed through its impact on its successors. This program has been pursued sporadically but never systematically.

Mongolic, Turkic, and Tungusic vocabularies contain layers of loans from earlier steppe languages. Identifying which loans are pre-Old-Turkic and pre-Common-Mongolic, and what their phonological shape implies about the source language, is a recoverable program. The same applies to Yeniseian languages (Ket and its now-extinct relatives), which Vovin and others have argued preserve traces of a Xiongnu substrate. The Yeniseian hypothesis has the virtue of being testable: if Xiongnu was Yeniseian or close to it, we expect specific phonological signatures in the Chinese transcriptions, and we expect specific kinds of substrate residue in the Turkic and Mongolic languages that displaced it.

Slavic, Hungarian, and Romance languages preserve possible Hunnic loans in ways that have been studied but not exhausted. Hungarian is particularly tantalizing because of the long historical association with steppe traditions, even though the modern consensus is that Hungarian itself is not a Hunnic descendant.

IX. Methodological Frontiers

Recovery is not only a question of finding more material. It is also a question of extracting more information from material we already have.

Computational reconstruction of Chinese transcriptions. The Chinese transcription of foreign names depends on the phonological values of the Chinese characters used, and these values change over time. Old Chinese reconstructions have improved dramatically since the work of Karlgren — Baxter-Sagart, Schuessler, Pulleyblank, and now various computationally informed proposals — but applying these reconstructions to the Xiongnu corpus is a separate operation from the core reconstruction work. A systematic re-transcription of every Xiongnu lexical item using current Old Chinese reconstructions, with explicit confidence intervals, would clarify which forms are securely reconstructed, which are ambiguous between reconstructions, and which are essentially unconstrained. Some of the strongest etymological arguments in the literature rest on transcription assumptions that are no longer current.

Multispectral imaging of palimpsests and degraded manuscripts. This is now a routine technique, but its application has been driven by texts that scholars already know they want — Galen, Archimedes, biblical material. A program targeted specifically at Byzantine historical manuscripts known to descend from chains that included Priscus, Olympiodorus, and Eunapius would be a distinct intervention.

Statistical analysis of Inner Asian onomastic patterns. Personal name structures often preserve information about phonotactics and morphology even when the names themselves are not etymologized. Statistical analysis of large Xiongnu and Hunnic name corpora — looking for recurrent prefixes, suffixes, vowel patterns, and consonant clusters — can yield phonological and morphological constraints without requiring that any individual name be successfully etymologized.

Genomic correlation. Population-genetic studies of the Xiongnu (Jeong et al. and successors) have shown that the confederation was genetically heterogeneous, which constrains but does not determine linguistic hypotheses. Linking specific genetic clusters to specific burial styles, name patterns, and material culture markers can produce indirect evidence about which parts of the linguistic record represent the politically dominant language and which represent client populations.

Machine-assisted philology. The processing of Han slip corpora, Bactrian documents, and Sogdian materials is currently bottlenecked by trained human reading time. Machine-assisted transcription, named-entity recognition tuned to steppe ethnonymics, and automatic cross-referencing across corpora can multiply the effective output of the small number of qualified scholars. This is not a replacement for philology but a force multiplier for it.

X. Why the Gap Persists

The evidentiary gap is not primarily a function of how much material exists. It is a function of how the institutions that govern recovery are organized.

Sinology, classical philology, Iranian studies, Indology, Armenology, Turkology, and Mongolic studies are separate disciplines with separate journals, separate funding streams, separate training pipelines, and separate canons. The Xiongnu and Hunnic problems sit at the intersection of all of them and are central to none. A scholar who tries to work across these fields acquires the combined methodological burden without acquiring the institutional support that would normally accompany it. The result is that each discipline contributes its own thin slice of evidence to the comparative pile, and no one is professionally responsible for assembling the pile or for noticing what is missing from it.

Archaeological recovery in Mongolia, Russia, and Central Asia operates under political and economic constraints that vary by decade. Permafrost thaw is destroying material faster than it is being recovered. Slip corpora published in Chinese journals are not always indexed in databases used by non-Sinologist linguists. Coin corpora published by numismatists are not always searched by historical linguists. Palimpsest imaging programs are funded for specific theological and classical priorities and rarely extend to historiographical material.

These are not insurmountable problems, but they are structural rather than evidentiary. The pessimism in the field about ever resolving the Xiongnu and Hunnic linguistic questions is partly justified by the genuine smallness of the current corpus and partly an artifact of how the work is organized. The corpus could grow substantially within a generation if growth were treated as a deliberate institutional objective rather than a hoped-for byproduct of unrelated research.

XI. Priority Programs

If one were to design a coordinated recovery program with realistic resources, the highest-yield interventions would be approximately as follows. A comprehensive editorial project to extract all non-Chinese-language items from the Han frontier slip corpora, with current Old Chinese reconstructions attached, would substantially expand the Xiongnu lexicon. A multispectral imaging campaign targeting Byzantine historical manuscript traditions descending from the Constantinian Excerpta tradition would offer a realistic chance of recovering Priscus material. A salvage archaeology program for the rapidly degrading frozen tombs of the Altai-Sayan zone would address an irreversible deadline. A coordinated publication and database initiative for unpublished Sogdian and Bactrian material would exploit corpora that already exist but are not accessible. A sustained comparative program that takes the substrate evidence in Turkic, Mongolic, and Yeniseian seriously as primary evidence rather than as confirmation of independently formed hypotheses would extract more information from material that is already in print.

None of these would produce a Rosetta Stone. Together, they would plausibly move the Xiongnu and Hunnic questions from the current condition — where every classification proposal is consistent with the evidence and none is decisively supported — to a condition where some proposals are decisively excluded and the remaining hypothesis space is small enough to be productively contested.

XII. Conclusion

The standard framing of the Xiongnu and Hunnic linguistic problems treats them as effectively closed: the material is gone, the transcriptions are too thin, the comparative work has been done and yielded indeterminacy. This framing is wrong in a specific way. It mistakes the current state of mobilized evidence for the current state of available evidence, and it mistakes the limits of single-discipline competence for the limits of what coordinated work could accomplish. The languages may never be fully recovered. They are not as far out of reach as the field’s habits of pessimism suggest, and the next generation of work — if it is organized to take recovery as an objective rather than as an accident — could materially change what we know.

Unknown's avatar

About nathanalbright

I'm a person with diverse interests who loves to read. If you want to know something about me, just ask.
This entry was posted in History, Musings and tagged , , , . Bookmark the permalink.

Leave a Reply