Abstract
The languages of New Guinea and its satellite archipelagos constitute the densest and least understood concentration of linguistic diversity on Earth. Roughly eight hundred and fifty languages — by some counts more, by others fewer, depending on how splitter-lumper questions are resolved — are spoken across an area smaller than the state of Texas, organized into perhaps twenty to forty proposed families plus a substantial residue of isolates. The label “Papuan” is itself a negative classification, denoting the non-Austronesian languages of the region rather than any positive genetic claim. Decades of careful comparative work have established some families with reasonable confidence (Trans-New Guinea in some formulation, Sepik, Torricelli, Lower Sepik-Ramu, and several smaller groupings), but vast portions of the classification remain contested, and the relationships among the established families remain largely opaque. This paper surveys computational methods that could materially improve this situation, with attention to what they can and cannot do, what kinds of data they require, and where the genuine bottlenecks lie. The argument is not that computation will resolve Papuan classification — the data limitations are too severe for that — but that computational methods can extract substantially more information from existing material, prioritize fieldwork where it would do the most good, and discipline the proliferation of weakly grounded family proposals.
I. The Shape of the Problem
Before considering what computational methods can contribute, the structure of the Papuan classification problem deserves careful statement, because methods that work well for Indo-European or Austronesian fail in ways specific to this region.
The Papuan languages are old in place. The human occupation of New Guinea is approximately fifty thousand years deep, and there is no reason to think the linguistic occupation is much shallower. By contrast, the time depth at which the comparative method is generally considered reliable is roughly six to eight thousand years, with some practitioners pushing to ten thousand under favorable conditions. The Papuan situation places most genuine genetic relationships beyond the standard reach of the comparative method, not because the relationships are absent but because the signal-to-noise ratio in surviving cognate sets has decayed past the threshold of confident reconstruction.
The languages have been in intense contact with each other for most of that period. Borrowing is endemic. Areal features cluster across genetic boundaries. Sprachbund effects have produced typological convergence across families that almost certainly are not related. The same areal pressures have produced morphological and phonological similarities that mimic genetic signal, generating false positives for relatedness when superficial methods are applied.
Documentation is uneven and in many cases thin. Some languages have full grammars, dictionaries, and text corpora produced by missionary linguists, SIL workers, or academic field linguists. Others have only short word lists, sometimes collected a century ago by patrol officers or anthropologists with no linguistic training, sometimes containing transcription errors that have propagated through subsequent comparative work. A non-trivial number of languages are represented in the comparative literature only by the same hundred-item Swadesh list, collected once, by one person, decades ago.
The languages exhibit typological features that make standard comparative methods less effective. Many have small phonemic inventories with extensive allophonic variation, making it difficult to identify regular sound correspondences. Many have complex serial verb constructions, switch-reference systems, and other morphosyntactic features whose comparative behavior is poorly understood. Many have extensive paradigmatic suppletion and irregular morphology that resists the kind of paradigm-based reconstruction that works well in Indo-European.
The Trans-New Guinea hypothesis, in its various formulations, attempts to group several hundred of these languages into a single family. Its strongest version (associated with Pawley, Ross, and others) is supported by a stable inventory of pronominal forms and a smaller set of putative cognate roots. Its weakest version is essentially a residual category for highland and adjacent languages that show some pronominal similarities. Whether Trans-New Guinea is one family, several families, or a contact-induced typological pool is a question that the standard comparative method has not been able to settle in fifty years of work, and there is no reason to expect that more of the same will settle it.
II. What Computation Can and Cannot Do
A clear-eyed account of what computational methods bring to this problem should begin by distinguishing the genuine contributions from the inflated claims that have accompanied the rise of computational phylogenetics in linguistics.
Computational methods are good at processing large amounts of data consistently, applying explicit criteria uniformly, quantifying uncertainty, and surfacing patterns that human pattern-recognition would miss or weight incorrectly. They are good at prioritization: given a thousand candidate cognate sets, which deserve closest scrutiny? Given five hundred languages, which sub-groupings show the strongest signal? Given a list of hypothesized sound correspondences, which are statistically robust and which are at chance level?
Computational methods are not good at substituting for absent data, distinguishing inheritance from borrowing without independent evidence, or producing cognate sets from raw word lists without prior expert judgment about what could plausibly be cognate. They are also not good — despite frequent claims — at identifying language families from typological features, because typological features cluster areally and the resulting trees reflect contact rather than descent.
The proper role of computational methods in Papuan linguistics is therefore not to replace traditional comparative work but to amplify it: to extract more information from existing data, to identify the highest-value targets for additional fieldwork, to test specific hypotheses with explicit statistical machinery, and to maintain a current and consistent picture of what is and is not supported by the available evidence.
III. Lexical Phylogenetics: Promise and Pitfalls
Bayesian phylogenetic methods, originally developed for biological systematics and adapted to linguistics by Gray, Atkinson, Bouckaert, and others, have produced striking results for Indo-European, Austronesian, Bantu, and several other families. Their application to Papuan languages has been limited and the results so far have been suggestive rather than decisive.
The basic procedure takes a matrix of cognate judgments — for each meaning slot in a word list, which languages share cognate forms — and infers a tree, with branch lengths corresponding to amounts of lexical replacement, that best accounts for the distribution of cognates. Bayesian methods quantify uncertainty by sampling from the posterior distribution of trees consistent with the data, producing not a single tree but a distribution of trees with associated probabilities.
For Papuan applications, several problems are acute. First, the cognate judgments are themselves the comparative-historical question; if we knew the cognates with confidence, we would already have the classification. The standard practice is to rely on cognate judgments produced by traditional comparative work, which means the phylogenetic analysis inherits all the uncertainty of those judgments and adds its own statistical machinery on top. Second, the deep time depth means that lexical replacement has progressed nearly to saturation in many cases, and the remaining signal is weak relative to the noise from chance resemblance and undetected borrowing. Third, the contact-heavy history means that the assumption of tree-like descent is violated to a degree that simple tree models cannot accommodate.
The honest assessment is that lexical Bayesian phylogenetics, applied to Papuan data, can probably resolve the structure within established families with moderate confidence and can probably distinguish closely related families from distantly related ones, but cannot resolve deep relationships across families on current data. The method is most useful as a discipline imposed on cognate judgments — forcing explicit, replicable, and statistically evaluated decisions about which forms are cognate — rather than as an oracle that produces classifications.
A more productive program is to use phylogenetic methods to test specific hypotheses against null models. Does the pronominal evidence for Trans-New Guinea produce a tree signal stronger than chance? Where exactly within the proposed Trans-New Guinea languages does the signal concentrate, and where does it dissipate? When borrowing is explicitly modeled rather than assumed away, how much of the apparent signal survives? These are questions that admit computational answers, even if the headline question of whether Trans-New Guinea exists as a single family does not.
IV. Automatic Cognate Detection
Several algorithms now exist for identifying potential cognate sets from word lists without prior expert judgment. The LexStat method developed by Johann-Mattis List, the SCA (Sound-Class-based Alignment) method, and various neural-network approaches have demonstrated reasonable performance on well-documented families.
For Papuan languages, automatic cognate detection faces a structural difficulty: the methods are trained or calibrated on languages where the genetic relationships are known and the regular sound correspondences are established. They identify cognates partly by sound similarity and partly by participation in regular correspondences. Where regular correspondences are not yet established — which is the case for most cross-family Papuan questions — the methods reduce to sound similarity, which conflates inheritance and borrowing and produces both false positives and false negatives at unacceptable rates.
The productive use of these methods is therefore not as a primary classification tool but as a discovery tool. Given a candidate language pair, run automatic cognate detection. Examine the candidate cognate sets it produces. Look for systematic patterns of sound correspondence among the candidates. Where systematic patterns appear, the candidate cognates become a hypothesis worth pursuing through traditional methods. Where patterns do not appear, the candidate cognates are most likely chance resemblance or scattered borrowing.
This workflow is not a replacement for the comparative method. It is a way of generating candidate correspondences faster than human inspection of word lists, and of doing so consistently across many language pairs simultaneously, so that the human comparativist can focus attention on the most promising material.
A related application is the systematic identification of likely loans. Once correspondence patterns are partially established, forms that should participate in the patterns but do not — and that have a sound shape consistent with a known contact language — can be flagged as likely loans and separated from the inheritance signal. This is something traditional comparative work does manually and inconsistently; computational methods can do it consistently across an entire corpus.
V. Phonological Reconstruction by Alignment
Multiple sequence alignment, borrowed from molecular biology, can be applied to phonological forms across languages to produce explicit reconstructions of proto-forms. The method aligns segments across languages — treating each segment as analogous to a nucleotide or amino acid — and infers the most parsimonious or most probable proto-segment at each position.
The application to Papuan languages is constrained by the same time-depth problem that affects everything else. Where the languages are close enough that alignment is feasible, the method can produce reconstructions that are at least as good as those produced by traditional methods, with the advantage that the reasoning is explicit and replicable. Where the languages are distant, the method produces alignments that are not meaningfully different from random.
The genuine contribution of alignment-based reconstruction is in mid-level cases: established families where the comparative work has been done in fragments by different scholars over decades, with no unified reconstruction. Computational alignment can produce a unified reconstruction that integrates all the existing evidence, identifies points where the existing fragments conflict, and generates explicit predictions that fieldwork can test. For families like Sepik, Torricelli, the various Trans-New Guinea sub-families, and others where partial reconstructions exist but no comprehensive synthesis has been produced, this is a tractable and high-value program.
VI. Typological Databases and Their Limits
The World Atlas of Language Structures (WALS), Grambank, PHOIBLE, and several specialized databases provide typological data on hundreds of Papuan languages. These data have been used to investigate areal patterns, to test correlations between typological features, and to generate typological clusters that have sometimes been interpreted as evidence for genetic groupings.
The genetic interpretation is a methodological error. Typological features cluster areally on time scales much shorter than the time depths at issue in Papuan classification. Two languages that share a typological profile may share it because they descend from a common ancestor, because they have been in contact, or because they have independently arrived at the same solution to a structural problem. Distinguishing these requires evidence that typological data alone does not provide.
The legitimate uses of typological data in Papuan classification are several. First, typological profiles can identify outliers — languages whose typology differs sharply from their geographic neighbors and whose ancestry may therefore lie elsewhere. Second, typological data can support hypotheses generated by lexical evidence: if a proposed family is also typologically coherent in non-trivial ways, that is weak additional support. Third, typological data can characterize the contact-induced convergence that comparative work needs to factor out. Fourth, typological data can identify rare or unusual features whose distribution is more diagnostic than common features: a feature found in only a handful of the world’s languages, occurring in a cluster of Papuan languages, is more likely to reflect a genuine historical connection than a feature found in half the world’s languages.
The Grambank project in particular has produced data of a quality and consistency that earlier typological work did not achieve, and computational analyses of Grambank data are beginning to produce results that deserve attention. The most promising line of work is not the inference of family trees from typological data — which the data cannot support — but the characterization of areal zones and the identification of features that pattern with established genetic groupings versus features that pattern areally.
VII. Pronoun Paradigms and the Trans-New Guinea Question
The Trans-New Guinea hypothesis rests, more than on any other single piece of evidence, on a recurrent pronominal pattern: a first person singular pronoun involving a nasal consonant (often *na or similar) and a second person singular pronoun involving a velar (often *ŋga or similar), with related forms in the plural. This pattern is observed across hundreds of languages in highland New Guinea and adjacent regions. Whether it constitutes evidence for genetic relationship has been debated for half a century.
This is a problem on which computational methods can make a specific contribution. The question is whether the observed distribution of pronominal forms is more concentrated, more systematic, and more phonologically coherent than would be expected by chance, given the size of the pronominal inventory, the number of languages involved, and the constraints on what pronouns typically look like cross-linguistically.
Several specific computational programs are tractable. The first is explicit null-model testing: given a database of pronominal forms from a typologically representative sample of the world’s languages, what is the probability that a region of the size and density of Trans-New Guinea would show the observed degree of pronominal similarity by chance? The second is internal structure analysis: within the languages that share the pronominal pattern, do the pronouns participate in regular sound correspondences with other lexical material, as would be expected if they are inherited, or do they appear as a kind of typological flag detached from systematic phonological history? The third is contact analysis: how does the distribution of pronominal forms correlate with geographic proximity, with established contact relationships, and with non-pronominal lexical similarity?
These questions admit explicit computational answers, and the answers — whatever they turn out to be — will discipline a debate that has been conducted largely on impressionistic grounds. The Trans-New Guinea hypothesis may emerge strengthened, weakened, or restructured into a smaller core with a larger periphery of areally affiliated languages. Any of these outcomes would be progress.
VIII. Sound Correspondences at Scale
The traditional comparative method works by identifying recurrent sound correspondences across languages and reconstructing the proto-segments that account for them. The method is well understood for cases where the family is small and the correspondences are clear. It scales poorly to cases involving hundreds of languages and partial documentation.
Computational methods can extract candidate correspondences from large word-list corpora consistently and quickly. Given a body of word lists for, say, two hundred languages with two hundred items each, the algorithms can identify segment-position pairs that recur across multiple language pairs, rank them by statistical strength, and produce an inventory of candidate correspondences for human evaluation. This is not the comparative method — the human evaluation step is essential — but it is a way of feeding the comparative method with material that has been pre-filtered for plausibility.
The high-value targets are the medium-density families where partial correspondences have been established by individual scholars but no comprehensive correspondence inventory exists. Computational methods can produce that inventory, identify gaps and inconsistencies in the existing work, and suggest specific lexical items where additional fieldwork would resolve open questions. For families like Lower Sepik-Ramu, Torricelli, the various Sepik sub-families, and the smaller Trans-New Guinea sub-families, this is a tractable program that would substantially improve the state of the field.
A related application is the identification of loanword strata. Languages in extended contact accumulate loans in waves, with each wave reflecting the phonological state of the donor language at the time of borrowing. Computational analysis of correspondence patterns can sometimes separate strata that correspond to different periods of contact, providing historical information about the relationships between the languages even when full genetic relationships cannot be established.
IX. Network Models and Reticulate Histories
The Papuan situation, with its endemic borrowing and areal convergence, violates the tree-model assumptions of standard phylogenetic methods. Network methods, which allow for reticulation — for languages to have multiple ancestors, for features to flow across branches, for the history to be a network rather than a tree — are a closer fit to the actual situation.
Several network methods have been developed in recent years. NeighborNet produces an unrooted network that displays both tree-like and reticulate signal. Bayesian methods that explicitly model borrowing alongside descent have been developed for specific cases. The TraitLab framework allows for testing whether tree models or network models better fit the data.
For Papuan applications, network methods are particularly valuable as a diagnostic. A language family for which the data produce a tree-like network is a family for which standard phylogenetic methods are appropriate. A language family for which the data produce a strongly reticulate network is one for which tree models will produce misleading results. Knowing which is which, before committing to a method, is itself useful.
Network methods are also useful for visualizing areal versus genetic signal. The classic problem in Papuan linguistics is that strong typological clustering may reflect either common descent or long contact, and the two are difficult to distinguish from typology alone. When typological data is combined with lexical data in a network framework, the relative contributions of the two signals can sometimes be separated, with features that pattern with the lexical signal more likely to be inherited and features that pattern against it more likely to be areal.
The honest limitation is that network methods, like tree methods, depend on the quality of the input data. They do not solve the Papuan documentation problem; they solve a different problem given that documentation problem is partially solved.
X. Field-Computational Integration
The most productive role for computational methods in Papuan linguistics is probably not as a stand-alone analytical program but as a feedback loop with fieldwork. Computational analysis of existing data identifies the highest-value targets for fieldwork — the languages whose better documentation would most improve the resolution of open classification questions. Fieldwork produces new data on those targets. The new data is fed back into the computational analysis, which produces a revised picture and identifies the next round of high-value targets.
This integrated workflow has not been pursued systematically for Papuan languages, in part because the institutional structures that fund and conduct fieldwork are separate from those that conduct computational analysis. The fieldworker chooses targets based on personal interest, accessibility, and the priorities of the host institution. The computational analyst works with whatever data has been produced. There is no mechanism for the analyst to communicate to the fieldworker that this particular language, in this particular region, would resolve this particular question if it were better documented.
A coordinated program that closed this loop would produce more progress per unit of fieldwork than the current uncoordinated arrangement. The specific targets would emerge from analysis: languages on contested boundaries between proposed families, languages whose minimal documentation creates the largest uncertainty in the regional picture, languages whose typological profile suggests an unexpected affiliation that better data could test. The fieldwork is hard regardless — New Guinea fieldwork is genuinely difficult and dangerous — but it could be allocated more effectively.
XI. Documentation Acceleration
Beyond classification, computational methods can accelerate documentation itself. Speech recognition adapted to low-resource languages, automated forced alignment of recordings to transcriptions, and machine-assisted glossing of texts can multiply the output of individual field linguists. The same recordings that would otherwise yield a few transcribed texts after months of work can yield substantially more if the routine portions of the work are automated.
These tools are not yet mature for languages with no prior documentation — speech recognition for an undocumented language is essentially unsolvable without bootstrapping data — but they are useful for languages with even minimal existing documentation. Once a few hundred utterances are transcribed, models can be trained to assist with additional transcription, and the rate of documentation accelerates. For the substantial number of Papuan languages that have partial documentation but no comprehensive corpus, this is a tractable acceleration.
The same techniques apply to legacy materials. Audio recordings made by missionaries and field linguists over the last fifty years exist in archive collections, often poorly catalogued and largely untranscribed. Machine-assisted transcription, paired with whatever paper documentation accompanies the recordings, could extract usable data from materials that would otherwise sit in archives indefinitely. For some languages, the recoverable archival material may exceed what current fieldwork could produce.
XII. Database Infrastructure
A practical bottleneck that computational methods cannot bypass but can be designed to address is the absence of adequate database infrastructure for Papuan linguistic data. The data exists, in many cases, but is scattered across journal articles, monographs, archived field notes, theses, and personal collections. The overhead of locating and integrating it is high, and much of what is theoretically available is effectively inaccessible.
Modern linguistic databases — Glottolog, CLDF (Cross-Linguistic Data Formats), Lexibank, Grambank — have begun to address this for the world’s languages collectively, with Papuan coverage growing but still partial. A targeted Papuan-specific database initiative, building on these frameworks but optimized for the specific needs of Papuan classification work, would multiply the productivity of every analytical program described above. Without it, every project starts by reassembling its own data; with it, projects can build on each other.
This is infrastructure work, not analysis, and infrastructure work is chronically underfunded relative to analysis. The pattern in which each new project produces its own dataset, uses it, and leaves it inaccessible to subsequent projects is wasteful in a way that the field has tolerated for a long time. A coordinated database, with proper data formats, version control, and documentation of cognate judgments and their justifications, would change the productivity baseline.
XIII. What Adjudication Would Look Like
A useful exercise is to imagine what a well-resolved Papuan classification would actually look like, and to ask how computational methods could move the field toward it.
A well-resolved classification would distinguish, for each proposed family, the core where the evidence is strong from the periphery where the evidence is weaker. It would explicitly model contact and borrowing, separating areal signal from genetic signal. It would quantify uncertainty: not “Trans-New Guinea is a family” or “Trans-New Guinea is not a family” but “the evidence supports a core Trans-New Guinea grouping of approximately N languages with this confidence level, an extended grouping of approximately M additional languages with lower confidence, and a residue of languages that are areally affiliated but cannot be confidently assigned.” It would produce explicit predictions that future fieldwork could test, and it would update systematically as new data arrives.
This is not what the current classification looks like. The current classification is a patchwork of proposals from different scholars working with different data and different methods, with no consistent uncertainty quantification and no shared infrastructure. Computational methods cannot fix this by themselves — the underlying comparative work is irreducibly hard — but they can provide the framework within which the comparative work accumulates rather than dispersing.
XIV. Realistic Expectations
The honest summary is that computational methods will not solve the Papuan classification problem in any decisive sense. The data limitations are too severe, the time depths too great, the contact effects too pervasive. What computational methods can do is more modest and still substantial: extract more information from existing data, discipline the proliferation of weakly grounded proposals, prioritize fieldwork where it will do the most good, accelerate documentation, and provide infrastructure within which future work can accumulate.
The pessimistic frame, in which Papuan classification is essentially intractable and computational methods will not change that, mistakes the current state of the field for its potential state. The inflated frame, in which computational phylogenetics will resolve the classification once enough data is fed in, mistakes the methods for oracles. The realistic frame is that a generation of integrated work — computational analysis tightly coupled with fieldwork, infrastructure built deliberately rather than as a byproduct, methods chosen for fitness to the specific Papuan problem rather than imported wholesale from better-behaved cases — could produce substantial improvement on the current picture without producing certainty.
What is certain is that the current pace of progress is slower than the pace of language loss. Many of the smaller Papuan languages have under a thousand speakers, intergenerational transmission is faltering in many communities, and the documentation that does not happen in the next generation will not happen at all. The classification problem is not the most urgent problem facing Papuan linguistics; documentation is. But the two are connected, and the same computational tools that improve classification can accelerate documentation, and the same coordinated infrastructure that supports one supports the other.
XV. Conclusion
The Papuan languages are a hard case for every analytical method that has been applied to them, and computational methods are no exception. They are also a high-value case, because the linguistic diversity of the region is unmatched and because what is lost when documentation fails cannot be recovered. The proper role of computational methods in this situation is as a force multiplier for the irreducibly human work of comparison, fieldwork, and analysis — not as a substitute for that work, and not as a source of conclusions that the data cannot support. Used in this role, computational methods could materially improve a situation that has been frustratingly static for decades. Used inappropriately — as oracles, or as ways of generating publications without the underlying philological work — they will produce a literature that looks more decisive than it is and that the next generation will have to clean up. The choice between these two trajectories is not technical. It is a choice about how the field organizes itself, what it rewards, and what infrastructure it is willing to build.
