Data & indices
LINGACON maintains two headline measures. ALPS counts what people have — the languages in their repertoire. The GLCI measures what those languages achieve — whether two people can actually understand one another. Both are built country by country from documented sources, under one set of definitions, and released openly with the papers that describe them.
ALPS
Average Languages a Person Speaks
ALPS is the population-weighted average number of languages spoken per person. "Speaking" is fixed at a single explicit threshold: conversational competence at CEFR A2 and above — enough to handle an everyday exchange, not a memorized phrase and not native-level mastery. Fixing that threshold is what makes national figures comparable; sources that use looser or stricter criteria are harmonized to it, and the adjustment is documented.
Two design choices distinguish the index from other attempts. First, it counts an unlimited number of languages per person. Most instruments stop at a first and second language, or cap the count at three or four — the Eurobarometer, for example, asks about a limited set of foreign languages, so anyone beyond that ceiling is silently truncated. A person who speaks six languages counts as six. Second, every national figure is audited against independent evidence rather than taken from a single source at face value.
The current global estimate is 1.59, across 235 countries and territories. ALPS is reported to two decimal places at global, regional, and national level.
GLCI
Global Linguistic Connectivity Index
The GLCI asks a different question: take two people at random — what is the probability that one can understand the other? Aggregated over all pairs, it becomes a measure of how linguistically connected a population is. The global value is 11.6%. Within a single country the figure rises to roughly 57%; between two different countries it falls to 8.3%.
Crucially, the index does not treat languages as isolated boxes. Someone who speaks Spanish understands a good deal of Portuguese; a Czech speaker follows much of Slovak; speakers of neighbouring Arabic varieties understand each other to a degree that depends on which varieties. LINGACON accounts for this through a mutual intelligibility coefficient (MIC) — a documented matrix of partial comprehension between related languages, built from the experimental literature rather than assumed. The matrix is asymmetric, because comprehension often is: speakers of one variety frequently understand another better than the reverse.
The full methodology, the coefficient matrix with its sources, and the sensitivity of the headline figure to alternative conventions are published with the GLCI paper.
Coverage and sources
The dataset covers 235 countries and territories — every jurisdiction with a resident population, not only UN member states, together accounting for more than 8 billion people, roughly 97% of humanity. Connectivity estimates are computed on the subset of jurisdictions with complete language-by-language detail, and the coverage convention is stated explicitly wherever a figure is reported.
National figures draw on population censuses, large-scale survey programmes such as the Eurobarometer and the Afrobarometer, national statistical offices, and documented language surveys — harmonized to LINGACON's common definitions. No single source is taken at face value: instruments that appear to ask the same question routinely produce different answers because they apply different thresholds, and reconciling them is a substantial part of the work rather than a footnote to it.
Access
Headline datasets are openly available for academic and scientific research, free of charge, with citation. They are published alongside the corresponding papers, together with the supplementary materials needed to reproduce every figure.
Beyond that, LINGACON holds a substantial body of proprietary data: language-by-language national detail, pairwise connectivity matrices, regional and sub-national breakdowns, and the audit layer behind each estimate. These are available on request under licence, on commercial terms, to companies, institutions, and organizations using them outside academic research — see Partnerships for data licensing and the planned API.
Researchers who need early access, additional country detail, or replication materials are welcome to write to research@lingacon.com.
