Method
How the words were classified, which texts were counted and what the analysis cannot establish.
Assigning the strata
A word's stratum is the language from which Romanian borrowed it directly. Dragoste is Slavonic even though Slavonic had taken it from Proto-Slavic; bulevard is French even though French had taken it from Dutch.
The assignment comes from Wiktionary's etymology categories, which the templates inside articles populate automatically, for 47,654 lexemes. Each word's chain is read from the first sentence of the etymology that contains a link. Disputed etymologies list their hypotheses in later sentences, and reading all of them would produce a route that does not exist.
For 13,635 published words, the result was compared with the etymological notes in the dictionaries aggregated by dexonline. The two sources indicate the same stratum in 86.7 per cent of cases and on the same source language in 91.1 per cent. Differences occur especially for nineteenth-century borrowings, where scholars may identify either Latin or French as the direct source. Each word entry flags these cases explicitly.
The two hundred most frequent lemmas account for nearly half the text, so misclassifying one of them can visibly change the percentages. Wiktionary keeps one article for homographs with different histories, and we read the first etymology section, which goes wrong in the automated result. The 14 cases in which the frequent sense was not the sense in that first section are corrected by hand and listed below.
The historical corpus: dated prose, 1825–1949
6,626 texts from Romanian Wikisource, with 10,998,608 words, grouped into five 25-year intervals. Each text gets a year from the page, where the page declares one, otherwise from the author's active period, computed from their Wikidata dates as birth plus thirty-five years, capped at the year of death.
Three precautions. Translations are excluded, because a foreign author's dates say nothing about the year of the Romanian text. No author may contribute more than 2,000,000 characters to one interval, so that we do not measure Eminescu's idiolect in place of the period's language. Repeated paragraphs are counted once, because Wikisource keeps both the parent page and its chapters.
The first interval, 1825–1849, is the weak one: 369,815 words from 18 authors, of whom the most prolific supplies 30.8 per cent of the text. Little prose was printed in Romanian before 1850, and it shows.
The contemporary corpora: two registers, read separately
Spoken: 306,494,203 words from the OpenSubtitles frequency list for Romanian, that is, film dialogue. Written: 15,566,007 words from 41,357 Romanian Wikipedia articles, taken in dump order.
The two corpora are analysed separately because they represent different registers.
Lemmatisation
The same lemma appears in several forms in running text: casă, casa, casei, case, casele, caselor. Lemmatisation groups them before frequencies are calculated.
Forms are tied to lemmas using the 2,273,343 inflected forms in the dexonline database. Normalisation brings the language's successive orthographies to one place: the cedilla in place of the comma, î for â, nineteenth-century diacritics, and texts typed without diacritics at all, among them the constitutions of 1952 and 1965.
A form leading to several lemmas is divided among them in proportion to what those lemmas collected from unambiguous forms. That is how si, written without diacritics, ends up almost entirely with și rather than with the musical note.
Licensing the dexonline data
The inflected forms and lemmas used for lemmatisation come from dexonline's official export dataset. Processed data derived from that dataset retain the terms of GNU GPL v2 or any later version. Copyright © 2004–2026 dexonline, https://dexonline.ro.
The original dump and definition texts are not redistributed. The public archive includes only processed form-to-lemma links, aggregate statistics and flags for etymological disagreements. You can consult the attribution, licence links and exact inventory in the public DEXONLINE-NOTICE.md file.
What the page does not say
First appearance in the corpus is not a word's first attestation. Philological attestations are found in the Dicționarul limbii române and cannot be extracted automatically. The page shows only that the word was already in use in the first interval where it appears; it may have existed earlier.
A word gets a curve of its own only for the 9,382 words with at least twelve occurrences in the dated prose, spread across at least three intervals. For the rest, the stratum's curve is shown.
Between 1949 and today's corpora there is a fifty-year hole, which the official texts cover only in their own register. The charts draw no continuous line across it.
Wikipedia is not written language in general but an encyclopedic register, which leans towards Latinisms and Gallicisms. Film dialogue is not speech but speech written by scriptwriters and translators. Both are named as such wherever they appear.
Stratum percentages are given as a share of the text whose donor could be named: 87.2 per cent of the spoken corpus and 78.8 per cent of the written one. The remainder are words with no etymology in Wiktionary's categories, not words with no etymology.
The hand corrections
Each with its reason. These are the heavy words where the automatic assignment had taken the rare homograph instead of the ordinary sense.
| Word | Stratum | Reason |
|---|---|---|
| mai | Inherited Latin | lat. magis; Wiktionary listează mai întâi sensurile «ciocan» și «lună», dar forma frecventă este adverbul comparativ |
| cap | Inherited Latin | lat. caput, moștenit; prima secțiune descrie sensul nautic, împrumutat din franceză |
| duce | Inherited Latin | lat. ducere; substantivul italian duce e omonimul rar |
| prin | Inherited Latin | per + in, amândouă latine |
| prim | Learned Latin | lat. primus, împrumut savant de secol XIX, nu maghiar |
| secol | Learned Latin | lat. saeculum |
| ultim | Learned Latin | lat. ultimus |
| mod | Learned Latin | lat. modus |
| persoană | Learned Latin | lat. persona |
| dar | Unknown | conjuncția nu are o origine stabilită; slavonul darŭ explică substantivul dar, nu legătura dintre propoziții |
| parte | Inherited Latin | lat. partem, moștenit |
| joc | Inherited Latin | lat. jocus, moștenit |
| trimite | Inherited Latin | lat. tramittere, moștenit |
| politic | French | fr. politique; nota rusească se referă la un derivat de secol XX |
Reproduction
The public archive contains the code, processed data, research notes and files needed to rebuild the page and check its figures.
Download the reproduction archive SHA-256
Regenerating every corpus starts from external dexonline, Wikisource, Wikipedia and OpenSubtitles dumps. They are not redistributed in the archive; REPRODUCERE.md lists the sources, expected filenames and run order. The commands below become runnable after downloading and extracting the package.
python3 tools/fetch_categories.py python3 tools/dex_extract.py python3 tools/ws_index.py python3 tools/corpus_hist.py python3 tools/corpus_now.py python3 tools/official.py python3 tools/shortlist.py python3 tools/fetch_entries.py python3 tools/dex_check.py python3 tools/audit_top.py python3 tools/analyse.py node build.mjs node qa_assertions.mjs