Method
How we classified the words, which texts we counted and where the analysis stops.
Assigning the strata
A word's stratum is normally the language from which Romanian borrowed it directly. Bulevard is placed under French even though French had borrowed it earlier from Dutch.
Words that come from older Slavic forms need more caution. Wiktionary separates Proto-Slavic from Church Slavonic, but an old form alone does not show whether Romanian borrowed a word through speech or writing. We group both under Early Slavic and preserve the source's wording in the word's history.
We classified 47,657 dictionary words (lemmas) using Wiktionary's origin categories and articles. The program reads the first sentence that names a source language. Alternative hypotheses stay separate: joining them would create an etymological route that the source does not support.
When the source article gives an explicit gloss for a link in the chain, the word card places it under the source form. This is a gloss supplied by Wiktionary, not an automatic translation; when it is absent, the page does not invent a meaning.
For 14,550 published words, we checked the result against the notes about word origins in the dictionaries collected by dexonline. The two sources indicate the same stratum in 88.0 per cent of cases and the same source language in 92.0 per cent. Differences occur especially in nineteenth-century borrowings, where scholars may identify Latin or French as the direct source. Each word entry flags these cases.
The two hundred most frequent dictionary words make up nearly half the text, so one wrong classification can visibly change the percentages. Wiktionary uses one article for words that are spelled alike but have different histories, and the program reads the first section about origin. The 36 cases where the frequent sense was not the one in that section were corrected by hand and are listed below.
Historical texts: dated prose, 1825–1949
We counted 6,626 Romanian texts from Wikisource, with 10,998,608 words, grouped into five 25-year intervals. Each text gets the year shown on its page. If no year is given, we estimate it from the author's active period: birth year plus thirty-five years, capped at the year of death.
Three precautions. We exclude translations because a foreign author's dates do not tell us when the Romanian text was written. No author may contribute more than 2,000,000 characters to one interval, so one author cannot dominate the result. Repeated paragraphs are counted once because Wikisource keeps both the parent page and its chapters.
The first interval, 1825–1849, has the least data: 369,815 words from 18 authors, with the most prolific supplying 30.8 per cent of the text. Little prose was printed in Romanian before 1850, which limits the comparison.
Contemporary texts: two kinds of language
OpenSubtitles 2018: 306,494,203 words of Romanian-subtitled film dialogue. Wikipedia: 15,566,007 words from 41,357 Romanian articles, taken in the order in which they appear in the source file.
We analyse them separately because they represent two different kinds of language.
How we join forms of the same word
The same word appears in several forms in a text: casă, casa, casei, case, casele, caselor. We bring them to the dictionary form before counting frequencies. This process is called lemmatisation.
We connect forms to dictionary words using the 2,273,343 inflected forms in the dexonline database. We also bring spelling variants together: the cedilla instead of the comma, î instead of â, nineteenth-century diacritics and texts written without diacritics, including the constitutions of 1952 and 1965.
If one form could belong to several dictionary words, we divided its occurrences between them according to the frequency of their unambiguous forms. That is why si, written without diacritics, ends up almost entirely with și rather than with the musical note.
Very frequent grammatical words
The main analysis counts every occurrence, including prepositions, articles, pronouns, conjunctions and other very frequent grammatical words. We also make a version that removes them to see how much the result changes.
We make the decision for the dictionary word, not for each occurrence. Some forms have both a grammatical role and a meaning of their own; when the roles cannot be separated safely, we keep or remove the form everywhere. This tests the sensitivity of the result; it is not a full grammatical analysis.
We remove 242 dictionary words. They account for 50.9 per cent of the occurrences analysed in OpenSubtitles and 42.1 per cent in Wikipedia. After filtering, 83.1 per cent of the remaining OpenSubtitles occurrences and 73.0 per cent in Wikipedia can be linked to a stratum.
We reviewed the first 200 forms, together with the exclusions and ambiguous forms that were kept, in descending order of their combined frequency in the two sources.
The cautious version keeps forms with several roles when those roles cannot be separated. In this version, inherited Latin accounts for 57.7 per cent in OpenSubtitles and 28.9 per cent in Wikipedia, while French accounts for 21.5 per cent and 48.6 per cent.
The extended version also removes putea, tot, mult, doar, chiar and cât. With 248 forms removed, inherited Latin reaches 55.1 per cent in OpenSubtitles and 27.6 per cent in Wikipedia, while French reaches 22.8 per cent and 49.5 per cent.
The complete inventory is published in data/function_lemmas.csv, while the frequency order, review flag and both variants appear in the audit report.
For each set of texts, percentages use the remaining occurrences that can be linked to a stratum. The filter is not applied to the historical statistics or constitutional texts.
What the Albanian parallel measures — and what it cannot measure
The count of Albanian-linked lemmas is calculated only over the site's public dictionary. The build searches for the language codes sq and aln in each etymological chain and sums those lemmas' frequencies in OpenSubtitles and Wikipedia.
The 24.3 per cent and 28.4 per cent figures use the mass of all searchable lemmas classified as substrate in the relevant corpus as their denominator. They are not percentages of the whole corpus and do not include lemmas excluded from public search.
A chain containing an Albanian form does not by itself establish whether Romanian borrowed from Albanian, Albanian from Romanian, or both from an older source. Likewise, the dataset does not encode the Italian dialect isoglosses invoked by Dan Ungureanu: those are comparisons among related forms, not donor languages.
Licensing the dexonline data
The inflected forms and dictionary words used to group the text come from dexonline's official export dataset. The processed and published data derived from it retain the terms of GNU GPL v2 or any later version. Copyright © 2004–2026 dexonline, https://dexonline.ro.
The original database files and definition texts are not redistributed. The public archive includes only processed links between forms and dictionary words, aggregate statistics and flags for disagreements about word origins. You can see the attribution, licence links and exact inventory in the public DEXONLINE-NOTICE.md file.
What the page cannot say
First appearance in the counted texts is not a word's first documented appearance. The earliest attestations in the history of Romanian are found in the Dicționarul limbii române and cannot be extracted automatically from the sources used here. The page shows only that the word was already in use in the first interval where it appears; it may have existed earlier.
A word gets its own graph only for the 9,385 words with at least twelve appearances in the dated prose, spread across at least three intervals. For the rest, the graph for its stratum is shown.
Between 1949 and today's corpora there is a fifty-year hole, which the official texts cover only in their own register. The charts draw no continuous line across it.
Wikipedia shows one kind of encyclopedic writing, not written language in general. Film dialogue is scripted or translated, not recorded speech. The page says explicitly what each source represents.
Stratum percentages are given as a share of the text whose donor could be named: 87.4 per cent of the OpenSubtitles list and 78.9 per cent of the Wikipedia articles analysed. For the remainder, Wiktionary's categories did not provide an origin clear enough to use.
The cautious version does not check the grammatical role of every occurrence. A form with several roles stays included or excluded everywhere, according to the public inventory.
Where dictionaries part company
For 14,550 published words, the dictionaries aggregated by dexonline also contain an etymological note. Classification agrees at stratum level in 88.0 per cent of cases; in 1,748 cases, the sources choose a different direct donor.
The differences are not randomly distributed. Most involve modern words that can be analysed either as direct Latin borrowings or as arriving through French, Italian or German, and the boundary between Early Slavic and later neighbouring Slavic languages.
There is also an internal Wiktionary alignment issue: for 803 of the 17,694 automatically checkable records, the category-assigned stratum does not match the first recognisable stage of the published chain. The usual cause is a single entry combining homonyms or competing etymologies. The interface marks these records with ?. They are not corrected automatically without sense-level disambiguation.
Learned Latin→Latin 403 cases
Examples abecedar abrevia acvilin ad-hoc admirabil alias ambiguu audient aulă aură auspiciu baptisteriu
French→Latin 238 cases
Examples abroga absolut act acvatic adjudeca afabil alega altitudine audio august aulic benign
Early Slavic→The later Slavic neighbours 180 cases
Examples babă baniță baștină bivoliță bleg bolovan borș boz bragă braniște breaz bujor
German→French 94 cases
Examples adresă africată automobil bandă bloc boxer cadavru clasă hinduism idiș infarct mașină
Italian→Latin 64 cases
Examples acut acvariu acvilă amic amiciție genera glorie incipient infern insuficient invada investiga
French→Italian 61 cases
Examples acont adio bancar bariton bas basorelief capitul caramel carmin cicerone ciment guelf
Learned Latin→French 51 cases
Examples abdomen administrator admira apex canicular cedru chimie dativ genitiv gratis gravitate imaginație
Inherited Latin→Unknown 50 cases
Examples acolea adică adulmeca borî brâncă buf buiestru butuc funigel gaie genune ghioc
German→Latin 39 cases
Examples abstract chirurg ie insurgent matroană personal reumatic tabelă tort pacient paralel parazit
French→Greek 37 cases
Examples acrostih aed caligrafie canapea gradat granitic granular gregorian icter pelican manie analog
Corrections checked by hand
These are frequent words for which the program chose the rare sense of a word spelled the same way. The reason for each correction is recorded.
| Word | Stratum | Reason |
|---|---|---|
| mai | Inherited Latin | lat. magis; Wiktionary listează mai întâi sensurile «ciocan» și «lună», dar forma frecventă este adverbul comparativ |
| cap | Inherited Latin | lat. caput, moștenit; prima secțiune descrie sensul nautic, împrumutat din franceză |
| duce | Inherited Latin | lat. ducere; substantivul italian duce e omonimul rar |
| prin | Inherited Latin | per + in, amândouă latine |
| prim | Learned Latin | lat. primus, împrumut savant de secol XIX, nu maghiar |
| secol | Learned Latin | lat. saeculum |
| ultim | Learned Latin | lat. ultimus |
| mod | Learned Latin | lat. modus |
| persoană | Learned Latin | lat. persona |
| dar | Unknown | conjuncția nu are o origine stabilită; slavonul darŭ explică substantivul dar, nu legătura dintre propoziții |
| parte | Inherited Latin | lat. partem, moștenit |
| joc | Inherited Latin | lat. jocus, moștenit |
| trimite | Inherited Latin | lat. tramittere, moștenit |
| politic | French | fr. politique; nota rusească se referă la un derivat de secol XX |
| fiecare | Internal formation | format în română din fie + care |
| întrece | Internal formation | format în română din în- + trece |
| scară | Inherited Latin | lat. scala; mențiunea franceză din dexonline explică unele sensuri după fr. échelle, nu originea cuvântului |
| vorbi | Internal formation | derivat în română din vorbă + -i; forma «vorbit» este participiul acestui verb |
| păi | Inherited Latin | interjecție moștenită din lat. post; articolul o lasă în afara categoriilor etimologice folosite la inventar |
| film | French | fr. film; clasificarea automată a păstrat și categoria germană, dar lanțul și DEX indică donatorul francez |
| mașină | French | fr. machine; germ. Maschine este etimon concurent, nu clasificarea proximă folosită aici |
| motiv | French | fr. motif; italiana și germana apar ca surse concurente în notele etimologice |
| telefon | French | fr. téléphone; germana apare ca etimon concurent în datele automate |
| corp | French | fr. corps pentru sensul modern; lat. corpus și germana apar ca trasee concurente |
| cameră | Italian | it. camera; clasificarea automată o trimitea greșit la engleză |
| scuză | Italian | it. scusa; franceza este menționată ca paralelă pentru excuse |
| profesor | French | fr. professeur; germ. Professor este etimon concurent |
| rezultat | French | fr. résultat; germ. Resultat este etimon concurent |
| birou | French | fr. bureau; rusa apare ca variantă concurentă, dar lanțul și DEX indică franceza |
| doctor | Learned Latin | lat. doctor; franceza și germana sunt forme concurente, iar clasificarea germană era o alegere automată greșită |
| spital | German | germ. Spital pentru forma română standard; greaca și latina apar în trasee alternative |
| religie | French | fr. religion; latina și germana apar ca surse concurente |
| ofițer | French | fr. officier; rusa și germana sunt etimoane concurente |
| bancă | French | fr. banc pentru sensul publicat; italiana explică un alt sens și o analiză paralelă |
| organizație | French | fr. organisation; germ. Organisation este etimon concurent |
| real | French | fr. réel în uzul modern; germana și latina apar în trasee concurente |
Rebuilding the analysis
The public archive includes the code, processed data, research notes and files needed to rebuild the page and check its figures.
Download the archive for rebuilding the analysis SHA-256
To rebuild the analysis from source, download the complete files separately from dexonline, Wikisource, Wikipedia and OpenSubtitles. They are not redistributed in the archive. REPRODUCERE.md lists the sources, expected filenames and order of steps. After downloading them, the commands below can be run in the stated order.
python3 tools/fetch_categories.py python3 tools/dex_extract.py python3 tools/ws_index.py python3 tools/corpus_hist.py python3 tools/corpus_now.py python3 tools/official.py python3 tools/shortlist.py python3 tools/fetch_entries.py python3 tools/dex_check.py python3 tools/audit_top.py python3 tools/function_filter.py python3 tools/analyse.py python3 tools/explore.py node build.mjs node qa_assertions.mjs