Romanian is the only Romance language that grew up surrounded by Slavs, under Ottoman suzerainty for centuries, first written in Church Slavonic and in Cyrillic letters, administered in turn from Buda and Vienna, schooled in Greek, and then deliberately rebuilt on a French model. Everything that passed through left words behind, and the words can still be sorted by who brought them.
See a sentence's layers
Type or paste Romanian text. The page lemmatises it, looks up each origin and shows how the composition changes when you count every occurrence, each lemma once, or remove very frequent grammatical words.
The analysis runs in your browser. The text is neither sent nor stored. If a form can belong to more than one lemma, it is marked ambiguous rather than guessed from context.
What are you counting?
Type a sentence to reveal its strata.
Three counts, three questions
The result depends on what is counted. Of the 47,657 lexemes whose origin can be classified, 59.9 per cent come from French and 4.7 per cent from inherited Latin. In the OpenSubtitles frequency list, counting every occurrence reverses the proportions: inherited Latin accounts for 77.7 per cent of classified occurrences and French for 10.0 per cent.
The inventory statistics use all 47,657 lemmas whose origin can be classified. Public search includes 18,721 of them, those occurring at least 0.5 per million words in the contemporary corpora.
That second count includes the function words repeated through nearly every sentence: prepositions, articles and determiners, pronouns, conjunctions, auxiliaries and particles. Much of Romanian’s highest-frequency grammatical scaffolding is inherited from Latin.
To estimate the influence of very frequent grammatical words, we remove 242 lemmas and recalculate the proportions over the remaining occurrences. This is a conservative lemma-level filter. Grammatical and lexical uses of the same lemma are not separated sentence by sentence.
In the unfiltered view, a stratum could be assigned to 87.4 per cent of the OpenSubtitles mass and 78.9 per cent of the Wikipedia mass analysed.
Classified distinct lemmasAll occurrences in dialogueDialogue · conservative filter
French59.9%10.0%21.5%
Early Slavic5.1%3.5%7.6%
Inherited Latin4.7%77.7%57.7%
Hungarian4.5%0.4%0.8%
The later Slavic neighbours4.4%0.7%1.4%
Ottoman Turkish3.5%0.4%0.8%
German3.1%0.5%1.0%
Greek2.6%0.5%1.2%
The conservative filter removes 242 lemmas, accounting for 50.9 per cent of the occurrences analysed. Percentages are recalculated over the remaining occurrences whose stratum can be classified.
The strata sit in the order they were laid down, oldest at the bottom. Band thickness represents each stratum's share. Choose the complete inventory, film dialogue or Wikipedia articles. For the two texts, you can view the result with or without very frequent grammatical words.
What are you comparing?
Very frequent grammatical words
This choice applies to the texts, not to the complete inventory.
Classified distinct lemmas
Unknown2.8%
Internal formation2.4%
English2.3%
French59.9%
Italian2.0%
Learned Latin2.5%
German3.1%
The later Slavic neighbours4.4%
Ottoman Turkish3.5%
Greek2.6%
Hungarian4.5%
Early Slavic5.1%
Inherited Latin4.7%
The substrate0.3%
Where the lemmas in the complete inventory come from. Each lemma is counted once.
Turkish words stayed in the kitchen
Turkish words are 1.9× more frequent in film dialogue than in Wikipedia articles. French words are 2.6× more frequent in Wikipedia articles than in film dialogue.
Turkish words entered mainly through trade and the household: names for food, household objects and colloquial expressions. French words entered through school, administration and law. This helps explain why the former occur more often in film dialogue and the latter in Wikipedia articles.
more in film dialoguemore in Wikipedia articles
The substrate1.33×
Inherited Latin1.28×
Early Slavic1.23×
Hungarian0.58×
Greek0.59×
Ottoman Turkish1.85×
The later Slavic neighbours1.09×
German0.48×
Learned Latin0.55×
Italian0.61×
French0.39×
English0.81×
Internal formation1.14×
Unknown4.11×
How concentrated is each stratum?
Some strata get half their occurrences from a handful of words; others need hundreds. In dialogue, half of inherited Latin’s occurrences come from just 14 words; French needs 282. In Wikipedia, French reaches halfway after 322 words.
The chart shows how many words from the public inventory are needed to reach half of a stratum’s occurrences. After removing very frequent grammatical words, inherited Latin in dialogue needs 38 words.
Choose the view
Dialogue · all words
A small number means a concentrated stratum: a few very frequent words account for much of its occurrences. The count uses the 18,721 words in the public inventory.
Which words contribute most to each stratum?
Aggregate percentages become easier to understand when you see the words that produce them. Choose a stratum and a source: the list shows the words that contribute most to its occurrences.
Choose the view
Choose the stratum
Frequencies are per million words. The percentage on the right shows how much of the stratum’s occurrences, in the selected source, is contributed by each word.
Direct donors and earlier stages
A page stratum normally identifies the language from which Romanian received a word directly. The chain can continue further back. We count only chains whose first stage confirms the assigned stratum. Of the 11,029 chains beginning with French, 2,061 record at least one earlier stage; 1,800 continue towards Latin.
Among the chains beginning with Ottoman Turkish, 262 of 536 continue to an earlier stage, including 105 towards Arabic and 105 towards Persian.
Choose the direct donor
These are stages published in the etymological chains, not a complete map of ultimate origin. Chains whose first stage does not match the assigned stratum are excluded, usually because one entry combines homonyms or competing analyses. Sources stop at different depths.
Two incompatible hypotheses about the formation of Romanian
Dan Alexe and Dan Ungureanu are often invoked together, but they do not tell the same story. Both reject an exclusively north-Danubian formation; from there, their routes diverge.
Dan Alexe
Balkan Latinity → structural contact with Albanian and South Slavic → movement north of the Danube
Romanian formed in the central Balkans through prolonged linguistic convergence with Albanian and South Slavic, and its speakers moved north of the Danube roughly between 800 and 1000.
What the argument centres on
Balkan structure and typology: the postposed article, Balkan subjunctive constructions, and Romanian–Albanian grammatical and lexical correspondences.
dialectalised Cisalpine Latin → Romania continua, probably the Balkans, after 400 → Romanian
Romanian's Romance ancestor carried a cluster of phonetic, morphological and lexical innovations concentrated in north-western Italy. Romanian itself, however, could only emerge later in a continuous Latin-speaking part of the Empire, probably the Balkans, after 400.
What the argument centres on
The mapped overlap between Romanian features and forms in western Lombardy, Piedmont and the Alps, together with late Romance vocabulary.
The words open in the site's current classification. Their presence here does not mean that the dataset adopts either author's interpretation.
What this dataset can see
Of the 18,721 searchable lemmas, 44 have an Albanian or Gheg Albanian stage in their published etymological chain; all are classified here as substrate. They account for 24.3 per cent of the searchable substrate mass in dialogue and 28.4 per cent in Wikipedia.
In dialogue, bucura alone accounts for 56.6 per cent of the mass of this 44-lemma group. This concentration matters: the size of a group and its weight are not the same thing.
This shows that the site's etymological sources contain a measurable Albanian-linked nucleus within the substrate vocabulary. It does not establish the direction of borrowing, a common source, the location of contact or Alexe's chronology.
The dataset can make Ungureanu's examples inspectable, but it cannot test his geographic hypothesis: these chains record donor languages, while his argument relies on dialect parallels and isoglosses rather than borrowings from Lombard or Piedmontese.
Alexe explicitly says that his Balkan model and Ungureanu's north-Italian model are incompatible, and faults Ungureanu for putting lexical resemblance ahead of grammatical structure.
Her 2018 review of Româna și dialectele italiene challenges the method, dialect attribution and absence of a convincing synthesis, and concludes that the proposed theses are not demonstrated.
His 2024 review calls the inventory of local dialect descriptions and dictionaries an acceptable method for a grounded redesign of Romanian language history.
In a public critique from 2024, the linguist argues that both authors turn difficult, weakly testable uncertainties into claims more definite than the sources and comparative-historical method allow.
The literary prose in the corpus stops in 1949, because copyright still protects more recent works. The next period uses nine public, dated constitutional texts from 1848 to 2003. They allow the official vocabulary to be compared across the full period.
The Early Slavic share falls from 3.5 per cent in 1848 to 0.9 per cent in 1866. The French share climbs from 8.2 per cent to 36.2 per cent and peaks under communism. Four centuries of suzerainty left hundreds of words from the market and household, but almost nothing in the constitutional vocabulary measured here.
FrenchInherited LatinEarly SlavicThe later Slavic neighbours
Percent of the words whose donor can be named
View as table
Year
Document
French
Inherited Latin
Early Slavic
The later Slavic neighbours
1848
Proclamația de la Islaz
8.2%
81.6%
3.5%
0.6%
1866
Constituția României
19.8%
71.4%
0.9%
0.2%
1923
Constituția României
24.3%
63.9%
1.1%
0.1%
1938
Constituția României
24.8%
63.1%
1.3%
0.1%
1948
Constituția Republicii Populare Române
34.0%
52.9%
2.3%
0.4%
1952
Constituția Republicii Populare Române
37.6%
49.4%
3.7%
1.0%
1965
Constituția Republicii Socialiste România
37.4%
48.5%
2.7%
1.1%
1991
Constituția României
36.0%
52.3%
1.5%
0.3%
2003
Constituția României, revizuită
36.2%
52.6%
1.2%
0.2%
The fourteen strata
The cards below include very frequent grammatical words. You can remove them only in the comparison above.
The substrate
before AD 106 · Whoever was here before Rome
The only stratum whose donor has no name, because the language these words come from was never written down. A word lands here by elimination: Latin cannot explain it, Slavic cannot explain it, and Albanian has a relative for it, since Albanian rests on the same pre-Roman substrate. There are about a hundred and twenty in the whole language, and they cluster where a shepherd works: landforms, trees, parts of an animal, the tools of a sheepfold. Since the proof is negative, disagreement is the norm: for 24 of the words here, dexonline simply says the origin is unknown.
Words
126
Film dialogue
0.3%
Wikipedia
0.2%
Most frequent
Inherited Latin
106 – seventh century · Rome, then a thousand years of talking
There are relatively few inherited Latin entries in the dictionary, but they account for 77.7 per cent of classified words in film dialogue and 60.8 per cent of those in Wikipedia articles. They were transmitted through speech: aqua became apă, noctem became noapte, oculus became ochi. The surprise is that the core Christian vocabulary is Latin rather than Slavonic: biserică comes from basilica and Dumnezeu from Domine Deus; cruce, înger and a boteza also have Latin origins. Slavic influence came later.
Words
2,238
Film dialogue
77.7%
Wikipedia
60.8%
Most frequent
Early Slavic
sixth to seventeenth centuries · Direct contact, then church and chancery
Dictionaries use labels including Proto-Slavic, Old Slavic and Church Slavonic. Church Slavonic was the written language of the church and chancery, based on an early South Slavic variety. The recorded form alone cannot show whether Romanian borrowed a word through speech or through writing, so this analysis counts them together. The layer includes everyday words such as nevoie, iubi, prieten, trăi and dragoste, as well as religious and administrative terms such as rai, duh, pravilă and hrisov. In the official texts analysed, its share is 3.5 per cent in 1848 and 0.9 per cent in 1866.
Words
2,422
Film dialogue
3.5%
Wikipedia
2.9%
Most frequent
Hungarian
eleventh to nineteenth centuries · Transylvania, eight centuries of neighbours and administration
The word for city is Hungarian and the word for village is Latin, which says plainly enough where power was kept and where life was kept. The stratum gave gând, chip, fel, marfă, hotar, meșter. It also gave a tool: the verbal suffix -ui, with which Romanian has been manufacturing verbs ever since, including from roots that were never Hungarian. A locui is Hungarian and a chibzui is Hungarian, but the verb-making machine stayed in the language and now works on its own account.
Words
2,143
Film dialogue
0.4%
Wikipedia
0.6%
Most frequent
Greek
tenth century to 1821 · Byzantium, then the Phanariot princes
Two channels more than five centuries apart. Byzantium sent words through the church; then, between 1711 and 1821, the Porte appointed Greek princes from Istanbul's Phanar quarter to Bucharest and Iași, with an administration and a school system in Greek. A hundred and ten years of Greek rule left a thinner stratum than anyone would guess: 2.6 per cent of the lexemes with a known donor. Like Hungarian, it left a tool behind, the ending -isi, out of which Romanian built a chivernisi and a economisi.
Words
1,224
Film dialogue
0.5%
Wikipedia
0.9%
Most frequent
Ottoman Turkish
fifteenth to nineteenth centuries · Trade and the household
Four centuries of Ottoman suzerainty left cafea, ciorbă, iaurt, dulap, cearșaf, geam, tavan, papuc, mahala, bacșiș, chef and haz. The Porte took tribute and appointed princes without directly administering Wallachia and Moldavia. The political relationship also left historical terms such as firman and pașalâc, while most of the words entered through trade and the household. In film dialogue they are 1.9× more frequent than in Wikipedia articles, and across the nine constitutional texts measured, from 1848 to 2003, they are effectively absent.
Words
1,652
Film dialogue
0.4%
Wikipedia
0.2%
Most frequent
The later Slavic neighbours
sixteenth to twentieth centuries · Bulgarians, Serbs, Ukrainians, Russians, Poles
Borders first, then empires. Russian arrives twice, with the occupation and administration after 1828 and again with the Soviet decades, and it is the only stratum in which communism is visible in the figures: in official text it climbs from 0.1 per cent in the 1938 constitution to 1.1 per cent in that of 1965, then falls back.
Words
2,116
Film dialogue
0.7%
Wikipedia
0.6%
Most frequent
German
twelfth to twentieth centuries · The Saxons, the Habsburg administration, the workshop
Saxons reach Transylvania from the twelfth century, Austrian administration after 1699, and technical vocabulary with industry. From here come șurub, ștecher, bormașină, chelner, șpriț, cartof and rucsac. German words appear more often in Wikipedia articles than in film dialogue.
Words
1,463
Film dialogue
0.5%
Wikipedia
1.0%
Most frequent
Learned Latin
1780–1900 · Scholars, on paper and on purpose
Unlike the Latin words inherited through speech, learned Latinisms were taken from books and introduced deliberately. The Transylvanian School used the language's Latin descent as a political argument, and later Timotei Cipariu and the Laurian–Massim dictionary of 1871–1876 tried to re-Latinise the vocabulary to the point of absurdity and were mocked for it. What survives are the doublets: the same Latin word arriving twice, with different meanings. From digitus, deget and digital; from gravis, greu and grav; from directus, drept and direct; from frigidus, frig and frigid. In each pair, the first word was inherited and changed through use; the second is the learned borrowing.
Words
1,191
Film dialogue
1.9%
Wikipedia
3.4%
Most frequent
Italian
1800–1930 · Music, art, banking
A small stratum, easy to recognise by domain: the score, the stage, the bank. The generation of 1848, educated in Italy and inclined to see Italian as an elder sister of the language, also brought in part of the political vocabulary of the moment.
Words
961
Film dialogue
0.8%
Wikipedia
1.3%
Most frequent
French
1830–1918 · Education, government and the prestige of Paris
French influence was adopted by Romanian elites in the nineteenth century. Young people studied in Paris, the civil code of 1864 followed the Napoleonic model, and the state was built to French patterns. The language adopted much of the vocabulary of politics, administration and abstraction along with them. Official documents show the scale of the change: French words account for 8.2 per cent of the 1848 proclamation, 19.8 per cent of the 1866 constitution, largely modelled on the Belgian one of 1831, and 36.2 per cent of the constitution now in force. The highest share appears in the 1952 Constitution, at 37.6 per cent, because many administrative terms used under Romanian communism are French in origin: planificare, industrializare, colectivizare.
Words
28,556
Film dialogue
10.0%
Wikipedia
26.1%
Most frequent
English
from 1960, mostly after 1990 · Prestige and the screen
English borrowings grew especially after 1990 through technology, media and popular culture. The page's dictionary contains 1,099 English words, or 2.3 per cent of the total.
Words
1,099
Film dialogue
0.4%
Wikipedia
0.4%
Most frequent
Internal formation
from the sixteenth century · Romanian, on its own
Words created in Romanian from existing elements, through compounding or derivation. They appear when the language needs a new name.
Words
1,121
Film dialogue
1.2%
Wikipedia
1.0%
Most frequent
Unknown
no date · Nobody can say
No established origin exists for 1,345 lexemes. Their share in film dialogue is strongly affected by one very frequent word, the conjunction dar.