The Layers of a Word

Stratigraphy

The Layers of a Word

Where 18,721 Romanian words come from

Romanian is the only Romance language that grew up surrounded by Slavs, under Ottoman suzerainty for centuries, first written in Church Slavonic and in Cyrillic letters, administered in turn from Buda and Vienna, schooled in Greek, and then deliberately rebuilt on a French model. Everything that passed through left words behind, and the words can still be sorted by who brought them.

See a sentence's layers

Type or paste Romanian text. The page lemmatises it, looks up each origin and shows how the composition changes when you count every occurrence, each lemma once, or remove very frequent grammatical words.

The analysis runs in your browser. The text is neither sent nor stored. If a form can belong to more than one lemma, it is marked ambiguous rather than guessed from context.

What are you counting?

Type a sentence to reveal its strata.

Three counts, three questions

The result depends on what is counted. Of the 47,657 lexemes whose origin can be classified, 59.9 per cent come from French and 4.7 per cent from inherited Latin. In the OpenSubtitles frequency list, counting every occurrence reverses the proportions: inherited Latin accounts for 77.7 per cent of classified occurrences and French for 10.0 per cent.

The inventory statistics use all 47,657 lemmas whose origin can be classified. Public search includes 18,721 of them, those occurring at least 0.5 per million words in the contemporary corpora.

That second count includes the function words repeated through nearly every sentence: prepositions, articles and determiners, pronouns, conjunctions, auxiliaries and particles. Much of Romanian’s highest-frequency grammatical scaffolding is inherited from Latin.

To estimate the influence of very frequent grammatical words, we remove 242 lemmas and recalculate the proportions over the remaining occurrences. This is a conservative lemma-level filter. Grammatical and lexical uses of the same lemma are not separated sentence by sentence.

In the unfiltered view, a stratum could be assigned to 87.4 per cent of the OpenSubtitles mass and 78.9 per cent of the Wikipedia mass analysed.

Classified distinct lemmasAll occurrences in dialogueDialogue · conservative filter
French 59.9% 10.0% 21.5%
Early Slavic 5.1% 3.5% 7.6%
Inherited Latin 4.7% 77.7% 57.7%
Hungarian 4.5% 0.4% 0.8%
The later Slavic neighbours 4.4% 0.7% 1.4%
Ottoman Turkish 3.5% 0.4% 0.8%
German 3.1% 0.5% 1.0%
Greek 2.6% 0.5% 1.2%

The conservative filter removes 242 lemmas, accounting for 50.9 per cent of the occurrences analysed. Percentages are recalculated over the remaining occurrences whose stratum can be classified.

How the proportions change

The strata sit in the order they were laid down, oldest at the bottom. Band thickness represents each stratum's share. Choose the complete inventory, film dialogue or Wikipedia articles. For the two texts, you can view the result with or without very frequent grammatical words.

What are you comparing?

This choice applies to the texts, not to the complete inventory.

Classified distinct lemmas

  1. Unknown 2.8%
  2. Internal formation 2.4%
  3. English 2.3%
  4. French 59.9%
  5. Italian 2.0%
  6. Learned Latin 2.5%
  7. German 3.1%
  8. The later Slavic neighbours 4.4%
  9. Ottoman Turkish 3.5%
  10. Greek 2.6%
  11. Hungarian 4.5%
  12. Early Slavic 5.1%
  13. Inherited Latin 4.7%
  14. The substrate 0.3%
Where the lemmas in the complete inventory come from. Each lemma is counted once.

Turkish words stayed in the kitchen

Turkish words are 1.9× more frequent in film dialogue than in Wikipedia articles. French words are 2.6× more frequent in Wikipedia articles than in film dialogue.

Turkish words entered mainly through trade and the household: names for food, household objects and colloquial expressions. French words entered through school, administration and law. This helps explain why the former occur more often in film dialogue and the latter in Wikipedia articles.

more in film dialoguemore in Wikipedia articles
The substrate 1.33×
Inherited Latin 1.28×
Early Slavic 1.23×
Hungarian 0.58×
Greek 0.59×
Ottoman Turkish 1.85×
The later Slavic neighbours 1.09×
German 0.48×
Learned Latin 0.55×
Italian 0.61×
French 0.39×
English 0.81×
Internal formation 1.14×
Unknown 4.11×

How concentrated is each stratum?

Some strata get half their occurrences from a handful of words; others need hundreds. In dialogue, half of inherited Latin’s occurrences come from just 14 words; French needs 282. In Wikipedia, French reaches halfway after 322 words.

The chart shows how many words from the public inventory are needed to reach half of a stratum’s occurrences. After removing very frequent grammatical words, inherited Latin in dialogue needs 38 words.

Choose the view

Dialogue · all words

A small number means a concentrated stratum: a few very frequent words account for much of its occurrences. The count uses the 18,721 words in the public inventory.

Which words contribute most to each stratum?

Aggregate percentages become easier to understand when you see the words that produce them. Choose a stratum and a source: the list shows the words that contribute most to its occurrences.

Choose the view

Choose the stratum

    Frequencies are per million words. The percentage on the right shows how much of the stratum’s occurrences, in the selected source, is contributed by each word.

    Direct donors and earlier stages

    A page stratum normally identifies the language from which Romanian received a word directly. The chain can continue further back. We count only chains whose first stage confirms the assigned stratum. Of the 11,029 chains beginning with French, 2,061 record at least one earlier stage; 1,800 continue towards Latin.

    Among the chains beginning with Ottoman Turkish, 262 of 536 continue to an earlier stage, including 105 towards Arabic and 105 towards Persian.

    Choose the direct donor

    These are stages published in the etymological chains, not a complete map of ultimate origin. Chains whose first stage does not match the assigned stratum are excluded, usually because one entry combines homonyms or competing analyses. Sources stop at different depths.

    Two incompatible hypotheses about the formation of Romanian

    Dan Alexe and Dan Ungureanu are often invoked together, but they do not tell the same story. Both reject an exclusively north-Danubian formation; from there, their routes diverge.

    Dan Alexe

    Balkan Latinity → structural contact with Albanian and South Slavic → movement north of the Danube

    Romanian formed in the central Balkans through prolonged linguistic convergence with Albanian and South Slavic, and its speakers moved north of the Danube roughly between 800 and 1000.

    What the argument centres on

    Balkan structure and typology: the postposed article, Balkan subjunctive constructions, and Romanian–Albanian grammatical and lexical correspondences.

    Words discussed by Alexe

    Official book preview

    Dan Ungureanu

    dialectalised Cisalpine Latin → Romania continua, probably the Balkans, after 400 → Romanian

    Romanian's Romance ancestor carried a cluster of phonetic, morphological and lexical innovations concentrated in north-western Italy. Romanian itself, however, could only emerge later in a continuous Latin-speaking part of the Empire, probably the Balkans, after 400.

    What the argument centres on

    The mapped overlap between Romanian features and forms in western Lombardy, Piedmont and the Alps, together with late Romance vocabulary.

    Examples used in his comparisons

    Official book preview

    The words open in the site's current classification. Their presence here does not mean that the dataset adopts either author's interpretation.

    What this dataset can see

    Of the 18,721 searchable lemmas, 44 have an Albanian or Gheg Albanian stage in their published etymological chain; all are classified here as substrate. They account for 24.3 per cent of the searchable substrate mass in dialogue and 28.4 per cent in Wikipedia.

    In dialogue, bucura alone accounts for 56.6 per cent of the mass of this 44-lemma group. This concentration matters: the size of a group and its weight are not the same thing.

    This shows that the site's etymological sources contain a measurable Albanian-linked nucleus within the substrate vocabulary. It does not establish the direction of borrowing, a common source, the location of contact or Alexe's chronology.

    The dataset can make Ungureanu's examples inspectable, but it cannot test his geographic hypothesis: these chains record donor languages, while his argument relies on dialect parallels and isoglosses rather than borrowings from Lombard or Piedmontese.

    How the dispute was received

    Vocabulary in official texts, 1848–2003

    The literary prose in the corpus stops in 1949, because copyright still protects more recent works. The next period uses nine public, dated constitutional texts from 1848 to 2003. They allow the official vocabulary to be compared across the full period.

    The Early Slavic share falls from 3.5 per cent in 1848 to 0.9 per cent in 1866. The French share climbs from 8.2 per cent to 36.2 per cent and peaks under communism. Four centuries of suzerainty left hundreds of words from the market and household, but almost nothing in the constitutional vocabulary measured here.

    French
    40.0% 0 1848: 8.2%1866: 19.8%1923: 24.3%1938: 24.8%1948: 34.0%1952: 37.6%1965: 37.4%1991: 36.0%2003: 36.2% 1848192319652003
    Inherited Latin
    90.0% 0 1848: 81.6%1866: 71.4%1923: 63.9%1938: 63.1%1948: 52.9%1952: 49.4%1965: 48.5%1991: 52.3%2003: 52.6% 1848192319652003
    Early Slavic
    4.0% 0 1848: 3.5%1866: 0.9%1923: 1.1%1938: 1.3%1948: 2.3%1952: 3.7%1965: 2.7%1991: 1.5%2003: 1.2% 1848192319652003
    The later Slavic neighbours
    2.0% 0 1848: 0.6%1866: 0.2%1923: 0.1%1938: 0.1%1948: 0.4%1952: 1.0%1965: 1.1%1991: 0.3%2003: 0.2% 1848192319652003

    Percent of the words whose donor can be named

    View as table
    YearDocumentFrenchInherited LatinEarly SlavicThe later Slavic neighbours
    1848Proclamația de la Islaz8.2%81.6%3.5%0.6%
    1866Constituția României19.8%71.4%0.9%0.2%
    1923Constituția României24.3%63.9%1.1%0.1%
    1938Constituția României24.8%63.1%1.3%0.1%
    1948Constituția Republicii Populare Române34.0%52.9%2.3%0.4%
    1952Constituția Republicii Populare Române37.6%49.4%3.7%1.0%
    1965Constituția Republicii Socialiste România37.4%48.5%2.7%1.1%
    1991Constituția României36.0%52.3%1.5%0.3%
    2003Constituția României, revizuită36.2%52.6%1.2%0.2%

    The fourteen strata

    The cards below include very frequent grammatical words. You can remove them only in the comparison above.

    The substrate

    before AD 106 · Whoever was here before Rome

    The only stratum whose donor has no name, because the language these words come from was never written down. A word lands here by elimination: Latin cannot explain it, Slavic cannot explain it, and Albanian has a relative for it, since Albanian rests on the same pre-Roman substrate. There are about a hundred and twenty in the whole language, and they cluster where a shepherd works: landforms, trees, parts of an animal, the tools of a sheepfold. Since the proof is negative, disagreement is the norm: for 24 of the words here, dexonline simply says the origin is unknown.

    Words
    126
    Film dialogue
    0.3%
    Wikipedia
    0.2%

    Most frequent

    Inherited Latin

    106 – seventh century · Rome, then a thousand years of talking

    There are relatively few inherited Latin entries in the dictionary, but they account for 77.7 per cent of classified words in film dialogue and 60.8 per cent of those in Wikipedia articles. They were transmitted through speech: aqua became apă, noctem became noapte, oculus became ochi. The surprise is that the core Christian vocabulary is Latin rather than Slavonic: biserică comes from basilica and Dumnezeu from Domine Deus; cruce, înger and a boteza also have Latin origins. Slavic influence came later.

    Words
    2,238
    Film dialogue
    77.7%
    Wikipedia
    60.8%

    Most frequent

    Early Slavic

    sixth to seventeenth centuries · Direct contact, then church and chancery

    Dictionaries use labels including Proto-Slavic, Old Slavic and Church Slavonic. Church Slavonic was the written language of the church and chancery, based on an early South Slavic variety. The recorded form alone cannot show whether Romanian borrowed a word through speech or through writing, so this analysis counts them together. The layer includes everyday words such as nevoie, iubi, prieten, trăi and dragoste, as well as religious and administrative terms such as rai, duh, pravilă and hrisov. In the official texts analysed, its share is 3.5 per cent in 1848 and 0.9 per cent in 1866.

    Words
    2,422
    Film dialogue
    3.5%
    Wikipedia
    2.9%

    Most frequent

    Hungarian

    eleventh to nineteenth centuries · Transylvania, eight centuries of neighbours and administration

    The word for city is Hungarian and the word for village is Latin, which says plainly enough where power was kept and where life was kept. The stratum gave gând, chip, fel, marfă, hotar, meșter. It also gave a tool: the verbal suffix -ui, with which Romanian has been manufacturing verbs ever since, including from roots that were never Hungarian. A locui is Hungarian and a chibzui is Hungarian, but the verb-making machine stayed in the language and now works on its own account.

    Words
    2,143
    Film dialogue
    0.4%
    Wikipedia
    0.6%

    Most frequent

    Greek

    tenth century to 1821 · Byzantium, then the Phanariot princes

    Two channels more than five centuries apart. Byzantium sent words through the church; then, between 1711 and 1821, the Porte appointed Greek princes from Istanbul's Phanar quarter to Bucharest and Iași, with an administration and a school system in Greek. A hundred and ten years of Greek rule left a thinner stratum than anyone would guess: 2.6 per cent of the lexemes with a known donor. Like Hungarian, it left a tool behind, the ending -isi, out of which Romanian built a chivernisi and a economisi.

    Words
    1,224
    Film dialogue
    0.5%
    Wikipedia
    0.9%

    Most frequent

    Ottoman Turkish

    fifteenth to nineteenth centuries · Trade and the household

    Four centuries of Ottoman suzerainty left cafea, ciorbă, iaurt, dulap, cearșaf, geam, tavan, papuc, mahala, bacșiș, chef and haz. The Porte took tribute and appointed princes without directly administering Wallachia and Moldavia. The political relationship also left historical terms such as firman and pașalâc, while most of the words entered through trade and the household. In film dialogue they are 1.9× more frequent than in Wikipedia articles, and across the nine constitutional texts measured, from 1848 to 2003, they are effectively absent.

    Words
    1,652
    Film dialogue
    0.4%
    Wikipedia
    0.2%

    Most frequent

    The later Slavic neighbours

    sixteenth to twentieth centuries · Bulgarians, Serbs, Ukrainians, Russians, Poles

    Borders first, then empires. Russian arrives twice, with the occupation and administration after 1828 and again with the Soviet decades, and it is the only stratum in which communism is visible in the figures: in official text it climbs from 0.1 per cent in the 1938 constitution to 1.1 per cent in that of 1965, then falls back.

    Words
    2,116
    Film dialogue
    0.7%
    Wikipedia
    0.6%

    Most frequent

    German

    twelfth to twentieth centuries · The Saxons, the Habsburg administration, the workshop

    Saxons reach Transylvania from the twelfth century, Austrian administration after 1699, and technical vocabulary with industry. From here come șurub, ștecher, bormașină, chelner, șpriț, cartof and rucsac. German words appear more often in Wikipedia articles than in film dialogue.

    Words
    1,463
    Film dialogue
    0.5%
    Wikipedia
    1.0%

    Most frequent

    Learned Latin

    1780–1900 · Scholars, on paper and on purpose

    Unlike the Latin words inherited through speech, learned Latinisms were taken from books and introduced deliberately. The Transylvanian School used the language's Latin descent as a political argument, and later Timotei Cipariu and the Laurian–Massim dictionary of 1871–1876 tried to re-Latinise the vocabulary to the point of absurdity and were mocked for it. What survives are the doublets: the same Latin word arriving twice, with different meanings. From digitus, deget and digital; from gravis, greu and grav; from directus, drept and direct; from frigidus, frig and frigid. In each pair, the first word was inherited and changed through use; the second is the learned borrowing.

    Words
    1,191
    Film dialogue
    1.9%
    Wikipedia
    3.4%

    Most frequent

    Italian

    1800–1930 · Music, art, banking

    A small stratum, easy to recognise by domain: the score, the stage, the bank. The generation of 1848, educated in Italy and inclined to see Italian as an elder sister of the language, also brought in part of the political vocabulary of the moment.

    Words
    961
    Film dialogue
    0.8%
    Wikipedia
    1.3%

    Most frequent

    French

    1830–1918 · Education, government and the prestige of Paris

    French influence was adopted by Romanian elites in the nineteenth century. Young people studied in Paris, the civil code of 1864 followed the Napoleonic model, and the state was built to French patterns. The language adopted much of the vocabulary of politics, administration and abstraction along with them. Official documents show the scale of the change: French words account for 8.2 per cent of the 1848 proclamation, 19.8 per cent of the 1866 constitution, largely modelled on the Belgian one of 1831, and 36.2 per cent of the constitution now in force. The highest share appears in the 1952 Constitution, at 37.6 per cent, because many administrative terms used under Romanian communism are French in origin: planificare, industrializare, colectivizare.

    Words
    28,556
    Film dialogue
    10.0%
    Wikipedia
    26.1%

    Most frequent

    English

    from 1960, mostly after 1990 · Prestige and the screen

    English borrowings grew especially after 1990 through technology, media and popular culture. The page's dictionary contains 1,099 English words, or 2.3 per cent of the total.

    Words
    1,099
    Film dialogue
    0.4%
    Wikipedia
    0.4%

    Most frequent

    Internal formation

    from the sixteenth century · Romanian, on its own

    Words created in Romanian from existing elements, through compounding or derivation. They appear when the language needs a new name.

    Words
    1,121
    Film dialogue
    1.2%
    Wikipedia
    1.0%

    Most frequent

    Unknown

    no date · Nobody can say

    No established origin exists for 1,345 lexemes. Their share in film dialogue is strongly affected by one very frequent word, the conjunction dar.

    Words
    1,345
    Film dialogue
    1.8%
    Wikipedia
    0.4%

    Most frequent