Method & Sources

How the frequency model was built — and what it is not.

Step one: noun-type frequency. The share of each of the sixteen noun types comes from the grammatical notes accompanying Česká čítanka (Kořánová, 2012). This gives, for example, the hrad / obchod pattern at 29.3% and žena / káva at 23.1%.

Step two: case distribution. Case shares come from a table in Anna Plocková’s bachelor’s thesis (2016), calculated from the SYN2010 corpus — a 100-million-word body of fiction, technical literature and journalism.

Step three: the proxy. That table counts adjectives, not nouns. To split noun-type frequency by case, the model assumes the distribution of forms across cases and genders is broadly the same for nouns as for adjectives. Each case share is divided by its gender subtotal, then multiplied by the noun-type frequency.

This is why every number here is presented as an estimated learning-priority model, not a prediction of the Czech you personally will hear. It is reliable enough to tell you what to learn first, and no more precise than that.

Known limitations

  • Approximate rows. Six noun types fall below 1% and cannot be measured exactly. They are shown as “<1%”, and cumulative totals from milestone 6 onward are marked with ≈. Combined, those approximate rows can add up past 100% — so they are never presented as exact totals.
  • Adjectives as a proxy. Multiple adjectives can describe one noun, so absolute counts differ; the distribution across cases is what is borrowed, not the volume.
  • Singular only. The roadmap covers singular declension. The same approach is replicable for plurals.
  • Vocative excluded. The fifth case describes no relationship between things, so it sits outside the six-case route.
IThe model at a glance

Share of forms by case

Share of nouns by type

  • Mihrad / obchod29.3%
  • Fžena / káva23.1%
  • Frůže / práce8.3%
  • Mapán / kamarád7.4%
  • Nstavení / cvičení6.6%
  • Nměsto / auto6.2%
  • Fkost / bolest5%
  • Mamuž / otec3.6%
  • Mistroj / počítač3.3%
  • Fpíseň / garáž2.8%
  • Mapředseda / kolega1%
  • Masoudce / průvodce<1%
  • Nmoře / ovoce<1%
  • Nkuře / dítě<1%
  • Ncentrum / centrum<1%
  • Ntéma / téma<1%
IIEditorial notes

Language review. Every sample sentence, trigger list and generated story here is intended for review by a native Czech linguist before wider publication. The manuscript itself notes that AI-generated Czech needs checking — that caution applies to this site.

Not a replacement. Existing textbooks and courses are not inadequate. This resource reorders and personalises what they contain, so that effort lands where it returns most.

IIIReferences
  • Danaher, D. S. (2015–2018). Case studies: nominative, genitive, dative, accusative, locative, instrumental.cokdybysme.net
  • Holá, L. (2016). Česky krok za krokem 1. Prague: Akropolis.
  • Holá, L., & Bořilová, P. (2009). Česky krok za krokem 2. Prague: Akropolis.
  • Janda, L. A., & Clancy, S. J. (2006). The Case Book for Czech. Slavica.
  • Kořánová, I. (2012). Česká čítanka + Grammatical Notes English. Prague: Akropolis.source of noun-type frequency
  • Kupka, P. (2010). Podstatná jména. Prague: Kupka Nakladatelství.
  • Lukeš, D., & Jitka, K. (2012). Yellow Pages of the Czech Language.
  • Naughton, J., & von Kunes, K. (2021). Czech: An Essential Grammar (2nd ed.). Routledge.
  • Plocková, A. (2016). Příznakovost a frekvence. Bachelor’s thesis, Masarykova univerzita.SYN2010 adjective distribution by case
  • Rejzek, J. (2012). Český etymologický slovník. Prague: Leda.

Based on Czech Cases: A Strategic Roadmap by Mitchell Burns — dedicated to every student and teacher of Czech who has struggled with the cases.