Where do the Vocab Frequency Lists Come From?

Apologies if this is addressed somewhere (tried searching, couldn’t find it), but I mean these lists specifically:

image

Like, where does the Bunpro team get it from? I’m also curious what is meant by “General” frequency. We already have default frequency (what Bunpro recommends for ease of learning), dictionary (I guess just pulled from an online dictionary), but where does that leave General?

10 Likes

My concern is that despite setting “Anime” as the priority, I’ve only seen Anime frequencies in the 500-1000 range when I was expecting it to start at 1 and descend from there.

Bumping this thread tho, also curious where the data is pulled from.

1 Like

The word frequency data we use is aggregate data from a list that you can find online. Much like the JLPT data, there isn’t really any “true” frequency list because whatever subset of anime/novels etc is used to make a frequency list will end up biasing the data.

@TangoTangoSIerra The data does start from 1 but in most cases the first couple hundred are going to be particles and super basic common words like いい, する, ある

7 Likes

@Jake
Thank you for your reply!
I understand, but would it really go as low as 1,000+? Or do I have a setting wrong somewhere?

I’m not really sure what you mean by go as low as 1000+. Each word takes up one spot so unless there are less than 1000 words in a frequency list, it would go beyond that pretty easily.

Sorry if I was unclear. What I mean is this: If I’m sorting by “Anime - Most Frequent” and the “first couple hundred are going to be particles and super basic common words like いい, する, ある” as you said, I am still seeing the frequency of the anime word as “Top 1,000” with the actual number being 1,139 (for example). Shouldn’t the frequency number be a lot higher? (lower number, technically)

my guess is you already know the other 1138 words in the list. For example I’m know few words and the words I get when studying the top 2k sorted by generaly frequency are much closer to the beginning of the list (see screenshot).

1 Like

Ah I see where the confusion might be. The frequency is global rather than based on the deck.

So a deck that only has rare words from anime might only ever show 1k+ or even 10k+ because there are tens of thousands of words in each frequency list.

3 Likes

@Jake are you able to share the raw data for these lists?

2 Likes

Hi,

I have to open this topic again, as I am a bit confused by the sorting order filter “general”.

I added the Core 2k deck to my queue and had already like half of it finished via other decks.

I thought I put the order to general to get the most frequent words in Japanese language first. But I get now 昭和 (Shōwa era) as vocabulary. First I did not expected it to be core 2k word and second not being so frequent compared to other words. This becomes even more evident when I switch to order it by “Netflix” my next word would be 音 which is a way more usual/frequent word than 昭和.

So what is “general” actually doing? And is 昭和 really one of the most basic and important 2k words?

昭和 isn’t a “basic” word. it’s just extremely frequent in printed media and online articles. year dates are usually given in both Gregorian (e.g. 1950) and Imperial Era (e.g. 昭和25年). every newspaper will have the era name at the top. every single wikipedia article about a person or historical event will probably contain era dates throughout.

1 Like

Really helpful thing to note. I think sometimes people confuse ‘core’ with being like the actual core of the language i.e. mega common words, instead of how it’s actually sourced. I’ll make a note to adjust the deck descriptions we have for the Core decks tomorrow to include something about this, in case anyone in the future is curious.

1 Like

Thx.

The description would be really helpful to understand the intentions and sources of the decks.

There are many core lists outside. In YouTube videos many people are talking about core lists. “Learn the 2k core list and you are able to understand daily conversation” at least vocabularywise… and so on.

Back to the initial question. 昭和 is more common in the general Japanese language than 音, because it shows up just way more often in NP, books, etc? Based on the filter “general”

it is more common in written non-fiction, and the frequency lists are primarily based on written language.

that doesn’t mean the core decks are useless for conversational japanese, but they are certainly skewed towards certain words and bookish expressions. it is difficult to build a deck around conversational language because, by their nature, it’s rare for full natural conversations to be written down like they are spoken. the closest to that would probably be a Netflix or Anime based frequency list. besides that, Kaishi 1.5k is always a great option for beginners because it was handpicked by a group of people as a basis for immediate immersion.

Is kaishi 1.5 a bunpro.jp deck?

it’s meant as an Anki deck, but there is a version of it in the Bunpro community decks.

1 Like