Hi!
I was interested in the similarity of words within the Bunpro JLPT vocabulary and I want to share my insights and tools with you. My main motivation is to better differentiate similar words in future and to understand their nuance differences. I am interested in your thoughts, feedback and other insights we could explore.
Similarity Graph
First, I wanted to get an overview on different similarity measures and how they “look like”. There can be many ways to measure similarity of words but I used the following two methods:
- “Bunpro Summary similarity”: two words are similar if the bunpro summary (definition) describes the same
- “Accepted answers similarity”: Two words are similar if they share some accepted answers
With these two, you can calcualte scores between each pair, take the 5 closest matches and place them on a map where similar words are closer to each other than non-similar ones. If one word links to another with a high enough score, I draw an edge.
You can play around with the map here. It may take a few seconds to load and might not work on your phone.
Missing Meanings in Vocabulary
Then I got interested in all study questions that use an English translation which isn’t listed in the vocabulary itself. For example, the word がっちり lists “solid”, “robust”, etc. but one of the sentences uses “thoroughly” which is missing in the listing.
I have a full list here. Notice that the list contains many false positives though it’s easy to recognize them.
Missing Accepted Answers
Lastly, I wanted to find all sentences that ask for a specific word but don’t allow another for any reason (wrong context, wrong grammar, missing accepted answer, etc.). For example, the sentence “____を買う。” (I will buy a ticket.) is looking for “チケット” while “乗車券” is rejected even though both could be translated to “ticket”.
The full list can be found here. Again, there will be quite a few false positives.
Next Steps
One can come up with other similarity measure but I think it will result in a similar image. Recreating the analysis on grammar might be interesting but I am not sure if it will convey the grammar differences accuratly.
I think the data shows that data quality on Bunpro is already very good. You can argue whether every single meaning needs to be listed in a vocabulary list or not. On the one hand, it adds quite a bit of overload to a word. On the other hand, it may open another view on the word itself. What’s more interesting is to add new explanations or alternative answers to some of the sentences that I identified.
Disclaimer
The analysis uses the official Bunpro N5 to N1 Vocab deck. I tried to link to Bunpro where it makes sense. Grammar decks have not been used so some JLPT words might be missing. A large part of the reports were generated by Claude.
