How WaniKani, two failed JLPT attempts, and a pile of unread ebooks led me to Yomumichi

I can only highly recommend you start Satori reader ASAP.
I jumped on it with an early N5 3 months ago and the boost to my vocabolary / knowledge is insane.
Now I when I learn a word in the N5 deck of bunpro I often have already at least read if not actually study it in the satori reader card system.
Start with spring, summer or kiki mimi radio. I have read those and while challenging they are perfect for new learners!
Any of their easy content will do, easy by name but you will learn a ton from them and kiki mimi radio has a great story so far, summer was great too!

1 Like

This looks really good. This is how I translate Japanese manga myself with pencil and paper, with the exact translation, then the literal translation, then the natural translation.

Having this in an app would be a big timesaver.

1 Like

Japanese definitions plus a short explanation, in Japanese, of how the word is working in that particular sentence: yes, I want that. Past a certain point, getting pulled back into English every few lines becomes its own kind of friction.

It also got me thinking about something broader: the lookup could be in Japanese while the rest of the interface stays in English, and further down the line advanced readers could move more of the interface and its explanations into Japanese too. That’s me extending your suggestion rather than something I had already planned, but I’ve written it down now.

The contextual part is something I’m already committed to. A lookup that gives you the most common dictionary meaning when a different one clearly fits the sentence isn’t much use. And if the explanation still doesn’t land, I want you to be able to ask Shiori about that word, grammar point or sentence while she already has the reading context. That’s future work rather than a beta promise, but it’s the direction I want to go.

That’s pretty much my own story. I put a lot of time into kanji, vocabulary and grammar, but the material at my level was rarely what I actually wanted to read. Then I’d open a book I cared about and the friction would wear me down before the story could properly pull me in.

Since you’re already on the beta list, I’d really value your feedback once it’s ready. Two things I’m curious about: what would you reach for first if difficulty and lookup friction weren’t in the way, and when you’ve tried native material before, what usually stopped or slowed you down?

Thanks for mentioning that. I’d much rather ask than assume what would actually be helpful for you.

If you’re comfortable sharing, what tends to make reading software work or not work for you? It could be how much information is visible at once, losing your place when a lookup opens, having to jump between applications, or something I haven’t considered at all.

There’s no need to write me an essay about it. Even a couple of notes while you’re actually using the beta would help me a lot.

Yomitan and Anki are genuinely powerful, and if you already have that workflow running the way you like it, I’m not going to pretend Yomumichi invented dictionary lookup or replaces all the specialist tools around it.

What I’m trying to get right is the book-reading experience as a whole. Bring a book in, keep its chapters, context and your reading position, and have furigana, contextual meanings, grammar and kanji lookup, the different meaning layers and eventually review features together in one place. If you love your current setup, Yomumichi might simply sit alongside it. If you never managed to get that entire chain working, hopefully it removes a lot of the hassle.

On the AI point, I’m with you. Dictionaries, grounded grammar sources and native explanations should remain the authority. The AI’s job is to connect reliable information to the sentence and book in front of you, not replace it with a confident-looking guess.

Your process maps almost one-to-one onto the three layers I’m building: interlinear, faithful and naturalized.

Breaking a sentence down that way works, but it eats a lot of time and interrupts the reading flow when you’re doing it by hand for every difficult line. I want to automate the repetitive part while keeping the Japanese visible so you can still compare every layer with the original.

If you haven’t joined the beta yet, I’d be glad to have you. Someone who already does this manually would be especially valuable in telling me whether the three layers are genuinely useful and trustworthy—or whether they’re just adding noise.

2 Likes

Yes, I’ve signed up and very happy to take part in providing feedback.

I’m away for a couple of weeks, but will definitely have time when I return.

1 Like

Following up on the copyright part, you may want to look into this more and potentially get advice from a lawyer before you invest more time/money into this and release it to the public. I don’t have any Japanese novels published this year on hand, but looking at the copyright section of one published in 2025 it explicitly says that scanning and translation, even for personal use, is prohibited. Japan also tends to be a lot more strict with copyrights than other countries.

It sounds like you wouldn’t necessarily be training your models based on scanned data? But if so, that’s another iffy area when it comes to copyrights, as a lot of books are starting to explicitly include AI training in their copyright statements (example below from an English book. Haven’t seen this in any Japanese books yet, but will keep an eye out).

Anyway, I don’t want this to turn into an overall discussion on AI, but did maybe want to caution you about doing legal due diligence before releasing to the public, and especially if you charge a subscription fee as you may open yourself up to unwanted legal trouble.

6 Likes

Thanks for following up! Just to clarify, Yomumichi isn’t scanning books, redistributing uploaded content, adding it to a shared catalog, or using it to train AI models. Users bring their own legally obtained EPUBs, and there’s no DRM removal or circumvention involved.

Other ebook and language-learning readers offer similar private-import and translation features, so my current understanding is that this approach should be fine as long as the content remains private and isn’t redistributed. But I’ll check everything properly before release to make sure the boundaries are correct.

I appreciate the warning—it’s definitely worth keeping in mind.

2 Likes

@Asher, I just got an email telling me the post got hidden… Do you know why?

Unhidden. Post doesn’t seem to be breaking any rules atm, but probably best to be careful of the language you use surrounding copyright, like @nminer suggested.

1 Like

Thank you! Appreciate it!

1 Like

Hmm… too much going on is a big thing. Like most people with ADHD, I am also dyslexic (which I learned to get through when younger, but am facing all over again in every new language I learn). This means delineation is important, and definitely not too many popups. I like colourful, especially if it makes sense.

Definitely keeping things on one page. So the text is in front of me - then I can mouse over something, right click to get options, hover for a quick definition (especially the reading, perhaps some colourful notations like the commonality of a word (eg, top 100~500, top 1000, top 2000, top 3000, uncommon, rare). ) A good example of this breakdown is the Longman3000.pdf for English (top 1, 2, and 3 thousand).

Um. It would be great if toggles were available - to choose how much is available. From brief to verbose (say, 3 to 5 levels). Because what I can absorb on any given day may be a TONNE or it may be one line.

Anyway! Thanks much! Keep at the project, can’t wait to see what you do!

1 Like

So the TLs are AI, the art is AI and your site has a heavy Claude coded style.
You also don’t plan on making it free nor open source.
Ironic considering open source is how your LLM was trained on it.
While I probably won’t be one of your clients, I really do want to warn you about copyright issues.

Not from the “you stole because you use AI” pov, but rather from how difficult it will be for you to defend your invention if someone was to make a clear copy. Putting it behind closed source doesn’t mean your code is protected, as you inputted snippets of it into one or several AI models. One could also take the generated art and use it themsleves, you won’t be able to attack them since the art isn’t copyrightable.
Same for your program if it depends on too much agentic ai, it might not be allowed to be registered.

Lastly, even if you don’t store AI TLs, the fact you allow them (that’s your concept, people can make TLs) still results in allowing authors to legally attack you to ensure your tool doesn’t step on their work. Which is something we saw with various tools that have to block or remove contents that at first sight did nothing bad.

Hence, I highly second the idea of seeking legal advice to be certain the various usages of AIs in your project, usages of your project, and quantity of AI usage during conception will not put you in a bad spot.

It looks very similar to Japanese Sentence Analyzer | Breakdown Grammar & Particles in principle. It is unclear what they use under the hood. You look to have taken a similar approach but wrapped it around a whole book-reader rather than a sentance analyser.

I am not sure how much benefit burning LLM tokens to parse the sentances gets you over and above using a morphological parser such as one of the dozens of Mecab variants.

1 Like

Why would they search legal advice? Have they ever claimed that they care about their IP being used somewhere else? I know this might be shocking for some, but most people simply dont care if there stuff is reused, or are willing to take the risk. After all, they created it with AI in the first place. Its not their art and they are aware of it I assume. Most people here seem to make such a giant deal over something OP probably does not even care about.

But to be fair the entire thing does seem incredibly AI coded. Like, to a ridiculous degree. He does not even try to put some personality in it by directing the AI to a unique art-style, or unique writing-style. No, OP just took the stock AI response and went with it. I am not against AI to steer artistic choices if its used as a tool but all of the AI here is just AI and nothing but AI. Even his responses here are 90% AI.

That’s the principle of an advice. They don’t have to take it, it’s not mandatory and is supposed to bring to their attention things they may not have considered or cared about until it becomes an issue.

Its not their art and they are aware of it I assume.

Also, my post wasn’t only about art, but nvm the part where Mersadajan said they were a dev and refused to open source (usually done to avoid people forking the project).

Hi. Great idea and keep it up. As a fellow developer and heavy ai user I understand you. It’s an awesome tool that speeds up development 10x in the right hands. Don’t listen to the haters and deniers :). Building software is really hard and it’s clear you’ve out hundreds of hours into yours. It’s respectable.

I like the idea behind the app, I think the reading niche is semi open, especially for books and novels.

As others said, syncing with other tools would be an advantage, so I can hide furigana on known words or sync my progress to anki/Bunpro.

I try to avoid having too many srs tools so that feature by itself doesn’t appeal to me.

But I really like how you split translations into chunks and literacy levels. That’s cool.

One thing to add - many people read on their phones or tablets (wouldn’t be surprised if more than half), so a responsive version is a must have from day 1 I think. Or a simple pwa.

As others said, Copyright and legal part is important, but first you have to launch and get paid users. Legal becomes a blocker only if you scale really well.

Just for reference, anyone who has ever created a self-study deck (anime, manga or novel based) and publicly uploaded it anywhere (or shared with anyone) is violating copyright law by reproducing parts or the whole of a copyrighted piece of art.

Yet they don’t sit in jail because small fish are not important and are generally hard to catch. Size and distribution power matters here a lot.

Good luck!

1 Like

I don’t know what new you actually did. I use jidoujisho with yomikan and read a lot. Ai is good but not for everything, sometimes you need yourself to do the work

Sounds like too much of a chore for a simple AI based reader. I feel like every Japanese learner who can vibecode a little, does something like this.

So you’re saying, don’t worry about breaking laws unless you start making lots of money. Oh and I’m going to waive the moral concern because other people have broken the law so its okay for you to do it too.

That makes a lot of sense, especially the fact that how much information you can absorb may change completely from one day to another.

I want the reader itself to remain as clean as possible so that the book stays in focus. I don’t want every sentence to be surrounded by colours, labels and controls while you’re simply trying to read. More information should appear on demand when you interact with a word or sentence, rather than constantly competing with the text.

For the quick lookup, I’m currently leaning towards showing the reading and main meaning that actually apply in that particular sentence. I don’t want to dump an entire dictionary entry into a popup when most of it is irrelevant to what you’re reading at that moment. If someone wants all the other readings, meanings and details, they could still open the full dictionary entry from there.

I hadn’t properly considered having three to five different detail levels yet. I’m not sure whether it will end up being that granular, but having at least a brief and more detailed mode could be really useful. Meaningful colours for things like frequency or information type could also help with delineation without making the page visually noisy.

I have ADHD myself, so keeping things focused, context-relevant and free from unnecessary clutter is one of my own requirements for the reader. Thanks for taking the time to explain this. This is exactly the sort of feedback that will help once you’re using the Beta.

You’re actually going in the right direction with that assumption. The tokenization of the Japanese sentences is not done by an LLM, so I’m not burning model tokens simply to split sentences into words.

I don’t want to go too deeply into the implementation details yet, but deterministic language tooling handles that part. The model-based work is reserved for the layers where context and judgment are more useful.

So the comparison with a sentence analyzer is fair at a surface level, but Yomumichi carries that kind of analysis into a complete book-reading experience with broader context around it.

Thank you. Honestly, this was really nice to read from another developer who uses AI seriously and understands what working with it is actually like.

AI speeds up my work enormously, but it doesn’t independently design, build and verify the application. There is still a huge amount of steering, reviewing, iterating, rejecting bad results and reworking things until they fit together properly. Anyone who has tried using AI for a larger software project knows that simply accepting whatever it produces doesn’t work.

I’m still a one-person team, I have a full-time job, and I’m building this in my free time. I’ve already put hundreds of hours into it and also spend a substantial amount on the tools and inference that allow me to work this way. AI is multiplying the work I’m already doing; it isn’t removing the work or the judgment behind it. So I really appreciate you recognizing that.

Your feature suggestions make a lot of sense too. I’m already looking into selectively hiding furigana based on known words and eventually syncing information with services such as Anki, Bunpro and WaniKani. I use these tools myself, so I also understand not wanting to maintain another completely isolated SRS. I can’t promise exactly what the first version of those integrations will look like yet, but interoperability is definitely valuable.

Mobile and tablet support are also important to me because I read on the go myself. A nice desktop reader is useful, but it cannot be the only good experience. I’m already thinking about how translations, lookups and the other tools should work on smaller screens without covering the book in controls and information. That is another reason I want the reader to stay clean and reveal most things only when the user asks for them.

Thanks again for the encouragement and the thoughtful feedback.

One separate clarification, since the way I write my replies was mentioned: yes, I use AI to help format some of them.

English isn’t my first language, and I often use speech-to-text to get all my thoughts down quickly. I then use AI to organize that into cleaner English, fix the grammar and make it more readable before I proofread and post it.

That doesn’t mean AI is deciding what I think or inventing my answers. The thoughts, opinions, product decisions and technical details are mine; the wording is assisted.

This is also similar to how I work as a developer. I direct multiple AI-assisted sessions, review what comes back, correct it and iterate. “AI-assisted” is completely accurate. Saying that AI generated my opinions or did the thinking for me isn’t.

BlockquoteI don’t really want to turn this into the much larger discussion about AI replacing artists. I’m a developer and I use AI heavily in my own work, but I don’t feel replaced by it. It enhances what I can do. I would see it the same way for an artist working on the project.

as an artist, i refuse to pretend the work of my peers wasn’t cannibalized for AI, and AI has no place in my workflow

looked like a neat project though, but I’ll be supporting real artists.

2 Likes