Unsubtitled: Japanese conversation practice

Hi folks! I wanted conversation practice with feedback on what I was getting wrong, and I wanted it to be more fun than grinding out drills. So I made something! I’ve been using it every day, and I think it’s ready to share with y’all :grinning_face_with_smiling_eyes:

It’s kind of like an interactive visual novel. You pick a situation and talk your way through it out loud: a ramen counter, the ward office, your first day on a new team. The characters answer and react to whatever you actually said, so you can have real conversations and practice things a drill would never make you say, and you come out of it knowing you could handle that situation for real.

I’ve built some learning tools into it too: every line you send comes back checked for grammar, vocab, and politeness for the situation, with a sentence on why. It’s really helped me clean up some mistakes I just couldn’t shake before.

There are switches for furigana, romaji, and the translation. You can turn them off one at a time as you outgrow them.

It’s designed to work well with your phone’s built-in voice dictation (the little mic icon on your keyboard), but you can also type responses if you don’t feel like talking out loud.

Please give it a try, and let me know what you think!

So, just to get this out of the way:

What’s your response to users skeptical about this being made with AI? What verification checks have been made to ensure the site is offering realistic/natural Japanese?

4 Likes

Good question - and it applies to every site, human or AI-assisted.

What verification checks has Bunpro made to ensure the site is offering realistic/natural Japanese?

Tongue-in-cheek, might want to consider why quality discussions focus only on AI. I’ve reported dozens of errors on Bunpro, IKnow, Duolingo, and other sites over the years, and all their staff has responded to them promptly. When errors are reported on Unsubtitled, I also resolve them promptly.

No content is 100% accurate, and it doesn’t matter, because it still helps us learn and progress.

Language learning is not a binary outcome; it’s a process, and mistakes are a part of that.

Currently, Unsubtitled has an error rate of around 3-6% (that’s the worst-case rate, including minor things like politeness mismatches), which is about par for the course with other apps at their start. I keep a close eye on it, doing full audits weekly, and continue to improve the generation and QA framework.

I am not touching the AI debate this time, I will just point out that it is really unlikely you received any feedback from Duolingo, they are know for not caring answering to people. Ialso gave them dozens for the past 2 years before quitting and never heard a single peep.
I don’t know for the others sites except Bunpro where indeed the staff is extremly responsive.

I know it’s tongue and cheek but Bunpro actually takes pride in using only native audio and working from Japan with Japanese native to create their content. The french translation is also handled by a French native working at their office.

OK touching the AI subject just a bit, please be transparent.
If you are using LLMs, don’t hide it. It’s not respectfull to your users if they think you used real japanese audio from native (which I doubt). Also all the AI art already make me belive it’s fully AI generated but most people don’t care so go for it.
Just be transparent on how the audio is created please :bowing_man:

1 Like

Congrats on shipping. The visual design honestly looks great, at least on the landing page.

I never liked the “AI companion” concept so this probably isn’t for me, but there does seem to be a market that absolutely loves it.

I guess what I really want is the bare bones ChatGPT experience, but it knows how to communicate in (insert langauge here) at my level and corrects and teaches me in a natural way, in whatever my use case is, without forcing me into a roleplay scenario.

1 Like

Yep, and that’s exactly my point. We have no idea how these sites work behind the scenes, and in most cases, no one has thought (or cared) to ask because the product experience is beneficial.

If you used it, you would see that it’s crystal clear that LLMs and audio generation are key elements - otherwise, it would be impossible to respond to your conversational input.

Anyway, I’m not here to get into a philosophical debate on how Van Gogh isn’t a real artist because he used brushes instead of finger painting. I would love some actual feedback on the experience rather than debating the value of technology.

Would be amazing if someone shared their thoughts after trying it :upside_down_face:

Comparing yourself to Van Gogh is not a geat start.

5 Likes

I mean most people have their own AI assistant so why should they try yours?
A big portion of the users here and on wanikani are anti AI so I don’t think you’ll get a lot of feedback but I might be wrong.
As usual any AI assisted website will quickly devolve into AI bashing so you better have some strong argument about how your tool is so much better than any already existing one.
In addition we see people promoting their AI tool/website (they swear that they use daily) nearly weekly as their are so easy to create now.

4 Likes

Man, yeah, tough crowd.

Didn’t think that asking for feedback would devolve into AI bashing.

Guess I should have made an English reading comprehension app instead, lmao.


Anyway, I made Unsubtitled for myself, and I find it helpful.

Hopefully someone else will too~~

1 Like

I hope too, really!
Sorry for being annoying and thank for not using chapgpt to answer back as a lot of AI enjoyer do :bowing_man:

Next project idea when you are done with unsubtitled :wink:

2 Likes

Yeah insulting someone’s English comprehension is a great idea too. Got any other genius insights into language learning Mr. Gogh?

2 Likes

My question was a good faith one that I brought up because it shows up in every thread about a tool like this.

Bunpro staff can answer questions about their quality control without getting defensive, and notably their quality control comes from people highly qualified to verify responses. If you’re the person doing quality control, are you qualified to be doing so?

I’m not categorically anti-AI, but quality control is an important feature of language learning tools to me :slight_smile:

7 Likes

One of the issues with LLMs is that they are inherently stochastic. You can modify a prompt to suggest that a specific grammar mistake not be made in the future, or that a specific topic not be discussed, but you cannot guarantee that the mistake will not be made.

2 Likes

I love Bunpro and their staff. Check my profile - I’ve been here for a long time.

The disconnect in this discussion is that we’re talking about how AI and professional human translators make mistakes in the same breath, but people focus only on AI authorship while humans get a free pass. I’m assuming most people don’t go to a bookstore, pick up a copy of Genki, and question the clerk about the authors’ and editors’ resumes before reading the first page.

Prefacing this by again saying that I love Bunpro and their staff, there is some irony: if the human quality control is so highly qualified, then why are people (specifically, learners) still finding and reporting errors to the professionals? Is finding a typo a critical breach of trust and a reason to throw away the site? I don’t think so. Despite anyone’s best efforts and pedigree, mistakes still slip through. The overall learning experience is what matters.

I completely agree that quality control is a crucial feature of language learning tools, which is why I’m open to discussing the error rate and the process behind it. I don’t believe Bunpro or any other language-learning tools have published theirs. My goal is for it to be completely error-free, and I’m proactively pursuing that, manually, and by using AI and automation to audit and review conversations for errors. Even so, given the current rate of one minor error per 20-30 clean lines, it’s a good learning tool for me.

I appreciate the discussion now that we’re a little bit back on track. I can discuss learning theory and AI all day long; my career and background are in applied learning, cognitive science, and memory, and I spent a few years at a major language-learning company.

But yeah, I’ll admit it’s frustrating that we’re having a philosophical discussion about AI quality instead of someone just saying “I gave it a try and it was fun, the Japanese was completely accurate so far” or “yeah, it was pretty good, but I ran into one furigana mismatch after talking for twenty minutes” or even just “I tried it and didn’t like it”.

Kind of like human conversations :wink:

I’ve intentionally tuned Unsubtitled so the conversations are more open-ended and freeform than your average ChatGPT convo, so it exposes you to other tangential vocab and related situations.

Could you share the prompt(s) and the model parameters?

It’s more involved than just a prompt and parameters, but sure, I can talk about the system.

The core is currently Opus 5, with low reasoning effort and a default temperature, producing structured output against a schema. Each prompt is assembled per scene: character, setting, the politeness that the setting calls for, and the learner’s CEFR level if they’ve set one.

The most important part is the eval framework. It replays anonymized conversations with the level attached, and every candidate model (about a dozen, mostly Anthropic and OpenAI) gets identical inputs generated from the same schema code the app runs on, so the test is accurate. Answers go head-to-head against the current live config and are judged by a separate model, which is never one of the candidates because it shows favoritism. A/B order flips on half the pairs, seeded off the pair so re-runs are stable, and the judge can return a tie or mark both bad. It doesn’t show the furigana; otherwise, the annotation quality bleeds into the verdict on the writing. Separately, there’s a set of checks that don’t need a model at all: furigana completeness, validated by the same functions the app calls at runtime; p50 and p95 latency; output tokens; and any turn that errored or came back unusable. Plus, a human review after all of that is done.

Model selection primarily focuses on the Japanese quality and correctness. The strongest GPT model I’ve tested is a few seconds faster and has comparable Japanese, but it leaves kanji unannotated at a high error rate, so that’s a no-go. Opus 5 is currently the best model by a long shot and often produces perfect results.

I run the whole process weekly to continuously improve the outputs. The error rate in evals is generally around 0-2%. The error rate with live conversations is generally around 3-6%, but that’s a trailing estimate by definition, and it’s getting more accurate as I continue to tune it.

What is your source of human review?