I co-built Fluentide, a Chinese listening app with 2,000+ learners. It turns breaking news into Chinese episodes that sound native but stay at each learner's level.
Our AI has the autonomy to write and publish every episode. Before a script is voiced, it has to pass checks on level, repetition and new words, and AI reviewers flag any Chinese a native speaker wouldn't say. I built many of those guardrails.
This summer, an episode for Mandarin learners still went live in Cantonese, which they can't understand, and none of our checks caught it.
The episode in Cantonese
I noticed it myself. Some of our episodes came out in Mandarin and some in Cantonese, and I traced it to our text-to-speech service. It takes the language as a separate setting from the voice, and one part of our code never sent it. The service guessed. The script itself was fine, and every check we had only looked at the script.
I locked the language setting and added a test that fails if it's ever left out again.
Easy words, Chinese nobody says
We asked for words at the learner's level. The AI kept to easy words but sometimes wrote phrases no native speaker would say, like 走海 ("to walk the sea"), or 有米饭吃 ("have cooked rice to eat") where anyone would say 有饭吃 ("have food"). Our level check only counted vocabulary. It passed both, and they made it into a beginner episode before we caught them.
They were easy to miss because of how our scripts are written. Each Chinese line sits next to its English translation, and the English tells you what the Chinese means, so you don't notice when it's wrong. A beginner would memorize a phrase like that as real Chinese. I built an AI reviewer that deletes the English and flags any line a native speaker wouldn't say.
A goose leg on repeat
New words should come back a few times so they stick. Our level check had a floor on reuse but no ceiling. A draft for our lowest level took that literally and repeated whole sentences: "It's a goose leg. It's a goose leg. It's not a goose leg…" It scored 99.8% known words and passed.
I added a ceiling on repetition and set it by measuring all our published scripts. It sits high enough that normal scripts pass, and only real padding gets flagged.
The exception the AI leaned on
We let the AI leave a word out of the new-word count if the episode explained that word. It explained so many that one beginner episode hit about 15 new words a minute, against our usual 9.
So we added a cap that counts every new word, explained or not. Over the cap, the AI is told to cut the words it uses only once first.
City guides with no facts
We also wanted almost every word to be one the learner already knows. The first batch of beginner city guides hit up to 99% by leaving out names and places. Half of them had no checkable facts at all.
Now a check counts facts like years, distances and prices, with the bar set at our own news episodes. Numbers are some of the first words beginners learn, so the facts came back to our city guides without making them harder.
Testing the checks on real HSK papers
A check can also be too strict and block good work. That's harder to spot, because a blocked script never reaches anyone to read. So I started testing our checks on work we knew was good. When our AI began writing practice exams for HSK, China's official Chinese proficiency test, every check ran on the official sample papers first. If a real exam question failed one of our checks, we fixed the check.
Our plan said the right answer is almost always a paraphrase of what you hear, so we were going to reject questions where it's said word for word. On the real HSK 4 paper, 9 of the 32 answers are said word for word. The rule would have failed 28% of the real exam. On the official papers, not one of 256 wrong answers is said word for word, so the check now rejects any question where a wrong answer is said word for word.
Whatever AI I work on next, the first thing I'll look for is its goose leg.