Type «ai roleplay» into Google in Jakarta and the box finishes it with bahasa Indonesia. In Tokyo, ロールプレイ AI completes to 日本語. In Moscow, «ии девушка» becomes «ии девушка на русском». Istanbul asks for yapay zeka sohbet and then for ücretsiz; Seoul asks which ai 채팅 앱 to use. Four scripts, four alphabets, and underneath them one question that is not «is this app good».
It is «will it be any good in MY language» — and the people typing it have already been burned. They have met the character who was charming for three turns and then started thinking in English. They have met the one who used the formal pronoun with a lover. They know the feeling of reading something that parses perfectly and sounds like nobody alive.
This piece is about what breaks, why roleplay is the worst possible case for it, and how to find out in ten minutes whether a platform is faking it. We make @talaforge_bot, so treat our own claims with the suspicion they deserve; they are at the bottom of the article, after the test you can run on us.
Why it breaks, mechanically
The failure has a name in the research literature. A Cohere Labs team published the Language Confusion Benchmark in 2024 — fifteen typologically diverse languages, one question: can a model reliably reply in the language it was addressed in. The answer was no. Even the strongest models failed to do it consistently, English-centric ones worst of all, and two conditions made it measurably worse: complex prompts and high sampling temperature.
Now read that with roleplay in mind. A character is a complex prompt by construction — a persona, a scene, a memory, style notes, and a conversation that grows every turn. And roleplay is run hot on purpose, because temperature is the thing that stops her repeating herself. Roleplay is not an edge case for language drift. It is the precise operating point where drift is worst, and the whole category lives there.
Underneath sits a quieter tax. Models read text as tokens, and the token vocabulary is learned mostly from English, which lands at roughly four characters per token. Everything else pays what practitioners call fertility: Chinese runs around 1.8× the tokens per word, Arabic and Hindi commonly 3–4×, and some languages far worse. You never see that number. You see its consequence — the same conversation window holds less of your story in your language than it would in English, so she runs out of room, and forgets, sooner.
Five ways it shows, in order of how fast you will notice
- Register collapse. Japanese keigo, Korean speech levels, Russian ты/вы, German du/Sie, Indonesian saya/anda against aku/kamu against gue/lo. Every one of those encodes a relationship. A character who cannot hold one, or who never switches, has forgotten who you are to her — and register is the single fastest thing to test.
- Agreement drift. In Russian, Arabic, Hebrew, Polish or Portuguese the verb or the adjective is gendered, so a woman writing about herself picks a different form than a man. A character who slips into the masculine about herself has broken the scene in one word, and English-trained models slip often.
- English bleed. Dialogue in your language, narration in your language, and the italic inner thought in English. It surfaces in the thoughts first, because that is the most stylised text in the reply and the least anchored by what you just wrote.
- Calqued idiom. Sentences that parse and that nobody says: English figures of speech carried over word for word, English sentence rhythm, English paragraph shape. This is what people mean when they say a product «reads like a dub».
- Mangled names and address terms. 先輩 rendered in Latin letters for a Japanese reader, 오빠 flattened into a first name, a Vietnamese anh/em pair replaced by a bare «you». The words that carry the relationship are exactly the words a translation layer drops, because English has no slot for them.
The ten-minute test
Run this on us and on whatever else you are considering. It costs nothing and it settles the question faster than any review.
- Minute one — write the first line in your language, in the register you actually use. Slang, particles, the informal pronoun. Not «hello».
- Minutes two to four — make her switch register. Ask her to be formal with you, then tell her to drop it. A platform with only one politeness setting fails here, and it fails silently.
- Minute five — complicate the scene. Three people in the room, an argument, something that happened earlier. Complexity is the documented trigger; go and trigger it.
- Minute six — read the italics. Her inner thought is where English arrives first, usually several turns before the dialogue cracks.
- Minutes seven and eight — check agreement and address. Does she use the right gendered form about herself? Does she call you what a person in your language would call you?
- Minutes nine and ten — come back tomorrow and ask what you talked about. Language quality that survives a night is language quality.
Fail three of those six and the language setting is a translation layer. The size of the character library does not compensate; it is one failure repeated a million times.
Where we honestly stand
@talaforge_bot is in 36 languages — every button, every menu, every error, and the memory she keeps about you is written back in your language rather than filed in English and translated on the way out.
Then the honest part, because the interface count is not a fluency guarantee and we refuse to sell it as one. How well a character writes in Vietnamese or Turkish depends on two things we control only partly: which model stands behind her, and how her card was written. A character authored in English, running on a model that was evaluated in English, will read in your language like a good translation — because that is exactly what it is.
What we did do is treat drift as a bug rather than a limitation. In May 2026 a character narrated in Russian, spoke in Russian, and thought in English, with fragments of Chinese in the italics. The fix was an explicit instruction pinned to the very end of the prompt — speech, narration and inner thought alike, stay in the language — because the last few hundred tokens of context anchor a model's surface-language choice hardest. That pin covers a dozen languages today. It does not cover all 36, and we are not going to imply that it does.
Where we lose: we have no library of millions of community-made characters, so your odds of finding one already written by a native speaker of your language are lower here than on the big platforms. If what you want is a character somebody in your country already wrote and tuned, the large catalogues are where those live, and a local app built by people who speak your language will sometimes beat us on idiom and slang outright. We would rather say that than pretend our shelf is their shelf.
Ten minutes, starting now
Open Telegram, open @talaforge_bot, and write the first line in your own language instead of in English. Free to start — energy is ⚡ and upgrades are Telegram Stars in the chat, no card, no account, no store in between. If she answers in the register you used, holds it when the scene gets complicated, and still remembers tomorrow, you have your answer. If she does not, you have it just as fast, and you spent ten minutes rather than a month.
And the memory is yours to open and edit, which matters more in your language than in English: when she does get a name or a form of address wrong, you can go in and fix the fact rather than argue with her about it for five turns.