Text-only roleplay hits the same wall every time: you have been building a scene for an hour and there is no way to see it. The character describes a room and you picture your own version of it. The character describes themselves and the picture drifts a little further from the last one every session.
Three things close that gap, and none of them needs configuring.
They draw the scene you are in
Ask for a picture and the character draws the moment you are actually in, not a generic portrait. The result arrives in the conversation, not in a separate gallery you have to go and check.
The part that matters is consistency: the same face every time, because the look is locked rather than re-rolled. Most image tools give you a new person on every generation, which is why character art from them never feels like the same character twice.
You pick which engine draws it. Different engines have genuinely different hands — one is better at faces, another at light and atmosphere — and the choice is per image rather than a setting you configure once.
They answer out loud
Replies can arrive as a real Telegram voice message — waveform, scrubbing, speed control, the same object as a voice note from a person. Not an audio file attachment you have to download.
Every character gets a voice that fits them. You can browse the available voices and preview one before committing, and the behaviour has three settings: off, on request, or automatic.
These are voice messages, not a phone call. There is no live conversation and nothing listens to you — you type, and the reply happens to arrive as audio.
They look at what you send
Send a photo into the chat and the character sees it. The view from your window, the outfit you were asking about, a drawing you made — and the reply is in character about what is actually in the image, not a generic acknowledgement.
This is the one that surprises people, because it changes the direction of the medium. Up to that point you have been describing your world in words. Now you can just show it.
Nothing to switch on
There is no settings page standing between you and any of this. Photos can arrive on their own as a scene moves, and the speaker button sits under every reply from the first message.
Both run on energy, topped up with Telegram Stars inside the chat. You can see how much is left at any moment.
- A render that fails is not charged.
- Either one can be turned off entirely if you want text only.
- You choose which engine draws, per image.
What happens to what you send
A photo you send is relayed to a third-party provider so the character can answer about it, the same way your text is relayed to generate a reply. This is written down in the policy rather than implied.
Nothing you send is used to build a public gallery, and images generated in your conversation land in your conversation.
Why the face staying the same is the hard part
Generating a beautiful image is easy now. Generating the same person twice is not, and it is the difference between illustrations and a character.
The look is fixed when the character is created and applied to every render afterwards, so a scene in a cafe and a scene on a street two weeks later are recognisably the same person. Nobody uploads a reference photo and nobody adjusts a setting — the consistency is a property of the character, not something you maintain.