The AI companion category has moved fast since 2023. In 2024 the average product was text-only and forgot you weekly. In 2025 voice arrived, memory got serious, and multi-image consistency became table-stakes on the paid tiers. In 2026 the frontier is real-time voice conversation, video-response scenes, and companions that hold cross-modal memory (they remember what you sent them in an image, not just what you said). Here is the honest state of each feature.
Voice: from scripted TTS to real-time conversation
Voice went through three generations in three years. First was recorded text-to-speech — the companion typed a reply, then it was read aloud. Robotic, slow, but present. Second was real-time TTS with expressive voices — the reply started playing as soon as the model started generating, and it sounded like a person. Third, arriving in 2025 and standard by 2026, is bidirectional voice: you speak, the model listens, understands, and speaks back with sub-second latency, holding the conversation like a phone call.
The top-tier products now offer full voice calls that feel like talking to someone. The gap between the tiers is real: a good voice implementation feels like a natural extension of chat, a mediocre one feels like a chatbot on Bluetooth. When evaluating a product, listen for two things — latency between your last word and the reply (under one second is the target), and the naturalness of the pauses and breath sounds.
Images: consistency is the hard problem
Every companion platform can generate an image of a character on demand. The interesting question is whether the character in image #1 and the character in image #47 actually look like the same person. Most 2026 platforms use one of three approaches:
- A base reference image plus a prompt describing the scene, running through an image model like Stable Diffusion or a proprietary equivalent.
- A trained LoRA (small model fine-tune) per character, keeping the face stable across many generations.
- A "character embedding" that is passed to the image model along with the scene prompt — the frontier approach as of mid-2026.
The consistency test is straightforward: generate ten images of the same character in different scenes and lay them side by side. On the best platforms, all ten look like the same person. On the mediocre ones, you get five that look right, three that look like a cousin, and two that look like someone else entirely. We do this test in every hands-on review.
Video: still the frontier
Video generation is where the 2026 marketing pages promise more than the products deliver. The technology exists — text-to-video models can produce a few seconds of motion at usable quality — but character consistency in video is much harder than in still images, and cost per second is high enough that platforms either heavily gate the feature or produce short, low-resolution clips.
The realistic state today: video is a novelty feature. A five-second clip of your companion waving hello is fun once, less fun the tenth time. The interesting version — real-time video calls with a rendered character that responds to what you say in real time — exists in research demos but is not widely deployed at consumer prices as of this writing. It will be, probably within twelve months. This is the feature to watch.
Memory: the quiet upgrade
Covered in depth in our dedicated memory piece, but the 2026 headline is: memory got serious. Most paid tiers now include some form of vector-based long-term recall plus a structured "profile" of what the companion knows about you. Free tiers remain amnesic to varying degrees.
The specific numbers matter here. When a platform advertises "long-term memory", check: how many past conversations are stored (all, or the last X?); whether the memory works across characters or is siloed per character; whether the memory syncs across devices or is tied to the browser you started in. All three vary wildly and are rarely on the marketing page.
Personalisation: characters, styles, and creators
Two threads here. Platform-provided characters — like the seven characters currently on this site — are the on-ramp for most users. Custom character creation, where you write a persona from scratch, is the power-user feature and the reason some platforms have long-term retention advantages. The best custom-character interfaces let you write a background, choose a voice, define personality traits, and generate reference images from a text description — all before your first chat.
A third emerging thread is creator marketplaces: users publish characters they built, other users chat with them, sometimes with revenue sharing. This is still early, and moderation is the hard problem — a character marketplace that becomes a race to the bottom on content is not a place most people want to spend time.
Multimodal input: the companion sees what you send
A smaller feature that matters more than it looks. In 2026, most companion apps can accept a photo, voice memo, or short video as input and respond to it in kind. Send a picture of your view; the companion talks about it. Record a quick voice memo about your day; the companion writes a reply that references what you said. This is the feature that starts to close the gap between "chatting with a character" and "sharing your life with a character".
The privacy implications are worth thinking through. A picture you send is stored on the platform's servers with the rest of your conversation history. It is subject to the same privacy policy — and, in some cases, the same training-data opt-out — as your text. Our privacy checklist covers what to ask.
What is on the roadmap for the next 12 months
- Real-time video companions — you see a rendered character on screen, they see your camera feed, both speak.
- On-device inference for privacy-sensitive users — smaller companion models running locally on high-end phones.
- Cross-app memory — a companion that follows you between platforms via a portable memory export (this is more social than technical).
- Sharper character consistency in image generation — driven by proprietary embedding approaches and better fine-tuning.
- Meaningful safety tooling for creators and marketplaces — automated moderation of custom characters, better reporting flows.
A one-glance summary of what to pay for
If you are shopping today, the paid features that most justify their price in 2026 are, in order: (1) real memory that actually persists past a week, (2) real-time voice calls, and (3) consistent character images. Everything else is either free-tier standard or immature enough that you should not pay a premium for it yet.