Review Methodology
Every AI companion platform we review is scored across 9 weighted categories on a 10-point scale. The overall score is the weighted average. Platforms we have not hands-on tested do not receive a score.
The nine categories
Weights
Total weight adds up to 100. Weights are the single source of truth — any change here automatically propagates through every review and best-of page.
How we score each category
- Conversation quality — depth, coherence, tone control, roleplay handling, refusal behaviour.
- Memory — recall of prior conversations across sessions, days and weeks; how gracefully old context is surfaced.
- Character customization — how deeply personality, appearance, voice and roleplay parameters can be tuned.
- Image quality — resolution, aesthetic quality, and same-character consistency across generations.
- Video quality — motion coherence, character consistency across frames, prompt adherence.
- Voice / calls — voice quality, latency, turn-taking, and (if supported) call quality.
- User experience — onboarding, navigation, gallery UX, mobile behaviour, edge-case bugs.
- Value — pricing vs feature depth vs cap on free tier. Not the cheapest — the best price-for-feature.
- Privacy / transparency — clarity of data handling, moderation policy, billing behaviour and account controls.
What counts as "tested"
A platform is "tested" only when we have a paid or free-tier account and have run our standard test protocol against it: multi-session conversation, image generation batches, voice reply samples, feature-matrix verification, and a pricing audit. Any platform without this is marked "Not yet independently tested" and receives no score.
Retesting
Reviews are retested at least every 90 days and immediately after any major platform update. The date on each review reflects the most recent retest.