The most common question on Twin Mode demos: "is the twin actually convincing or is this another deepfake-grade output that gets caught immediately?"
Fair question. This post is the data.
We tracked twin fidelity for 24 creators across 12 weeks of active use. Voice match scores. Likeness scores. Audience detection rates. Real numbers from real renders.
The four fidelity dimensions we track
Voice match score. Pairwise similarity between the twin's voice and the creator's authentic voice samples. Scale 0 to 100. Measured by ElevenLabs' internal model.
Visual likeness score. Pairwise similarity between the twin's face and the creator's authentic face samples. Scale 0 to 100. Measured by a third-party face-similarity model we license.
Audience detection rate. Percentage of audience members shown a twin video and an authentic video back-to-back who correctly identified which was which. Lower is better (means twin is harder to detect).
Failure rate. Percentage of twin renders that get rejected in the creator's approval queue. Lower is better.
Week 1 baseline
After enrollment and the first 4 renders:
- Voice match: 91.2 (good but not great)
- Visual likeness: 88.4 (slightly distinguishable from authentic by careful audience)
- Audience detection rate: 38% (audience catches the twin 38% of the time, near chance for an alert audience)
- Failure rate: 11% (about 1 in 9 renders rejected)
The day-1 numbers are honest. The twin is convincing but not perfect. A careful viewer can sometimes spot it. The creator catches about 1 in 9 renders that miss.
Week 4 progress
After about 28 renders of training:
- Voice match: 94.6 (+3.4)
- Visual likeness: 92.1 (+3.7)
- Audience detection rate: 24% (-14)
- Failure rate: 6% (-5)
The system has learned. Voice tightens. Likeness sharpens. Audience detection drops below the level where a casual viewer would notice. Failures drop in half.
Week 8 mid-training
After about 56 renders:
- Voice match: 96.8 (+5.6 from baseline)
- Visual likeness: 94.7 (+6.3)
- Audience detection rate: 14% (-24)
- Failure rate: 3.5% (-7.5)
By week 8 the twin is hard for even attentive audiences to detect. The failure rate has dropped to where most creators stop manually checking renders before approval.
Week 12 stabilization
After about 84 renders:
- Voice match: 97.4 (+6.2)
- Visual likeness: 95.8 (+7.4)
- Audience detection rate: 11% (-27)
- Failure rate: 2.8% (-8.2)
The numbers stabilize. Marginal improvements after week 12 are small. The twin has reached its operational ceiling for this creator with this enrollment quality.
What drives the improvement
Three things, in order of impact.
Volume of training data. The biggest gain comes from the system having seen more of your face, voice, and gesture patterns. Renders 1 to 50 generate the bulk of the lift.
Failure feedback. Every rejected render teaches the system what to avoid. By week 8, the failure pattern catalog is dense enough that new renders bypass most common misses.
Audience signal. The system tracks which twin renders get the best audience engagement and which get the worst. The high-engagement patterns get reinforced. The low-engagement patterns get downweighted.
What does not improve over time
Two honest limitations.
Enrollment quality is the ceiling. A bad enrollment (poor lighting, monotone voice, limited expression range) caps the fidelity ceiling. No amount of training overcomes a weak enrollment. If your week-1 numbers are below the baseline above, re-enroll. The re-enrollment is free for the first 60 days.
Edge-case content is harder. Highly emotional content (genuine laughter, sudden anger, tears) is the hardest fidelity surface for any twin. The twin will hit 97% on a calm explainer and 85% on a tearful storytelling moment. The gap closes slowly.
The Twin Mode primer covers the broader scope of what the twin does and does not do.
What this means for your output quality
The detection rate of 11% by week 12 means that even an attentive audience catches the twin only about 1 in 9 times. For most content, that is well below the threshold where engagement or trust gets affected.
For sensitive content (medical, financial, legal, political), we recommend using authentic footage only, regardless of twin quality. The fidelity is high enough to mislead, which is exactly why we restrict twin use on those topics. The ethics framework details this.
For everyday content (educational, brand storytelling, product demos, brand explainers, recurring segments), the twin is operationally indistinguishable from authentic footage by week 12.
The compounding implication
These numbers are why the twin is a switching cost.
A creator at week 12 in ENCORE has a twin operating at 97% voice match and 96% likeness with 11% detection. A creator who switches to a competitor on day 1 has a twin operating at 91% voice match and 88% likeness with 38% detection.
The competitor's day-1 twin is materially worse than the ENCORE week-12 twin. To match the quality elsewhere, the creator would need to invest another 12 weeks of training data.
This is why creators in the private beta describe the twin as "the thing I would not switch away from even if you raised the price."
The full switching-cost argument is here.
What to do with this
If you are evaluating Twin Mode against alternatives, the question to ask the alternatives is: "what does the twin look like at week 12, with data?"
Most cannot answer. We can.
Reserve your spot → or see the Studio tier features if you are ready to start training.