Upload a clear sample
Use 30–120 seconds of one speaker, recorded without music, echo, or background noise. These conditions help the model separate your vocal identity from room sound and competing voices, producing a cleaner profile with more consistent pronunciation, pacing, and character.