Introduction
You generate the confession scene and the voice is perfect: low, hesitant, breaking on the last word. You generate the next scene and she sounds like a radio announcer. Same character, same generator, different day. The audience does not hear "the model varied"; they hear "that's a different person," and the series loses the one thing that makes a character a character.
An AI character voice generator produces a voice. Making it a character voice is a workflow problem: the voice has to be defined once, attached to the character, and directed per line the same way every time. This guide shows how to do that in ArcLoop — from the voice definition on the asset, through directing lines with state and scene, to choosing the model for the job and reviewing the voice against the face.
Where the voice lives
In ArcLoop a character asset holds more than the face. Bind the character voice to the asset alongside the Main image and references, and every shot that references @Mira pulls the same voice, the way it pulls the same face. The voice becomes part of identity, not a per-clip setting.
That has one practical consequence: define the voice before you need it. A character who gets her voice chosen in episode four sounds retrofitted, because the first three episodes were guessed. Cast the voice when you build the asset, the same session as the turnaround.
If you are choosing a voice quickly from a library, the AI Character Voice Generator tool page walks through picking a voice, pasting lines, and directing the performance; this article is the version that lives inside a series.
Rules for a voice that holds
- One profile per character, written down. Perceived age, pitch, pace, texture, where the voice breaks. Keep it on the asset's description; the voice prompt guide covers the fields.
- Direct the state, not the adjective. "Sad" is an average. "Trying to stay quiet, anger breaking on the last words" is a performance.
- Every line carries its scene. Where she is, who she is talking to, what she is hiding. The model performs the context.
- Same model for the same character across a season. Switching engines mid-series is the fastest way to a different voice.
- Review voice with the face, never alone. A line that sounds right in isolation can be wrong for the expression it plays over.
How to build and use the character voice, step by step
- Write the voice profile on the character asset in My Assets: age, tone, pace, texture, emotional range, one signature habit ("clips the ends of words when lying").
- Bind the character voice to the asset. Pick from the library or shape one; it is now part of
@Mira. - Pick the engine by the job. Seed Audio 1.0 for performed scenes where the emotion has to move inside the line; ElevenLabs for stable, reusable delivery with tags like
[whispers]or[crying]. The dubbing comparison has the split. Choose once per character. - Write each line with its state and scene on the shot card: who she is talking to, what she wants, where the voice breaks.
- Generate the shots' video first, so the face and timing exist.
- Run Generate Voiceover in Edit. The lines are placed on the timeline against their shots.
- Review with the picture. Watch the shot with the voice; if the delivery fights the expression, change the direction on the line, not the voice profile.
- Reuse across episodes. The asset carries the voice; new episodes only need new lines. For multi-character scenes, the multi-character voiceover guide covers keeping the voices apart.
Example 1: Voice profile on the asset
@Mira voice profile: female, reads 26–28, low-medium pitch, slightly dry texture, unhurried pace with short pauses before the words that cost her something. Emotional range: controlled to breaking; never shouts, gets quieter when angry. Signature: clips the ends of words when she is lying, and laughs only on an exhale. Engine: Seed Audio 1.0 for the whole season; do not switch engines for one-off lines.
Example 2: A directed line for a confrontation
Line for @Mira, Shot 9, kitchen, facing @Daniel across the counter, she has already seen the boarding pass. State: calm on the surface, anger under it, trying not to give him the satisfaction. Pace: slower than normal; a full pause before the last word. Line: "How was Osaka." Delivery note: not a question — flat, the period audible, no rise at the end, and no breath before she says it.
Example 3: The same character, a different scene, same voice
Line for @Mira, Shot 3 of EP4, hospital corridor, alone, on the phone to @Sora. State: exhausted, relieved, close to laughing at herself. Pace: quicker, breath audible, the dry texture stays. Line: "I know. I know. I'm coming home." Delivery note: the third phrase is almost a whisper; keep the pitch where it was in EP1 — same person, worse night.
The most common ways character voices break
- Voice chosen per clip. Different voice each generation. Bind it to the asset.
- Adjective direction. "Emotional" produces a random emotion. Write the state and the scene.
- Engine switched mid-season. The character changes throats in episode five. One engine per character.
- Lines written for the page. Long sentences the voice has to rush. Cut to what a caption holds.
- Reviewed without the face. Sounded fine, played over the wrong expression. Review on the timeline.
FAQ
Can two characters share a generated voice? They should not. Even with different direction, the audience keys on timbre. Give each recurring character its own profile and library pick.
What if I want the voice in several languages? Keep the profile and the direction; generate the line in each language with the same engine so the character reads the same across locales.
How much direction is too much? If the note is longer than the line, cut the note. State, scene, pace, one delivery detail is enough.
Does the voice work for narration too? Yes, but narration is a different character. Give it its own profile so the narrator never sounds like the lead.
Next step
Open your lead character's asset in ArcLoop and write the voice profile in one paragraph: age, pitch, pace, texture, where it breaks, one habit. Bind the voice, pick the engine, and direct one line with its state and scene. Generate it, play it against the shot, and adjust the direction until the voice and the face are the same person.





