← All work

Live — this is the real project running in the page.

VoiceEnroll

The enrollment moment, and nothing else.

Role
Sole designer and engineer
Stack
React · TypeScript · Vite
Dates
2025

Picking the voice your assistant will speak in is a small moment that gets designed badly almost every time. It is usually a settings row with names and no sound, or a wall of samples with no way to hold two in mind at once. You end up choosing a label rather than a voice.

This is that moment and nothing else: preview, switch, confirm. No account step, no personality quiz, no progress bar.

The decision

Selection and audition are the same gesture. Clicking a card plays it and selects it; clicking another instantly switches both. There is no separate play button, because a separate play button makes auditioning and choosing two different actions and quietly encourages you to choose without listening.

The card animates while it is speaking, which sounds decorative and is not. In a silent room you know which card is talking; with the volume low, or on a phone in a pocket, the animation is the only signal that the tap did anything.

Confirm is a distinct, deliberate step. The preview is reversible and free; the commitment is separate and explicit.

What it cost

Two voices is not enough to prove the interaction scales. At a dozen, the grid needs grouping and the instant-switch behaviour starts to feel frantic rather than fluid. The design as it stands is honest about being a prototype of a moment, not a shipping enrollment flow.

The audio samples were also missing from the repository entirely — the project referenced two MP3s that had never been committed, so the demo had been silent since it was written. I generated them through the same Azure Speech resource that drives MAI-Voice-Lab, choosing two voices that genuinely differ along the axis the interface copy claims: one clear and articulate, one warm and conversational. A voice picker where both options sound the same is not a picker.

What I’d change next

The two cards describe themselves as “Professional” and “Friendly”, which is the label problem I opened with, moved one level down. Better would be to let the sample itself carry the distinction — the same sentence, chosen so the difference between the voices is audible in the first two seconds — and drop the adjectives entirely.