Suno has opened the beta of Speech, a model that turns text into spoken narration with an original musical score. Both elements are generated together as one continuous audio track, allowing the delivery and music to align around pauses, transitions or changes in vocal intensity.
Text, voice characteristics and musical direction
The prompt has three parts: the text to be read, a description of the voice and instructions for the accompanying music. Users can describe the voice through an accent or historical period, as well as its tone and vocal register. For the music, they specify a genre, mood or dramatic direction.
Described uses include dramatized readings of poems and stories, or personal messages accompanied by music. Other possibilities include character performances and guided recordings such as meditations. Short-form media is another suggested use where variation between generations is acceptable.
Music can follow pauses and dramatic delivery
Generating both elements together allows a pause in speech to coincide with a musical transition, or an increase in vocal intensity to align with a swell in the accompaniment or a rhythmic accent. AlphaSignal suggests that successful short recordings could need less subsequent editing when precise, repeatable timing is not required.
Suno demonstrated the approach with a passage from the Odyssey rewritten in Gen Z slang. A second example delivered a request for a roommate to wash the dishes in Victorian English. In both company demonstrations, the music followed the reading’s pace and dramatic progression.
Accents and timing remain inconsistent
Suno warns that the beta produces inconsistent results. An accent can shift within a single recording, for example from British toward Australian pronunciation. Pauses can last longer than requested, while repeated generations can produce different pacing and interpretations.
The beta’s advertised capabilities do not include controls for exact recording length or word-level timing. Speech also does not advertise control over the placement of musical cues. Deterministic regeneration and voice cloning from a reference recording are not listed either.
Available to all users in the Suno app
Speech is available to all users directly in the Suno app. The wider release follows a month-long closed beta with a smaller group of users. The feature uses the same workflow as the app’s song generator.
Generation uses Suno’s existing credit-based plans; no separate pricing for Speech has been published.
Access remains limited to the app. Suno has not announced an API, SDK or other programmatic access, leaving developers without an announced route for submitting scripts automatically or integrating generation into a production workflow.



