Suno Speech combines spoken narration and music in a single audio track

Suno Speech combines spoken narration and music in a single audio track

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
4. 10. 2026
2 minutes reading · 2 views
Listen to the article
Audio version of the article
Suno Speech combines spoken narration and music in a single audio track

Suno has opened the beta of Speech, a model that turns text into spoken narration with an original musical score. Both elements are generated together as one continuous audio track, allowing the delivery and music to align around pauses, transitions or changes in vocal intensity.

Text, voice characteristics and musical direction

The prompt has three parts: the text to be read, a description of the voice and instructions for the accompanying music. Users can describe the voice through an accent or historical period, as well as its tone and vocal register. For the music, they specify a genre, mood or dramatic direction.

Described uses include dramatized readings of poems and stories, or personal messages accompanied by music. Other possibilities include character performances and guided recordings such as meditations. Short-form media is another suggested use where variation between generations is acceptable.

Music can follow pauses and dramatic delivery

Generating both elements together allows a pause in speech to coincide with a musical transition, or an increase in vocal intensity to align with a swell in the accompaniment or a rhythmic accent. AlphaSignal suggests that successful short recordings could need less subsequent editing when precise, repeatable timing is not required.

Suno demonstrated the approach with a passage from the Odyssey rewritten in Gen Z slang. A second example delivered a request for a roommate to wash the dishes in Victorian English. In both company demonstrations, the music followed the reading’s pace and dramatic progression.

Accents and timing remain inconsistent

Suno warns that the beta produces inconsistent results. An accent can shift within a single recording, for example from British toward Australian pronunciation. Pauses can last longer than requested, while repeated generations can produce different pacing and interpretations.

The beta’s advertised capabilities do not include controls for exact recording length or word-level timing. Speech also does not advertise control over the placement of musical cues. Deterministic regeneration and voice cloning from a reference recording are not listed either.

Available to all users in the Suno app

Speech is available to all users directly in the Suno app. The wider release follows a month-long closed beta with a smaller group of users. The feature uses the same workflow as the app’s song generator.

Generation uses Suno’s existing credit-based plans; no separate pricing for Speech has been published.

Access remains limited to the app. Suno has not announced an API, SDK or other programmatic access, leaving developers without an announced route for submitting scripts automatically or integrating generation into a production workflow.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Google moves federated learning computation from devices to serversGoogle moves federated learning computation from devices to servers
Google Research has introduced a federated learning system based on trusted execution environments. Gboard already uses it for English and Japanese next-word prediction; Google says it delivers more accurate models and faster training.
3 min read
4. 10. 2026
Shopify introduces Canvas for building online stores through AI chatShopify introduces Canvas for building online stores through AI chat
Canvas connects online store creation with the AI agent Sidekick. It displays changes in real time and lets merchants test interactive elements, animations and layouts across different screen sizes.
1 min read
4. 10. 2026
Hugging Face and Liquid AI bring model training to coding agents without changing their codeHugging Face and Liquid AI bring model training to coding agents without changing their code
The open stack enables reinforcement learning inside Claude Code, Codex and OpenCode. In Hugging Face and Liquid AI’s experiment, LFM2.5-2.6B’s success rate across four environments rose from 42.2% to 54.2%.
4 min read
3. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok