The Decoder reports that Suno, best known for generating music, has added a feature called Speech. It produces spoken text and matching background music together as a single audio track. A user types an idea or a written text and describes the voice and the music style they want. The model then generates both the voice and the sound. The article does not give a price or say where the feature is available.

[1]
Two white lines cross an indigo block; the upper line ends in an ochre dot and the lower line stops sooner.
Two white lines share one block. One ends in a dot and the other stops short. That is speech and music placed in the same piece, not at the same length. An illustration, not an audio waveform., AI-generated illustration, not a news photograph

Suno product chief Jack Brody says the company tested Speech with a small group for a month. Suno says the feature can be used for poems, meditations, and bedtime stories. The beta still has bugs. The example in the article is that a British accent can sometimes sound Australian.

[1]

Suno has not said how it trained the model. The article places that silence next to criticism of music generators over copyright: major record labels have already sued Suno, and a Munich court recently ruled against the company, rejecting fair use as a justification for using copyrighted data. That is background the article supplies. It is not a new account of training data released with Speech. The source line names Suno. This is a short piece from The Decoder. Another outlet described the same feature, and that account is not filed separately.

[1]

要点

  • Speech puts spoken audio and matching background music on one track. The user describes the voice and the music.
  • Jack Brody says a small group tested it for a month. Uses include poems, meditations, and bedtime stories. A British accent can sound Australian.
  • Suno has not said how the model was trained. The label lawsuits and the Munich ruling against fair use are background in the article.