The model is capable of generating music that remains consistent over several minutes and operates on a hierarchical sequence-to-sequence modeling task. The generated music is at 24 kHz and outperforms prior systems in terms of audio quality and consistency with the text description.
Comments
Sign in to join the conversation.
No comments yet. Be the first to say what you thought.