
Click the blue text to follow us

Written by | Cheng Shushu Edited by | Li Xinma Cover image | Doubao AI
Once upon a time, AI music was merely a novel toy in the hands of tech enthusiasts—users could generate a melody by inputting a few keywords, akin to “opening a blind box.” However, the results were often filled with randomness and experimental flair, making it difficult to truly enter the realm of professional music.
But this situation is rapidly changing.
Recently, the AI music startup Suno is seeking a new round of financing with a valuation exceeding $2 billion, quadrupling from before; it is reported that its annual recurring revenue has surpassed $100 million, injecting solid commercial confidence into this emerging sector. Meanwhile, streaming giant Spotify announced a partnership with three major record labels and industry organizations to jointly develop “responsible and artist-centered” AI music products, marking a shift from passive observation to active collaboration in the traditional music industry. Furthermore, news that ElevenLabs, a leader in AI voice technology, has received strategic investment from NVIDIA, and that OpenAI is reportedly about to officially enter the field, indicates that top tech companies are paying attention to this area.
This series of intense capital and industry dynamics points to a clear trend: driven by technological iteration, capital support, and industry collaboration, AI music is no longer just a demo in a laboratory or a topic of online discussion. It is rapidly integrating into the complete industry chain from creation to consumption, accelerating its growth from a flashy “toy” to a genuine “business.”
01.
From “Opening a Blind Box” toProductivity
The core limitation of early AI music was its “one-time generation” or “blind box attribute”: the melody obtained after inputting keywords was often a “one-off deal,” which could neither be modified nor adjusted, and it was difficult to guarantee sound quality and professionalism, serving only as a fun experience, making it hard to enter formal creative scenarios.
However, since 2025, new generation tools launched by players like Udio and Suno have broken this predicament through upgraded editing functions, breakthroughs in sound quality, and reconstruction of creative logic, allowing AI music to enter a stage of “modifiable, precisely controllable, and deeply polished” craftsmanship.
First, the implementation of visual editing tools has achieved “paragraph-level refinement.”
On April 1st of this year, the AI music generation platform Udio, developed by the Australian company TopazLabs, launched the new “UdioStyles” feature, allowing users to upload content they own or control, thereby generating new music that mimics the “sound characteristics” of existing tracks. At the same time, it released an updated version of the existing AI model, v1.5Allegro, which improved output speed by 30% without sacrificing quality and consistency, greatly enhancing creative efficiency and helping creators convert inspiration into musical works more quickly.
Two months later, Udio quickly launched the visual editing tool Sessions, directly filling the gap of “difficult to modify” in AI music. This tool can automatically identify musical structures such as verses, choruses, and bridges from audio waveforms, allowing creators to move, expand, or replace different parts of the song. More importantly, the modified sections can automatically adapt to the original music in terms of tonality and rhythm, avoiding issues of disconnection.
Secondly, high-quality models combined with professional workstations achieve “detail-level control.”
Also in June, the American AI music generator startup Suno acquired the AI audio workstation WavTool and launched the V5 model and its self-developed digital audio workstation (DAW) SunoStudio in September. The V5 model brought a leap in sound quality, generating music that approaches the natural texture of real recordings; while the self-developed digital audio workstation SunoStudio revolutionizes traditional DAWs, combining “generation + editing” to change the previous model of AI generating content that cannot be edited.

Image source: Screenshot from SunoStudio official tutorial
Some users without formal music training only need to input the music style, general lyrics, emotional tone, specific directional prompts, or reference clips, or even hum a melody and upload a recording, and SunoStudio can synthesize a complete music product within minutes, simultaneously generating separate audio tracks for each instrument. By modifying editing instructions in specific segments of the audio track, AI can regenerate new musical sections.
Musicians can use their professional knowledge to “conduct” AI to provide more creative materials, such as editing, layering, and reorganizing materials based on their needs on the separate audio tracks, generating multiple AI versions for selection or combination, helping creators capture inspiration during bottlenecks and significantly shortening production cycles.
Meanwhile, minimalist interactive tools have achieved “demand-level precision.”
The ElevenMusic from the UK-based AI voice generator company ElevenLabs has lowered the professional threshold for AI music generation; its main interface retains only an input box, and the operation is entirely conversational. Users only need to input descriptive prompts, such as music style, emotional atmosphere, and instrument configuration, and the system can generate various types of music accordingly. Even more astonishing is that users can choose whether the music includes vocals, specific instruments, and other detailed elements, greatly enriching the freedom of creation. Currently, this AI supports the generation of songs in multiple languages, including English, Spanish, German, and Japanese.
The collective evolution of these tools has made AI music generation content modifiable, combinable, and embeddable, allowing it to truly become a productivity tool in the hands of creators, rather than just a flashy demo.
02.
Competing for “AI Music“
As the technological foundation takes shape, a global commercial race around AI music has fully ignited. From tech giants to startups, from overseas to domestic players, various forces are laying out strategies from multiple dimensions, including technology, products, and ecosystems.
On the international stage, competition is becoming increasingly fierce. The aforementioned technology-driven startups are consolidating their leading positions with “high quality + strong implementation.”
Suno and Udio, as benchmarks in the field, have achieved a closed loop of “technological breakthrough – commercial validation”: Suno has built a “sound quality + controllability” technological moat with its V5 model and SunoStudio, and its annual revenue of $150 million, with a fourfold growth over three years, validates the feasibility of subscription-based and enterprise-level music licensing business models; Udio, through its Styles library and Sessions editing tool, focuses on enhancing “professional creative efficiency,” becoming a favored “quick demo tool” for short video creators and independent musicians, with its commercialization progress and user stickiness leading the way.
Next are the tech giants entering the fray, cutting in with “resource integration + vertical scenarios.”
Google released the second-generation Lyria model in May this year, avoiding the red ocean of “general music generation” and instead focusing on “advertising music”—leveraging its advertising ecosystem resources, the second-generation Lyria can quickly adapt to the style needs of different industry advertisements, directly addressing the customized needs of commercial clients.
OpenAI has also been reported to have quietly initiated the research and development of AI music generation technology within its internal team. To provide high-quality training data for music generation models, OpenAI is collaborating with some students from the Juilliard School, who are professionally annotating music scores.
The domestic market is also showing vibrant innovative energy. Currently, domestic players in the AI music large model space can be divided into three categories:
The first category is represented by major companies like ByteDance and Alibaba. Among them, ByteDance’s Sponge Music has rapidly acquired users through a free strategy and platform ecosystem; Alibaba’s Tongyi Laboratory has released the InspireMusic model, taking the “tool empowerment” path, open-sourcing the InspireMusic full-chain toolkit to enable small and medium developers and enterprises to access AI music generation capabilities, aiming to seize the B-end market through “ecosystem co-construction.”

Image source: Screenshot from Sponge Music webpage
The second category is represented by emerging large model manufacturers like TianGong SkyMusic under Kunlun Wanwei. As the first music SOTA model in China, TianGong SkyMusic relies on the technical foundation of the “TianGong 3.0” super large model, focusing on “rapid generation + multi-style adaptation,” targeting high-frequency demand scenarios such as “micro-short drama music” and “game soundtrack segments”; its subsequent MurekaO1 model topped the industry SOTA leaderboard, attracting professional creative teams to collaborate, attempting to establish a voice in the “professional-grade AI music” field.
The third category is represented by vertical track unicorns like TianPuYue under Quwan Technology. As the world’s first multimodal music scoring large model, TianPuYue not only supports text-to-music and audio-to-music but also pioneers image-to-music and video-to-music functions, launching three months earlier than the international leader Suno. Since its launch, TianPuYue has fully integrated with Quwan’s ChaoYa App, directly reaching millions of music enthusiasts, achieving deep binding of “product – scenario – user” and rapidly accumulating users and data.
03.
Valuation Surge and Copyright Reefs
As the industry develops vigorously, hidden problems are quietly surfacing.
The core capability of AI music models relies on learning and mimicking a vast amount of music works, but most of this training data consists of commercially protected works (such as songs released by record companies and original works by independent musicians). Currently, there is a widespread issue of “data source opacity” in the industry: most AI music companies have not disclosed the authorization status of their training data, nor have they paid corresponding copyright fees to original creators.

Image source: Doubao AI
This model of “unauthorized training” has raised alarms among global copyright holders—the German Music Copyright Association has publicly questioned the legality of Suno’s training data, stating that “using copyrighted music to train AI without authorization is essentially an infringement on the labor results of creators.”
More complex is the definition of the creative subject. In traditional music creation, the logic of “the creator is the copyright owner” is clear and straightforward, but AI music cannot follow this logic: a user inputs lyrics and emotional prompts through SunoStudio, and the AI automatically generates a complete song with vocals, beats, and bass lines; another creator uploads their hummed melody, which is expanded into a symphonic version by ElevenMusic. In these works, the creativity comes from humans, the execution is done by algorithms, and the material sources are from training data—so who should the copyright belong to: the user, the platform, or the original musicians who were never credited?
Currently, both the “data infringement” in the model training phase and the “ambiguous ownership” of generated works have yet to form a globally unified solution.
In the face of copyright dilemmas, some leading players are beginning to proactively build copyright cooperation ecosystems. Spotify’s collaboration with three major record labels, Merlin, and Believe focuses on establishing an “AI music copyright distribution mechanism”: if AI-generated works use authorized data from copyright holders, they will pay royalties to original creators based on traffic; ElevenLabs has also reached agreements in advance with independent music organizations Merlin and copyright company Kobalt to ensure the compliance of training data and plans to launch an “AI music copyright traceability system” to record the sources of training data through technical means, achieving “transparent distribution.”
The formulation of industry policies and standards is also accelerating. The EU’s “Artificial Intelligence Act” has included “copyright labeling of AI-generated content” in regulatory requirements, clearly stating that AI companies must disclose the sources of training data for generated works; the National Internet Information Office of China has also made “training data compliance” a core review criterion in AI model filing.
For the AI music industry, compliance is not the end, but the starting point for the next round of innovation. When technological breakthroughs and copyright regulations achieve co-evolution, and when capital enthusiasm and legal frameworks find a balance, AI music can truly complete its transformation from “toy” to “business.”
He Xiaosheng
Great courage appears timid; great wisdom appears foolish.

“He Li” cares about what you care about ❤
“He Li” has opened the [Message] function, welcome to leave your paw prints~
Come leave a messageabout the issues or content you care about, we will discuss and interpret them in subsequent content.
Recommended Read


