• Dominic@beehaw.org
    link
    fedilink
    English
    arrow-up
    5
    ·
    edit-2
    1 year ago

    There are a few reasons why music models haven’t exploded the way that large-language models and generative image models have. Maybe the strength of the copyright-holders is part of it, but I think that the technical issues are a bigger obstacle right now.

    • Generative models are extremely data-inefficient. The Internet is loaded with text and images, but there isn’t as much music.

    • Language and vision are the two problems that machine learning researchers have been obsessed with for decades. They built up “good” datasets for these problems and “good” benchmarks for models. They also did a lot of work on figuring out how to encode these types of data to make them easier for machine learning models. (I’m particularly thinking of all of the research done on word embeddings, which are still pivotal to large language models.)

    Even still, there are fairly impressive models for generative music.

    • i_am_not_a_robot
      link
      fedilink
      English
      arrow-up
      4
      ·
      1 year ago

      Example of music generation: MusicLM. The abstract mentions having to create a new dataset to get these results.