Skip to content

What AI Music Slop Is, and What Sonic Multiplicities Is Instead

What slop is

The dominant form of AI music in 2026 is slop. The word names a kind of system. It says nothing about whether the output sounds good.

A slop system runs a short pipeline. Train a model on a corpus of recorded music, usually scraped without consent. Collapse that corpus into a latent space where genres, timbres, and formal conventions become regions you can sample from. Take a text prompt (upbeat corporate background, acoustic guitar, medium tempo). Render audio that fits it. The audio sounds like music. Nobody performed it, and it answers a description of a thing instead of the thing.

Four things hold for every system built this way, however good it sounds:

  • No human is in the room, and none ever was. No body, no nervous system, no years spent learning an instrument.
  • Everything it knows came from other people's recordings. It has never played a note itself. Calling what gradient descent does to an audio file "listening" is generous.
  • The music exists because someone typed a prompt. It does not shift when the performer shifts, because there is no performer, or when the room changes, because there is no room.
  • It cannot fail. It cannot miss an entrance, lose the thread, or startle you. It draws from a distribution whose center is the average of music that already exists.

The output is competent and dead. It fills time the way a screensaver fills a screen: active on the surface, absent underneath.

More training data will not change this. A system built to turn text into audio produces slop because that is what the architecture yields. Some of it sounds pleasant and stays slop, because the path from prompt in to audio out, with nothing performed in between, is the path from a vending machine to its can.


What Sonic Multiplicities is

Andrew Grathwohl built Sonic Multiplicities (SM) in SuperCollider and has run it in live performance since 2010. It pairs a solo improviser with a real-time listening and response system. Garrett Semmelink has played violin through it since 2010, Tyler Dinner guitar since 2019, and ten to fifteen others across its history. SM does not generate music. It joins music that is already happening.

Its corpus is its own past: the SuperCollider logs from real performances, and nothing else. It has never ingested another artist's catalog. The corpus is not fixed. Every show adds to it, so the system knows more about its own playing after each performance than it did before. When it answers Semmelink's violin, it answers through the memory of every previous night it heard that violin. Dinner's guitar draws on a different memory. The relationship is built over years of playing together.

The listening runs on a state-space model, a diagonal S4D network that replaced an earlier LSTM, trained on those history logs. It reads live audio. There is no text anywhere in the loop. It recognizes the gestures a player is making, matches them against everything it has done before, and answers. Music goes in and music comes out, with the system's own experience in between.

SM cannot play alone. Take away the performer and it produces nothing. It has no demo mode and no text-to-audio path, and that dependency is the design. Because SM only plays when a person is playing, everything it produces is tied to a specific event, room, and player, and none of it survives the occasion that made it.

The response system runs several coordinated engines. The site calls them datacoustics, the five engines, and the performance-completion loop. They analyze, answer, and complete musical gestures as they happen. What SM contributes is shaped by what the player just did and by what its history suggests could come next, so the result reads as a conversation rather than a string of generated phrases.

There is no score and no arrangement. No fixed progressions, no templates. Two parties improvise in real time, each carrying its own history into the room, and the encounter comes out different every time.


Principality

The thread tying these choices together is the principality. Through accumulated performance, SM builds a sonic world with its own internal logic: its own laws for what opposition, imitation, tension, resolution, and silence mean. The world grows with every show and belongs to the history of what happened in those rooms.

No two principalities are alike, because no two histories are. The performances that shaped one instance are specific to it. So are the players, the rooms, and the nights when something broke open or changed direction. All of it accumulates into a system unlike what it was before and unlike any other instance that lived through different nights.

A slop system has none of this. It has a distribution, and every output is another sample from the same space. Using it does not change it. Nothing accumulates and nothing is remembered. SM's larger aim, a network of principalities where performers visit each other's spaces and distinct machine-musician identities negotiate instead of dissolving, only works if each one is specific down to the root. That specificity is earned in live performances with the people who played them, over time. Prompt engineering cannot manufacture it.


Why the distinction gets missed

"AI music" now covers both kinds of system at once. A journalist writing about AI and music in 2026 will likely place Suno and SM on the same spectrum, one commercial and one experimental, both filed under "AI-generated." That framing hides the distinction that matters most.

Critics who dismiss all of it as soulless have slop pegged. The absence of a performer, the regression to the mean, the inability to surprise: those are properties of prompt-driven systems, and the critics describe them well. Their vocabulary just stops there, with no way to name a system that works on other principles. When they say "AI music" they mean slop, and about slop they are correct.

The boosters defending "AI creativity" err the other way. They hold up systems like SM as proof that AI can be a creative partner, then generalize to claims that only fit slop: infinite content from a text box, session musicians replaced, production democratized. SM shows none of that. It shows that a system can grow a musical identity out of its own performance history, which is the more interesting claim, and the one the argument should be about.

What separates the two is occasion. A slop system answers a prompt and finishes when the file is delivered. SM answers a person and keeps what happened. Play it today and it is already different from yesterday, because yesterday's performance changed it.


What this means

SM rejects the premise that music should be generated at all. The site puts it bluntly: uninstall your DAW. The digital audio workstation treats music as something you produce, and that is the wrong tool for what SM does. SM treats music as a practice, and a practice needs presence, time, and a thing that is made in answer to something happening.

The slop model treats music as content: specify it, generate it, consume it. The transaction closes on delivery. No performer, no history, no occasion, no stakes, no reason for the file to exist beyond the request.

SM treats music as something lived. It starts when a player starts and ends when they stop, and the music is whatever happens between. It carries the whole history of the system and the whole present of the room. You cannot reproduce it, because that player, that system state, and that evening will not line up again. It does not need reproducing. Its worth is in having happened.

Tyler Dinner, who has played guitar with SM since 2019, put it well: "It's more like a ghost, a presence."

A presence has been somewhere. It carries what it has been through and answers from there. A vending machine has been nowhere. It dispenses.

Sonic Multiplicities has been in active performance since 2010. It does not accept prompts.