I measured my music taste, then asked AI to make music for it. It couldn't.

The idea was simple and, in hindsight, naive. Take everything Spotify knows about what I listen to, distill it into a description of my taste, and use that description to have a music model generate pieces I would actually want to hear. Music that is not confined to a genre, because my listening isn't. One Saturday in August, one Claude Code session, a few dollars in API credits and one annual subscription I regretted within the hour later, I have a very good description of my taste and a firm conclusion that the second half does not work yet. This note is about both, because the failed half is the more interesting one.

Getting the data out

Spotify's developer API in 2026 gives a hobby app much less than it used to. No genre labels, no popularity, none of the old audio features (tempo, energy, valence), no previews. What it does give: my 1,580 liked songs with artists, albums, release years and ISRCs (the industry recording identifier), my playlists, and, it turned out, the ranked "top items" endpoints page through the whole listening history: 24,280 distinct tracks and nearly 2,000 artists over the long term. That history became the behavioral signal next to the explicit likes.

Genre labels came from MusicBrainz, one lookup per artist, about fifty minutes. Useful as a secondary layer, but I objected to it as the primary one: the artist is the wrong unit. The same artist writes very different songs, and what I like is the song. So the profile had to be built from the audio.

Measuring the audio itself

This is the part I am pleased with. Deezer's public API resolves an ISRC to a 30-second preview MP3, no account needed. That covered 1,551 of the 1,580 tracks. Every preview was then analyzed locally on my own machine, about 1.7 seconds per track on an RTX 3090: pretrained Essentia models for a 400-genre probability vector, danceability, vocal presence, electronic versus acoustic, mood axes, timbre, instrument and mood tags; signal features for tempo, key and mode, loudness, dynamic range; and a CLAP audio embedding, a 512-number fingerprint of how a clip sounds, made deterministic by averaging three windows. Spotify's own downloaded files were ruled out on purpose: they are DRM-protected and decrypting them would be unlawful. Deezer previews are meant for public playback and were only analyzed, never redistributed.

The tracks were then weighted (a liked track I play a lot counts more than one I liked once) and clustered by how they sound. Ten clusters came out, and they were nameable, which is what mattered:

ClusterWeightTempo
Ambient / experimental electronic16%113
House / deep house / nu-disco12%118
Vocal synth-pop / electro-pop / French pop12%120
Trip-hop / downtempo / future jazz12%103
Instrumental techno, deep and minimal11%125
High-energy 90s-style rave / breakbeat10%132
Quiet acoustic-electronic songwriter / downtempo10%106
Classical (piano-led, Romantic to Ligeti, baroque, opera)7%101
Rock / post-punk / new wave5%124
Jazz (hard bop, cool, modal)3%118

Beyond the clusters, the profile found what runs through everything: minor keys in anything with a beat (reversed in classical and jazz), pulse without aggression, "deep" as the dominant mood tag, 120 BPM as home, piano showing up in seven of ten clusters, vocals as texture rather than the point and, when present, often French. It found bridge tracks that sit between two clusters, outliers, and drift over time: in 2025 classical led my new likes for the first time, and right now I am exploring more than liking, mostly classical piano and Berlin techno.

Then I calibrated it, which took one round of questions. Two cluster names were jargon and got reworded. The rock cluster was confirmed as real taste and not vacation-playlist noise. Classical, at 7% by count, I raised to a core: it is the most developed and most sophisticated music I know, and a like-count underweights it badly. Classical-into-electronic crossings, I said, mostly do not work for me (an opera sample over a beat is fun once) unless the integration is structural, Anna Meredith being the model. French lyrics are a real preference. And two strands the clustering was too coarse to see, but which I know are there: Bartók's folk-derived dance forms, and Arabic and Maghrebi modal music, usually inside electronic production. Both point the same way: unfamiliar modal melody and irregular rhythm as a source of freshness, inside a production idiom that is already mine.

Reading that profile back was slightly uncanny. It is more accurate than anything I could have written about myself, and it was derived from measurements, not from asking me.

Attempt one: a commercial music model via its API

From the profile, the agent wrote seventeen prompts. Not "make some deep house" but measured properties: tempo, minor or major, vocal treatment (none, voice as instrument, or French lyrics), instrument set, mood vocabulary ("deep, hypnotic, restrained, not aggressive"), and length. Three kinds: inside a cluster, on a bridge between two, and structural crossings such as piano and string quartet inside minimal techno with a 7/8 cell, or an oud-and-ney maqam line over a deep groove.

These went to ElevenLabs Music: seventeen pieces of 75 to 90 seconds, about 25 minutes of audio, roughly four dollars. Ten prompts were first rejected by the terms-of-service filter for naming living artists or specific works; rewritten descriptively, all went through. Every result was then run through the same analysis as my likes and compared: requested versus measured tempo and mode, genre labels, nearest liked tracks.

The measurements said: tempo adherence excellent (13 of 17 exact), electronic idioms correct, vocals unreliable (one "song" came out instrumental), acoustic and folk idioms drifting toward film music. My ears said: competent and empty. Nothing happens in these pieces. They sound like the genre they were asked for and like nothing in particular. The embedding-based "closeness to your taste" number turned out too noisy for generated audio to be more than a hint, which is worth knowing on its own.

Attempt two: compose it ourselves

A hypothesis: maybe the emptiness comes from leaving every decision to the model. So the agent wrote a piece as code. All decisions explicit: 124 BPM, D minor, a seven-note piano cell cycling against a 4/4 techno floor so that the two grids drift apart and realign every seven bars, a second piano in canon, a harmonic turn, the cell in augmentation in the strings, a sixteen-bar breath where the kick drops out, subtractive ending. Eight and a half minutes, rendered with a free General MIDI soundfont and synthesized drums. The structure document said: judge the composition, not the timbre.

I judged the composition. Very little of interest. The relentless beat through most of the piece is annoying; the development is the only slightly interesting element, and only once; no inclination to listen again. The lesson, recorded in the project so nobody relitigates it: a process is not a piece. Structure alone does not clear the bar. If this is ever retried it needs real melodic and harmonic invention, a beat that is an event rather than a floor, and a shorter form.

Attempt three: the best-known consumer model

Suno has no public API, so the agent drove its web app through the Chrome extension: the six best briefs (four instrumental, the folk-dance quartet, and a French synth-pop song with lyrics written for it, female voice), two takes each, twelve songs. Verdict: cheap restaurant muzak. Generic drivel. I cancelled both subscriptions the same evening and asked Suno for a refund of the annual fee under the EU right of withdrawal, having used sixty credits.

To be fair about what I tested: one-shot text-to-music, in August 2026, driven by good prompts derived from a real profile. Not fine-tuning, not stems, not a producer iterating for a week. Within that scope the conclusion is firm: the outputs are competent and empty, and my bar is human-level creativity. Not worth more money or more evenings unless something fundamental changes.

What survived: the profile, and sixty-one records

The profile and the analysis pipeline are sound and reusable; any audio can be measured the same way and placed against my clusters. And the profile turned out to be excellent for the thing I should perhaps have asked for first: recommending music that people made.

The agent produced sixty-one items, each an album or a track with a one-line reason and a Spotify link, chosen against the profile and then checked against my whole history so that everything is graded as new to me, brushed past, or a deeper cut of something I know. It built them into a private playlist. The reasoning was specific in a way I have never gotten from a recommendation engine: Nik Bärtsch's "ritual groove music" as "your ambient/techno corner done acoustically, pulse without pressure"; Taraf de Haïdouks as "the real thing behind the Romanian Folk Dances"; the Massive Attack remix of Nusrat Fateh Ali Khan as "the best-known proof that this crossing can work"; Barker's kickless techno "given what you said about relentless beats"; Michael Gordon's "Timber," an hour of six percussionists on wooden 2x4s, with the note "the piece I should have written today instead."

So did they land? Mostly. Michael Gordon's "Timber" is very good, full stop. The Nusrat Fateh Ali Khan remix and Nik Bärtsch both go back on; I will listen to them again. Barker was a hit with a footnote: the track on the playlist, "Paradise Engineering," turned out to be one I had already heard a few times on Spotify without ever liking it, so it was not new to me, but it was a good recommendation all the same. Taraf de Haïdouks was fun once, but not something I will play repeatedly, and I do not hear the connection to Bartók the agent was reaching for. That is roughly the hit rate I would want from a good friend with a big record collection, and much better than anything an algorithm has ever pushed at me.

What I make of it

The measuring works. The describing works, better than expected. The recommending works, because it is really the describing pointed at a catalog of things humans made. The generating does not, and the interesting part is that the failure is not technical. Tempo, mode and idiom were all hit; the models can do what they are asked. What they cannot do is want anything. Every piece is an average of the genre it was asked for, and an average has nothing to say.

Tools: Claude Code throughout, with Spotify's and Deezer's public APIs, MusicBrainz, Essentia and a CLAP model running locally, ElevenLabs Music via API, and Suno via its web app driven through the Chrome extension. Total cost of the generation experiments: about four dollars of API credits, plus one refund request.