• Welcome to ASR. There are many reviews of audio hardware and expert members to help answer your questions. Click here to have your audio equipment measured for free!

AI and Future of Music Production

You don't like watching generated AI beautiful young women dancing in open fields created by old men?
Let's not kid ourselves
 
You don't like watching generated AI beautiful young women dancing in open fields created by old men?
Let's not kid ourselves

Not sure who's kidding here, but this stuff is the 21st century version of Victoriana kitsch (or if you prefer different cultural references poshlost or prolefeed).

As a descriptive term, kitsch originated in the art markets of Munich, Germany in the 1860s and the 1870s, describing cheap, popular, and marketable pictures and sketches. In Das Buch vom Kitsch (The Book of Kitsch), published in 1936, Hans Reimann defined it as a professional expression "born in a painter's studio".

Kitsch is regarded as a modern phenomenon, coinciding with social changes in recent centuries such as the Industrial Revolution, urbanization, mass production, modern materials and media such as plastics, radio and television, the rise of the middle class and public education—all of which have factored into a perception of oversaturation of art produced for the popular taste.

I think this new generative 'art' or cultural production shares many of the same characteristics, motivations and resonances.

Kitsch is less about the thing observed than about the observer. According to Roger Scruton, "Kitsch is fake art, expressing fake emotions, whose purpose is to deceive the consumer into thinking he feels something deep and serious."

Tomáš Kulka, in Kitsch and Art ... proposes three essential conditions:

Kitsch depicts a beautiful or highly emotionally charged subject;
The depicted subject is instantly and effortlessly identifiable;
Kitsch does not substantially enrich our associations related to the depicted subject.

Straight-up nostalgic and melancholic kitsch eschews irony, and is incapable of it. Pop-art kitsch embraced irony and differentiates by making it fundamental. Anyway now that this is becoming Brian's blog, I've gone back to hiding their posts, some occasional curious glancing notwithstanding.
 
That's exactly what it is.
These "arts" aren't 1) replacing human race 2) doing anything we can't do already.
 
Why not to acknowledge the fact AI generated this by bringing some unreal elements to transitions. Like her turning to flower fields and vice versa. I might be tripping a bit with this kind of imagination but isnt AI here for us what can’t we do in real sober states ? Not to derail your efforts to accomplish the realness of your ai AV creations, just a thought to more experimental approach (of course she can still have two eyes and 10 fingers on her hands).

It is harder to make it look as real as possible than to just let it do something crazy. That is the challenge. I clearly stated in the description what I was doing and that AI was involved. Not trying to hide anything or deceive anyone.
 
I've been working on the lyrics for this one for a while. Finished them today and spent then spent 7 hours trying to convince AceStep 1.5 xl-sft to generate a version I liked.

 
This is a piano song generated a while back with an older version of AceStep 1.5 running in Gradio on my pc. I used a reference track I played on one of my synthesizers. I had AceStep convert the audio to a solo piano, so this is either AI augmented or edited. The image was generated in Fooocus and then I had Grok resize it and output it to fill the frame.

 
An original recording cleaned up using AI with the cover mode in AceStep 1.5. This one is 94% of the original recording from 1985. Only 6% AI.

 
I've been working on this one for a while. I think it is my best one yet. Just finished it.

I quite like that (EDIT - on second listen, and following the lyrics - Like it a lot). I'd like to know how much of it is yours, vs the AI (to see if I can allow myself to like it :p)

Lyrics yours (presumably - beautiful - I think I can allow myself to like it just from them)
Melody - (I hope so also)
Choice of instruments?
Selection of voice?
Arrangement? (my definition here which may or may not be correct: score for each instrument) - guessing we're heading into AI territory now - even if "directed" by you. Or is it all yours and the AI simply performs?

(only answer if you are happy to do so. Feel free to just ignore otherwise)
 
Last edited:
I quite like that (EDIT - on second listen, and following the lyrics - Like it a lot). I'd like to know how much of it is yours, vs the AI (to see if I can allow myself to like it :p)

Lyrics yours (presumably - beautiful - I think I can allow myself to like it just from them)
Melody - (I hope so also)
Choice of instruments?
Selection of voice?
Arrangement? (my definition here which may or may not be correct: score for each instrument) - guessing we're heading into AI territory now - even if "directed" by you. Or is it all yours and the AI simply performs?

(only answer if you are happy to do so. Feel free to just ignore otherwise)

Thanks. I like that one too. :)

People seem to have two different things in mind when they call something AI slop.

The first I agree with completely. There are people who use an online service like Suno or one of the others and have it generate everything including the lyrics with a veru simple prompt like: "A pop song, female singer". They generate thousands of "songs" very quickly without even changing the prompt and upload all of them to Spotify or elsewhere without ever listening to them or doing anything else. No effort, all automated. That is 100% AI slop, designed to get as much ad money as possible.

The second are like several people here who say unless all the parts are done by human musicians, it is AI slop. I disagree. A good song needs an idea, lyrics, the right instruments, the right type of voice, melody, rhythm and speed.

The good AI generated / augmented songs take musical knowledge and quite a bit of time. They have a human songwriter who may or may not play some of the parts and may or may not sing it themselves.

They can hire people (if there are any available in their area and they have funds to pay them) to play guitar, drums, etc. to bring their idea (the song) to life or they can use an AI agent to play the parts or sing. The songwriter may or may not tell the human players what to play and might just describe it with "start with a slow or fast piano part then add X and Y. We need guitars, bass and drums and a kazoo and it should sound like ........ Add a fancy guitar solo here. They start playing and stop because it isn't what the songwriter had in mind. He gives more information to influence them in a certain direction and they try again. This repeats until it sounds right, but during the process, the songwriter may have different ideas that are even better and change everything up. Get rid of the kazoo and add a french horn. Hey, we need more drums or louder drums. More or less or no distortion on the guitar part. Turn off the phase shifter. Less or more reverb on the vocals. This music part sounds good, but the words don't fit. Ok change the lyrics a bit so it fits, but still has the same meaning or feel.


I am happy to answer, but it will be long too:

I started with the phrase "A dramatic sigh". That popped into my head one day and I wrote it down on my list. I started the lyrics with the idea of making a funny / sarcastic song. How can a sigh be dramatic? But then I got to thinking and changing it a little at a time till it changed completely and ended up pretty close to the end result. With these lyrics, it didn't all come to me at once like sometimes happens. I worked on the lyrics for a while each day as ideas came to me. I finished them yesterday. I would say I spent about 4 or 5 hours total on the lyrics before I did anything else. For the melody I was thinking about a simple piano line. I started working with AceStep 1.5 at 3:02pm yesterday. I took a previous workflow that was close to what I needed and started with 90 beats per minute. I set the key signature to E minor starting out. I started with a simple prompt like slow piano, cello, female singer, mid range vocals. I added my lyrics with markup to further influence the AI. I set the denoise steps to 50 starting out for faster generations to see what I needed to change. After each generation, I listened to the full song that was generated and then made changes to fix what was wrong according to what I wanted it to sound like. That involved small changes to the original prompt and to the markup on the lyrics. At 50 denoise steps my pc can generate a ~4 minute song in about 45 seconds. When the results started getting closer to what I wanted, I changed the denoise steps to 80 (higher quality). Still making other small changes to the overall prompt, to the CFG settings (controls how close it follows the instructions vs being more random / "creative"). There are two CFG settings. The temperature setting, controls how "calm and reasonable" the AI should be. Too low, generates monotone. Too high generates complete garbage. There are also two random seeds for each generation. Those can be set to iterate or be completely random. Random works best as their influence is smaller than everything else. There are settings for Sampler type and scheduler type that change how instruments and vocals sound. There is a sampler aura flow setting that can influence things too.

The version you heard was generated at 4:21 am. Between 3:02pm and 4:21am, I only took breaks to eat, use the restroom and tend to my pets.

The final prompt that worked was:
"slow melancholic ballad, sad and introspective, emotional release, cinematic intimacy, mature female vocal in her 50s, warm husky slightly breathy timbre, rich lower register with emotional depth, lived-in vulnerable voice, subtle natural vibrato, soft and world-weary delivery, very slow tempo 55-68 bpm, piano and cello only, sparse piano with heavy sustain pedal and delicate arpeggios, warm resonant cello with sustained notes and gentle pulses, minimalistic chamber feel, quiet theatrical atmosphere, somber reflective mood, high emotional intimacy, soft dynamics with subtle swells only in bridge, warm natural reverb, no drums no percussion, experienced singer conveying lifetime of quiet strength, very clear phrasing with natural pauses and breaths between lines, deliberate pacing, gentle breaths where punctuation indicates"

I referred to markup on the lyrics. Commas are a small pause / breath. The ... is a slightly longer pause. The [headers] in the markup affect the way the AI does voices and instruments. [Chorus] means a little more dramatic. You can also refer to the headers in the prompt to change things. The AI sometimes decides to ignore bits of the markup, depending on the CFG and temperature settings. One generation can be great on the instruments but the voice is wrong. Another generation can make the piano sound like a kid's toy or the cello sound more like a fog horn. Very frustrating when the voice is perfect and the rest is bad. Over the time period, I generated and made changes for 170 generations until I got the one you heard. For some of my songs I used a weak or strong reference track. For this one, I was a songwriter and director / producer. The process was fun and frustrating, and I am happy with the end result.

As for whether it qualifies as AI slop or it is ok to enjoy it, that is up to everyone else to decide. The idea was mine. I wrote and rewrote the lyrics until they felt right. I changed the voice type several times. Made changes to the piano and cello tone. Increased and decreased the BPM and tried different key signatures until it felt right. Some generations, the voice sounded like a very heavy smoker and others sounded too young, too autotuned or had too much country twang for this song. I listened to every bad generation, every almost right generation and then finally the one I wanted. I ended up with the beats per minute set to 65 and the key signature set to D minor.

Here are the final lyrics and markup.

[Verse 1]
The clock hands bleed into the wall
I’ve traced the cracks... I’ve watched them fall,
These shelves hold ghosts of who I used to be,
Remembered dreams... I couldn’t keep...
The curtain’s thin, the stage light’s low
I’ve played my part, I’ve let it go
But something lingers in the quiet space,
Between the final bow... and grace

[Short piano and cello] (The AI chose to ignore this insertion)

[Chorus]
So let it out... let it fly,
Just a dramatic sigh
For every tear I swallowed down,
For every mask I wore without a sound,
It’s not a break... it’s not a cry,
Just a dramatic sigh...
The weight of all these years I’ve held,
Exhaled... into the air

[Verse 2]
I gave my all to shifting tides,
Chased the light and paid the price...
The mirror shows the lines I’ve earned,
The battles fought, the lessons learned,
To stand with empty hands and know,
Every scar’s a quiet show,
In scripts I never got to write
Still played the role... through fading light

[Chorus]
So let it out... let it fly,
Just a dramatic sigh...
For every vow I couldn’t keep,
For every promise buried deep,
It’s not a fall, it’s not a lie,
Just a dramatic sigh...
The weight of all these years I’ve held
Exhaled... into the air

[Bridge]
They call it weakness... call it defeat,
But I know the truth of how it feels,
To carry oceans in your chest,
And finally... choose to rest...
No grand finale... no trumpet’s sound,
Just breath returning to the ground,
The drama wasn’t in the fight,
It’s in the surrender... of the night

[Outro]
The stage goes dark,
The footlights fade...
I leave the ghost of what I made,
One last breath...
One final line...
Just a dramatic sigh...

[Fade]
 
I am happy to answer, but it will be long too:
Thanks for the detailed response.

You've created the first (part) AI generated song I've heard that I am 100% comfortable with calling real music.

I think for me the lyrics - the meaning - is critical. If you'd told me they had been AI generated I'd have been horrified given the emotional resonance they had with me.

Then - given the lyrics, a stochastic music generation would never have worked. It is obvious you had a vision for how you wanted it to sound to match the depth of the lyrics - and for me that is what matters. Human vision. Human creation. After that, it doesn't matter much to me (at leat in terms of the song) if you are directing musicians to create your vision or bullying an AI into generating it. It is the vision that counts.

I'll put to one side my disquiet at the impact of AI on the livelihood of musicians by accepting that in all likelihood you are not in a position to pay musicians to help create your music, and that here, AI is an enabler for you to realise your vision that would not happen any other way.

Great work.
 
Last edited:
These are transient fads

AI was ordained to make funny cat videos.

Is the only thing that lasts.
... and porn, of course. Would the internet have achieved critical mass  sans porn?
 
I think for me the lyrics - the meaning - is critical. If you'd told me they had been AI generated I'd have been horrified given the emotional resonance they had with me.

AI doesn't have feelings or lived experience. That is still necessary for real creativity.

I'll put to one side my disquiet at the impact of AI on the livelihood of musicians by accepting that in all likelihood you are not in a position to pay musicians to help create your music, and that here, AI is an enabler for you to realise your vision that would not happen any other way.

I live in a very rural area in southeast Oklahoma. There are very few musicians near me and they play because they want to, in gospel groups and similar. Most of them would not be capable of playing how or what I want.

I'm not depriving any musicians of pay with what I'm doing. I am using the AI music generator as a tool to build what I imagine. I am capable of playing most instruments myself and could build up these songs a track at a time over a month or so. After that, the result might be better or worse and I still wouldn't have a singer for the vocals.
 
This song is based on a strange dream I was having as I woke up this morning. The cover image is a good representation of what I was seeing in the dream. I had some of the lyrics in my head when I woke up and the main idea so these lyrics came easy. Most songs require over a hundred generations to get what I'm looking for. This one only took 7. I went back and forth between Gemini and Grok to get the image like I wanted. They had a hard time understanding what I wanted.

 
@Brian Hall Amazing stuff! Axo1989 on the other AI thread pointed me here.

Below is my local music generation setup: I have two clustered DGX Sparks running the Ace-Step 1.5 XL music generation models and saving the results to my Plex server. I fine-tune these on my organic FLAC collection to give the model some direction. This gives me an effectively infinite supply of private and personalized music. I am now working on adding a reinforcement learning interface where I can upvote/downvote generations and have the model continually learn from my preferences.

IMG_0703.jpeg


The beefy hardware is not because each song takes long to generate, but I need 10-100 generations to get something that works! You must be an expert prompter because you need so few generations for your creations.
 
Back
Top Bottom