I quite like that (EDIT - on second listen, and following the lyrics - Like it a lot). I'd like to know how much of it is yours, vs the AI (to see if I can allow myself to like it

)
Lyrics yours (presumably - beautiful - I think I can allow myself to like it just from them)
Melody - (I hope so also)
Choice of instruments?
Selection of voice?
Arrangement? (my definition here which may or may not be correct: score for each instrument) - guessing we're heading into AI territory now - even if "directed" by you. Or is it all yours and the AI simply performs?
(only answer if you are happy to do so. Feel free to just ignore otherwise)
Thanks. I like that one too.
People seem to have two different things in mind when they call something AI slop.
The first I agree with completely. There are people who use an online service like Suno or one of the others and have it generate everything including the lyrics with a veru simple prompt like: "A pop song, female singer". They generate thousands of "songs" very quickly without even changing the prompt and upload all of them to Spotify or elsewhere without ever listening to them or doing anything else. No effort, all automated. That is 100% AI slop, designed to get as much ad money as possible.
The second are like several people here who say unless all the parts are done by human musicians, it is AI slop. I disagree. A good song needs an idea, lyrics, the right instruments, the right type of voice, melody, rhythm and speed.
The good AI generated / augmented songs take musical knowledge and quite a bit of time. They have a human songwriter who may or may not play some of the parts and may or may not sing it themselves.
They can hire people (if there are any available in their area and they have funds to pay them) to play guitar, drums, etc. to bring their idea (the song) to life or they can use an AI agent to play the parts or sing. The songwriter may or may not tell the human players what to play and might just describe it with "start with a slow or fast piano part then add X and Y. We need guitars, bass and drums and a kazoo and it should sound like ........ Add a fancy guitar solo here. They start playing and stop because it isn't what the songwriter had in mind. He gives more information to influence them in a certain direction and they try again. This repeats until it sounds right, but during the process, the songwriter may have different ideas that are even better and change everything up. Get rid of the kazoo and add a french horn. Hey, we need more drums or louder drums. More or less or no distortion on the guitar part. Turn off the phase shifter. Less or more reverb on the vocals. This music part sounds good, but the words don't fit. Ok change the lyrics a bit so it fits, but still has the same meaning or feel.
I am happy to answer, but it will be long too:
I started with the phrase "A dramatic sigh". That popped into my head one day and I wrote it down on my list. I started the lyrics with the idea of making a funny / sarcastic song. How can a sigh be dramatic? But then I got to thinking and changing it a little at a time till it changed completely and ended up pretty close to the end result. With these lyrics, it didn't all come to me at once like sometimes happens. I worked on the lyrics for a while each day as ideas came to me. I finished them yesterday. I would say I spent about 4 or 5 hours total on the lyrics before I did anything else. For the melody I was thinking about a simple piano line. I started working with AceStep 1.5 at 3:02pm yesterday. I took a previous workflow that was close to what I needed and started with 90 beats per minute. I set the key signature to E minor starting out. I started with a simple prompt like slow piano, cello, female singer, mid range vocals. I added my lyrics with markup to further influence the AI. I set the denoise steps to 50 starting out for faster generations to see what I needed to change. After each generation, I listened to the full song that was generated and then made changes to fix what was wrong according to what I wanted it to sound like. That involved small changes to the original prompt and to the markup on the lyrics. At 50 denoise steps my pc can generate a ~4 minute song in about 45 seconds. When the results started getting closer to what I wanted, I changed the denoise steps to 80 (higher quality). Still making other small changes to the overall prompt, to the CFG settings (controls how close it follows the instructions vs being more random / "creative"). There are two CFG settings. The temperature setting, controls how "calm and reasonable" the AI should be. Too low, generates monotone. Too high generates complete garbage. There are also two random seeds for each generation. Those can be set to iterate or be completely random. Random works best as their influence is smaller than everything else. There are settings for Sampler type and scheduler type that change how instruments and vocals sound. There is a sampler aura flow setting that can influence things too.
The version you heard was generated at 4:21 am. Between 3:02pm and 4:21am, I only took breaks to eat, use the restroom and tend to my pets.
The final prompt that worked was:
"slow melancholic ballad, sad and introspective, emotional release, cinematic intimacy, mature female vocal in her 50s, warm husky slightly breathy timbre, rich lower register with emotional depth, lived-in vulnerable voice, subtle natural vibrato, soft and world-weary delivery, very slow tempo 55-68 bpm, piano and cello only, sparse piano with heavy sustain pedal and delicate arpeggios, warm resonant cello with sustained notes and gentle pulses, minimalistic chamber feel, quiet theatrical atmosphere, somber reflective mood, high emotional intimacy, soft dynamics with subtle swells only in bridge, warm natural reverb, no drums no percussion, experienced singer conveying lifetime of quiet strength, very clear phrasing with natural pauses and breaths between lines, deliberate pacing, gentle breaths where punctuation indicates"
I referred to markup on the lyrics. Commas are a small pause / breath. The ... is a slightly longer pause. The [headers] in the markup affect the way the AI does voices and instruments. [Chorus] means a little more dramatic. You can also refer to the headers in the prompt to change things. The AI sometimes decides to ignore bits of the markup, depending on the CFG and temperature settings. One generation can be great on the instruments but the voice is wrong. Another generation can make the piano sound like a kid's toy or the cello sound more like a fog horn. Very frustrating when the voice is perfect and the rest is bad. Over the time period, I generated and made changes for 170 generations until I got the one you heard. For some of my songs I used a weak or strong reference track. For this one, I was a songwriter and director / producer. The process was fun and frustrating, and I am happy with the end result.
As for whether it qualifies as AI slop or it is ok to enjoy it, that is up to everyone else to decide. The idea was mine. I wrote and rewrote the lyrics until they felt right. I changed the voice type several times. Made changes to the piano and cello tone. Increased and decreased the BPM and tried different key signatures until it felt right. Some generations, the voice sounded like a very heavy smoker and others sounded too young, too autotuned or had too much country twang for this song. I listened to every bad generation, every almost right generation and then finally the one I wanted. I ended up with the beats per minute set to 65 and the key signature set to D minor.
Here are the final lyrics and markup.
[Verse 1]
The clock hands bleed into the wall
I’ve traced the cracks... I’ve watched them fall,
These shelves hold ghosts of who I used to be,
Remembered dreams... I couldn’t keep...
The curtain’s thin, the stage light’s low
I’ve played my part, I’ve let it go
But something lingers in the quiet space,
Between the final bow... and grace
[Short piano and cello] (The AI chose to ignore this insertion)
[Chorus]
So let it out... let it fly,
Just a dramatic sigh
For every tear I swallowed down,
For every mask I wore without a sound,
It’s not a break... it’s not a cry,
Just a dramatic sigh...
The weight of all these years I’ve held,
Exhaled... into the air
[Verse 2]
I gave my all to shifting tides,
Chased the light and paid the price...
The mirror shows the lines I’ve earned,
The battles fought, the lessons learned,
To stand with empty hands and know,
Every scar’s a quiet show,
In scripts I never got to write
Still played the role... through fading light
[Chorus]
So let it out... let it fly,
Just a dramatic sigh...
For every vow I couldn’t keep,
For every promise buried deep,
It’s not a fall, it’s not a lie,
Just a dramatic sigh...
The weight of all these years I’ve held
Exhaled... into the air
[Bridge]
They call it weakness... call it defeat,
But I know the truth of how it feels,
To carry oceans in your chest,
And finally... choose to rest...
No grand finale... no trumpet’s sound,
Just breath returning to the ground,
The drama wasn’t in the fight,
It’s in the surrender... of the night
[Outro]
The stage goes dark,
The footlights fade...
I leave the ghost of what I made,
One last breath...
One final line...
Just a dramatic sigh...
[Fade]