Type a sentence, get a track. That pitch sells itself right up until you need a 40-second bed for a product demo and every result Stable Audio hands back sounds like the trailer for a 2011 superhero film. The tool isn’t the problem. The prompt usually is, along with a workflow that treats generation as the finish line instead of step two.
What follows is the loop I run most weeks: how to set up a session, what to actually type, how to rescue a near-miss with audio-to-audio, and how to get the file into a timeline without it sounding like library music nobody would license.
Settings First, Because They Constrain Everything
Head to audio.stability.ai and you face a few choices before a single word of your prompt matters. Two of them shape results more than anything you type.
- Mode. Text-to-audio builds from nothing. Audio-to-audio takes a clip you upload and restyles it, which is the fastest way to save a generation that nailed the mood but got the instruments wrong.
- Length. Clips under 20 seconds hold together better and loop cleanly. Push past a minute and the model starts inventing new sections you never asked for, which is charming in a demo and infuriating in an edit.
Free accounts get a monthly credit allowance and unused credits don’t roll over, so line up a batch of prompts before you start spending. If you’d rather self-host, the open-weight Stable Audio Open runs on a decent consumer GPU and caps out around 45 seconds per clip. That ceiling rules out full songs but leaves plenty of room for stings, transitions, and background beds.
A Prompt Formula That Beats Guessing
Generic input gives generic output. “Epic music” returns something you’ve heard in a thousand mattress ads. Fill five slots in every prompt instead: genre, lead instrument, production texture, mood or use case, and tempo in BPM.
Compare these two:
Weak: “chill background music”
Workable: “Lo-fi hip-hop instrumental, dusty Rhodes piano, brushed drums, vinyl crackle, 85 BPM, late-night study mood”
Two more that reliably produce usable material:
“Minimal ambient techno, soft analog pulse, deep sub bass, no percussion fills, 120 BPM, corporate explainer bed”
“Sparse fingerpicked acoustic guitar, close-miked, natural room reverb, 70 BPM, bittersweet documentary outro”
Notice that none of them ask for a genre and stop. Specificity is what stops the model reaching for its default soundtrack voice.
Walkthrough: Building a 45-Second Loop for a YouTube Segment
Here’s the exact sequence, using a talking-head video that needs music underneath narration.
1. Write four prompts that differ by one variable. Keep genre and tempo identical, and change only the texture: warm tape saturation in one, clean digital in the next, roomy and distant in the third, dry and close in the fourth.
2. Generate all four at 35 to 40 seconds. Ignore anything shorter than 30 seconds, since you need trim room at both ends.
3. Audition them with your voiceover already playing. This is the step people skip. A loop that sounds lovely on its own often fights a mid-range voice. The winner is usually the one you barely notice.
4. Download the WAV and trim the head and tail. The first 200 to 400 milliseconds frequently contain a fade-in the model adds on its own. Cut it.
5. Set your loop point on a downbeat, not on silence. Zoom in and cut where the waveform crosses zero, or you’ll get a click every pass.
6. Do the boring processing. High-pass at 30 Hz to clear rumble, scoop 200 to 400 Hz if it’s muddy, and normalise to about -14 LUFS integrated if the video is heading to YouTube.
7. Duck it under speech. Sidechain compression knocking off 3 to 6 dB works, but plain volume automation of 8 to 12 dB during dialogue is easier to control and sounds cleaner.
Rescuing Weak Generations With Audio-to-Audio
Audio-to-audio is the feature people underuse. Hum four bars into your phone, upload it, and describe what you want it to become: “warm analog synth pad, slow attack, cinematic, 70 BPM.” Lower strength settings preserve your melody while rebuilding the sound. Push the strength higher and the source becomes a suggestion rather than a structure.
The practical version of this: you keep the rhythm and contour you already like, and you stop rolling the dice on whether the model invents a melody you can hum back. It’s the difference between a slot machine and a sketch pad.
Where Stable Audio Fits in a Wider Stack
Stable Audio is best at instrumentals. The moment you need a sung hook, generate the backing here and handle the vocal separately. A workflow like cloning a voice with OpenVoice pairs well, because you get a clean instrumental bed with no competing frequencies where the vocal sits.
If you want actual song structure, with verses and a chorus that arrives where you expect, that’s a different job and a different set of prompts. The techniques in this guide to generating full songs from text prompts cover arranging beyond the loop stage.
It’s also worth having a second tool on hand for comparison. Riffusion’s approach to building a full track in under 20 minutes leans harder into structure, while something like Loudly AI suits creators who want fewer knobs and faster output. Different projects, different trade-offs.
And if you’d rather call Stable Audio Open from your own code than click around a web app, the pattern for running AI models through Replicate translates directly.
Five Prompts Worth Stealing
- “Warm lo-fi drum break, dusty kick, snappy snare, vinyl noise, 92 BPM, no melody”
- “Cinematic strings swelling slowly, single cello line, wide reverb, 60 BPM, restrained and hopeful”
- “Retro arcade synth arpeggio, square wave lead, gated reverb drums, 140 BPM, energetic”
- “Muted electric piano chords, soft ride cymbal, upright bass walking line, 100 BPM, jazz café background”
- “Sub-heavy ambient drone, granular textures, no percussion, evolving pad, 50 BPM, tense documentary”
Three Problems You’ll Hit in Week One
Everything sounds washed out
Long reverb tails are the model’s comfort zone. Add “dry, close-miked” or “tight room” to your prompt, and cut below 200 Hz and above 10 kHz on the result. Most of the mush lives at the edges.
Your loop clicks
Ninety percent of the time it’s a loop point that lands mid-note or on a non-zero crossing. Find a downbeat, cut there, and crossfade 10 to 20 milliseconds. If the clip has a big reverb tail, fade the final 300 milliseconds out entirely rather than trying to loop it.
Every generation sounds the same
You’re probably reusing the same prompt skeleton. Shift the tempo by five BPM, swap one instrument, or change the era cue. Small deltas produce surprisingly different results, and they’re cheaper than regenerating the same prompt five times hoping for variance.
Keep a Prompt Log, Not a Folder of Random WAVs
The producers getting the most out of Stable Audio aren’t writing better sentences than you. They’re keeping records. A plain spreadsheet with five columns does the job: prompt, mode, length, a rating out of five, and where the file ended up being used.
After 20 or 30 generations you’ll see patterns you’d otherwise miss. Certain phrasings reliably produce clean low end. Certain BPM ranges loop more gracefully. Certain textural words, “dusty” and “brushed” among them, pull the output toward usable and away from trailer cliché.
That log turns the tool from a novelty into a library you can search. Six months in, you’ll have 40 clips tagged by mood and tempo, ready to drop into a project in the time it takes to open the app.

