Twenty-three generations in, I had a folder full of blurry hands and melting coffee cups. The prompt wasn’t the problem. Everything around the prompt was.
FLUX rewards a specific way of working: pick the right variant, structure the prompt in layers, set numbers that match the variant, then repair what’s broken instead of starting over. That’s the whole game. Here’s the workflow in order, with the settings I actually use.
Step 1: Match the Variant to Your Hardware
FLUX.1 ships in three flavours and the differences aren’t cosmetic. Schnell runs in four steps and finishes a 1024px image on a mid-range card in a couple of seconds. Dev takes 20 to 30 steps and gives noticeably better lighting and anatomy. Pro is API-only and mostly matters when you need commercial rights without thinking about the licence.
On 8GB of VRAM, start with schnell in a quantised build (GGUF Q4 or FP8). On 16GB, or on a rented cloud GPU, go straight to dev. The open-source release that reset expectations for what you could run at home is the same architecture across all three, so a prompt you tune on schnell transfers to dev with only minor edits to the light and texture language.
Step 2: Write the Prompt in Layers
FLUX weighs the start of a prompt more heavily than the end. Subject first, mood last. I build every prompt from the same five layers, in the same order.
The five layers
- Subject and count: “one elderly fisherman”, never “fishermen”
- Action or pose: “mending a net, seated”
- Setting and time: “on a stone harbour wall at first light”
- Light: “cold blue ambient with warm lantern spill from the left”
- Camera: “35mm, shallow depth of field, slight grain”
Assembled: one elderly fisherman mending a net, seated on a stone harbour wall at first light, cold blue ambient with warm lantern spill from the left, 35mm, shallow depth of field, slight grain.
What layering actually changes
My first attempt at that scene was “old fisherman working, harbour, morning”. It came back as a stock photo: clean clothes, flat overcast light, no story. Same model, same seed, thirty extra words, and now there’s a lantern and a reason for the shadow across his face. The model didn’t get smarter. I stopped making it guess.
Step 3: Use Numbers, Not Vibes
Two settings do most of the heavy lifting, and both have a point past which you’re just burning time.
- Schnell: 4 steps, guidance 0. Pushing to 8 steps doesn’t add detail, because the model is distilled for four.
- Dev: 20-28 steps, guidance 3.0-4.0. Once you pass 30 steps you’re rendering noise that isn’t there.
- Resolution: generate at 1024×1024 or 1024×768. Go wider than 1440px at generation time and limbs start duplicating.
- Upscaling: do it afterwards with a dedicated upscaler at 2x, then a light inpaint pass over faces and hands.
A test worth running once: same prompt, same seed, guidance at 2.0, 3.5 and 5.0, three images side by side. The tipping point is usually between 3.5 and 4.5, where colour goes oversaturated and skin turns waxy. Once you’ve seen it on your own subject matter you never have to guess again.
Step 4: Fix Hands and Text Instead of Re-rolling
Re-rolling is the most expensive habit in image generation. Four seeds at 1024px cost roughly the same wall time as one seed at 2048px, and one of those four is usually usable.
Hands
Hands fail when they’re small and doing something precise. Two fixes work reliably. Compose so hands are either large in frame or out of frame entirely, and if a hand has to thread a needle or hold a pen, generate the wide pose first at 1024, then inpaint just the hand region at 0.45-0.55 denoise with a prompt describing only the hand. Cropping tighter before you inpaint makes the result blend better; the seam shows up when you give the model too much surrounding context.
Text in images
FLUX handles a short word on a sign better than most open models, but packaging, menus and multi-line copy still wobble at small sizes. When a shot lives or dies on legible type, either add the type in post or move that job to a model built specifically around rendering readable text. Fighting a general model for clean lettering is a waste of an afternoon.
Step 5: Get Consistency With LoRAs, Not Luck
Keeping the same face or product across forty images with prompt wording alone will eat a weekend. A LoRA does it in one line: append the trigger word and set the weight. Start at 0.7. Move to 0.85 only if the likeness is slipping, and stay below 0.9, where output tends to go flat and plasticky.
Adapters are small, usually a couple of hundred megabytes, and most are hosted publicly on the model hub where a million AI models live, so trying three before committing to one costs you ten minutes and no money.
ControlNet and IP-Adapter cover the other half of consistency: composition and palette. Reach for a depth or pose map when a client sends a reference layout, and IP-Adapter when they send a mood board and say “like this, but with our product”.
Step 6: Batch, Log, Then Refine
The single biggest speed gain isn’t a setting. It’s bookkeeping. Every prompt I keep gets one line in a plain text file: variant, seed, steps, guidance, LoRA weights, and a one-word verdict. Six weeks later that file is worth more than any downloaded prompt pack.
- Generate four seeds per prompt, then stop and choose one.
- Lock the seed before changing a single word, so you can see what the change actually did.
- Inpaint the broken region rather than regenerating the frame.
- Upscale once, at the very end.
If you’re still deciding whether FLUX is the right tool for the work in front of you, don’t read benchmarks. Push your five hardest real briefs through two or three models and compare side by side, the way you would test AI image generators against your own real work. Your briefs are the only benchmark that matters.
Step 7: Decide Where FLUX Lives
Local gives you unlimited runs for the price of electricity, and ComfyUI is the least painful way to run it. A workflow you build once becomes the template for everything after it. Cloud rentals make sense for a fortnight of heavy production. The API suits pipelines where images need to appear behind a queue.
One shortcut worth keeping: use a general chat assistant to expand a one-line brief into a layered prompt, roughly the way a conversational model can restructure messy notes into something usable, and then edit the result yourself. Raw model output leans toward adjective soup, and FLUX will faithfully render every unnecessary adjective you leave in.
What a Realistic Run Looks Like
A twelve-image product set: about 25 minutes of prompting and batching, 15 minutes of picking and inpainting, 10 minutes of upscaling. A thirty-panel storyboard: roughly three hours, most of it spent on panel one, because panel one is where you find the LoRA weight, the guidance value and the prompt structure that the other twenty-nine panels inherit.
The point of learning the settings is that they eventually stop being decisions. Once guidance sits at 3.5 and every prompt starts with the subject, your attention moves to the brief instead of the sliders, and the folder stops filling up with melting coffee cups.

