Type “fisherman mending a net at dawn, cold blue light, film grain” into a box, wait eight seconds, and four images appear. One is usable. That gap between usable and finished is where most of the real work in artificial intelligence drawing happens.
Since 2022 these tools have gone from party trick to default first step for a lot of visual jobs: mood boards, storyboards, indie game assets, tattoo roughs, architectural massing studies, book jackets. They have also stayed weirdly bad at things any ten-year-old can do. Knowing which is which saves whole afternoons.
Four different jobs hiding behind one phrase
“AI drawing” gets used for at least four separate tasks, and mixing them up is the quickest route to frustration.
- Text to image. Midjourney, DALL·E, Stable Diffusion, Firefly, Ideogram. You describe, it renders.
- Image to image and inpainting. You supply the base picture and edit a masked region. Most paid studio work lives here.
- Guided generation. ControlNet, pose skeletons, depth maps, edge detection. You lock the composition and let the model handle surface.
- Grunt work. Line cleanup, colour flats, background removal, upscaling, denoising. Unglamorous, and quietly the biggest time saver of the lot.
If you need one character in one exact pose, prompting from a blank canvas is the wrong instrument. If you need forty thumbnail directions before lunch, it is exactly the right one.
What is actually happening when it paints
Diffusion models learn by un-ruining pictures. Training adds noise to an image until nothing is left but static, then the network practises reversing that, one step at a time, with your caption attached as a hint. Generation runs the same trick backwards from pure noise.
The trick that made this affordable: Stable Diffusion 1.5 does not paint at 512×512 pixels. It works in a compressed latent space of 64×64 and decodes at the end, which drops the memory needed to something a mid-range gaming card can manage. Its training slice of the LAION-5B dataset leaned on roughly two billion English image-caption pairs. That scale explains why it understands “brutalist” and “shot on Portra 400”, and also why it inherited every lighting habit, framing cliché, and stock-photo watermark in the crawl.
Why prompts behave the way they do
Your text is chopped into tokens and squeezed into a mathematical summary that the model nudges its denoising toward. Word order matters less than you would expect, and early tokens carry more weight than later ones. That is why “red silk dress, white background” and “white background, red silk dress” can produce noticeably different results, and why cramming in fifteen modifiers usually muddies an image instead of refining it.
The workflow that actually holds up
Professionals rarely generate a finished picture. They generate scaffolding, then draw.
- Block the composition with a rough sketch or a crude 3D proxy, then push it through a structure model so the AI respects your staging.
- Generate twenty variants at low resolution. Cheap, fast, disposable.
- Pick one, mask the parts that work, regenerate the parts that do not.
- Upscale, then paint over the failures by hand. Hands, jewellery, small type, and any straight edge that carries meaning.
A 4× upscale from 512px to 2048px takes seconds on a decent card and is usually still not enough for print. Add detail at the upscale stage, fix, upscale again. Slower, much better.
Consistency is the hard part
Keeping the same character across twelve images is the problem that eats the most time. Fine-tuned adapters, fixed seeds, reference images, and IP-Adapter style conditioning all help. None of them fully solve it. For comic work, a lot of artists settle on generating reference sheets, then drawing the panels by hand with the output taped beside the drawing board.
Where AI drawing still falls over
Ask the models to do any of the following and watch them stumble.
- Two people touching on purpose. Hands still merge and fingers still multiply, though far less than in 2023.
- Text longer than a few words. Logos, signage, and packaging labels remain a lottery.
- Spatial logic across a series. A door on the left in panel one drifts right by panel four.
- Physics. Reflections, shadows, and liquids follow rules the model approximates rather than knows.
- Anything specific about your own house, your dog, or your mother’s face without a reference-driven model.
There is also the sameness problem. Prompt a thousand people with the same words and you get a recognisable house style: glossy skin, teal-and-orange grading, shallow depth of field, centred composition. It is a tell, and art directors now spot it instantly. That is one reason the smarter workflows lean hard on editing rather than straight generation.
Copyright, likeness, and fights nobody has settled
The legal ground keeps shifting. In the United States, the Copyright Office has ruled that images generated purely from a prompt cannot be copyrighted, because there is no human author at the wheel. The same filing that protected Kris Kashtanova’s Zarya of the Dawn covered the text and the arrangement, not the generated art. Getty Images is suing Stability AI on both sides of the Atlantic over training data, and several artist class actions are grinding through discovery.
Training-data consent has become a selling point. Adobe Firefly was trained on Adobe Stock and licensed material, and offers indemnification for enterprise customers. That clause alone decides the question for plenty of studios.
Artists have pushed back with software as well as lawyers. Glaze and Nightshade add near-invisible perturbations that make an image hostile to training, or quietly poison the model that scrapes it. Whether they hold up long term is unclear. What they signal is not: people who draw for a living want a say in this.
This sits inside a wider argument about who gets to use machines to make images, and at whose expense. The Pope’s letter on AI and human dignity is worth reading even if you share none of the premises, because it frames the questions regulators keep circling back to.
Likeness is the next battleground. Non-consensual sexual imagery, celebrity deepfakes, and covert recording all ride the same rails as a harmless character sketch, which is why privacy law is moving quickly. Wearable cameras have already triggered bans, a pattern worth watching in Norway’s move against camera-enabled glasses.
Local, cloud, or plugin
Three practical routes, and the differences are less about quality than about terms.
- Local. Stable Diffusion variants run on 8 to 12GB of VRAM for 1.5-class models, and 16GB or more for the newer, heavier ones. Free, uncensored, fiddly.
- Cloud. Midjourney, Firefly, Ideogram. Fewer knobs, better defaults, subscriptions, clearer commercial terms.
- Plugin. Photoshop, Krita, Blender, and Figma all have AI panels built in now. Handy, because you never leave the file you are working in.
Expect a fraction of a cent to a few cents per image in the cloud, and an hour or two of setup for a local install. Neither path is expensive. The decision hinges on licensing, not cost.
Where this is heading
Two things are happening at once. The models are getting better at following instructions, so requests like “a bicycle with the chain on the wrong side” are now tests they mostly pass. And the money is consolidating around a handful of players with the compute to keep training. The commercial pressure shaping that race is its own story, and the reality of pitching AI products to investors explains why so many of these tools look identical after eighteen months.
Meanwhile the architecture that draws a fox in a waistcoat turns out to be manipulable. Researchers keep finding ways to steer assistants off their rails with clever phrasing, and attacks that exploit chatbot personalities have close cousins in image pipelines, where hidden instructions can leak data or override your settings.
Not every organisation is saying yes. A council in England spent years and a serious amount of public money deciding it wanted nothing to do with one particular AI contractor, a small but telling case documented in the English county that said no to Palantir. Public bodies are asking the same questions artists ask: whose work trained this, who owns the output, and can we walk away if we need to.
Before you publish anything made this way, do three dull things. Read the licence for the model you used, since commercial rights differ wildly between them. Keep the C2PA provenance metadata intact if the tool attaches it, because clients and platforms are starting to check. And disclose the tools in your credit line, the same way you would credit a photographer or an assistant. None of that is exciting, and all of it is the difference between a quick win and a ruined deadline.

