Guide · Creating
AI image, video and audio
Generative media tools produce images, video, speech and music from a description or a reference. Producing one impressive result is easy and is not the skill. The skill is producing a set that belongs together: same character, same style, same voice, across twenty assets, on a deadline.
How long does it take to learn?
A striking single image on your first evening. A consistent set you would hand to a client, several weeks, because consistency is a different problem from quality and nothing about the first prepares you for the second.
The fastest path, in order
- 01
Learn what each family of tools is for
Image, video, speech and music are separate systems with separate limits. Choosing the wrong one and prompting harder is the most common way an afternoon disappears.
- 02
Solve consistency before you solve quality
Reference images, seeds, style guides and character references. A merely good image that matches the other nineteen is worth more than a spectacular one that does not, and every real brief is the second problem.
- 03
Learn the tells
Hands, text in images, reflections, physics in video, breath in speech. Knowing where each system fails is how you check work before a client does, and it is faster than looking for something wrong in general.
- 04
Build a repeatable pipeline
Prompt, reference, generate, select, fix, export, in the same order every time, written down. Improvisation is fine for one asset and falls apart at twenty, which is where deadlines actually live.
- 05
Keep the last mile human
Colour, crop, timing, mix. Generated media almost always needs a real editing pass, and the people whose output looks professional are doing this step rather than finding a better prompt.
- 06
Get the rights question straight
What a tool's licence allows commercially, what it says about training data, what your client will ask. Settle this before the work rather than after, because it is the question that can invalidate a finished project.
Tools worth your time
| Tool | What it is for |
|---|---|
| Midjourney | Still the strongest for style, with reference features for holding a look across a set. |
| Flux and Stable Diffusion | Open models you can run and fine-tune, which is where real consistency control lives. |
| Runway | Video generation plus the editing tools around it, built for people delivering work. |
| Sora, Veo and Kling | Text-to-video at the frontier: strong results, expensive iteration, weak fine control. |
| ElevenLabs | Speech and voice cloning, the current default for narration and dubbing. |
| Suno and Udio | Music generation, useful for scratch tracks and beds, licence terms worth reading closely. |
This is the fastest-moving corner of AI, and the leader changes every few months. The transferable parts are the workflow and the eye, which is why the emphasis belongs on consistency, review and the editing pass rather than on any tool's current interface.
Mistakes that cost people weeks
Optimising the single image
Forty minutes on one hero shot, then discovering the other nineteen assets cannot match it. Establish the look across three images first, then go deep on any of them.
Skipping the human edit
Straight-out-of-the-model work has a texture people recognise even when they cannot name it. A colour pass and a crop is usually the entire difference between looks generated and looks made.
Ignoring the licence until delivery
Commercial terms differ sharply between tools and tiers, and clients increasingly ask. Read it at the start, when it is a five-minute question rather than a finished project.