A caption tells you what is in a picture. A prompt tells you how it was made. This is the difference, the five layers a good reading has to extract, and how WEDO AI Studio turns a reference into a prompt for your own product.
You have a reference — a shot you saved, a competitor's ad, a frame from a film, something a client sent with "make it look like this". Image to prompt reads that picture and writes the prompt that would produce its treatment: how it is lit, what lens it was seen through, how the colour is graded, what the surfaces are doing, how the frame is composed.
That is a different job from captioning. A caption says "a woman in a red coat on a wet street at night". A prompt says the light is a single practical behind her throwing a rim along the shoulder, the street is wet enough to act as a second light source, it is a 35mm at f/2 so the background lights bloom into soft circles, and the grade is teal shadows with the reds left hot. The first sentence gets you a stock photo. The second gets you the look.
Direction, quality, count. One hard source or a broad soft one; whether there is fill or the shadows are left to fall; where the practicals are; what is being blocked. This is the layer that carries almost all of a photograph's character, and the layer most people never notice consciously.
Wide-angle edge stretch, telephoto compression, how much of the background is dissolved and how the out-of-focus highlights are shaped. Read from perspective cues rather than guessed.
Where the shadows are pushed, where the highlights are pushed, which colour has been allowed to stay saturated and which has been drained. Naming the grade is what makes two images from the same prompt feel like one campaign.
Wet, matte, brushed, powdered, worn. Whether there is dust, condensation, fingerprints, fabric slub. These details are what separate photographic output from plastic output.
Crop ratio, where the subject sits, where the negative space is and whether it is doing a job — leaving room for a headline is a different decision from leaving room for air.
Worth separating, because people ask for one and mean the other.
A bottle of olive oil on a wooden table with some herbs, natural light, rustic style
True, and useless. Regenerating from this gives you a generic food-blog photo, not the image you were looking at.
Single soft window source from camera-right, roughly 60 degrees off axis, no fill, so the
right side of the bottle glows and the left falls into a deep warm shadow that still holds
detail. Shot on a 60mm macro at f/2.8, focus on the bottle shoulder, the table foreground
dissolving within 15cm. Unfinished oak with visible open grain and a faint oil ring;
a scatter of thyme, slightly wilted, not styled straight. Grade is warm and low contrast
with the greens desaturated and the amber of the oil left as the only fully saturated
colour. Vertical 4:5, bottle on the right third, negative space left for type.Reading a reference for its technique is what art directors have always done — lighting setups and lens choices are not ownable. Reproducing a specific photograph, an identifiable person, a copyrighted character or a competitor's protected creative is a different thing, and not something to do with a prompt tool. Use references to learn a treatment and then apply it to your own subject. That is both the legally sound path and, in practice, the one that produces work that looks like yours.
You can also attach reference images anywhere in the tool — up to six on the image generation page — and the engine is instructed to read what it can actually see rather than paraphrase your text back at you.
It is a tool that analyses a picture and writes the AI prompt that would recreate its treatment — the lighting direction and quality, the lens and depth of field, the colour grade, the surface materials and the composition. It is not the same as image captioning, which describes what is in the picture rather than how it was made.
A description names the subject; a prompt names the decisions. 'A woman in a red coat on a wet street' will get you a generic stock photo. Naming the single rim light behind her, the 35mm at f/2, the wet ground acting as a second source and the teal-shadow grade gets you the look you were pointing at.
Yes, and it is worth treating as a separate job from style reference. In product-match mode you upload the reference and state what your actual product is, and the prompt is rebuilt around your item while the lighting, lens, grade and composition of the reference are preserved.
Learning a technique from a reference — a lighting setup, a lens choice, a colour grade — is standard creative practice and those things are not ownable. Reproducing a specific copyrighted photograph, an identifiable person or a protected character is a different matter and not something we would recommend doing with any tool. Use references to learn a treatment, then apply it to your own subject.
Any of them. The output is written in the conventions of whichever target model you select — Midjourney parameters, a flowing natural paragraph for Nano Banana, Gemini Image or GPT Image, or a nested JSON prompt if you need structured output for an API or batch job.
Yes, for stills. Export a frame and upload it like any other image. What it reads is the treatment of that frame; for the motion side of a clip you would want the video prompt tool instead, which handles camera movement, beat timing and motion physics.
Upload any reference image and get the lighting, lens, grade and composition written out as a prompt you can rerun on your own product. 30 free generations.
Try WEDO AI Studio free →