Prompting for Different Modalities: Text, Code, and Images
We break down what prompting techniques and vocabulary work best for different modalities.
A prompt that gets you a great blog post will not get you great code, and a prompt that gets great code will fall flat if you're actually trying to describe an image. Each modality, text, code, and images, rewards a different vocabulary and a different level of specificity. Here's what actually works for each.
Prompting for Text
Text generation is the most forgiving modality, but it's also where vague requests hurt you most because there's no objectively "wrong" output to catch. Precision here comes from four things: audience, tone, length, and structure.
Instead of "Write about our new product," specify: "Write a 150-word announcement for our email newsletter, aimed at existing customers, in a warm and confident tone, ending with a single clear call to action." Every one of those four details removes a decision the model would otherwise have to guess at, and every guess is a chance for the output to miss what you actually wanted.
Prompting for Code
Code prompting rewards a completely different kind of specificity: naming your exact tools, versions, and constraints. "Write a function to sort a list" is technically answerable a hundred different ways. "Write a Python function using only the standard library that sorts a list of dictionaries by the 'date' key in descending order, and handles missing keys by treating them as the oldest possible date" gives the model almost no room to guess wrong.
Three things matter most for code prompts:
Name your stack. Language, framework, and version. "Using React 18 with TypeScript" produces meaningfully different code than a generic "using React."
State your constraints up front. No external dependencies, must run in a browser, needs to handle edge case X. Constraints stated after the fact often get ignored in a first draft.
Ask for the reasoning on anything non-trivial. For genuinely tricky logic, asking the model to briefly explain its approach before writing the code catches faulty assumptions before they're baked into fifty lines of code you now have to unwind.
Prompting for Images
Image prompting is the modality where vocabulary matters most, because there's no back-and-forth clarification, you get one shot per generation. The most useful structure is: subject, style, composition, and lighting, in that order.
"A cat" produces something generic. "A ginger tabby cat sitting on a windowsill, soft afternoon light, shallow depth of field, photographed in the style of a warm lifestyle photograph" gives the model concrete visual anchors for subject, lighting, and style all at once.
A few vocabulary notes that consistently help:
Named art styles and camera terms (35mm film, isometric, watercolor, cinematic lighting) do far more work than adjectives like "nice" or "professional looking."
Negative space matters. If you need room for text overlay, say so directly: "with empty space in the upper third for a headline."
Iterate on one variable at a time. Change the lighting description, regenerate, compare. Changing five things at once makes it impossible to tell which change helped.
Key Takeaways
Text rewards specificity about audience, tone, length, and structure.
Code rewards specificity about your stack, your constraints, and asking for reasoning on anything tricky.
Images reward concrete visual vocabulary: named styles, lighting terms, and composition details, since there's no clarifying question round.
Across all three, the same rule holds: the model fills every gap you leave with its own guess. The fewer gaps, the closer the first result lands to what you actually pictured.