Modifying Prompts
Inference parameters: randomness and diversity
The story: A model writes one word at a time, picking the next word the way you'd pick a snack from a vending machine. Some snacks are popular, some rarely chosen. Three dials on the machine change how you pick:
- An adventure dial. Turned down, you always take the most popular snack. Turned up, you're more willing to try something unusual.
- A "popular enough" rule: only choose from the most popular snacks until their combined share of sales reaches a chosen percentage. Set it low and only the top sellers qualify. Set it high and more snacks qualify.
- A "top K" rule: only choose from the K best-selling snacks. K = 10 keeps it safe. K = 500 opens up almost everything.
In AI/AWS terms: Which parameters are available depends on the model.
| Dial | Parameter | What it does | Lower value | Higher value |
|---|---|---|---|---|
| Adventure | Temperature | Reshapes the probability distribution of the next token | More deterministic, focused, repeatable | More diverse, creative, random |
| "Popular enough" share | Top P | Samples only from the smallest set of tokens whose probabilities add up to P | Only the most likely tokens | A wider range of tokens |
| Top K sellers | Top K | Samples only from the K most probable tokens | Focused, coherent (for example K = 10) | More varied (for example K = 500) |
For the exam: Low temperature = consistent and deterministic. High temperature = creative and varied. Top P is a probability cutoff; Top K is a count cutoff.
Inference parameters: length
The story: Two ways to stop a speech: a timer that cuts the speaker off after five minutes no matter what, and a code word ("thank you, everyone") that ends it the moment it's said, even if there's time left.
In AI/AWS terms:
- Maximum length: caps the number of generated tokens to avoid runaway output (the timer).
- Stop sequences: tokens that end generation as soon as they appear, whatever the maximum length (the code word).
For the exam: Maximum length caps tokens. Stop sequences end output as soon as they appear.
Prompting best practices
The story: How to brief a new assistant well:
- Speak plainly and keep it short.
- Tell them what it's for.
- Say what shape you want: "a short list", "one paragraph", "under 100 words".
- Put the actual request at the end, after the background.
- Frame it as a question: who, what, where, when, why, how.
- Show a sample of what a good answer looks like.
- For a big job, split it into smaller steps, or ask them to talk through it step by step.
- Try different ways of asking.
- Keep a standard form for requests you make often.
In AI/AWS terms:
- Be clear and concise, in natural language.
- Include context, such as what the output is for.
- Use directives for the response type: summary, full sentence, list, length.
- Put the requested output at the end of the prompt.
- Start with a question: who, what, where, when, why, how.
- Give an example response in brackets.
- Break up complex tasks into subtasks or separate prompts, or ask the model to think step by step.
- Experiment with different prompts.
- Use prompt templates for consistent, reliable inputs.
For the exam: Put the requested output at the end, break complex tasks up, and use templates for consistency.
Scenario update
The story: For the finance report, the company turns the adventure dial up for more creative ideas, allows a long answer, says the report is for small and medium finance businesses, and lists the sections it wants.
In AI/AWS terms: Temperature 0.9 and top p 0.999 for more creative output, maximum length 5,000, the finance industry and SMB audience as context, and a list of report sections as the directive.
For the exam: Raising temperature and top p makes output more creative.