AIF-C01 notes
Essentials of prompt engineering

Modifying Prompts

Inference parameters: randomness and diversity

The story: A model writes one word at a time, picking the next word the way you'd pick a snack from a vending machine. Some snacks are popular, some rarely chosen. Three dials on the machine change how you pick:

  • An adventure dial. Turned down, you always take the most popular snack. Turned up, you're more willing to try something unusual.
  • A "popular enough" rule: only choose from the most popular snacks until their combined share of sales reaches a chosen percentage. Set it low and only the top sellers qualify. Set it high and more snacks qualify.
  • A "top K" rule: only choose from the K best-selling snacks. K = 10 keeps it safe. K = 500 opens up almost everything.

In AI/AWS terms: Which parameters are available depends on the model.

DialParameterWhat it doesLower valueHigher value
AdventureTemperatureReshapes the probability distribution of the next tokenMore deterministic, focused, repeatableMore diverse, creative, random
"Popular enough" shareTop PSamples only from the smallest set of tokens whose probabilities add up to POnly the most likely tokensA wider range of tokens
Top K sellersTop KSamples only from the K most probable tokensFocused, coherent (for example K = 10)More varied (for example K = 500)

For the exam: Low temperature = consistent and deterministic. High temperature = creative and varied. Top P is a probability cutoff; Top K is a count cutoff.

Inference parameters: length

The story: Two ways to stop a speech: a timer that cuts the speaker off after five minutes no matter what, and a code word ("thank you, everyone") that ends it the moment it's said, even if there's time left.

In AI/AWS terms:

  • Maximum length: caps the number of generated tokens to avoid runaway output (the timer).
  • Stop sequences: tokens that end generation as soon as they appear, whatever the maximum length (the code word).

For the exam: Maximum length caps tokens. Stop sequences end output as soon as they appear.

Prompting best practices

The story: How to brief a new assistant well:

  • Speak plainly and keep it short.
  • Tell them what it's for.
  • Say what shape you want: "a short list", "one paragraph", "under 100 words".
  • Put the actual request at the end, after the background.
  • Frame it as a question: who, what, where, when, why, how.
  • Show a sample of what a good answer looks like.
  • For a big job, split it into smaller steps, or ask them to talk through it step by step.
  • Try different ways of asking.
  • Keep a standard form for requests you make often.

In AI/AWS terms:

  • Be clear and concise, in natural language.
  • Include context, such as what the output is for.
  • Use directives for the response type: summary, full sentence, list, length.
  • Put the requested output at the end of the prompt.
  • Start with a question: who, what, where, when, why, how.
  • Give an example response in brackets.
  • Break up complex tasks into subtasks or separate prompts, or ask the model to think step by step.
  • Experiment with different prompts.
  • Use prompt templates for consistent, reliable inputs.

For the exam: Put the requested output at the end, break complex tasks up, and use templates for consistency.

Scenario update

The story: For the finance report, the company turns the adventure dial up for more creative ideas, allows a long answer, says the report is for small and medium finance businesses, and lists the sections it wants.

In AI/AWS terms: Temperature 0.9 and top p 0.999 for more creative output, maximum length 5,000, the finance industry and SMB audience as context, and a list of report sections as the directive.

For the exam: Raising temperature and top p makes output more creative.

On this page