AIF-C01 notes
Optimizing foundation models

Model Evaluation

Three ways to grade written text

The story: A teacher grades a student's writing against a model answer, three ways:

  • For a summary: "Did you mention all the important points from the model answer?" It's fine if you added a bit extra, but missing key points costs you.
  • For a translation: "How much of what you wrote matches the model answer word for word?" And writing just three words to avoid mistakes is penalized.
  • For anything: "Did you mean the same thing, even if you used different words?" "Car" and "automobile" both get credit.

In AI/AWS terms: Three metrics compare generated text with human-written reference text:

Grading styleMetricMeasuresFocusBest for
All the important points?ROUGEOverlap of words, n-grams, or sequences with the referenceRecall: how much of the important information was capturedSummarization (also translation)
Word-for-word match, no cheating by being shortBLEUn-gram precision against reference translations, with a brevity penalty for overly short outputPrecisionMachine translation
Same meaning?BERTScoreCosine similarity of contextual BERT embeddingsSemantic similarity: recognizes paraphrases and synonymsAny task where meaning matters more than exact words

For the exam: ROUGE = recall, summarization. BLEU = precision, translation, brevity penalty. BERTScore = meaning.

ROUGE variants

The story: One teacher counts how many single words, or pairs of words, match the model answer. Another checks whether the key points appear in the same order, because a summary with the right points in a jumbled order is harder to follow.

In AI/AWS terms:

  • ROUGE-N: n-gram overlap (ROUGE-1 for single words, ROUGE-2 for pairs). Measures fluency and coverage of key ideas.
  • ROUGE-L: longest common subsequence. Measures coherence and order.

For the exam: ROUGE-N = n-gram overlap. ROUGE-L = longest common subsequence.

Limitations

The story: The word-matching teachers mark down a student who says "automobile" when the model answer says "car", even though that's correct. The translation teacher also can't really tell whether the grammar flows. So the school pairs them with the meaning-focused teacher.

In AI/AWS terms:

  • ROUGE and BLEU depend on exact matches, so they penalize valid paraphrases.
  • BLEU struggles to judge fluency and grammar.
  • BERTScore is often used alongside them for a fuller picture.

For the exam: ROUGE and BLEU penalize paraphrases. BERTScore handles them.

AnyCompany results

The story: The fashion store's personal stylist got better at descriptions and advice, and the shop's numbers went up with it.

In AI/AWS terms:

  • Conversion rate rose 15%, with ROUGE averaging 0.85.
  • Average order value rose 20%, with BLEU at 0.78.
  • Customer retention rose 25%, with BERTScore averaging 0.90.

For the exam: Good text metrics matter because they move business metrics like conversion, order value, and retention.

On this page