Model Evaluation
Three ways to grade written text
The story: A teacher grades a student's writing against a model answer, three ways:
- For a summary: "Did you mention all the important points from the model answer?" It's fine if you added a bit extra, but missing key points costs you.
- For a translation: "How much of what you wrote matches the model answer word for word?" And writing just three words to avoid mistakes is penalized.
- For anything: "Did you mean the same thing, even if you used different words?" "Car" and "automobile" both get credit.
In AI/AWS terms: Three metrics compare generated text with human-written reference text:
| Grading style | Metric | Measures | Focus | Best for |
|---|---|---|---|---|
| All the important points? | ROUGE | Overlap of words, n-grams, or sequences with the reference | Recall: how much of the important information was captured | Summarization (also translation) |
| Word-for-word match, no cheating by being short | BLEU | n-gram precision against reference translations, with a brevity penalty for overly short output | Precision | Machine translation |
| Same meaning? | BERTScore | Cosine similarity of contextual BERT embeddings | Semantic similarity: recognizes paraphrases and synonyms | Any task where meaning matters more than exact words |
For the exam: ROUGE = recall, summarization. BLEU = precision, translation, brevity penalty. BERTScore = meaning.
ROUGE variants
The story: One teacher counts how many single words, or pairs of words, match the model answer. Another checks whether the key points appear in the same order, because a summary with the right points in a jumbled order is harder to follow.
In AI/AWS terms:
- ROUGE-N: n-gram overlap (ROUGE-1 for single words, ROUGE-2 for pairs). Measures fluency and coverage of key ideas.
- ROUGE-L: longest common subsequence. Measures coherence and order.
For the exam: ROUGE-N = n-gram overlap. ROUGE-L = longest common subsequence.
Limitations
The story: The word-matching teachers mark down a student who says "automobile" when the model answer says "car", even though that's correct. The translation teacher also can't really tell whether the grammar flows. So the school pairs them with the meaning-focused teacher.
In AI/AWS terms:
- ROUGE and BLEU depend on exact matches, so they penalize valid paraphrases.
- BLEU struggles to judge fluency and grammar.
- BERTScore is often used alongside them for a fuller picture.
For the exam: ROUGE and BLEU penalize paraphrases. BERTScore handles them.
AnyCompany results
The story: The fashion store's personal stylist got better at descriptions and advice, and the shop's numbers went up with it.
In AI/AWS terms:
- Conversion rate rose 15%, with ROUGE averaging 0.85.
- Average order value rose 20%, with BLEU at 0.78.
- Customer retention rose 25%, with BERTScore averaging 0.90.
For the exam: Good text metrics matter because they move business metrics like conversion, order value, and retention.