Evaluating an FM
Why evaluate
The story: A driver who drives beautifully but delivers to the wrong addresses is no use to your business. You test them against what the business actually needs.
In AI/AWS terms: Evaluation checks whether the model meets the business objectives.
For the exam: Evaluation asks whether the model meets business objectives, not only technical scores.
Three evaluation methods
The story: Judging a cooking contest. Having expert chefs taste every dish is the best judgment, but slow and expensive. Having everyone cook the same standard dish lets you compare contestants fairly, but it might not be the dish your restaurant serves. A machine that measures salt and temperature is instant, but can't tell you whether the dish is delicious.
In AI/AWS terms:
| Contest judge | Method | Strengths | Weaknesses |
|---|---|---|---|
| Expert chefs | Human evaluation | The gold standard. Judges coherence, relevance, factuality, and quality | Slow and expensive at scale |
| Standard dish | Benchmark datasets | Standardized comparison across models and over time | May not match your specific use case |
| Measuring machine | Automated metrics | Quick and scalable for fast iteration | Miss nuance and may not match human judgment |
For the exam: Human evaluation is the gold standard. Benchmarks give standardized comparison. Automated metrics are fast but miss nuance.
Benchmarks to know
The story: Standard school exams: a general reading comprehension exam, a harder advanced version, a reading-and-answer-questions exam, and a translation exam.
In AI/AWS terms:
- GLUE: language understanding (classification, Q&A, inference)
- SuperGLUE: harder GLUE tasks
- SQuAD: question answering
- WMT: machine translation
For the exam: GLUE and SuperGLUE = language understanding. SQuAD = question answering. WMT = translation.
Metrics to know
The story: Grading student work against a model answer:
- For a summary, did the student cover the key points of the model answer?
- For a translation, how many of the student's word sequences match the model answer?
- Did the student say the same thing in different words? Give them credit for meaning.
- How surprised is a reader by each next word? Less surprise means more natural writing.
- For a sorting task, how well did they balance getting things right and not missing any?
In AI/AWS terms:
| Grading question | Metric | Measures | Best for |
|---|---|---|---|
| Covered the key points? | ROUGE | Overlap between generated text and a reference, focused on recall | Summarization |
| Matching word sequences? | BLEU | n-gram precision against a reference | Machine translation |
| Same meaning? | BERTScore | Semantic similarity using BERT embeddings and cosine similarity | Any generation where meaning matters more than exact wording |
| How surprised? | Perplexity | How well the model predicts the next token (lower is better) | Language model quality |
| Balanced sorting? | F1 score | Balance of precision and recall | Classification and entity recognition |
Automated metrics give a first read. Combine them with human evaluation for a complete picture.
For the exam: ROUGE = summarization. BLEU = translation. BERTScore = meaning. Perplexity: lower is better.