AIF-C01 notes
Developing generative AI solutions

Evaluating an FM

Why evaluate

The story: A driver who drives beautifully but delivers to the wrong addresses is no use to your business. You test them against what the business actually needs.

In AI/AWS terms: Evaluation checks whether the model meets the business objectives.

For the exam: Evaluation asks whether the model meets business objectives, not only technical scores.

Three evaluation methods

The story: Judging a cooking contest. Having expert chefs taste every dish is the best judgment, but slow and expensive. Having everyone cook the same standard dish lets you compare contestants fairly, but it might not be the dish your restaurant serves. A machine that measures salt and temperature is instant, but can't tell you whether the dish is delicious.

In AI/AWS terms:

Contest judgeMethodStrengthsWeaknesses
Expert chefsHuman evaluationThe gold standard. Judges coherence, relevance, factuality, and qualitySlow and expensive at scale
Standard dishBenchmark datasetsStandardized comparison across models and over timeMay not match your specific use case
Measuring machineAutomated metricsQuick and scalable for fast iterationMiss nuance and may not match human judgment

For the exam: Human evaluation is the gold standard. Benchmarks give standardized comparison. Automated metrics are fast but miss nuance.

Benchmarks to know

The story: Standard school exams: a general reading comprehension exam, a harder advanced version, a reading-and-answer-questions exam, and a translation exam.

In AI/AWS terms:

  • GLUE: language understanding (classification, Q&A, inference)
  • SuperGLUE: harder GLUE tasks
  • SQuAD: question answering
  • WMT: machine translation

For the exam: GLUE and SuperGLUE = language understanding. SQuAD = question answering. WMT = translation.

Metrics to know

The story: Grading student work against a model answer:

  • For a summary, did the student cover the key points of the model answer?
  • For a translation, how many of the student's word sequences match the model answer?
  • Did the student say the same thing in different words? Give them credit for meaning.
  • How surprised is a reader by each next word? Less surprise means more natural writing.
  • For a sorting task, how well did they balance getting things right and not missing any?

In AI/AWS terms:

Grading questionMetricMeasuresBest for
Covered the key points?ROUGEOverlap between generated text and a reference, focused on recallSummarization
Matching word sequences?BLEUn-gram precision against a referenceMachine translation
Same meaning?BERTScoreSemantic similarity using BERT embeddings and cosine similarityAny generation where meaning matters more than exact wording
How surprised?PerplexityHow well the model predicts the next token (lower is better)Language model quality
Balanced sorting?F1 scoreBalance of precision and recallClassification and entity recognition

Automated metrics give a first read. Combine them with human evaluation for a complete picture.

For the exam: ROUGE = summarization. BLEU = translation. BERTScore = meaning. Perplexity: lower is better.

On this page