AIF-C01 notes
Optimizing foundation models

Evaluate Results

Human evaluation versus benchmark datasets

The story: Testing a new chef two ways. You can have real diners eat and tell you how it felt: was the food right for the occasion, was it creative? Or you can give every chef the same written test with fixed correct answers and a timer, which makes it easy to compare chefs, or the same chef this year versus last year.

In AI/AWS terms:

Human evaluation (diners)Benchmark datasets (written test)
KindQualitativeQuantitative and objective
MeasuresUser experience, contextual appropriateness, creativity and flexibilityAccuracy, speed and efficiency, scalability
Best forIterative tuning to meet user expectationsInitial testing and comparing models or model versions

For the exam: Comparing models or model versions → benchmark datasets. User experience and creativity → human evaluation.

Building a benchmark dataset

The story: Senior chefs write a set of tough test questions about your cuisine, along with the correct answers and which cookbook page each answer comes from. The new chef answers, and their answers are marked against the senior chefs' answers. To save time, a trusted senior chef can do the marking.

In AI/AWS terms:

  1. Subject matter experts (SMEs) write relevant, challenging questions about the topic or documents.
  2. SMEs provide the correct answers and the context they come from.
  3. The model answers the questions, and its answers are scored against the SMEs' answers.

Scoring can be automated with LLM as a judge: another LLM compares the model's answers to the benchmark answers (the trusted marker).

For the exam: SMEs write the benchmark questions and answers. LLM as a judge automates the scoring.

Combined approach

The story: The restaurant gives the written test before the chef starts, and collects diner reviews once they're cooking, so the chef keeps getting better.

In AI/AWS terms: Benchmarks show technical capability, and humans show real-world usefulness. AnyCompany tests against benchmarks before production and collects human ratings after, so the model keeps improving.

For the exam: Use both: benchmarks before production, human ratings after.

On this page