Evaluate Results
Human evaluation versus benchmark datasets
The story: Testing a new chef two ways. You can have real diners eat and tell you how it felt: was the food right for the occasion, was it creative? Or you can give every chef the same written test with fixed correct answers and a timer, which makes it easy to compare chefs, or the same chef this year versus last year.
In AI/AWS terms:
| Human evaluation (diners) | Benchmark datasets (written test) | |
|---|---|---|
| Kind | Qualitative | Quantitative and objective |
| Measures | User experience, contextual appropriateness, creativity and flexibility | Accuracy, speed and efficiency, scalability |
| Best for | Iterative tuning to meet user expectations | Initial testing and comparing models or model versions |
For the exam: Comparing models or model versions → benchmark datasets. User experience and creativity → human evaluation.
Building a benchmark dataset
The story: Senior chefs write a set of tough test questions about your cuisine, along with the correct answers and which cookbook page each answer comes from. The new chef answers, and their answers are marked against the senior chefs' answers. To save time, a trusted senior chef can do the marking.
In AI/AWS terms:
- Subject matter experts (SMEs) write relevant, challenging questions about the topic or documents.
- SMEs provide the correct answers and the context they come from.
- The model answers the questions, and its answers are scored against the SMEs' answers.
Scoring can be automated with LLM as a judge: another LLM compares the model's answers to the benchmark answers (the trusted marker).
For the exam: SMEs write the benchmark questions and answers. LLM as a judge automates the scoring.
Combined approach
The story: The restaurant gives the written test before the chef starts, and collects diner reviews once they're cooking, so the chef keeps getting better.
In AI/AWS terms: Benchmarks show technical capability, and humans show real-world usefulness. AnyCompany tests against benchmarks before production and collects human ratings after, so the model keeps improving.
For the exam: Use both: benchmarks before production, human ratings after.