AIF-C01 notes
Developing generative AI solutions

Evaluating an FM

11 exam-style questions on this lesson.

  1. Question 1 of 11Which evaluation method is considered the gold standard for judging coherence, relevance, and factuality?
  2. Question 2 of 11A team wants a standardized way to compare several models against each other and track them over time. Which method fits?
  3. Question 3 of 11What is a weakness of automated metrics?
  4. Question 4 of 11Which metric is best suited to evaluating text summarization?
  5. Question 5 of 11Which metric is best suited to evaluating machine translation?
  6. Question 6 of 11A team wants a metric that gives credit when the generated text means the same as the reference but uses different words. Which metric fits?
  7. Question 7 of 11What does a lower perplexity indicate for a language model?
  8. Question 8 of 11Which benchmark is designed for question answering?
  9. Question 9 of 11A team finds GLUE too easy for modern models. Which benchmark offers harder language understanding tasks?
  10. Question 10 of 11Which metric fits a named entity recognition task where you need a balance of precision and recall?
  11. Question 11 of 11Why does evaluation matter in the generative AI lifecycle?