Developing generative AI solutionsFull notesSummaryRingkasanStoriesPracticeEvaluating an FM11 exam-style questions on this lesson.Question 1 of 11Which evaluation method is considered the gold standard for judging coherence, relevance, and factuality?AAutomated metricsBPerplexityCHuman evaluationDBenchmark datasetsCheck answerQuestion 2 of 11A team wants a standardized way to compare several models against each other and track them over time. Which method fits?AHuman evaluation onlyBGuardrailsCA/B testing in production onlyDBenchmark datasetsCheck answerQuestion 3 of 11What is a weakness of automated metrics?AThey can miss nuance and may not match human judgmentBThey can't be computed for generated textCThey are slow to compute at scaleDThey are expensive to run on large test setsCheck answerQuestion 4 of 11Which metric is best suited to evaluating text summarization?APerplexityBROUGECR squaredDBLEUCheck answerQuestion 5 of 11Which metric is best suited to evaluating machine translation?AMSEBROUGECBLEUDF1 scoreCheck answerQuestion 6 of 11A team wants a metric that gives credit when the generated text means the same as the reference but uses different words. Which metric fits?ABLEUBROUGECAccuracyDBERTScoreCheck answerQuestion 7 of 11What does a lower perplexity indicate for a language model?AThe model predicts the next token betterBThe model is more random and creativeCThe model produces more toxic outputDThe model uses fewer tokens per answerCheck answerQuestion 8 of 11Which benchmark is designed for question answering?AGLUEBSQuADCWMTDBLEUCheck answerQuestion 9 of 11A team finds GLUE too easy for modern models. Which benchmark offers harder language understanding tasks?ASQuADBWMTCSuperGLUEDROUGECheck answerQuestion 10 of 11Which metric fits a named entity recognition task where you need a balance of precision and recall?ABERTScoreBPerplexityCBLEUDF1 scoreCheck answerQuestion 11 of 11Why does evaluation matter in the generative AI lifecycle?AIt checks whether the model meets the business objectivesBIt trains the model on the evaluation dataCIt replaces the need for a deployment stageDIt removes the need for monitoring laterCheck answerImproving the Performance of an FM13 exam-style questions on this lesson.Deploying the Application6 exam-style questions on this lesson.