Lyra
Learn
AI Learning Platform
π Midnight
π Comfort
π₯ Ember
π Paper
β Contrast
Exams
Sign in to track progress
β Back to the lesson
Module 23 Β· Quiz
Evaluating AI Systems
1. What is a golden set in the context of evaluating AI systems?
A collection of random questions for testing
A curated set of real questions with expected outcomes
A database of model training data
A framework for judging AI outputs
2. What are retrieval metrics focused on during the evaluation of a RAG system?
They measure the end-user satisfaction
They assess the quality of the summarized output
They determine if the correct information chunks were retrieved
They evaluate the overall accuracy of the AI model
3. How can bias in LLM scoring of answers be mitigated according to the lesson?
By always using the same model for judging
By avoiding human calibration entirely
By using randomized ordering in comparisons
By allowing judges to choose their preferred answers
4. What role do regression evaluations serve in the evaluation process?
They allow for one-time checks of model accuracy
They help track changes in quality continuously
They discard older sets of questions
They are only required before new model releases
Submit answers
Continue: LLMOps — Release Discipline β
Review this lesson
Retake quiz