LLM-as-judge
The only practical way to score answer quality automatically. But using a judge without validating it is weighing things on a broken scale.
1 steps
60 XP
A free account is needed
Start the lesson →
Sources
- Zheng, L. et al. 2023 · Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena · NeurIPS 2023
- Wang, P. et al. 2023 · Large Language Models are not Fair Evaluators (konum yanlılığı) · ACL 2024
- Panickssery, A. et al. 2024 · LLM Evaluators Recognize and Favor Their Own Generations · NeurIPS 2024
- Cohen, J. 1960 · A Coefficient of Agreement for Nominal Scales (kappa) · Educational and Psychological Measurement, 20(1)
ML Academy · an interactive machine learning course that runs in your browser ·
All lessons