Building an eval set
How do you know your prompt got better? If you are not measuring, you do not. And measuring with 10 examples is not much better than not measuring at all.
2 steps
120 XP
Open without an account
Start the lesson →
Sources
- Wilson, E. B. 1927 · Probable Inference, the Law of Succession, and Statistical Inference · JASA, 22(158)
- Liang, P. et al. 2023 · Holistic Evaluation of Language Models (HELM) · TMLR
- Zheng, L. et al. 2023 · Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena · NeurIPS 2023
- Alpaydın, E. 1999 · Combined 5×2cv F Test for Comparing Supervised Classification Learning Algorithms · Neural Computation, 11(8)
ML Academy · an interactive machine learning course that runs in your browser ·
All lessons