Teaching obedience: instruction tuning
Pretraining teaches the model the operations but leaves no way of telling it which one you want. Instruction tuning opens that route, and adds no new ability.
4 steps
225 XP
A free account is needed
Start the lesson →
Sources
- Wei, J. et al. 2022 · Finetuned Language Models Are Zero-Shot Learners (FLAN) · ICLR 2022
- Sanh, V. et al. 2022 · Multitask Prompted Training Enables Zero-Shot Task Generalization (T0) · ICLR 2022
- Ouyang, L. et al. 2022 · Training Language Models to Follow Instructions with Human Feedback (InstructGPT) · NeurIPS 2022
- Zhou, C. et al. 2023 · LIMA: Less Is More for Alignment · NeurIPS 2023
ML Academy · an interactive machine learning course that runs in your browser ·
All lessons