Multimodal models: bringing image and text into one space
CLIP's core is genuinely trained here. Two of the three results do not come out as expected.
4 steps
225 XP
A free account is needed
Start the lesson →
Sources
- Radford, A. et al. 2021 · Learning Transferable Visual Models From Natural Language Supervision · ICML 2021 · CLIP
- Liang, W. et al. 2022 · Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning · NeurIPS 2022
- Oord, A. van den, Li, Y. & Vinyals, O. 2018 · Representation Learning with Contrastive Predictive Coding · arXiv:1807.03748 · InfoNCE
ML Academy · an interactive machine learning course that runs in your browser ·
All lessons