Why tokenisers behave differently
Two models split the same text into different numbers of pieces. Where that difference comes from is written in the corpus the tokeniser was trained on.
3 steps
175 XP
A free account is needed
Start the lesson →
Sources
- Sennrich, R. et al. 2016 · Neural Machine Translation of Rare Words with Subword Units · ACL 2016
- Kudo, T. & Richardson, J. 2018 · SentencePiece: A simple and language independent subword tokenizer · EMNLP 2018
- Singh, A. K. & Strouse, D. J. 2024 · Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs · arXiv:2402.14903
- Petrov, A. et al. 2023 · Language Model Tokenizers Introduce Unfairness Between Languages · NeurIPS 2023
ML Academy · an interactive machine learning course that runs in your browser ·
All lessons