ML Academy · Track 3 · Large Language Models

Why tokenisers behave differently

Two models split the same text into different numbers of pieces. Where that difference comes from is written in the corpus the tokeniser was trained on.

3 steps 175 XP A free account is needed
Start the lesson →

Sources

ML Academy · an interactive machine learning course that runs in your browser · All lessons