Comment on: Show HN: How LLMs Work – Interactive visual guide based on Karpathy's lecture
by thesz
The page does very poor job tokenizing phrase "Noinceolik fiyulnabmed fyvaproldge" into "Noinceolik fiyulnabm ed fyvaproldge", factoring only "ed" suffix. As if made up words such as "noinceolik" are so common they are part of 100K token vocabulary.The actual application of GPT-5 tokenizer at [1] to my made up phrase results in 14 tokens, only two of them are four characters long and there are tokens containing spaces.[1] https://gpt-tokenizer.dev/I will read along, though.
View Discussion ↗
Discussion Thread
Parent Entity
Points: 207 • Comments: 49
Posted: Apr 24, 2026
Other Comments / Reviews
SaaS Metrics