Foundry

"Token" care of business: understanding the fuel that drives the LLM engine

"Token" care of business: understanding the fuel that drives the LLM engine

"Token" care of business: understanding the fuel that drives the LLM engine

Udith Vaidyanathan

CEO & Co-Founder, LogicFlo AI

If you were a six-year-old kid, you'd be learning a language by understanding each of the letters and then how they group together to form words and attaching meaning to those words (A for apple, B for ball, C for cat,…)

Over time, you'd understand sentences and "rules" (grammar) that influence sentence structure, and eventually, when you're an adult, you read a sentence and break that down into individual words. You know the meaning of those words that you have memorized, and that gives you a good sense of what is being said.

So your "training" would involve starting with the smallest possible elements, which are letters, building up to words, sentences, paragraphs, and eventually your neurons feel comfortable understanding information that is spread across a dense set of pages across large tomes, until you become proficient enough that you begin to understand articulations of complex abstract concepts, emotions, and poetry (To be or not to be, that is the question!).

But if you were a Large Language Model and your overlords didn't have the patience to handhold you, character by character, through the entire corpus of human information ever generated, then this is probably a really inefficient way to make you learn. It would take eons of time, not to mention guzzle electricity and water enough to consume the entire planet.

Alternatively, you could try, for example, teaching the LLM every single word ever made, like a dictionary, and just say, "Memorize the dictionary, so now you know the meaning of every single word that has ever existed in the past."  The problem with that would be that a language doesn't have a fixed number of words. You'd have to teach the LLM the hundreds of thousands that have existed in the past and then every surname, typo, product name and neologism coined since breakfast. Any dictionary you freeze in advance meets a word it has never seen and your model blames you for being a terrible teacher and asking questions out-of-syllabus.

You'd also never have taught the model the concepts of context and analogies that makes language beautiful and complex at the same time (Shit didn't hit the fan literally, no matter how vehemently the teenagers in your life might testify to the contrary)

So the human overlords found a slick way to find the best of both worlds. They taught the model the characters and then they told it how many times the character appears across all the information generated by humans (training data).

I, A, …

Then for the most common characters, they measured what the most common following characters are

IT, IN, AT, AM, TO,…

And do this over and over again so that not only would the model know that "I" or "A" is common, but it would know that all the prepositions(TO, IN, BY) , articles (A, AN, THE) and conjuctions (AND, OR, NOT) are extremely common.

This is called Byte Pair Encoding - BPE where you build frequency distributions of characters or bytes and  then do that for larger and larger groupings of characters and feed it to the model to build a statistical map of the entire language that has been sampled over the entire corpus of information ever generated - not just a dictionary. You have in essence "modeled" the entire language including common parlance and popular usage - without explicitly teaching the model about words, sentences and grammar.

This brings us to the most important supposedly magical thing of all - "inference" or model output.

This is when your shiny model is ready to start performing miracles for you,  and you feed it a prompt, it just breaks down the prompt into these tokens - some maybe one character, some may be a full word and some, an arbitrary slicing of the word that has chunks of these repeating units. This is also why there is no one single answer for "how many tokens is this word/sentence" it could be anything based on the statistical distribution of the training data.

For example - unbelievable likely gets broken down to "un", "bel", "ie","v","able". This isn't exactly how you would do it as a proficient adult, but then again, you don't have the ability to maintain the frequencies of all possible combinations of characters across every book you have ever read. You have, at some point assigned some rules (every time there is an "un" it negates what comes after that). The model does something close to that, but it will not have a rule book. It will just make a probabilistic guess that "un" negates the next thing that comes purely because of the number of times it has seen it happen (it will likely have to have seen the word unique, uniform etc. to make sure that it doesn't assign a probability of 100%, so you see why training data distribution and hygiene is so incredibly important).


When the model responds, what it is actually doing is predicting the next token based on the statistics it has seen before. Nobody ever taught it the difference between "true" and "sounds like the kind of thing that would be true," so when it is uncertain it does not stop and tell you - it generates the most probable continuation.

You can now see why hallucinations happen - the way in which a model reads a sentence is vastly different from the way you just read this one. It just keeps guessing without caring about the meaning. It just so happens that it has seen, read and memorized more information than you can ever fathom and so its guesses, generally turn out to be right (and often mind blowing)

It helps to notice that hallucinations do not all have the same cause. Sometimes the model simply never learned the fact; sometimes it learned a wrong or half-formed pattern; sometimes it misreads what you gave it; and sometimes it has been effectively trained to guess by evaluations that reward a confident answer over an honest "I don't know" (here's an interesting paper about LLM hallucinations).

Also this problem of teaching "meaning" to models is also being solved in many elegant ways that make them truly genius-like but I will deep dive into these challenges in subsequent articles.

25 Water st, New York, NY 10004

1111B S Governors Ave STE 39697
Dover, DE 19904

700 Soldier's Field Rd,
Boston, Massachussetts, MA 02163

21/13 Sri Krupa, 3rd Seaward Road, Valmiki Nagar,
Thiruvanmiyur, Chennai 600041

LogicFlo Inc• Copyright © 2026

25 Water st, New York, NY 10004

1111B S Governors Ave STE 39697
Dover, DE 19904

700 Soldier's Field Rd,
Boston, Massachussetts, MA 02163

21/13 Sri Krupa, 3rd Seaward Road, Valmiki Nagar,
Thiruvanmiyur, Chennai 600041

LogicFlo Inc• Copyright © 2026

25 Water st, New York, NY 10004

1111B S Governors Ave STE 39697
Dover, DE 19904

700 Soldier's Field Rd,
Boston, Massachussetts, MA 02163

21/13 Sri Krupa, 3rd Seaward Road, Valmiki Nagar,
Thiruvanmiyur, Chennai 600041

LogicFlo Inc• Copyright © 2026

25 Water st, New York, NY 10004

1111B S Governors Ave STE 39697

Dover, DE 19904

700 Soldier's Field Rd, Boston,

Massachussetts, MA 02163

21/13 Sri Krupa, 3rd Seaward Road,

Valmiki Nagar, Thiruvanmiyur, Chennai 600041

LogicFlo Inc• Copyright © 2026

25 Water st, New York, NY 10004

1111B S Governors Ave STE 39697
Dover, DE 19904

700 Soldier's Field Rd,
Boston, Massachussetts, MA 02163

21/13 Sri Krupa, 3rd Seaward Road, Valmiki Nagar,
Thiruvanmiyur, Chennai 600041

LogicFlo Inc• Copyright © 2026