Tokens per word
One of the most-asked questions about LLMs. Enter any word count to convert it to tokens instantly.
Short answer
About 1.3 tokens per English word — roughly 0.75 words per token.
Rule of thumb for ordinary English prose. Code, non-English text, and heavy punctuation tokenize less efficiently.
Estimated tokens
≈ 1,300
1,000 words × 1.3 tokens/word
| Words | Tokens (approx.) |
|---|---|
| 1 word | 1.3 |
| 10 words | 13 |
| 50 words | 65 |
| 100 words | 130 |
| 500 words (1 page) | 650 |
| 1,000 words | 1,300 |
| 10,000 words | 13,000 |
| 100,000 words (a novel) | 130,000 |
Based on 1.3 tokens per English word. Actual counts vary with vocabulary, formatting, and the model's tokenizer — use the token calculator for exact numbers.
About 1.3 tokens on average for English. Common words like “the” are often a single token; rare or long words split into several.
Tokenizers split text into subword pieces based on frequency, not dictionary words. Frequent words get one token; uncommon ones are broken into fragments, which pushes the average above 1.
Yes. English is the most token-efficient at ~1.3 per word. Many other languages cost noticeably more per word because tokenizers were trained mostly on English text.
Code usually costs more than prose — roughly 1.5 to 2+ tokens per word equivalent. Braces, indentation, camelCase names, and long ID strings all split into extra pieces.
It is the same idea from the other side. An average English word plus its following space is about 5 characters, and 5 ÷ 4 ≈ 1.25 — consistent with 1.3 tokens per word.
Paste your text into the free token calculator on the homepage. It runs a real tokenizer in your browser and shows the exact count.