Tokens & RAG
Token Counter and Visualizer
Count cl100k_base tokens and inspect byte-safe token boundaries, IDs and exact bytes.
Processed on our CPU server.Submitted input is processed for this run and is not stored.
256 KiB limitInput
Result
A useful result starts here.
Paste your input or load an example,
then run the tool.
Supported formats & limitations
- Pinned cl100k_base ordinary-text encoding using tokenizer v0.8.1; vocabulary is compiled locally. Special-token literals count as ordinary text, not reserved token IDs. No universal model/chat-wrapper/billing claim.
- 32 KiB UTF-8 text maximum. Consecutive whitespace/non-whitespace runs capped at 2048 bytes to bound BPE work. Empty text counts as zero.
- Visualization is limited to 4,096 tokens; the first 1,000 complete visual spans are drawn. The full bounded token IDs/hex/offsets remain in the report.
- Byte offsets are zero-based, half-open. Tokens can split UTF-8 characters; visual spans group tokens until complete code points and retain individual IDs and hex bytes.
- Server CPU processing with no input history or outbound requests. MCP access and public deployment remain separate launch tasks.
Execution budget: 5s; output limit: 4096 KiB. Browser worker startup has a separate 3s allowance.