Tokens & RAG
Chunking Playground
Split text by Unicode characters, recursive separators or token boundaries, with source offsets and overlaps.
Processed on our CPU server.Submitted input is processed for this run and is not stored.
256 KiB limitInput
Result
A useful result starts here.
Paste your input or load an example,
then run the tool.
Supported formats & limitations
- Pinned cl100k_base ordinary-text encoding using tokenizer v0.8.1; vocabulary is compiled locally. Special-token literals count as ordinary text, not reserved token IDs. No universal model/chat-wrapper/billing claim.
- 32 KiB UTF-8 text maximum. Consecutive whitespace/non-whitespace runs capped at 2048 bytes to bound BPE work. Empty text counts as zero.
- Byte offsets are zero-based, half-open. Tokens can split UTF-8 characters; visual spans group tokens until complete code points and retain individual IDs and hex bytes.
- Size 1–4096 units; overlap must be nonnegative and smaller. At most 200 chunks; increase size if the cap is reached.
- Character/recursive units are Unicode code points, not grapheme clusters. Recursive separators prefer paragraph, newline, space, then code-point fallback, retaining separators.
- Token mode uses whole-source token offsets. Cuts group split code points and may exceed requested size; actual overlap can differ and is reported. Chunks are not independently retokenized. No embeddings or semantic splitting.
- Server CPU processing with no input history or outbound requests. MCP access and public deployment remain separate launch tasks.
Execution budget: 5s; output limit: 4096 KiB. Browser worker startup has a separate 3s allowance.