Zero telemetry
No network calls, no analytics, no server. Your prompts stay local.
A small tool that answers a specific question: where are my prompt tokens going, and am I wasting them?
Each block in the flamegraph represents a part of your prompt. Width is proportional to token count, so the biggest budget eaters stand out immediately.
Accepts nested dict, list and str structures. Common top-level keys like system_prompt, tools, rag_context and chat_history become the top row.
The HTML report is a single file with no external dependencies. Open it in any browser, share it, or archive it with your experiment.
Pass cost_per_token and the report shows the estimated price of each category, not just the raw token count.
The report renders the tree as stacked, colored bars. Move the cursor over a block to see its name, tokens and share of the total.
Nodes too thin to draw fold into striped · N more · buckets so the graph stays readable. Pass --no-aggregate (or aggregate=False) to draw every node — see the unaggregated demo.
detect_waste() scans the prompt tree and returns concrete findings you can act on before calling the API.
Each report contains total_tokens, wasted_tokens and a list of Findings. A finding has a kind, path, human message and the number of tokens_wasted.
| Kind | What it means | Action |
|---|---|---|
duplicate | The same text appears in several leaves. | Dedupe your RAG chunks or history. |
too_many_tools | More than 5 tool definitions. | Only declare tools the model is likely to call. |
huge_rag | RAG context is over 50% of the prompt. | Trim, rerank or chunk your documents. |
long_history | Chat history is over 30% of the prompt. | Summarize or truncate old turns. |
large_system_prompt | System prompt is over 35% of the prompt. | Make it shorter or split instructions. |
Compare two versions of the same prompt and see where tokens were added, removed or changed. Useful for A/B testing system prompts or RAG chunking strategies.
The diff report uses green for added, red for removed and orange for changed nodes. Unchanged nodes stay neutral.
Diffs render to HTML, SVG and Markdown just like regular flamegraphs, so you can embed them in pull requests or documentation.
Choose the output that fits your workflow.
| Format | Best for | Command |
|---|---|---|
| HTML | Interactive exploration in the browser | --format html or default |
| SVG | Embedding in documentation or presentations | --format svg |
| Markdown | Paste into GitHub issues, PRs or wiki | --format md |
| JSON | Machine-readable output for CI and downstream tooling | --format json |
| Terminal | Quick look from the shell | --terminal |
Use the default word-punctuation heuristic, install tiktoken for OpenAI-style counts, or pass your own callable.
Fast, dependency-free estimator that counts word-like tokens and punctuation. Perfect for quick checks and offline usage.
Install with pip install prompt-flamegraph[tiktoken] then pass tokenizer="tiktoken" for accurate counts matching OpenAI models.
Any Callable[[str], int] works. Plug in your own tokenizer, a Hugging Face tokenizer, or a model-specific counter.
Use --tokenizer tiktoken, --tokenizer words or any registered name from the command line.
Details that make the tool pleasant to use.
No network calls, no analytics, no server. Your prompts stay local.
Modern Python, no compatibility hacks, clean type annotations where it matters.
Run --demo to see a sample report in seconds, or pipe JSON directly.
Pass --model gpt-4o to derive the tokenizer encoding, per-token pricing and context window. --list-models shows every preset.
--update-models refreshes LiteLLM community pricing into a local cache; --offline (or PROMPT_FLAMEGRAPH_OFFLINE) keeps it bundled-only. Prices are community estimates.
normalize() and from_messages() take raw OpenAI/Anthropic payloads; from_langchain(), from_litellm_messages() and profile_any() adapt framework objects by duck-typing — langchain is never imported.
--budget TOKENS exits with code 3 when the prompt is too big. Example workflow: .github/workflows/prompt-budget.yml.example.
Open source, weak copyleft — free to import into proprietary code.
Try the examples, read the API docs or install from PyPI.