Everything it does

A small tool that answers a specific question: where are my prompt tokens going, and am I wasting them?

Interactive flamegraph

Each block in the flamegraph represents a part of your prompt. Width is proportional to token count, so the biggest budget eaters stand out immediately.

Built for structured prompts

Accepts nested dict, list and str structures. Common top-level keys like system_prompt, tools, rag_context and chat_history become the top row.

Standalone output

The HTML report is a single file with no external dependencies. Open it in any browser, share it, or archive it with your experiment.

Cost overlay

Pass cost_per_token and the report shows the estimated price of each category, not just the raw token count.

Hover and explore

The report renders the tree as stacked, colored bars. Move the cursor over a block to see its name, tokens and share of the total.

Aggregation buckets

Nodes too thin to draw fold into striped · N more · buckets so the graph stays readable. Pass --no-aggregate (or aggregate=False) to draw every node — see the unaggregated demo.

Token waste detection

detect_waste() scans the prompt tree and returns concrete findings you can act on before calling the API.

Detections included

  • Duplicate text across leaves
  • Too many tools declared at once
  • Oversized RAG context
  • Long chat history
  • Large system prompt

WasteReport output

Each report contains total_tokens, wasted_tokens and a list of Findings. A finding has a kind, path, human message and the number of tokens_wasted.

KindWhat it meansAction
duplicateThe same text appears in several leaves.Dedupe your RAG chunks or history.
too_many_toolsMore than 5 tool definitions.Only declare tools the model is likely to call.
huge_ragRAG context is over 50% of the prompt.Trim, rerank or chunk your documents.
long_historyChat history is over 30% of the prompt.Summarize or truncate old turns.
large_system_promptSystem prompt is over 35% of the prompt.Make it shorter or split instructions.

Prompt diff

Compare two versions of the same prompt and see where tokens were added, removed or changed. Useful for A/B testing system prompts or RAG chunking strategies.

Color-coded changes

The diff report uses green for added, red for removed and orange for changed nodes. Unchanged nodes stay neutral.

Same export formats

Diffs render to HTML, SVG and Markdown just like regular flamegraphs, so you can embed them in pull requests or documentation.

Export formats

Choose the output that fits your workflow.

FormatBest forCommand
HTMLInteractive exploration in the browser--format html or default
SVGEmbedding in documentation or presentations--format svg
MarkdownPaste into GitHub issues, PRs or wiki--format md
JSONMachine-readable output for CI and downstream tooling--format json
TerminalQuick look from the shell--terminal

Pluggable tokenizers

Use the default word-punctuation heuristic, install tiktoken for OpenAI-style counts, or pass your own callable.

Default: words

Fast, dependency-free estimator that counts word-like tokens and punctuation. Perfect for quick checks and offline usage.

tiktoken / cl100k

Install with pip install prompt-flamegraph[tiktoken] then pass tokenizer="tiktoken" for accurate counts matching OpenAI models.

Custom callable

Any Callable[[str], int] works. Plug in your own tokenizer, a Hugging Face tokenizer, or a model-specific counter.

CLI override

Use --tokenizer tiktoken, --tokenizer words or any registered name from the command line.

Other nice things

Details that make the tool pleasant to use.

Zero telemetry

No network calls, no analytics, no server. Your prompts stay local.

Python 3.10+

Modern Python, no compatibility hacks, clean type annotations where it matters.

CLI with helpful defaults

Run --demo to see a sample report in seconds, or pipe JSON directly.

Model-aware profiling

Pass --model gpt-4o to derive the tokenizer encoding, per-token pricing and context window. --list-models shows every preset.

Dynamic pricing

--update-models refreshes LiteLLM community pricing into a local cache; --offline (or PROMPT_FLAMEGRAPH_OFFLINE) keeps it bundled-only. Prices are community estimates.

Adapters for anything

normalize() and from_messages() take raw OpenAI/Anthropic payloads; from_langchain(), from_litellm_messages() and profile_any() adapt framework objects by duck-typing — langchain is never imported.

CI budget gate

--budget TOKENS exits with code 3 when the prompt is too big. Example workflow: .github/workflows/prompt-budget.yml.example.

LGPL-3.0-or-later

Open source, weak copyleft — free to import into proprietary code.