Documentation
Install, import, call. The API is intentionally small so you can start profiling in seconds.
Installation
Core package has no required dependencies. Optional extras add tiktoken and rich.
bash
# Base install pip install prompt-flamegraph # Accurate OpenAI-style token counts pip install prompt-flamegraph[tiktoken] # Prettier terminal output pip install prompt-flamegraph[rich]
Python API
All public functions are exported from prompt_flamegraph.
| Function | Return | What it does |
|---|---|---|
profile_prompt(data, ...) | str | Build a tree and write a standalone HTML flamegraph. |
diff_prompts(v1, v2, ...) | str | Build a diff tree and write a standalone HTML diff. |
build_tree(data, name="prompt", tokenizer=None) | Node | Build a token tree without rendering. |
detect_waste(tree) | WasteReport | Find duplicates, oversized categories and too many tools. |
count_tokens(text, tokenizer=None) | int | Count tokens in a single string. |
get_tokenizer(tokenizer=None) | Tokenizer | Resolve a tokenizer name or callable. |
normalize(payload) | dict | Convert a raw OpenAI/Anthropic payload into a prompt dict. |
from_messages(messages) | dict | Build a prompt dict from a chat messages list. |
from_langchain(obj) | dict | Convert LangChain-style objects by duck-typing. |
from_litellm_messages(msgs) | dict | Convert a LiteLLM message list. |
profile_any(obj, ...) | str | Auto-detect the input shape and profile it. |
python
profile_prompt(
data,
output="prompt_flamegraph.html", # None to skip writing
title=None, # window title
tokenizer=None, # "tiktoken" | "words" | callable
model=None, # e.g. "gpt-4o": derives tokenizer + pricing
cost_per_token=None, # e.g. 1.5e-6
detect_waste=True,
aggregate=True, # fold thin nodes into "· N more ·" buckets
width=1200,
height=720,
) -> str
diff_prompts(
v1,
v2,
output="prompt_diff.html",
title=None,
tokenizer=None,
cost_per_token=None,
width=1200,
height=720,
) -> str
build_tree(
data,
name="prompt",
tokenizer=None,
model=None,
) -> Node
detect_waste(tree: Node) -> WasteReport
Node
Internal tree node. Useful when you want to inspect the tree yourself.
name: stringtokens: intchildren: list of Nodestext: string or Noneis_leaf: property
WasteReport
Returned by detect_waste().
total_tokens: intwasted_tokens: intwaste_ratio: float propertyfindings: list of Finding objects
Finding
One waste observation.
kind: stringpath: stringmessage: stringtokens_wasted: int
Tokenizer values
Strings or callables accepted anywhere a tokenizer is expected.
None: auto-loads tiktoken if installed, else words"tiktoken"/"cl100k_*""words": built-in heuristic- Any
Callable[[str], int]
CLI options
Run prompt-flamegraph --help for the full list.
| Option | Description | Default |
|---|---|---|
input | JSON file, raw JSON string or OpenAI/Anthropic payload; - reads stdin. | Required unless --demo |
-o, --output | Output file (inferred from format). | prompt_flamegraph.<ext> |
-t, --title | Title in the generated report. | Prompt Flamegraph / Prompt Diff |
--format | html, svg, md or json. | html |
--tokenizer | tiktoken, words or a callable name. | auto |
--cost | Cost per token, e.g. 1.5e-6. | None |
--diff FILE | Diff input against another file. | None |
--terminal | Print a bar chart in the terminal. | False |
--no-waste | Disable waste detection for HTML. | False |
--demo | Use the built-in sample prompt. | False |
--width, --height | Graph dimensions in pixels. | 1200x720 |
--model MODEL | Derive tokenizer, pricing and context window from a model preset (e.g. gpt-4o). | None |
--list-models | List bundled and cached model presets. | False |
--update-models | Refresh LiteLLM community pricing into the local cache. | False |
--offline | Bundled models only; also PROMPT_FLAMEGRAPH_OFFLINE. | False |
--budget TOKENS | Exit with code 3 when total tokens exceed the budget. | None |
--no-aggregate | Draw every node instead of folding thin ones into buckets. | False |
Next steps
Try the examples or read the feature deep-dive.