Documentation

Install, import, call. The API is intentionally small so you can start profiling in seconds.

Installation

Core package has no required dependencies. Optional extras add tiktoken and rich.

bash
# Base install
pip install prompt-flamegraph

# Accurate OpenAI-style token counts
pip install prompt-flamegraph[tiktoken]

# Prettier terminal output
pip install prompt-flamegraph[rich]

Python API

All public functions are exported from prompt_flamegraph.

FunctionReturnWhat it does
profile_prompt(data, ...)strBuild a tree and write a standalone HTML flamegraph.
diff_prompts(v1, v2, ...)strBuild a diff tree and write a standalone HTML diff.
build_tree(data, name="prompt", tokenizer=None)NodeBuild a token tree without rendering.
detect_waste(tree)WasteReportFind duplicates, oversized categories and too many tools.
count_tokens(text, tokenizer=None)intCount tokens in a single string.
get_tokenizer(tokenizer=None)TokenizerResolve a tokenizer name or callable.
normalize(payload)dictConvert a raw OpenAI/Anthropic payload into a prompt dict.
from_messages(messages)dictBuild a prompt dict from a chat messages list.
from_langchain(obj)dictConvert LangChain-style objects by duck-typing.
from_litellm_messages(msgs)dictConvert a LiteLLM message list.
profile_any(obj, ...)strAuto-detect the input shape and profile it.
python
profile_prompt(
    data,
    output="prompt_flamegraph.html",  # None to skip writing
    title=None,                        # window title
    tokenizer=None,                    # "tiktoken" | "words" | callable
    model=None,                        # e.g. "gpt-4o": derives tokenizer + pricing
    cost_per_token=None,               # e.g. 1.5e-6
    detect_waste=True,
    aggregate=True,                    # fold thin nodes into "· N more ·" buckets
    width=1200,
    height=720,
) -> str

diff_prompts(
    v1,
    v2,
    output="prompt_diff.html",
    title=None,
    tokenizer=None,
    cost_per_token=None,
    width=1200,
    height=720,
) -> str

build_tree(
    data,
    name="prompt",
    tokenizer=None,
    model=None,
) -> Node

detect_waste(tree: Node) -> WasteReport

Node

Internal tree node. Useful when you want to inspect the tree yourself.

  • name: string
  • tokens: int
  • children: list of Nodes
  • text: string or None
  • is_leaf: property

WasteReport

Returned by detect_waste().

  • total_tokens: int
  • wasted_tokens: int
  • waste_ratio: float property
  • findings: list of Finding objects

Finding

One waste observation.

  • kind: string
  • path: string
  • message: string
  • tokens_wasted: int

Tokenizer values

Strings or callables accepted anywhere a tokenizer is expected.

  • None: auto-loads tiktoken if installed, else words
  • "tiktoken" / "cl100k_*"
  • "words": built-in heuristic
  • Any Callable[[str], int]

CLI options

Run prompt-flamegraph --help for the full list.

OptionDescriptionDefault
inputJSON file, raw JSON string or OpenAI/Anthropic payload; - reads stdin.Required unless --demo
-o, --outputOutput file (inferred from format).prompt_flamegraph.<ext>
-t, --titleTitle in the generated report.Prompt Flamegraph / Prompt Diff
--formathtml, svg, md or json.html
--tokenizertiktoken, words or a callable name.auto
--costCost per token, e.g. 1.5e-6.None
--diff FILEDiff input against another file.None
--terminalPrint a bar chart in the terminal.False
--no-wasteDisable waste detection for HTML.False
--demoUse the built-in sample prompt.False
--width, --heightGraph dimensions in pixels.1200x720
--model MODELDerive tokenizer, pricing and context window from a model preset (e.g. gpt-4o).None
--list-modelsList bundled and cached model presets.False
--update-modelsRefresh LiteLLM community pricing into the local cache.False
--offlineBundled models only; also PROMPT_FLAMEGRAPH_OFFLINE.False
--budget TOKENSExit with code 3 when total tokens exceed the budget.None
--no-aggregateDraw every node instead of folding thin ones into buckets.False