Tool-output optimization
Reduce repetitive tool output before it enters a model's context, with reversible transforms and measurable limits.
On this page
Token X-ray and per-tool recommendations
Open Dashboard → Token X-ray to rank observed tool patterns by estimated opportunity. Leash records bounded pattern, count and optimization metadata, not raw tool input/output in the control plane. Use Spend insights for model/provider/cache breakdowns and Burn Reports for imported session logs.
Ranked suggestions and request-size changes are estimates. Confirm task outcomes and provider-reported usage before treating them as billing savings. Compression stays lossless on supported tool data and leaves user instructions unchanged.
Inspect your own tool output first
Use the local optimizer to compare the original text with a smaller representation. It can compact JSON whitespace, encode eligible repeated-row JSON as a self-described table, and encode repeated log lines as counted runs. The output is only changed when the transformed representation is smaller.
No model call, signup or provider key is required for the browser workshop or CLI analysis. Inspect the transformed output before choosing to enable it on real traffic.
leash optimize ./tool-output.json
cat ./tool-output.log | leash optimize -
# Return the diagnostics and optimized output as JSON:
leash optimize ./tool-output.json --json
# Write a new private file; existing files are never overwritten:
leash optimize ./tool-output.json --output ./compact-output.txtAutomatic compression, with an off switch
Compression is on by default at the local proxy and gateway. Set LEASH_TOOL_COMPRESSION=off and restart to disable it. Alternatively set toolCompression = "off" in local config.toml; the environment setting takes precedence.
Supported surfaces are string-valued OpenAI Chat tool messages, Responses function_call_output outputs, and Anthropic user tool_result content. Structured multimodal outputs, error tool results and signed-thinking requests are skipped. Requests with explicit cache_control, prompt_cache_key or prompt_cache_retention are skipped.
LEASH_TOOL_COMPRESSION=lossless leash start
# Disable it again and restart:
LEASH_TOOL_COMPRESSION=off leash startWhat lossless means here
JSON compaction preserves values, numeric lexemes and structure while removing formatting whitespace. Eligible table encodings can reconstruct the canonical JSON values. Log-run encoding preserves repeated line text and newline structure. These are structural transforms, not AI summaries or truncation.
The proxy evaluates bounded work: at most 1 MiB of tool text across up to 32 outputs, with a JSON depth limit. Unsupported or oversized input stays unchanged. Budget reservations and loop fingerprints use the original request, preserving conservative admission.
Separate measured text reduction from cost claims
The report shows original and optimized character counts, reduction percentage, chosen strategy and an estimated token count using characters divided by four. That token estimate is a heuristic, not the provider's tokenizer or billed usage.
The proxy exposes x-leash-tool-compression, x-leash-tool-chars-saved and x-leash-tool-estimated-tokens-saved response headers. Actual provider usage remains the source for cost metering. Leash does not claim a universal 30× saving: a large compression ratio on a repetitive fixture cannot establish the same reduction on a customer's total bill.