History on repeat
Every turn re-sends the entire growing transcript. You pay for the same tokens again, and again, and again.
Tidbit sits between your client and the model, and trims the waste out of what you send. Point your client at it once. No SDK, no code changes, and the replies come back untouched.*
Works with✓ Claude Code✓ Claude Desktop✓ the Anthropic API
The same context gets re-billed turn after turn, and it adds up quietly, in more ways than one. The longer and more complex the task, the more there is to save.
Every turn re-sends the entire growing transcript. You pay for the same tokens again, and again, and again.
Repeated context can bill at a deep discount, but only when it's flagged correctly and reused in time. Most setups miss it, so every turn pays full price.
One giant file read or log dump rides along inside every single request that follows it.
The usual fixes don't fit: SDK rewrites mean code changes in every app, quota walls stall the team mid-task, and nobody can see where money actually went.
Two different things get smaller: what you send, and what you pay when context repeats. We measure them separately, and every mark below names the tracked run it came from.
One large tool-heavy request, 197,521 tokens. The recorded answer check passed. Evidence: large-trace-lossless.
The same stable context sent four times, with the cache-filling send inside this measurement. Evidence: cache-reuse-4.
A pre-warmed 24-turn window, with the cache-filling send outside this measurement. Evidence: cache-session-24.
Every number above is regenerated from a committed result file. These runs used our most conservative setting. See all four tracked runs →
The product view reports observed workload outcomes.
Illustrative trend; measured results vary by workload.
See it work on your very next request.
The same optimization core, wherever you already work: the CLI, the desktop app, or any tool built on the API.
Run one script and the CLI routes through Tidbit. Same commands, same workflow, calls optimized from then on.
A one-click Mac installer wires the desktop app through Tidbit: no terminal, no config. It is signed, verifies each of its own updates before the swap, and refuses downgrades. The coding work you do in the app is optimized on the way out.
Already building on the API? Point any tool at the Tidbit gateway URL with your key. SDKs, scripts, and agents all keep working.
On a Claude subscription, Tidbit is free. On the API, Tidbit is free for the duration of early access. Your Console shows the savings we measure and attribute to Tidbit's own work.
On any Claude plan, the full gateway and Console, no charge.
No card to start.
On a Claude subscription, Tidbit is free, including the full gateway and Console. On the Anthropic API, Tidbit is free for the duration of early access. Your Console shows the savings we measure and attribute to our own work. Savings produced by your own client caching will never be billed.
No. Tidbit sits in front of your existing setup as a gateway. Point your client at Tidbit and keep your current workflow, with no SDK swap and no library migration.
We ran a small number of live workloads twice over: once straight to Anthropic, once through Tidbit on our lossless profile, the setting where nothing is dropped from your request. On those runs the answer check passed both ways, and the input and output token counts matched across the repeated-context turns we measured. That is a handful of tracked runs on particular workloads, so read it as evidence about those runs and not as a general quality or latency guarantee. The method and the raw numbers are on the benchmarks page.
Every number we publish comes from a recorded benchmark run rather than a projection. Each one carries an evidence ID, a fingerprint of the source data, and a derivation checked into the repository, so it can be reproduced from the recorded run. The savings in your Console are a separate thing, measured from your own account activity.
Tidbit does not use request content to train a model and does not sell data. On the API, you keep your own key and Tidbit uses it only to forward your request. On a Claude subscription, your client keeps signing in the way it already does and we never hold a key for it. Retention windows and processor terms are in the privacy policy.
No. Your traffic always rides your own credential, never a shared or pooled account, and our own key pays only for our benchmark runs. There is no rate-limit circumvention and no terms-of-service workaround. The efficiency comes from sending fewer tokens, not from evading limits.
Claude Code, Claude Desktop, and services that call the Anthropic API. Tidbit forwards your request to Anthropic rather than choosing a model for you, so it works with the Claude models your client already uses.
You can point your client back at Anthropic at any time and carry on, because removal is immediate in every mode. If Tidbit cannot resolve a credential it refuses the request outright rather than quietly sending it somewhere you did not choose. We do not have a live status page yet, so until that exists the trust page and direct email are where we report incidents.
Removal is one step. Point your client back at the direct endpoint and the path is exactly as it was before, with nothing left behind in your application code.
We add cache hints and trim re-sent history. What you actually wrote reaches Claude untouched.
Your traffic goes to Anthropic and nowhere else. No third-party models, no detours, and nothing you send is collected or used for training.
We only optimize the input you send; the model's response is never altered.
Your key is used to forward the request to Anthropic, and nothing more.
Keep every workflow, and watch the savings add up live. Free to start, no card required.
No card. No SDK. Switch it off whenever you want.