The Token Archive
Pricing

Simple plans. You keep the provider savings.

We charge a fixed plan. Token $ you save stay yours.

Free

Hobbyists and evaluation

$0forever
  • 10,000 requests / 30d
  • All core optimization
  • 1 API key · 1 user
  • 7-day logs
  • Discord / community support
  • BYOK — no pass-through billing
Start Free
Most popular

Pro

Indie founders in production

$99per month

Typical savings $2,000–$5,000 / month

  • 500,000 requests / 30d
  • 5 API keys
  • Cache TTL up to 24h
  • 30-day logs
  • Failover routing
  • Hard spend caps
  • Request tracing
  • Manual model override
  • Email support (48h)
Start 14-day trial

Team

Startups and agencies

$499per month

Typical savings $20,000–$50,000 / month

  • 5,000,000 requests / 30d
  • Unlimited keys · 10 users
  • 7-day cache TTL
  • 90-day logs
  • Per-project breakdown · per-user caps
  • LangChain / LlamaIndex
  • A/B + shadow traffic · eval suite
  • PII redaction · custom routing policies
  • Live chat (4h)
Start 14-day trial

Enterprise

Compliance and control

Custom$2k–$50k / month
  • Unlimited / negotiated volume
  • SSO · RBAC · audit logs
  • SOC 2 · HIPAA / GDPR
  • On-prem / VPC · data residency
  • White-label
  • 99.9% SLA · dedicated CSM
Contact sales

Volume is requests per 30 days — not tokens. The plan fee is a flat SaaS fee for the proxy and dashboard; provider spend stays on your own BYOK key, and output tokens usually cost 3–5× more per million than input.

Live today: OpenAI + Anthropic proxy, streaming, BYOK, exact-match cache, heuristic routing, code_ast compression, failover, budget headers, request/key/user limits, dashboard spend and savings. Embedding-based semantic cache, email/Slack alerts, evals, PII redaction, SSO and compliance controls are included in their tier as they ship — they are not shipping yet and not billed separately.

Hosted platform keys, when used, run on gpt-4o-mini under a separate cost cap. That cap is an infrastructure limit, not the marketing request gate.

Compare plans
CapabilityFreeProTeamEnterprise
Volume & limits
Requests / 30 days10,000500,0005,000,000Negotiated
API keys15UnlimitedUnlimited
Users1110Unlimited
Cache TTL1hup to 24hup to 7dCustom
Log retention7 days30 days90 daysCustom
Core proxy
OpenAI + AnthropicYesYesYesYes
StreamingYesYesYesYes
Exact-match cacheYes (shared global)YesYesYes
Semantic (embedding) cacheas it shipsYesYesYes
Heuristic routingAuto onlyManual overrideCustom policiesCustom policies
code_ast compressionYesYesYesYes
Pro controls
Manual model overrideYesYesYes
Failover routingYesYesYes
Hard spend capsYesYesYes
Request tracingYesYesYes
Email / Slack alertsas it shipsYesYesYes
Team collaboration
Per-project breakdownYesYes
Per-user capsYesYes
Per-endpoint cache TTLas it shipsYesYes
LangChain / LlamaIndexas it shipsYesYes
A/B + shadow trafficas it shipsYesYes
Eval suiteas it shipsYesYes
PII redactionas it shipsYesYes
Enterprise & compliance
SSOas it shipsYes
RBACas it shipsYes
Audit logsas it shipsYes
SOC 2as it shipsYes
On-prem / VPCas it shipsYes
Data residencyas it shipsYes
White-labelas it shipsYes
Service
Uptime SLANoneNone99.5%99.9%
SupportCommunityEmail 48hLive chat 4hDedicated CSM

Rows marked “as it ships” are included in that tier when the capability ships. Free hits 10,000 requests or needs a second key → upgrade to Pro. Need team visibility → Team. Need SSO or compliance → contact sales.

FAQ

What do you bill for?

Plan fee only. The $99 Pro / $499 Team fee is a flat SaaS fee for the compression proxy and dashboard. Upstream usage is billed by your provider on your BYOK key, and output tokens usually cost 3–5× more per million than input tokens. Set hard limits in your provider dashboard so users cannot overcharge the account.

How is volume measured?

Requests per 30 days — not tokens. Free is 10,000, Pro is 500,000, Team is 5,000,000, Enterprise is negotiated.

Which capabilities are live today?

OpenAI + Anthropic proxying, streaming, BYOK, exact-match cache, heuristic routing, code_ast compression, failover, budget headers, request/key/user limits, and the spend/savings dashboard. Embedding-based semantic cache, email/Slack alerts, evals, PII redaction, SSO and compliance controls are included in their tier as they ship — not shipping today.

When should I upgrade?

Free hits 10k requests or needs a second key → Pro. Need team visibility, per-project breakdown or per-user caps → Team. Need SSO, RBAC or compliance → contact sales.

Will quality drop?

Modes control aggression. Start lossless; raise it only if you want more savings. Typical result is 30–60% fewer input tokens — not a guaranteed cut to your total provider bill.

What clients work?

Anything with a custom OpenAI base URL. Not Copilot or Cursor Tab.

Can I try first?

Free tier. No card.