Pricing
Start Free. Scale When You're Ready.
$99 buys the default monitor — weekly, 150 of your own prompts, 2 candidates, 5 judges rotated from a fixed pool, plus catalog and price triggers, 10 runs a month. Model inference is included.
Free
forever free
One sponsored comparison on your own prompts.
Get Started Free- 1 sponsored comparison, judged and scored
- Standard & Advanced models
- Directional result — 1 judge on a handful of examples
- Reports guaranteed available for 7 days
- 1 seat
- Response caching
- Standing monitors
- Custom eval criteria
- API & MCP access
- Premium & Frontier models
Pro
per workspace
A standing monitor on the model you run in production.
Upgrade to Pro- Up to 2 monitors · 10 runs/mo each
- Weekly, plus catalog, price and drift triggers
- 150 of your own prompts, 2 candidates per run
- 5 judges rotated from a frozen pool of 8
- Model inference included — no metering, no markup
- Premium models as candidates
- Unlimited seats
- Open suite API and production traffic ingest
- Custom eval criteria
- Reports guaranteed available for 90 days
- Prompt CI
Team
per workspace
Every prompt change validated before it ships.
Upgrade to Team- Everything in Pro
- Prompt CI — deploy webhook and GitHub Action
- Deploy triggers bypass the daily debounce
- Prompt-as-variable runs
- Up to 6 monitors · 30 runs/mo each
- Reports guaranteed available for 180 days
- Choose your own judges
- Bring your own provider keys
- Frontier models as candidates
Enterprise
stated allowance per contract
For organizations that need frontier models, their own judges, and dedicated support.
Start with a Free Managed TrialTalk to sales →- Everything in Team
- Frontier models as candidates
- Choose your own judges
- Bring your own provider keys — judges and candidates both
- Run allowance stated in your contract
- Reports guaranteed available for 365 days
- Audit logs
- Dedicated support & SLA
- Free managed trial included
All paid plans billed monthly. No contracts — cancel anytime.
What a run includes
You buy monitors, runs, and triggers. PeerLM pays the model providers. There is no second, metered vocabulary to reconcile against the sticker price.
150 of your prompts
Each cycle replays a sample of your own production traffic — not a synthetic benchmark. PeerLM rejects a run it cannot execute at the size you asked for rather than quietly shrinking it.
5 judges, frozen pool of 8
PeerLM selects the judges. A model from the vendor under test can't judge it. Five judges rotate from a fixed pool so consecutive cycles compare like with like.
Inference included
Generations and judging are bundled. We show provider list price as a range so you can see what the cycle costs us — not what you owe on top of the plan.
Runs are the ceiling
Each monitor gets a fixed number of runs per calendar month. A run either happens or it doesn't. Unused runs do not roll over.
Add a premium model to the judge pool for $49 per monitor per month. It sits in the rotation — it does not become the sole judge.
Monitors — standing comparison on production traffic
Connect production traffic to a Monitor and continuously compare candidates for quality and cost. Standing monitors start on Pro. Free includes one sponsored comparison to try it out.
Sources
SDK, file upload, Langfuse, Helicone, Cloudflare AI Gateway, and OpenTelemetry. Sync traffic into each Monitor.
Triggers
Weekly heartbeat, plus catalog, price, and drift. Team adds deploy triggers for Prompt CI — those bypass the daily debounce.
Living verdict
Switch, route, or hold — with quality retained, projected savings, Switch Confidence (SCS), and per-task routing.
Routing export
Export recommended model routing as JSON, YAML, LiteLLM proxy config, Portkey gateway config, or a code snippet.