Test prompt and model changes against your current version. See which cases went from passing to failing, and compare the outputs before you ship.
No credit card required
Example: support triage release check
A prompt edit removed the escalation rule. Now an outage gets routed as a routine technical issue.
Pass rate · baseline v16 vs candidate v17 · 100 test cases · same model, same dataset
Regression detected
22 cases passed the baseline but fail the candidate.
Prompt changes · v16 → v17
model gpt-5.4 · temperature 0.2 · unchanged
Cases that flipped · pass → fail
"The checkout page is down and we are losing orders every minute."
urgent ✓technical ✗
JudgeMissed the urgency signal on a production outage with revenue loss.
"I was charged twice and my card closes tomorrow."
urgent ✓billing ✗
JudgeDouble charge with deadline pressure should escalate as urgent.
"Can you send me the invoice for last month?"
billing ✓The message is about billing. ✗
JudgeReturned prose instead of the required single lowercase category.
Recommendation: Block the release. Restore the escalation rule in v17, then rerun the suite.
Test across the models your team uses.
Start with a prompt, a set of test cases, and your team's provider API key.
Bring your current prompt, test inputs, and expected outputs. Include the edge cases you need to keep working.
Add your team's provider API key and choose the models to test. Configure a model to score the outputs against your expected answers.
Run your current version as the baseline, then test your prompt or model change on the same cases. Open the comparison to see what improved and what failed.
Check the overall pass rate, then look at individual failures. Use the outputs and prompt changes to decide what needs fixing before release.
Inside the report
An LLM judge compares each output to your expected output and returns pass or fail with a one-sentence explanation.
See which cases passed with your current version but fail after the change. Review the judge's explanation for each failure.
Review both outputs alongside the prompt diff, model, and settings to investigate what changed.
Send your team one link with the comparison and failed examples. Reviewers can open it without an account.
Every plan includes evaluations with LLM judge scoring, baseline comparisons, and unlimited shared regression reports. Live runs use encrypted organization provider keys so model choice and spend stay under your control.
Choose by volume
Plans scale by projects, daily evaluation runs, dataset size, and provider controls.
Baseline comparisons, failure evidence, and shareable release decisions.
Bigger test suites, more projects, and higher daily volume for weekly AI releases.
Higher volume, organization provider keys, and a security review for your team.
Projects
Prompts
Evaluation runs
Test cases per dataset
Playground calls
LLM judge scoring
Shared report links
Model access
Support
Billing note
Live evaluations use your organization's provider keys. PromptLens records provider-reported or estimated usage so every comparison shows quality, latency, and model cost.
Common questions about PromptLens and how it compares to the tools you're already using.
Run your test cases on the next prompt or model change. Review what failed before your users find it.