Catch AI regressions before they hit production.

Test prompt and model changes against your current version. See which cases went from passing to failing, and compare the outputs before you ship.

View sample report

No credit card required

PromptLens

Example: support triage release check

A prompt edit removed the escalation rule. Now an outage gets routed as a routine technical issue.

94%72%−22 points

Pass rate · baseline v16 vs candidate v17 · 100 test cases · same model, same dataset

Regression detected

22 cases passed the baseline but fail the candidate.

Prompt changes · v16 → v17

[system]
You are a support triage assistant.
Categorize each ticket as billing,
technical, urgent, or general.
Escalate outages and revenue loss
as "urgent", even when technical.
Reply with the single most
specific category.

model gpt-5.4 · temperature 0.2 · unchanged

Cases that flipped · pass → fail

"The checkout page is down and we are losing orders every minute."

urgenttechnical

JudgeMissed the urgency signal on a production outage with revenue loss.

"I was charged twice and my card closes tomorrow."

urgentbilling

JudgeDouble charge with deadline pressure should escalate as urgent.

"Can you send me the invoice for last month?"

billingThe message is about billing.

JudgeReturned prose instead of the required single lowercase category.

Recommendation: Block the release. Restore the escalation rule in v17, then rerun the suite.

Test across the models your team uses.

OpenAI
Anthropic
Google
Meta
DeepSeek
Mistral

Run your first regression check.

Start with a prompt, a set of test cases, and your team's provider API key.

01

Add your prompt and test cases

Bring your current prompt, test inputs, and expected outputs. Include the edge cases you need to keep working.

02

Connect your models

Add your team's provider API key and choose the models to test. Configure a model to score the outputs against your expected answers.

03

Run both versions

Run your current version as the baseline, then test your prompt or model change on the same cases. Open the comparison to see what improved and what failed.

See the outputs behind every score.

Check the overall pass rate, then look at individual failures. Use the outputs and prompt changes to decide what needs fixing before release.

Inside the report

Understand each score

An LLM judge compares each output to your expected output and returns pass or fail with a one-sentence explanation.

Find cases that stopped passing

See which cases passed with your current version but fail after the change. Review the judge's explanation for each failure.

Compare outputs side by side

Review both outputs alongside the prompt diff, model, and settings to investigate what changed.

Share the regression report

Send your team one link with the comparison and failed examples. Reviewers can open it without an account.

Start free. Upgrade when your suites grow.

Every plan includes evaluations with LLM judge scoring, baseline comparisons, and unlimited shared regression reports. Live runs use encrypted organization provider keys so model choice and spend stay under your control.

Free

$0/month

Baseline comparisons, failure evidence, and shareable release decisions.

Pro

Recommended
$99/month

Bigger test suites, more projects, and higher daily volume for weekly AI releases.

Teams

Custom

Higher volume, organization provider keys, and a security review for your team.

Projects

Free3
Pro10
TeamsCustom

Prompts

Free10
ProUnlimited
TeamsUnlimited

Evaluation runs

Free50 per day
Pro100 per day
TeamsCustom

Test cases per dataset

FreeUp to 50
ProUp to 200, plus comprehensive datasets
TeamsCustom

Playground calls

Free50 per hour
Pro500 per hour
TeamsCustom

LLM judge scoring

FreeIncluded
ProIncluded
TeamsIncluded

Shared report links

FreeUnlimited
ProUnlimited
TeamsUnlimited

Model access

FreeOrganization provider keys
ProOrganization provider keys
TeamsProvider policy controls

Support

FreeEmail
ProPriority email
TeamsDedicated

Billing note

Live evaluations use your organization's provider keys. PromptLens records provider-reported or estimated usage so every comparison shows quality, latency, and model cost.

No card required to start

Frequently asked questions

Common questions about PromptLens and how it compares to the tools you're already using.

Catch regressions before production.

Run your test cases on the next prompt or model change. Review what failed before your users find it.