Capability · Improve
See what’s working. Coach what isn’t. Ship the gain.
Replay real talks, test changes, and measure the result. See on a chart whether the new agent works better.
Replays. Evals. Experiments.
These three tools help each new agent work better on talks that match your business.
Replay each talk from phone, chat, kiosk, or email. Sort by result, language, or handoff. See what led to a sale.
Use saved test cases to check each new version. A version only goes live when it passes the same key cases as the current one.
Test two prompts, tones, voices, or handoff rules. Fairshift sends more talks to the stronger choice and ends the test when there is enough proof.
Real cases. Sealed rubrics. Hard gates.
Each saved case has a question, a good result, clear scoring rules, and a tag. Fairshift tests every new version. A version cannot go live if it fails a key saved case.
1id: booking_weekend2tags: [booking, hours, hospitality]3language: en4input: |5 do you accept walk-ins on saturday afternoon?6expected:7 intent: booking_enquiry8 must_mention:9 - "saturday"10 - "weekend hours"11 must_not:12 - "i don't know"13 - external_url_pattern14 tone: friendly_professional15 knowledge_used:16 - policy:hours.weekend17rubric:18 intent_match: 0.419 content_match: 0.420 tone_match: 0.1521 no_hallucination: 0.0522gate:23 baseline: 0.8624 required: regression_blockFind a better version with fewer talks.
Fairshift sends more talks to the choice that looks better. It still tests the other choice. The test ends when one choice has enough proof, often after 600 to 2,000 talks.
When it goes wrong, the answer is one click away.
Each reply has a step-by-step record. See every model call, tool, source, and wait time. Open the talk and click a reply to see what happened.
1└── turn_14 · 740ms · ok2 ├── intent_classify · 110ms · booking_enquiry · weekend3 │ └── route:core_local · 92ms4 ├── knowledge_retrieve · 240ms · 2 chunks5 │ ├── policy.hours.weekend · 0.946 │ └── policy.walkins · 0.917 ├── plan · 180ms · answer + offer_booking8 │ ├── route:frontier_a · 142ms · 1.2k tokens9 │ └── tool_choice · skip10 ├── compose · 160ms · 71 tokens out11 │ └── route:core_local · 134ms12 └── emit_audit · 50ms · okWhat Improve does. And what it does not.
Improve helps your team watch and test live agents. It is not a general AI test tool or a report for your customers.
Improve handles
- Replays of every conversation across every channel
- Eval suites with regression gates per agent version
- Bayesian A/B experiments on prompt, tone, voice, escalation
- Per-call traces. Every model call, tool, retrieval, latency
- Insights. Pattern detection across recent conversations
- Multi-channel attribution reporting
Improve does not
- Generic LLM benchmarks. Wrong tool
- Customer-facing dashboards your end users see
- Predictive lead scoring (use Convert)
- Marketing analytics (use your existing tools)
- Performance monitoring of your app (use APM)
- Replace your data warehouse
Better agents close more. See it on the pipeline.
The agent gets better. You watch the results.
Thirty minutes. We walk through real replays, your eval suite, and a live experiment. You see the loop running before the call ends.