What benchmark scores cannot tell you
Public benchmarks evaluate models on standardized problems using correct answers or evaluator scores. They do not tell you how a model performs in each feature of your service.
A model that excels at summarizing reviews may be average at answering support questions. Test each feature separately with real traffic.
How LLM A/B testing works
First, split users of the same AI feature across variants. Assignment is based on the device. Send a device ID with each call to keep a user on the same variant until you change the split.
Next, define a success metric, such as a purchase or copying a summary. ABTO shows each variant's outcomes alongside cost per call, latency, and error rate. When outcomes are similar, you can choose the cheaper variant for each feature.
Start new models with a small share of traffic
A new variant is saved at 0% and receives no users until you increase its share. We recommend starting at 10%, checking for problems, then increasing the share once the difference is clear.
If something goes wrong, set the new variant to 0% and the existing one to 100% to roll back in seconds. A feature can have up to five variants.
Compare prompts as well as models
User behavior also helps you determine whether a revised prompt is better. Before exposing it to real users, you can compare candidate models and prompts using the same inputs.