LLM A/B testing: choose AI models by how your users respond

Your users' responses—not a benchmark score—tell you which LLM fits your service. LLM A/B testing measures those responses by model. ABTO lets you start with a few lines of code.

Start freeRead the integration guide

What benchmark scores cannot tell you

Public benchmarks evaluate models on standardized problems using correct answers or evaluator scores. They do not tell you how a model performs in each feature of your service.

A model that excels at summarizing reviews may be average at answering support questions. Test each feature separately with real traffic.

How LLM A/B testing works

First, split users of the same AI feature across variants. Assignment is based on the device. Send a device ID with each call to keep a user on the same variant until you change the split.

Next, define a success metric, such as a purchase or copying a summary. ABTO shows each variant's outcomes alongside cost per call, latency, and error rate. When outcomes are similar, you can choose the cheaper variant for each feature.

Start new models with a small share of traffic

A new variant is saved at 0% and receives no users until you increase its share. We recommend starting at 10%, checking for problems, then increasing the share once the difference is clear.

If something goes wrong, set the new variant to 0% and the existing one to 100% to roll back in seconds. A feature can have up to five variants.

Compare prompts as well as models

User behavior also helps you determine whether a revised prompt is better. Before exposing it to real users, you can compare candidate models and prompts using the same inputs.

Related guides

Learn more