Zero of twelve models named Convert first on the direct prompt; zero named Statsig. Convert was named by ten of the twelve models and Statsig by ten and Convert carries 17 labels and Statsig 21, so the shares are not directly comparable.
Named in one category this edition.
Named in three categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the experimentation and personalization page.
Across every category in the September 2026 Edition, Convert and Statsig were named in the same answer twenty-three times, of the 53 answers naming Convert and the 79 naming Statsig. In those answers Statsig took the first choice two times and Convert two.
| Model | Direct | Paraphrase | Comparative | Budget-constrained | Scale-constrained | Negative |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | ||||||
| GPT-5.4 mini | ||||||
| Gemini 3.5 Flash | ||||||
| Perplexity Sonar | ||||||
| Grok 4.1 Fast | ||||||
| Mistral Small | ||||||
| DeepSeek V4 Flash | ||||||
| Llama 4 Maverick | ||||||
| Qwen 3.7 Flash | ||||||
| Kimi K2 | ||||||
| GLM 4.7 FlashX | ||||||
| MiniMax M2.5 |
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of four in this category shown.
“I would recommend Convert.com if you want a strong balance of reliability, privacy, and straightforward pricing” Perplexity Sonar · paraphrase prompt · first choice
“I would recommend Convert.com as the best A/B testing and personalization tool for a mid-sized B2B company” Llama 4 Maverick · paraphrase prompt · first choice
“Privacy-first (excellent GDPR), affordable (~$299+/month), up to 50 goals/test.” Grok 4.1 Fast · negative prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Five of five in this category shown.
“Proceed with awareness of roadmap ... some product teams have expressed concern over its long-term roadmap” Gemini 3.5 Flash · negative prompt · soft negative
“Major ownership instability ... per-event pricing penalizes experimentation culture” Kimi K2 · negative prompt · soft negative
“But require devs (no visual editors), best if you have engineering resources.” Grok 4.1 Fast · budget prompt · soft negative
“I'd recommend starting with Statsig if you have technical resources” GLM 4.7 FlashX · budget prompt · first choice
“I'd suggest evaluating Statsig or VWO first” Kimi K2 · scale prompt · first choice
Comparisons are drawn for the top eight products in each category, each against each. The output is the models' output; nothing here is a recommendation by the index.