One of twelve models named Optimizely CMS first on the direct prompt; zero named Statsig. Optimizely CMS was named by twelve of the twelve models and Statsig by ten and Optimizely CMS carries 47 labels and Statsig 21, so the shares are not directly comparable.
By Optimizely, San Francisco, United States, founded 2010. Named in eight categories this edition.
Named in three categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the experimentation and personalization page.
Across every category in the September 2026 Edition, Optimizely CMS and Statsig were named in the same answer forty-eight times, of the 195 answers naming Optimizely CMS and the 79 naming Statsig. In those answers Statsig took the first choice three times and Optimizely CMS twelve.
| Model | Direct | Paraphrase | Comparative | Budget-constrained | Scale-constrained | Negative |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | ||||||
| GPT-5.4 mini | ||||||
| Gemini 3.5 Flash | ||||||
| Perplexity Sonar | ||||||
| Grok 4.1 Fast | ||||||
| Mistral Small | ||||||
| DeepSeek V4 Flash | ||||||
| Llama 4 Maverick | ||||||
| Qwen 3.7 Flash | ||||||
| Kimi K2 | ||||||
| GLM 4.7 FlashX | ||||||
| MiniMax M2.5 |
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Six of eight in this category shown.
“Major Structural / Commercial Red Flags ... Optimizely ... Extreme pricing opacity ... Security incidents” Kimi K2 · negative prompt · hard negative
“High Cost, Declining Support, Rigid Platform... Be very cautious | Optimizely (for SMB/mid-market)” DeepSeek V4 Flash · negative prompt · hard negative
“Platforms to Avoid or Be Very Cautious About ... 1. Optimizely (Now part of Episerver)” Mistral Small · negative prompt · hard negative
“Deepest experimentation infrastructure - considered the enterprise standard for A/B testing... The most mature experimentation platform” Mistral Small · comparative prompt · first choice
“The category reference point for statistical rigor... Choose Optimizely if: You need the most rigorous experimentation program” Kimi K2 · comparative prompt · first choice
“Optimizely — best-in-class statistical rigor + feature experimentation; needs some engineering support; enterprise pricing” DeepSeek V4 Flash · scale prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Five of five in this category shown.
“Proceed with awareness of roadmap ... some product teams have expressed concern over its long-term roadmap” Gemini 3.5 Flash · negative prompt · soft negative
“Major ownership instability ... per-event pricing penalizes experimentation culture” Kimi K2 · negative prompt · soft negative
“But require devs (no visual editors), best if you have engineering resources.” Grok 4.1 Fast · budget prompt · soft negative
“I'd recommend starting with Statsig if you have technical resources” GLM 4.7 FlashX · budget prompt · first choice
“I'd suggest evaluating Statsig or VWO first” Kimi K2 · scale prompt · first choice
Comparisons are drawn for the top eight products in each category, each against each. The output is the models' output; nothing here is a recommendation by the index.