Zero of fourteen models named Mail-Tester first on the direct prompt; one named Validity Everest. Mail-Tester was named by eleven of the fourteen models and Validity Everest by twelve and Mail-Tester carries 20 labels and Validity Everest 27, so the shares are not directly comparable.
Named in one category this edition.
Named in one category this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all fourteen models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the email deliverability and testing tools page.
| Model | Direct | Paraphrase | Comparative | Budget-constrained | Scale-constrained | Negative |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | ||||||
| GPT-5.4 mini | ||||||
| Gemini 3.5 Flash | ||||||
| Perplexity Sonar | ||||||
| Grok 4.1 Fast | ||||||
| Mistral Small | ||||||
| DeepSeek V4 Flash | ||||||
| Llama 4 Maverick | ||||||
| Qwen 3.7 Flash | ||||||
| Kimi K2 | ||||||
| GLM 4.7 FlashX | ||||||
| MiniMax M2.5 | ||||||
| GPT-6 Luna | ||||||
| Muse Glimmer 30B |
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Six of eight in this category shown.
“Mail-Tester scores against a single seed address, so it can't tell you where your email lands across a multi-ISP seed network” Claude Haiku 4.5 · negative prompt · hard negative
“These are inaccurate and misleading... do not base business decisions on them.” DeepSeek V4 Flash · negative prompt · hard negative
“Avoid: Relying solely on these web-based scores. They are useful for a quick syntax check but should not be your primary deliverability testing tool.” GLM 4.7 FlashX · negative prompt · soft negative
“Mail-Tester for occasional pre-send checks of a message and its sending setup. It allows up to 3 free tests per day.” GPT-6 Luna · budget prompt · first choice
“The best email deliverability tool for a company with a limited budget is Mail-Tester or Unspam for pre-send testing” Llama 4 Maverick · budget prompt · first choice
“Start with free tools → Google Postmaster Tools + MXToolbox + Mail-Tester free tier” Kimi K2 · budget prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Six of eight in this category shown.
“its pricing often starts at $20,000 to $65,000/year, making it massive overkill and too expensive for most mid-sized operations” Gemini 3.5 Flash · paraphrase prompt · soft negative
“enterprise suites like *Validity Everest* (which can cost upward of $15,000/year) to be too complex and cost-prohibitive” Gemini 3.5 Flash · direct prompt · soft negative
“The gold standard for enterprise, but likely overkill for mid-sized companies on both features and budget.” DeepSeek V4 Flash · paraphrase prompt · soft negative
“The gold standard for larger mid-market and enterprise brands. Extremely deep seed-testing network and comprehensive reputation monitoring, though highly priced.” Gemini 3.5 Flash · scale prompt · first choice
“My default pick: Validity Everest” GPT-6 Luna · direct prompt · first choice
“Everest is the largest and most robust platform in the space... generally cost-prohibitive for SMBs.” Gemini 3.5 Flash · comparative prompt · alternative
Comparisons are drawn for the top eight products in each category, each against each. The output is the models' output; nothing here is a recommendation by the index.