Seven of fourteen models named Human Interest first on the direct prompt; one named Fidelity Advantage 401. Human Interest was named by thirteen of the fourteen models and Fidelity Advantage 401 by thirteen and Human Interest carries 42 labels and Fidelity Advantage 401 36, so the shares are not directly comparable.
Named in one category this edition.
Named in one category this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all fourteen models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the 401(k) providers page.
Across every category in the October 2026 Edition, Human Interest and Fidelity Advantage 401 were named in the same answer fifty-six times, of the 105 answers naming Human Interest and the 107 naming Fidelity Advantage 401. In those answers Fidelity Advantage 401 took the first choice nine times and Human Interest twelve.
| Model | Direct | Paraphrase | Comparative | Budget-constrained | Scale-constrained | Negative |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | ||||||
| GPT-5.4 mini | ||||||
| Gemini 3.5 Flash | ||||||
| Perplexity Sonar | ||||||
| Grok 4.1 Fast | ||||||
| Mistral Small | ||||||
| DeepSeek V4 Flash | ||||||
| Llama 4 Maverick | ||||||
| Qwen 3.7 Flash | ||||||
| Kimi K2 | ||||||
| GLM 4.7 FlashX | ||||||
| MiniMax M2.5 | ||||||
| GPT-6 Luna | ||||||
| Muse Glimmer 30B |
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Six of seven in this category shown.
“slam poor support, payroll delays/errors, and communication issues” Grok 4.1 Fast · negative prompt · hard negative
“Human Interest, Slavic401k, Guideline, and Ubiquity/Online 401(k) in various user reviews or forum posts reporting high costs, poor service” Perplexity Sonar · negative prompt · soft negative
“notable player, but it has historically been more focused on smaller businesses than true mid-market complexity” GPT-5.4 mini · direct prompt · soft negative
“Best Tech-Forward & Automated: Human Interest ... If you use Rippling, Gusto, or BambooHR and want the lowest administrative burden: Go with Human Interest.” Gemini 3.5 Flash · direct prompt · first choice
“Human Interest is recommended as the best fit for mid-market employers that want managed 401k administration with structured onboarding and conversion support.” Claude Haiku 4.5 · direct prompt · first choice
“Human Interest: Best Overall Value & Ease of Use ... go with Human Interest or Guideline” Gemini 3.5 Flash · budget prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of four in this category shown.
“Fidelity – Highly rated for its broad investment lineup, transparent fees, strong digital tools, and education resources. It's a top choice” Mistral Small · direct prompt · first choice
“1. Fidelity Investments - Best for: Large and mid-sized companies; top-rated mobile app, zero-fee index funds (e.g., FZROX).” Grok 4.1 Fast · comparative prompt · first choice
“If it is Fidelity, Vanguard, Schwab, or Empower, you are in safe territory regarding platform security and customer service.” Qwen 3.7 Flash · negative prompt · first choice
Comparisons are drawn for the top eight products in each category, each against each. The output is the models' output; nothing here is a recommendation by the index.