Two of twelve models named Pave first on the direct prompt; zero named Comprehensive.io. Pave was named by eleven of the twelve models and Comprehensive.io by eleven and Pave carries 25 labels and Comprehensive.io 12, so the shares are not directly comparable.
Named in three categories this edition.
Named in three categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the compensation benchmarking data page.
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Two of two in this category shown.
“Use modern, integration-based salary benchmarking software that pulls live, anonymized data directly from HR and payroll systems (e.g., Pave, Ravio, or Figures)” Gemini 3.5 Flash · negative prompt · first choice
“Start with Pave's free Market Data Lite (if you're under 200 employees) — it's the most accessible, current, and comprehensive free option.” DeepSeek V4 Flash · budget prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of four in this category shown.
“three free salary benchmarking tools cover most needs at zero cost: Pave (tech roles, under 200 employees), Comprehensive.io (US tech daily data), and BLS OEWS” Claude Haiku 4.5 · budget prompt · first choice
“Free daily-refreshed tech salary data from 6,000+ U.S. companies (no participation needed) ... Great for startups/lean teams.” Grok 4.1 Fast · budget prompt · first choice
“Pave and Comprehensive.io are currently leading the market for affordable or free options” Qwen 3.7 Flash · budget prompt · first choice
Comparisons are drawn for the top three products in each category. The output is the models' output; nothing here is a recommendation by the index.