HR AI Index
Index Compensation and total rewards September 2026 Edition

Compensation benchmarking data

Asked as “salary benchmarking data provider”, and as “compensation survey data”, on behalf of a mid-market B2B company. 50 first choices recorded across the direct, paraphrase, budget and scale prompts, twelve models each.
Standing
Contested
24% of first choices, contested.

By buyer segment

The same question asked on behalf of a different buyer. Each standing is computed within its segment; they sit side by side and are never added together.

01The standing

Share is the count of first choices across the direct, paraphrase, budget and scale prompts, over all twelve models, for a mid-market B2B company. Ordered by share.
ProductFirst-choice shareNegative rateLabelsQuadrant
01Payscale24%21%28accepted challenger
02Pave20%0%25accepted challenger
03Comprehensive.io12%0%12accepted challenger
04Mercer6%43%28criticized challenger
05Radford4%37%27criticized challenger
06Salary.com2%6%16accepted challenger
Show the two products at 0%, ordered by negative rate
08WTW0%30%10criticized challenger
07Ravio0%0%22accepted challenger
Bars are the share of first choices, 0 to 100Every product with at least 10 labels here. Every product name links to its vendor page.

Recommended versus criticized

Every product with at least 10 labels here, on both axes. The 30% line names a quadrant, not the verdict above: that one needs more than 40%.

Criticized challengerCriticized default
Negative label rate →
01
02
03
04
05
06
07
08
Accepted challengerEndorsed leader
0%First-choice share → · lines at 30% share and 25% negative40%
Key
01Payscale24%
02Pave20%
03Comprehensive.io12%
04Mercer6%
05Radford4%
06Salary.com2%
07Ravio0%
08WTW0%

02What they warned about

Zero of twelve models held their first choice under the paraphrase. Claude Haiku 4.5, GPT-5.4 mini, Gemini 3.5 Flash, Perplexity Sonar, Grok 4.1 Fast, Mistral Small, DeepSeek V4 Flash, Llama 4 Maverick, Qwen 3.7 Flash, Kimi K2, GLM 4.7 FlashX and MiniMax M2.5 changed. A high negative share on a product with few labels is a warning. A low share on a product with many labels is salience, not sentiment.
Mercer
43%
12 of 28 labels negative · 8 of 12 models · 2 hard negative
“**Avoid** over-investing in Mercer/WTW unless you have global complexity” Kimi K2, paraphrase prompt
Radford
37%
10 of 27 labels negative · 7 of 12 models · 2 hard negative
“What to Avoid on a Limited Budget ... enterprise-grade pricing with far more data than most small/mid-market companies need” Kimi K2, budget prompt
Glassdoor
75%
6 of 8 labels negative · 5 of 12 models · 2 hard negative
“Use these for market signals only. Do not use them for making concrete compensation decisions” Qwen 3.7 Flash, scale prompt
UserBenchmark
100%
4 of 4 labels negative · 4 of 12 models · 4 hard negative
“**UserBenchmark** - Has received numerous complaints about: Severe bias toward certain hardware manufacturers” MiniMax M2.5, negative prompt

03What they cite

Citations exist only for the models that return a source list: twelve of the twelve in this edition, and all six flagship models on the expanded tier.

Sites the answers cite

62 of 72 answers in this category came back with a source list, from 12 of 12 models: citations where the model returns them, or the search results it consulted. 6 of those lists are Google grounding redirects that name no site and are left out of the counts. 771 links across 252 sites, every framing counted. Ranked by the number of answers carrying the site or page.

vendor site · Ravio37 answers · 97 citations · 10 models
vendor site · Comprehensive30 answers · 35 citations · 10 models
21 answers · 52 citations · 10 models
vendor site · SalaryCube19 answers · 32 citations · 10 models
vendor site · HRTechSaaS15 answers · 15 citations · 9 models
14 answers · 19 citations · 9 models
13 answers · 22 citations · 8 models
vendor site · Salary.com12 answers · 13 citations · 7 models
vendor site · CaptivateIQ12 answers · 12 citations · 7 models
vendor site · Ficstar12 answers · 12 citations · 8 models
10 answers · 24 citations · 9 models
vendor site · Pin9 answers · 9 citations · 9 models

Pages the answers cite

The ten pages named in the most answers, by full address. A page here is one the models returned with a recommendation, not one the index endorses.

04How they answered

Six framings of the same buying question, each sent to every model in a fresh session with search on. One row per model, so a row shows whether it held its answer under rewording, what it named when cost was the constraint, and what it argued against. Computed from the raw judge labels.
ModelDirect“What is the best salary benchmarking data provider for a mid-market B2B company?”Paraphrase“Which compensation survey data would you recommend to a mid-sized B2B company?”Comparative“What are the top benchmarking providers and how do they differ?”Budget-constrained“What is the best salary benchmarking data provider for a company with a limited budget?”Scale-constrained“We are a 500 person company evaluating a salary benchmarking data provider. What should we look at?”Negative“Which benchmarking providers should I avoid or be cautious about?”
Claude Haiku 4.5Payscale
Three alternativesLattice, SalaryCube, beqom
Mercer Compensation Database, RadfordChanged
Two alternativesCulpepper, Willis Towers Watson
against: Payscale, Salary.com
no first choiceComprehensive.io, Pave
Two alternativesBLS OEWS, BambooHR
no first choicenothing named
GPT-5.4 miniPayscale
Two alternativesRadford, Salary.com
against: Mercer, WTW
Mercer, RadfordChanged
Two alternativesAon, WTW
Mercer, WTW
One alternativeSHRM
Payscale
One alternativeSalary.com
against: Korn Ferry, Mercer, Radford, Wage data from government sources
no first choicenothing named
Gemini 3.5 FlashPave
Six alternativesCarta Total Comp, Figures, Payscale, Radford, Ravio, Salary.com
Pave, PayscaleChanged
Five alternativesPavilion, Ravio, RepVue, Salary.com, The Bridge Group
against: Radford
Aon Radford, Mercer
Three alternativesERI Economic Research Institute, Pave, Willis Towers Watson [WTW]
Pave
Five alternativesComprehensive.io, ERI (Economic Research Institute) Salary Assessor, Pequity, Salary.com, U.S. Bureau of Labor Statistics (BLS) - OEWS
against: Mercer, Payscale, Radford, Willis Towers Watson
Pave
Seven alternativesFigures, Mercer, Option Impact, Ravio, Syndio, Trusaic, WTW
against: Aon Radford, Glassdoor, Payscale
against: Glassdoor, Hays, Payscale, Robert Half
Perplexity SonarPayscale
Two alternativesRavio, SalaryCube
against: Mercer
CompData, Mercer, Salary.comChanged
Three alternativesKorn Ferry Pay, RepVue/Bridge Group-style data, SHRM CompAnalyst Market Data
no first choiceComprehensive.io
One alternativePave
no first choicenothing named
Grok 4.1 FastPayscale
Two alternativesPave, Salary.com
against: Mercer/ Radford
CompData SurveysChanged
Four alternativesCulpepper, Pavilion GTM Compensation Benchmarks, Payscale, WorldatWork
against: Mercer, Radford
Deloitte
Six alternativesBCG, Bain & Company, Computer Economics, Gartner, KPMG, PwC
Comprehensive.io, U.S. Bureau of Labor Statistics (BLS) Occupational Employment and Wage Statistics
Three alternativesPave, Ravio, SalaryExpert
against: Glassdoor, Levels.fyi, Mercer, Payscale, Radford
no first choicenothing named
Mistral SmallSalaryCube
Three alternativesPave, Payfactors, Salary.com
Pavilion's B2B Tech Salary BenchmarksChanged
Four alternativesCompensation-IQ, Mercer, Ravio, Willis Towers Watson
Deloitte
Six alternativesBain & Company, Boston Consulting Group, Computer Economics, Gartner, KPMG, PwC
Bureau of Labor Statistics (BLS) OEWS
Three alternativesComprehensive.io, Pave, SalaryExpert
no first choiceagainst: Anthropic, OpenAI, UserBenchmark
DeepSeek V4 FlashPayscale
Four alternativesERI, Pave, Ravio, Salary.com
against: Mercer, Radford
Payscale / CompAnalystChanged
Two alternativesRadford, SalaryCube
against: Mercer Benchmark Database
no first choicePave
Four alternativesBLS Occupational Employment Statistics, Comprehensive.io, Payscale, Ravio
against: Glassdoor, Levels.fyi, Radford/Mercer
no first choiceagainst: 3DMark, Cinebench 2024, Geekbench 6, MMLU, Open LLM Leaderboard, PCMark, PassMark PerformanceTest, UserBenchmark
Llama 4 MaverickPayscaleno first choiceChangedno first choiceBLS OEWS, Comprehensive.io, Pave
Two alternativesFicstar, Ravio
no first choicenothing named
Qwen 3.7 FlashPayscale
Four alternativesPave, Payfactors, Salary.com, SalaryCube
against: Mercer, Radford, WTW
Mercer, Robert HalfChanged
Four alternativesPayscale, Radford, Salary.com, Willis Towers Watson
Gartner
Three alternativesComputer Economics, Deloitte, ISG
Comprehensive.io, Pave
Five alternativesBLS, Levels.fyi, Payscale, Ravio, Salary.com
against: Mercer, Towers Watson
no first choice
Four alternativesCompAnalyst, Option Impact, Pave, Radford
against: Deloitte, Glassdoor, Korn Ferry, Levels.fyi, LinkedIn, Mercer, WTW
against: PassMark, UserBenchmark
Kimi K2Payscale, SalaryCube
Three alternativesMercer, Radford, Salary.com
Payscale / CompAnalystChanged
Two alternativesCompData Surveys, Salary.com CompAnalyst
against: Mercer, Radford, Willis Towers Watson
no first choice
Seven alternativesComputer Economics, Deloitte, Forrester, Gartner, IDC, KPMG, McKinsey & Company
Comprehensive.io, U.S. Bureau of Labor Statistics
Four alternativesCulpepper, Local/Regional HR Consulting Firms, Payscale, Ravio
against: Betts, Kelly Services, Radford, Robert Half
no first choicenothing named
GLM 4.7 FlashXPave
Eight alternativesAon Radford, CaptivateIQ, Ficstar, Mercer, OpenComp, Payscale Ascent, Ravio, Salary.com CompAnalyst
CompDataChanged
Two alternativesMercer, Radford
no first choicePave
Five alternativesBLS, Comprehensive.io, Ravio, Salary.com CompAnalyst, StartupCFO’s Startup Salary Benchmark
against: Glassdoor, Levels.fyi, Payscale
no first choicenothing named
MiniMax M2.5Payscale
One alternativeComprehensive.io
against: Mercer, Radford
CompAnalyst, PayscaleChanged
Two alternativesMercer, Radford
no first choiceBLS Occupational Employment and Wage Statistics
Four alternativesFigures, Glassdoor, Pave Market Data Lite, Payscale
no first choiceagainst: UserBenchmark
Bold is the first choiceAlternatives are counted; the count opens them.What the answer argued against

05The record

One row per call: the version string exactly as returned, whether the model searched, sources cited, and latency. Full answer text is in the free responses file. Download the record
Seventy-two rows: every prompt, every model, every answer.
PromptModelVersion stringTime (UTC)SearchedSourcesLatency
Direct recommendationClaude Haiku 4.5claude-haiku-4-5-202510012026-09-17 20:12yes98 s
Direct recommendationGPT-5.4 minigpt-5.4-mini-2026-03-172026-09-17 17:59yes56 s
Direct recommendationGemini 3.5 Flashgemini-3.5-flash2026-09-17 22:02yes1353 s
Direct recommendationPerplexity Sonarsonar2026-09-17 20:41yes207 s
Direct recommendationGrok 4.1 Fastspacexai/grok-4.1-fast-non-reasoning via vertex2026-09-17 20:41yes3419 s
Direct recommendationMistral Smallmistral/mistral-small via mistral2026-09-17 20:35yes54 s
Direct recommendationDeepSeek V4 Flashdeepseek/deepseek-v4-flash via fireworks2026-09-17 18:49yes1522 s
Direct recommendationLlama 4 Maverickmeta/llama-4-maverick via bedrock2026-09-17 21:33yes52 s
Direct recommendationQwen 3.7 Flashalibaba/qwen3.7-flash via alibaba2026-09-17 19:51yes525 s
Direct recommendationKimi K2moonshotai/kimi-k2 via novita2026-09-17 19:21yes1428 s
Direct recommendationGLM 4.7 FlashXzai/glm-4.7-flashx via zai2026-09-17 19:07yes1164 s
Direct recommendationMiniMax M2.5minimax/minimax-m2.5 via minimax2026-09-17 19:42yes1123 s
ParaphraseClaude Haiku 4.5claude-haiku-4-5-202510012026-09-17 21:13no04 s
ParaphraseGPT-5.4 minigpt-5.4-mini-2026-03-172026-09-17 20:22no04 s
ParaphraseGemini 3.5 Flashgemini-3.5-flash2026-09-17 18:50yes1021 s
ParaphrasePerplexity Sonarsonar2026-09-17 21:36yes207 s
ParaphraseGrok 4.1 Fastspacexai/grok-4.1-fast-non-reasoning via vertex2026-09-17 18:45yes2914 s
ParaphraseMistral Smallmistral/mistral-small via mistral2026-09-17 20:12yes55 s
ParaphraseDeepSeek V4 Flashdeepseek/deepseek-v4-flash via fireworks2026-09-17 21:42yes1016 s
ParaphraseLlama 4 Maverickmeta/llama-4-maverick via bedrock2026-09-17 21:27yes52 s
ParaphraseQwen 3.7 Flashalibaba/qwen3.7-flash via alibaba2026-09-17 19:12no029 s
ParaphraseKimi K2moonshotai/kimi-k2 via novita2026-09-17 21:40yes1929 s
ParaphraseGLM 4.7 FlashXzai/glm-4.7-flashx via zai2026-09-17 18:30yes23107 s
ParaphraseMiniMax M2.5minimax/minimax-m2.5 via minimax2026-09-17 21:35yes527 s
ComparativeClaude Haiku 4.5claude-haiku-4-5-202510012026-09-17 19:55yes98 s
ComparativeGPT-5.4 minigpt-5.4-mini-2026-03-172026-09-17 19:20no08 s
ComparativeGemini 3.5 Flashgemini-3.5-flash2026-09-17 18:33yes1634 s
ComparativePerplexity Sonarsonar2026-09-17 21:41yes208 s
ComparativeGrok 4.1 Fastspacexai/grok-4.1-fast-non-reasoning via vertex2026-09-17 20:25yes2917 s
ComparativeMistral Smallmistral/mistral-small via mistral2026-09-17 21:24yes57 s
ComparativeDeepSeek V4 Flashdeepseek/deepseek-v4-flash via fireworks2026-09-17 20:50yes1315 s
ComparativeLlama 4 Maverickmeta/llama-4-maverick via bedrock2026-09-17 19:59yes52 s
ComparativeQwen 3.7 Flashalibaba/qwen3.7-flash via alibaba2026-09-17 18:55yes1037 s
ComparativeKimi K2moonshotai/kimi-k2 via novita2026-09-17 21:46yes1345 s
ComparativeGLM 4.7 FlashXzai/glm-4.7-flashx via zai2026-09-17 21:04yes25164 s
ComparativeMiniMax M2.5minimax/minimax-m2.5 via minimax2026-09-17 21:08yes520 s
Budget constrainedClaude Haiku 4.5claude-haiku-4-5-202510012026-09-17 19:03yes98 s
Budget constrainedGPT-5.4 minigpt-5.4-mini-2026-03-172026-09-17 18:05no03 s
Budget constrainedGemini 3.5 Flashgemini-3.5-flash2026-09-17 17:56yes1423 s
Budget constrainedPerplexity Sonarsonar2026-09-17 20:27yes204 s
Budget constrainedGrok 4.1 Fastspacexai/grok-4.1-fast-non-reasoning via vertex2026-09-17 20:34yes3215 s
Budget constrainedMistral Smallmistral/mistral-small via mistral2026-09-17 21:18yes54 s
Budget constrainedDeepSeek V4 Flashdeepseek/deepseek-v4-flash via fireworks2026-09-17 20:47yes913 s
Budget constrainedLlama 4 Maverickmeta/llama-4-maverick via bedrock2026-09-17 21:42yes74 s
Budget constrainedQwen 3.7 Flashalibaba/qwen3.7-flash via alibaba2026-09-17 18:03yes517 s
Budget constrainedKimi K2moonshotai/kimi-k2 via novita2026-09-17 19:31yes934 s
Budget constrainedGLM 4.7 FlashXzai/glm-4.7-flashx via zai2026-09-17 20:17yes21100 s
Budget constrainedMiniMax M2.5minimax/minimax-m2.5 via minimax2026-09-17 18:10yes515 s
Scale constrainedClaude Haiku 4.5claude-haiku-4-5-202510012026-09-17 18:12no05 s
Scale constrainedGPT-5.4 minigpt-5.4-mini-2026-03-172026-09-17 21:30no09 s
Scale constrainedGemini 3.5 Flashgemini-3.5-flash2026-09-17 22:05yes1258 s
Scale constrainedPerplexity Sonarsonar2026-09-17 20:44yes207 s
Scale constrainedGrok 4.1 Fastspacexai/grok-4.1-fast-non-reasoning via vertex2026-09-17 19:04yes1310 s
Scale constrainedMistral Smallmistral/mistral-small via mistral2026-09-17 20:33no09 s
Scale constrainedDeepSeek V4 Flashdeepseek/deepseek-v4-flash via fireworks2026-09-17 19:07yes1419 s
Scale constrainedLlama 4 Maverickmeta/llama-4-maverick via bedrock2026-09-17 20:35yes53 s
Scale constrainedQwen 3.7 Flashalibaba/qwen3.7-flash via alibaba2026-09-17 21:33no029 s
Scale constrainedKimi K2moonshotai/kimi-k2 via novita2026-09-17 20:46yes1430 s
Scale constrainedGLM 4.7 FlashXzai/glm-4.7-flashx via zai2026-09-17 21:18yes1450 s
Scale constrainedMiniMax M2.5minimax/minimax-m2.5 via minimax2026-09-17 18:11yes1045 s
Negative framingClaude Haiku 4.5claude-haiku-4-5-202510012026-09-17 17:58no02 s
Negative framingGPT-5.4 minigpt-5.4-mini-2026-03-172026-09-17 18:10yes56 s
Negative framingGemini 3.5 Flashgemini-3.5-flash2026-09-17 18:25yes2223 s
Negative framingPerplexity Sonarsonar2026-09-17 21:20yes156 s
Negative framingGrok 4.1 Fastspacexai/grok-4.1-fast-non-reasoning via vertex2026-09-17 18:24yes3918 s
Negative framingMistral Smallmistral/mistral-small via mistral2026-09-17 19:02yes56 s
Negative framingDeepSeek V4 Flashdeepseek/deepseek-v4-flash via deepinfra2026-09-17 19:30yes2859 s
Negative framingLlama 4 Maverickmeta/llama-4-maverick via bedrock2026-09-17 20:41yes53 s
Negative framingQwen 3.7 Flashalibaba/qwen3.7-flash via alibaba2026-09-17 21:07yes2257 s
Negative framingKimi K2moonshotai/kimi-k2 via novita2026-09-17 21:02yes1535 s
Negative framingGLM 4.7 FlashXzai/glm-4.7-flashx via zai2026-09-17 19:40yes2675 s
Negative framingMiniMax M2.5minimax/minimax-m2.5 via minimax2026-09-17 19:14yes1538 s

Normalization in this category

Every judgment call made between the raw labels and the numbers above, listed so it is visible and reversible.

Category-scoped readings
None. Every name in this category resolved on its own.
Unresolved, counted raw
AnandTech
Anthropic
BLS Occupational Employment and Wage Statistics (OEWS)
Betts
Cinebench 2024
CompTelecom
CrystalDiskMark
ERI (Economic Research Institute) Salary Assessor
ERI Economic Research Institute
Gamers Nexus
Geekbench 6
HELM
Kelly Services
LinkedIn job adverts
Local/Regional HR Consulting Firms
MMLU
Mercer Compensation Database
NERA
Open LLM Leaderboard
Oxera
PC-Kombo
PCMark (UL Solutions)
PassMark
PassMark PerformanceTest
Pavilion (formerly Revenue Collective)
Pavilion GTM Compensation Benchmarks
Pavilion's B2B Tech Salary Benchmarks
Pay Transparency Law Postings
Primate Labs
RepVue/Bridge Group-style data
S&P Global Mobility
SWE-bench Verified
StartupCFO’s Startup Salary Benchmark
Strategy&
TechPowerUp
TechSpot
Towers Watson
Trusted salary benchmarking surveys (imercer)
U.S. Bureau of Labor Statistics (BLS) - OEWS
Wage data from government sources
Willis Towers Watson [WTW]
WorldatWork
WorldatWork Research
Discontinued, still offered
No shut-down product was recommended here.
← Online course librariesCompensation management →