AI Indexes
HR AI Index
October 2026 Edition · The permanent record of this edition. The unqualified address always carries the latest edition.
Index › Compensation and total rewards › October 2026 Edition

Compensation benchmarking data

Asked as “salary benchmarking data provider”, and as “compensation survey data”, on behalf of a mid-market B2B company. 58 first choices recorded across the direct, paraphrase, budget and scale prompts, fourteen models each.
Standing · first-choice share
19%
Contested · Payscale 19%
19Pave19Payscale10SalaryCube52others

19% of first choices, contested.

Since September 2026=heldSince September 2026: 20% → 20%, ±0 points. Inside the 13-point floor: within noise. Read over the models both editions asked.Pave held the lead, ±0 points on 20%, inside the 13-point floor.

By buyer segment

The same question asked on behalf of a different buyer. Each standing is computed within its segment; they sit side by side and are never added together.

The standing

Share is the count of first choices across the direct, paraphrase, budget and scale prompts, over all fourteen models, for a mid-market B2B company. Ordered by share.
ProductFirst-choice shareNegative rateLabelsQuadrantSince September 2026
01Pave19%4%27accepted challenger=heldSince September 2026: 20% → 20%, ±0 points. Inside the 13-point floor: within noise. Read over the models both editions asked.20% → 20%
02Payscale19%14%35accepted challenger▼−2Since September 2026: 24% → 22%, −2 points. Inside the 13-point floor: within noise. Read over the models both editions asked.24% → 22%
03SalaryCube10%8%12accepted challenger▲+6Since September 2026: 4% → 10%, +6 points. Inside the 13-point floor: within noise. Read over the models both editions asked.4% → 10%
04Salary.com CompAnalyst9%4%24accepted challenger▲+4Since September 2026: 2% → 6%, +4 points. Inside the 13-point floor: within noise. Read over the models both editions asked.2% → 6%
05Ravio7%4%26accepted challenger▲+8Since September 2026: 0% → 8%, +8 points. Inside the 13-point floor: within noise. Read over the models both editions asked.0% → 8%
06Comprehensive.io7%0%10accepted challenger▼−4Since September 2026: 12% → 8%, −4 points. Inside the 13-point floor: within noise. Read over the models both editions asked.12% → 8%
07Radford5%50%28criticized challenger▲+2Since September 2026: 4% → 6%, +2 points. Inside the 13-point floor: within noise. Read over the models both editions asked.4% → 6%
08Mercer Total Remuneration Survey3%52%33criticized challenger▼−4Since September 2026: 6% → 2%, −4 points. Inside the 13-point floor: within noise. Read over the models both editions asked.6% → 2%
09WTW2%58%12criticized challenger=heldSince September 2026: 0% → 0%, ±0 points. Inside the 13-point floor: within noise. Read over the models both editions asked.0% → 0%
Show the four products at 0%, ordered by negative rate
13Glassdoor0%100%10criticized challenger=heldSince September 2026: 0% → 0%, ±0 points. Inside the 13-point floor: within noise. Read over the models both editions asked.0% → 0%
10Figures0%0%13accepted challenger=heldSince September 2026: 0% → 0%, ±0 points. Inside the 13-point floor: within noise. Read over the models both editions asked.0% → 0%
11Gartner0%0%12accepted challenger=heldSince September 2026: 0% → 0%, ±0 points. Inside the 13-point floor: within noise. Read over the models both editions asked.0% → 0%
12APQC0%0%10accepted challenger=heldSince September 2026: 0% → 0%, ±0 points. Inside the 13-point floor: within noise. Read over the models both editions asked.0% → 0%

The floor is 13 points of share, measured: how far the models move a leader on their own when the same questions are asked twice with nothing changed. A larger change is movement; a smaller one is noise, and both are shown. Movement is read over the twelve models both editions asked; GPT-6 Luna, Muse Glimmer 30B joined this edition and are in the standing but not yet in the comparison. How the floor is measured

Bars are the share of first choices, 0 to 100Every product with at least 10 labels here. Every product name links to its product page.
All twenty-eight head-to-head pages: the top eight products, each against each

Recommended versus criticized

Every product with at least 10 labels here, on both axes. The 30% line names a quadrant, not the verdict above: that one needs more than 40%.

Criticized challengerCriticized default
Negative label rate →
01
02
03
04
05
06
R07
08
09
10
11
12
13
Accepted challengerEndorsed leader
0%First-choice share → · lines at 30% share and 25% negative30%
Key
01Pave19%
02Payscale19%
03SalaryCube10%
04Salary.com CompAnalyst9%
05Ravio7%
06Comprehensive.io7%
07Radford5%
08Mercer Total Remuneration Survey3%
09WTW2%
10Figures0%
11Gartner0%
12APQC0%
13Glassdoor0%

What they warned about

Three of fourteen models held their first choice under the paraphrase. GPT-5.4 mini, Gemini 3.5 Flash, Perplexity Sonar, Grok 4.1 Fast, Mistral Small, Llama 4 Maverick, Qwen 3.7 Flash, Kimi K2, GLM 4.7 FlashX, GPT-6 Luna and Muse Glimmer 30B changed. A high negative share on a product with few labels is a warning. A low share on a product with many labels is salience, not sentiment.
Mercer Total Remuneration Survey
52%
17 of 33 labels negative · 9 of 14 models · 4 hard negative
“**Avoid:** Traditional survey providers like Mercer, Korn Ferry, or Willis Towers Watson” GLM 4.7 FlashX, paraphrase prompt
Radford
50%
14 of 28 labels negative · 9 of 14 models · 1 hard negative
“"Why you should avoid traditional providers (Mercer, Radford, Willis Towers Watson)"” Muse Glimmer 30B, negative prompt
Glassdoor
100%
10 of 10 labels negative · 6 of 14 models · 3 hard negative
“Avoid providers that rely heavily on employee self-reporting (like Glassdoor data) or scraped job postings, which are often inflated or inaccurate.” Qwen 3.7 Flash, scale prompt
Willis Towers Watson
56%
5 of 9 labels negative · 5 of 14 models · 4 hard negative
“**Avoid:** Traditional survey providers like Mercer, Korn Ferry, or Willis Towers Watson” GLM 4.7 FlashX, paraphrase prompt

What they cite

Citations exist only for the models that return a source list: fourteen of the fourteen in this edition, and all six flagship models on the expanded tier.

Sites the answers cite

74 of 84 answers in this category came back with a source list, from 14 of 14 models: citations where the model returns them, or the search results it consulted. 1009 links across 241 sites, every framing counted. Ranked by the number of answers carrying the site or page. 28 of the 252 answers across every segment cited this index's own page for the category; the method page measures whether that reading tilts an answer.

vendor site · Ravio49 answers · 134 citations · 12 models
28 answers · 66 citations · 11 models
vendor site · Comprehensive28 answers · 35 citations · 11 models
28 answers · 28 citations · 12 models
vendor site · Pin24 answers · 24 citations · 12 models
vendor site · SalaryCube23 answers · 49 citations · 11 models
vendor site · Deel23 answers · 25 citations · 10 models
22 answers · 42 citations · 11 models
16 answers · 36 citations · 10 models
vendor site · Figures16 answers · 19 citations · 9 models
vendor site · G215 answers · 20 citations · 8 models
vendor site · CaptivateIQ14 answers · 15 citations · 10 models

Pages the answers cite

The ten pages named in the most answers, by full address. A page here is one the models returned with a recommendation, not one the index endorses.

Search against answers

Each company's standing in the answers beside its site's footprint in Google search, one row a site: the products the models named on it with their shares, and the share they add up to; monthly searches on Google, and DataForSEO's estimate of AI search demand (modeled from search signals, directional, not a count of queries to any assistant), for the most-searched of the company's and its products' names (the name is in each row's hover text); estimated monthly organic visits to the site; and its best position in Google's top ten for “benchmarking providers”, “best salary benchmarking data provider”, “salary benchmarking data provider”. US estimates from DataForSEO and Google's Ads Transparency Center. A small company's site, or a mid-sized company's site for its flagship, is marked company; a product on a large parent's site (Google, Microsoft) has no site figures. A column with no figures for this category is left out, and an empty cell means none were seen, not none exist. Two measurements side by side: neither is read as the cause of the other.
Company and productsShareOwn site citedName searches, GoogleAI search demand, est.Organic visitsPaid search a month, est.Google ads, last 30 daysBest Google position
Pave pave.com
Pave 19%, Pave Market Data Lite 2%
21%21327,10028,4237,645$917
Payscale payscale.com
Payscale 19%, Payscale Ascent 2%
21%2299,9002991,408,989$26333
SalaryCube salarycube.com10%373002,82218
Salary.com CompAnalyst salary.com9%9026031,069,118$1,64270
Ravio ravio.com7%1651,9001818,4191
Comprehensive.io comprehensive.io7%555002,0213
Radford5%18,1005,850
Mercer mercer.com
Mercer Total Remuneration Survey 3%, Mercer Benchmark Database 2%
5%10649,50023,68198,058 company184
CompData Surveys compdata.ai3%3
Lattice lattice.com2%66290,50013,900122,123$1,328120

Google's top ten for the three searches hold 25 sites; 5 of them are among the sites the models cited here (captivateiq.com, comprehensive.io, g2.com, ravio.com, salarycube.com). The first result for “benchmarking providers” is thehackettgroup.com.

How they answered

Six framings of the same buying question, each sent to every model in a fresh session with search on. One row per model, so a row shows whether it held its answer under rewording, what it named when cost was the constraint, and what it argued against. Computed from the raw judge labels.
ShowHide
ModelDirect“What is the best salary benchmarking data provider for a mid-market B2B company?”Paraphrase“Which compensation survey data would you recommend to a mid-sized B2B company?”Comparative“What are the top benchmarking providers and how do they differ?”Budget-constrained“What is the best salary benchmarking data provider for a company with a limited budget?”Scale-constrained“We are a 500 person company evaluating a salary benchmarking data provider. What should we look at?”Negative“Which benchmarking providers should I avoid or be cautious about?”
Claude Haiku 4.5Payscale
Four alternativesComprehensive, Lattice, Pave, SalaryCube
PayscaleHeld
Three alternativesRavio, Salary.com CompAnalyst, Willis Towers Watson
against: Mercer Total Remuneration Survey
no first choiceComprehensive.io, Pave
Three alternativesBureau of Labor Statistics, Payscale, Salary.com CompAnalyst
no first choiceagainst: UserBenchmark
GPT-5.4 miniPayscale
Three alternativesMercer Total Remuneration Survey, Radford, WTW
no first choiceChangedno first choiceMercer MarketPricer, Payscaleagainst: Culpepper Compensation Surveys, Korn Ferry Pay, Mercer Total Remuneration Survey, Radford, WTWno first choicenothing named
Gemini 3.5 FlashPave
Five alternativesFigures, Payscale, Ravio, RepVue, Salary.com CompAnalyst
against: Aon Radford, Mercer Total Remuneration Survey
Pave, PayscaleChanged
Four alternativesCarta Total Compensation, Figures, Ravio, Salary.com CompAnalyst
against: Radford
Mercer Total Remuneration Survey, Willis Towers WatsonPave
Five alternativesBureau of Labor Statistics (BLS) - OEWS, Comprehensive.io, Figures, OpenComp, Ravio
against: Glassdoor, Mercer Total Remuneration Survey, Radford, WTW
no first choiceagainst: Glassdoor, Levels.fyi, Payscaleagainst: Glassdoor, Korn Ferry Pay, MGMA, Payscale, Radford, Sullivan Cotter
Perplexity SonarRavio
Five alternativesCompUp, Pave, Payfactors, Payscale Ascent, Salary.com CompAnalyst
CompData Surveys, Xactly's/Benchmarkit's B2B sales compensation surveyChanged
Two alternativesBDO private company executive compensation survey, Mercer SME/market surveys
against: Payscale
Mercer Total Remuneration SurveyPave
Four alternativesBLS OEWS, Comprehensive.io, Payscale, Ravio
no first choiceagainst: Korn Ferry Pay, Radford, Willis Towers Watson
Grok 4.1 FastSalaryCube
Two alternativesPayscale, Salary.com CompAnalyst
against: Korn Ferry Pay, Lattice, Mercer Total Remuneration Survey, Ravio, WTW
Salary.com CompAnalystChanged
Two alternativesCulpepper Compensation Surveys, Payscale
against: Radford, WTW
no first choiceComprehensive.io
Five alternativesPave, Payscale, Ravio, Salary.com CompAnalyst, U.S. Bureau of Labor Statistics (BLS) Occupational Employment Statistics
against: Glassdoor
no first choiceagainst: Glassdooragainst: GSM8K, HumanEval, LMSYS Chatbot Arena, MLPerf, MMLU, SPEC, SWE-bench, TPC, Terminal-Bench, WebArena
Mistral SmallRavio, SalaryCube
Two alternativesCompUp, Figures
Mercer Benchmark Database, RadfordChanged
Two alternativesBureau of Labor Statistics (BLS) – Occupational Employment and Wage Statistics, Payscale
The Conference Board
Six alternativesAPQC, Computer Economics, Deloitte, Gartner, McKinsey & Company, Oxera
Payscale, Radfordno first choiceagainst: UserBenchmark
DeepSeek V4 FlashPayscale
Two alternativesPave, Salary.com CompAnalyst
against: Mercer Total Remuneration Survey, Radford, SalaryCube
PayscaleHeld
Two alternativesBetts Recruiting Compensation Guide, Pavilion's B2B GTM Compensation Benchmarks
against: Mercer Total Remuneration Survey, Radford
no first choice
Seven alternativesAvasant, Computer Economics, Deloitte, Forrester, Gartner, ISG, KPMG
Comprehensive.io, Pave
Two alternativesBLS Occupational Employment & Wage Statistics, Ravio
against: Aon's Radford McLagan Compensation Database, Mercer Total Remuneration Survey, Willis Towers Watson
no first choiceagainst: Paveagainst: Aon's Radford McLagan Compensation Database, Brightmine, Mercer Total Remuneration Survey, Radford, WTW
Llama 4 Maverickno first choiceMercer Total Remuneration Survey, PayscaleChanged
Two alternativesravio.com, salarycube.com
no first choice
Two alternativesDeloitte, McKinsey & Company
Pave
Two alternativesBLS OEWS, Comprehensive.io
no first choiceagainst: Glassdoor, Hays
Qwen 3.7 FlashLattice
Four alternativesComprehensive, Deel HR, Pave, SalaryCube
against: Mercer Total Remuneration Survey, Radford
PayscaleChanged
Three alternativesKorn Ferry Pay, Radford, Robert Half
against: Aon's Radford McLagan Compensation Database, Mercer Total Remuneration Survey, WTW
Mercer Total Remuneration Survey
Two alternativesAon's Radford McLagan Compensation Database, Radford
Salary.com CompAnalyst
Three alternativesPave, Ravio, U.S. Bureau of Labor Statistics (BLS) Occupational Employment and Wage Statistics
against: Mercer Total Remuneration Survey, Willis Towers Watson
no first choice
Three alternativesLattice, Payscale, Ravio
against: Glassdoor, Gusto, Mercer Total Remuneration Survey, Radford, WTW, Workday
nothing named
Kimi K2Payscale, SalaryCube
Four alternativesFigures, Pave, Ravio, Salary.com CompAnalyst
against: Radford
Radford, Salary.com CompAnalystChanged
Seven alternativesBridge Group, Mercer Total Remuneration Survey, Pave, Payfactors, Payscale, SalaryCube, Workleap Compensation
no first choicePave, Ravio, U.S. Bureau of Labor Statistics
Four alternativesDeel Compensation, Option Impact, Pequity, SalaryExpert
against: Glassdoor
no first choice
Five alternativesFigures, Lattice, Pave, Payscale, Salary.com CompAnalyst
against: GAIA, LINPACK, LMSYS Chatbot Arena, OSWorld, SWE-bench, Terminal-Bench, WebArena, Whetstone
GLM 4.7 FlashXPave, Payscale Ascent
Two alternativesFigures, Salary.com CompAnalyst
against: Mercer Total Remuneration Survey
RavioChanged
Three alternativesBureau of Labor Statistics, Lattice, SalaryCube
against: Glassdoor, Indeed, Korn Ferry Pay, Mercer Total Remuneration Survey, Willis Towers Watson
no first choice
Seven alternativesAPQC, Accenture, Deloitte, Forrester, Gartner, Grant Thornton, McKinsey & Company
BLS, Comprehensive.io, Pave
Five alternativesFigures, Levels.fyi, Payscale, Ravio, Salary.com CompAnalyst
against: Glassdoor, Indeed
no first choicenothing named
MiniMax M2.5SalaryCube
Three alternativesCompUp, Lattice, Salary.com CompAnalyst
SalaryCubeHeld
Five alternativesMercer Total Remuneration Survey, Pave, Payscale, Radford McLagan, Ravio
no first choiceBureau of Labor Statistics
Three alternativesComprehensive.io, Indeed, Pave
against: Payscale, Salary.com CompAnalyst
no first choice
Four alternativesCompAnalyst, Figures, Payscale, Ravio
against: Mercer Total Remuneration Survey, Radford
nothing named
GPT-6 LunaPave
Three alternativesMercer Total Remuneration Survey, Payscale, Radford
Mercer Total Remuneration Survey, WTWChanged
Two alternativesAon Radford, Payscale
Aon's Radford McLagan Compensation Database, Mercer Total Remuneration Survey, WTW
Four alternativesAPQC, Gartner, ISG, The Hackett Group
Salary.com CompAnalyst
One alternativeBureau of Labor Statistics' free wage data
no first choicenothing named
Muse Glimmer 30BSalaryCube
Two alternativesPave, Salary.com CompAnalyst
CompData Surveys, Salary.com CompAnalystChanged
Four alternativesMercer Total Remuneration Survey TRS / Management Benchmarking Database / Comptryx, Payscale Knowledge / Payscale Salary Survey, Radford, Aon, SalaryCube
Deloitte
Three alternativesAPQC, Computer Economics, Gartner
Pave Market Data Lite
Two alternativesBLS Occupational Employment and Wage Statistics OEWS, Comprehensive.io
no first choiceagainst: Mercer Total Remuneration Survey, Payscale, Radford, UserBenchmark, Willis Towers Watson
Bold is the first choiceAlternatives are counted; the count opens them.What the answer argued against

The record

One row per call: the version string exactly as returned, whether the model searched, sources cited, and latency. Full answer text is in the free responses file. Download the record
Eighty-four rows: every prompt, every model, every answer.
PromptModelVersion stringTime (UTC)SearchedSourcesLatency
Direct recommendationClaude Haiku 4.5claude-haiku-4-5-202510012026-10-01 08:26yes97 s
Direct recommendationGPT-5.4 minigpt-5.4-mini-2026-03-172026-10-01 11:17yes36 s
Direct recommendationGemini 3.5 Flashgemini-3.5-flash2026-10-01 10:51yes1626 s
Direct recommendationPerplexity Sonarsonar2026-10-01 10:08yes204 s
Direct recommendationGrok 4.1 Fastspacexai/grok-4.1-fast-non-reasoning via vertex2026-10-01 10:51yes198 s
Direct recommendationMistral Smallmistral/mistral-small via mistral2026-10-01 11:14yes115 s
Direct recommendationDeepSeek V4 Flashdeepseek/deepseek-v4-flash via deepinfra2026-10-01 10:09yes1826 s
Direct recommendationLlama 4 Maverickmeta/llama-4-maverick via bedrock2026-10-01 08:08yes52 s
Direct recommendationQwen 3.7 Flashalibaba/qwen3.7-flash via alibaba2026-10-01 10:01yes855 s
Direct recommendationKimi K2moonshotai/kimi-k2 via novita2026-10-01 10:06yes1420 s
Direct recommendationGLM 4.7 FlashXzai/glm-4.7-flashx via zai2026-10-01 09:36yes12310 s
Direct recommendationMiniMax M2.5minimax/minimax-m2.5 via minimax2026-10-01 09:27yes539 s
Direct recommendationGPT-6 Lunagpt-6-luna2026-10-01 07:29yes318 s
Direct recommendationMuse Glimmer 30Bmeta/muse-glimmer-30b via togetherai2026-10-01 07:51yes2253 s
ParaphraseClaude Haiku 4.5claude-haiku-4-5-202510012026-10-01 07:53yes1710 s
ParaphraseGPT-5.4 minigpt-5.4-mini-2026-03-172026-10-01 12:08no06 s
ParaphraseGemini 3.5 Flashgemini-3.5-flash2026-10-01 09:11yes1222 s
ParaphrasePerplexity Sonarsonar2026-10-01 09:17yes203 s
ParaphraseGrok 4.1 Fastspacexai/grok-4.1-fast-non-reasoning via vertex2026-10-01 09:06yes1813 s
ParaphraseMistral Smallmistral/mistral-small via mistral2026-10-01 11:11no06 s
ParaphraseDeepSeek V4 Flashdeepseek/deepseek-v4-flash via deepinfra2026-10-01 07:45yes2228 s
ParaphraseLlama 4 Maverickmeta/llama-4-maverick via bedrock2026-10-01 07:42yes51 s
ParaphraseQwen 3.7 Flashalibaba/qwen3.7-flash via alibaba2026-10-01 11:53no037 s
ParaphraseKimi K2moonshotai/kimi-k2 via novita2026-10-01 08:44yes1326 s
ParaphraseGLM 4.7 FlashXzai/glm-4.7-flashx via zai2026-10-01 09:34yes2139 s
ParaphraseMiniMax M2.5minimax/minimax-m2.5 via minimax2026-10-01 11:17yes725 s
ParaphraseGPT-6 Lunagpt-6-luna2026-10-01 08:01yes533 s
ParaphraseMuse Glimmer 30Bmeta/muse-glimmer-30b via togetherai2026-10-01 09:57yes1234 s
ComparativeClaude Haiku 4.5claude-haiku-4-5-202510012026-10-01 09:18yes179 s
ComparativeGPT-5.4 minigpt-5.4-mini-2026-03-172026-10-01 09:07no06 s
ComparativeGemini 3.5 Flashgemini-3.5-flash2026-10-01 09:47yes1823 s
ComparativePerplexity Sonarsonar2026-10-01 09:15yes216 s
ComparativeGrok 4.1 Fastspacexai/grok-4.1-fast-non-reasoning via vertex2026-10-01 08:49yes2111 s
ComparativeMistral Smallmistral/mistral-small via mistral2026-10-01 11:05yes56 s
ComparativeDeepSeek V4 Flashdeepseek/deepseek-v4-flash via deepinfra2026-10-01 07:53yes2442 s
ComparativeLlama 4 Maverickmeta/llama-4-maverick via bedrock2026-10-01 11:13yes510 s
ComparativeQwen 3.7 Flashalibaba/qwen3.7-flash via alibaba2026-10-01 09:08yes1632 s
ComparativeKimi K2moonshotai/kimi-k2 via novita2026-10-01 11:53yes2536 s
ComparativeGLM 4.7 FlashXzai/glm-4.7-flashx via zai2026-10-01 11:48yes2332 s
ComparativeMiniMax M2.5minimax/minimax-m2.5 via minimax2026-10-01 07:56yes550 s
ComparativeGPT-6 Lunagpt-6-luna2026-10-01 09:41yes613 s
ComparativeMuse Glimmer 30Bmeta/muse-glimmer-30b via togetherai2026-10-01 08:18yes1431 s
Budget constrainedClaude Haiku 4.5claude-haiku-4-5-202510012026-10-01 08:00yes1610 s
Budget constrainedGPT-5.4 minigpt-5.4-mini-2026-03-172026-10-01 08:14yes35 s
Budget constrainedGemini 3.5 Flashgemini-3.5-flash2026-10-01 08:24yes1919 s
Budget constrainedPerplexity Sonarsonar2026-10-01 10:05yes203 s
Budget constrainedGrok 4.1 Fastspacexai/grok-4.1-fast-non-reasoning via vertex2026-10-01 09:10yes1912 s
Budget constrainedMistral Smallmistral/mistral-small via mistral2026-10-01 08:16yes53 s
Budget constrainedDeepSeek V4 Flashdeepseek/deepseek-v4-flash via deepinfra2026-10-01 12:01yes2431 s
Budget constrainedLlama 4 Maverickmeta/llama-4-maverick via bedrock2026-10-01 12:09yes52 s
Budget constrainedQwen 3.7 Flashalibaba/qwen3.7-flash via alibaba2026-10-01 09:25yes524 s
Budget constrainedKimi K2moonshotai/kimi-k2 via novita2026-10-01 10:07yes919 s
Budget constrainedGLM 4.7 FlashXzai/glm-4.7-flashx via zai2026-10-01 09:01yes2329 s
Budget constrainedMiniMax M2.5minimax/minimax-m2.5 via minimax2026-10-01 11:16yes832 s
Budget constrainedGPT-6 Lunagpt-6-luna2026-10-01 08:57yes314 s
Budget constrainedMuse Glimmer 30Bmeta/muse-glimmer-30b via togetherai2026-10-01 09:11yes2227 s
Scale constrainedClaude Haiku 4.5claude-haiku-4-5-202510012026-10-01 08:55no06 s
Scale constrainedGPT-5.4 minigpt-5.4-mini-2026-03-172026-10-01 09:50no09 s
Scale constrainedGemini 3.5 Flashgemini-3.5-flash2026-10-01 12:17no051 s
Scale constrainedPerplexity Sonarsonar2026-10-01 08:04yes206 s
Scale constrainedGrok 4.1 Fastspacexai/grok-4.1-fast-non-reasoning via vertex2026-10-01 09:35yes137 s
Scale constrainedMistral Smallmistral/mistral-small via mistral2026-10-01 08:28no06 s
Scale constrainedDeepSeek V4 Flashdeepseek/deepseek-v4-flash via deepinfra2026-10-01 11:50yes1639 s
Scale constrainedLlama 4 Maverickmeta/llama-4-maverick via bedrock2026-10-01 08:07yes52 s
Scale constrainedQwen 3.7 Flashalibaba/qwen3.7-flash via alibaba2026-10-01 11:57yes847 s
Scale constrainedKimi K2moonshotai/kimi-k2 via novita2026-10-01 10:54yes1925 s
Scale constrainedGLM 4.7 FlashXzai/glm-4.7-flashx via zai2026-10-01 09:10yes1893 s
Scale constrainedMiniMax M2.5minimax/minimax-m2.5 via minimax2026-10-01 07:46yes533 s
Scale constrainedGPT-6 Lunagpt-6-luna2026-10-01 08:09yes319 s
Scale constrainedMuse Glimmer 30Bmeta/muse-glimmer-30b via togetherai2026-10-01 11:55yes1221 s
Negative framingClaude Haiku 4.5claude-haiku-4-5-202510012026-10-01 09:40yes208 s
Negative framingGPT-5.4 minigpt-5.4-mini-2026-03-172026-10-01 09:40yes57 s
Negative framingGemini 3.5 Flashgemini-3.5-flash2026-10-01 10:55yes1222 s
Negative framingPerplexity Sonarsonar2026-10-01 07:33yes394 s
Negative framingGrok 4.1 Fastspacexai/grok-4.1-fast-non-reasoning via vertex2026-10-01 08:23yes249 s
Negative framingMistral Smallmistral/mistral-small via mistral2026-10-01 09:20yes55 s
Negative framingDeepSeek V4 Flashdeepseek/deepseek-v4-flash via deepinfra2026-10-01 10:57yes2133 s
Negative framingLlama 4 Maverickmeta/llama-4-maverick via bedrock2026-10-01 08:41yes52 s
Negative framingQwen 3.7 Flashalibaba/qwen3.7-flash via alibaba2026-10-01 11:04yes1030 s
Negative framingKimi K2moonshotai/kimi-k2 via novita2026-10-01 11:29yes2527 s
Negative framingGLM 4.7 FlashXzai/glm-4.7-flashx via zai2026-10-01 11:44no024 s
Negative framingMiniMax M2.5minimax/minimax-m2.5 via minimax2026-10-01 08:57yes1328 s
Negative framingGPT-6 Lunagpt-6-luna2026-10-01 08:29no03 s
Negative framingMuse Glimmer 30Bmeta/muse-glimmer-30b via togetherai2026-10-01 07:39yes1320 s

Noise floor in this category

Flips between the edition run and its calibration repeat. Six prompts per model is a small sample; the index-wide floor is the number to trust.

ShowHide
Claude Haiku 4.5
3 of 3 flipped
GPT-5.4 mini
4 of 4 flipped
Gemini 3.5 Flash
2 of 4 flipped
Perplexity Sonar
3 of 5 flipped
Grok 4.1 Fast
2 of 3 flipped
Mistral Small
3 of 4 flipped
DeepSeek V4 Flash
2 of 4 flipped
Llama 4 Maverick
1 of 2 flipped
Qwen 3.7 Flash
4 of 4 flipped
Kimi K2
3 of 4 flipped
GLM 4.7 FlashX
4 of 4 flipped
MiniMax M2.5
3 of 3 flipped
GPT-6 Luna
4 of 4 flipped
Muse Glimmer 30B
3 of 4 flipped

Normalization in this category

Every judgment call made between the raw labels and the numbers above, listed so it is visible and reversible.

ShowHide
Category-scoped readings
Aon read as Aon's Radford McLagan Compensation Database
Bain read as Bain & Company
Carta Total Comp read as Carta Total Compensation
Culpepper read as Culpepper Compensation Surveys
Korn Ferry read as Korn Ferry Pay
Korn Ferry (Hay Group) read as Korn Ferry Pay
McKinsey read as McKinsey & Company
Mercer read as Mercer Total Remuneration Survey
Salary.com read as Salary.com CompAnalyst
Salary.com (CompAnalyst) read as Salary.com CompAnalyst
Salary.com (CompAnalyst/PayFactors) read as Salary.com CompAnalyst
Salary.com by CompAnalyst read as Salary.com CompAnalyst
Unresolved, counted raw
APEC
ARC-AGI
Artificial Analysis
Avasant
BDO private company executive compensation survey
BLS Occupational Employment and Wage Statistics OEWS
Betts Recruiting (Comp Engine)
Betts Recruiting Compensation Guide
Bureau of Labor Statistics (BLS) - OEWS
Bureau of Labor Statistics (BLS) – Occupational Employment and Wage Statistics (OEWS)
Bureau of Labor Statistics' free wage data
Comptryx
CrystalDiskMark
Deel Compensation
GAIA
GSM8K
Grant Thornton
HumanEval
Kienbaum
LINPACK (1979)
LinkedIn Job Postings
MGMA
MLCommons / MLPerf
MMLU
Mercer SME/market surveys
Mercer Total Remuneration Survey TRS / Management Benchmarking Database / Comptryx
OSWorld
PCMark
PassMark
Pavilion's B2B GTM Compensation Benchmarks
Payscale Knowledge / Payscale Salary Survey
RSM US
Radford McLagan
Radford, Aon
SPEC benchmarks
Sales Incentive Compensation Association (SICA)
Squirrel
Sullivan Cotter
U.S. Bureau of Labor Statistics (BLS) Occupational Employment Statistics
Whetstone (1964)
Xactly's Sales Compensation Report
Xactly's/Benchmarkit's B2B sales compensation survey
ravio.com
salarycube.com
Discontinued, still offered
No shut-down product was recommended here.
← 401(k) providersCompensation management →