HR AI Index
Index Research › Repeat measurement · September 2026 Edition
Repeat measurement · September 2026 Edition

Asked the same question twice, the models moved their own first choice 70% of the time. The leader still held in 7 of 10 categories.

Ten categories were asked again one days after the edition run, with nothing changed: the same prompts, the same model versions, the same settings, one buyer segment, six framings, twelve models, 720 answers. Whatever differs between the two runs is noise, and the noise is what this page measures. That is the reason the index asks every question six ways of every model rather than once: on its own, one ask moved 70% of the time; seventy-two, read together, are what the index reports.

Model by model

For each model, the share of its questions where the top pick differed between the two runs. The pairs are one question asked in each run; a pair flips when the first choice is not the same product.
MiniMax M2.589%34 of 38
Qwen 3.7 Flash83%40 of 48
GPT-5.4 mini76%29 of 38
GLM 4.7 FlashX76%29 of 38
Kimi K272%29 of 40
DeepSeek V4 Flash70%33 of 47
Gemini 3.5 Flash69%33 of 48
Mistral Small68%25 of 37
Claude Haiku 4.567%24 of 36
Llama 4 Maverick55%17 of 31
Grok 4.1 Fast55%24 of 44
Perplexity Sonar51%18 of 35

MiniMax M2.5 changed its first choice most often, 89% of its 38 pairs; Perplexity Sonar least, 51% of 35. Pooled over every model and question, 70%.

A flip is the same model, the same question, a different first choice a few days apart. It says how much one answer can be trusted on its own, and nothing about why the model answered as it did.

Category by category

The leader's share of first choices in the edition run and in the repeat, and whether the leader was the same product both times. Sorted by the size of the move.
CategoryLeader in the edition runShare, run oneShare, repeatMoveLeader
Contingent workforce man SimpleVMS35%22%-13 pointsheld
Compensation benchmarkin Payscale24%33%+9 pointsheld
AI recruiting assistants GoPerfect19%11%-8 pointschanged: Workable
Employee onboarding BambooHR29%25%-4 pointschanged: Rippling
Benefits administration PlanSource23%20%-2 pointschanged: Gusto
Learning management syst TalentLMS48%50%+2 pointsheld
Performance management Lattice41%43%+2 pointsheld
Employee engagement surv Culture Amp40%42%+2 pointsheld
Global payroll Deel22%24%+2 pointsheld
Coaching platforms Growthspace26%27%+1 pointheld

In 7 of the 10 categories the same product led both runs. Where the leader changed, the two products were within 3 points of each other in the edition run.

The share floor

The bar a change has to clear before the index calls it a change, in the unit of the change itself.
9 pointsthe floor: the 90th percentile of the moves above
2 pointsmedian move of a leader's share on a repeat
13 pointsthe largest move, Contingent workforce man

From the next edition on, a product's change in share counts as movement only when it is larger than 9 points, and a new leader is reported only when it clears the old one by more than that. Nine repeats in ten move a leader less. The floor is measured again with every edition and the method page carries the rule: how the floor is measured.

Cite this

HR AI Recommendation Index, September 2026 Edition: repeat measurement. hr-ai-index.com/research/repeat-measurement/. Published under CC BY 4.0. Every figure on this page is computed from the published edition and changes with it; the edition and its date are the citation.

The output is the models' output. Nothing here says the models can be steered, and nothing here is a recommendation by the index.