HR AI Index
Index Performance and talent management Performance management › Small Improvements vs Betterworks
Performance management · September 2026 Edition

Small Improvements vs Betterworks

Zero of twelve models named Small Improvements first on the direct prompt; one named Betterworks. Small Improvements was named by ten of the twelve models and Betterworks by twelve and Small Improvements carries 11 labels and Betterworks 31, so the shares are not directly comparable.

Small Improvements

accepted challenger

Named in two categories this edition.

Betterworks

accepted challenger

Named in seven categories this edition.

First-choice share7%4%Of first choices across the direct, paraphrase, budget and scale prompts, 0 to 100.
Negative rate9%10%Negative labels as a share of the product's labels, 0 to 100.
Rank in category#5#6A position in a field of 14; printed, not drawn.
Labels1131A count; the two differ.
The two percentage rows are drawn on one 0 to 100 track, Small Improvements reading right to left. Rank and label count are printed, not drawn.Lattice was named alongside these two in twelve of the twelve direct answers. Lattice vs Small Improvements · Lattice vs Betterworks · PerformYard vs Small Improvements

Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the performance management page.

By framing

How many of the twelve models made each the first choice, per way of asking, and how many argued against it.
Small ImprovementsFirst choices, of twelve modelsBetterworks
Direct01
Paraphrase01
Comparative00
Budget-constrained301 against Small Improvements
Scale-constrained00
Negative003 against Betterworks
Bars are first choices, 0 to 12 each sideModels that argued againstA model can name both, so the two sides of a row do not sum to twelve.

Every model, every framing

The seventy-two answers behind the chart above, one cell each: where Small Improvements and Betterworks stood in it.
ModelDirectParaphraseComparativeBudget-constrainedScale-constrainedNegative
Claude Haiku 4.5
GPT-5.4 mini
Gemini 3.5 Flash
Perplexity Sonar
Grok 4.1 Fast
Mistral Small
DeepSeek V4 Flash
Llama 4 Maverick
Qwen 3.7 Flash
Kimi K2
GLM 4.7 FlashX
MiniMax M2.5
Small Improvements Betterworks first choice named as an alternative argued againstblank: not namedEach cell is one answer, Small Improvements on the left and Betterworks on the right.

The direct prompt

The plain question, one answer per model, grouped by where Small Improvements and Betterworks stood in it.

Betterworks first, Small Improvements not the choice

1 of 12 modelsSmall Improvements was named in the answer but not as the choice, or not at all.
Llama 4 MaverickBetterworks, HiBob, Lattice

Neither was the first choice, one was named

5 of 12 modelsThe answer put something else first and named one of the two as an alternative.
Claude Haiku 4.5Lattice alternatives: 15Five, Betterworks, Culture Amp, HiBob, Leapsome, PerformYard
Gemini 3.5 FlashLattice alternatives: 15Five, Betterworks, HiBob, Leapsome
Perplexity SonarLattice alternatives: 15Five, Betterworks, HiBob, Leapsome
DeepSeek V4 FlashLattice alternatives: 15Five, Betterworks, Culture Amp, Leapsome, PerformYard
Qwen 3.7 FlashLattice alternatives: 15Five, Betterworks, HiBob

Neither was named

6 of 12 modelsThe answer made no first choice from these two in this category.
GPT-5.4 miniLattice alternatives: Mitratech Perform
Grok 4.1 FastLattice alternatives: 15Five, HiBob, Leapsome, Paylocity
Mistral SmallLattice alternatives: 15Five, Culture Amp
Kimi K2Lattice alternatives: 15Five, Culture Amp, Leapsome, PerformYard
GLM 4.7 FlashXLattice alternatives: 15Five, Culture Amp, HiBob, Leapsome, Rippling
MiniMax M2.5Lattice alternatives: Engagedly, Paylocity

Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.

By buyer segment

The same question asked on behalf of a different buyer. Each standing is computed within its segment and they are never added together. The figures above are the mid-market standing, which is the one the category orders by.
Small business
Small Improvements leads by six points.
Small Improvements6%#7 of 10
Betterworks0%#– of 10
The full small business standing →
Mid-marketThe figures above
Small Improvements leads by two points.
Small Improvements7%#5 of 14
Betterworks4%#6 of 14
The full mid-market standing →
Enterprise
The order flips: Betterworks leads at enterprise.
Betterworks25%#2 of 9
Small Improvements0%#– of 9
The full enterprise standing →

What the models said about Small Improvements

Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Four of four in this category shown.

“Cheapest credible option (but has seat minimums)” Kimi K2 · budget prompt · soft negative
“The best performance management software for a company with a limited budget is Small Improvements, which is the cheapest credible performance appraisal software” Llama 4 Maverick · budget prompt · first choice
“Often cited as the best value for money, Small Improvements strikes a balance between robust features and affordability.” Qwen 3.7 Flash · budget prompt · first choice
“Often cited as the cheapest credible option at $3/user/month, making it ideal for small businesses” Mistral Small · budget prompt · first choice

What the models said about Betterworks

Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Two of two in this category shown.

“Betterworks | Extremely expensive, poor integration across multiple systems, unresponsive customer service” GLM 4.7 FlashX · negative prompt · hard negative
“The best performance management software for a mid-market B2B company is Lattice, HiBob HRIS, or Betterworks.” Llama 4 Maverick · direct prompt · first choice
Also compared

Comparisons are drawn for the top eight products in each category, each against each. The output is the models' output; nothing here is a recommendation by the index.