HR AI Index
Index Talent acquisition Technical hiring assessm › TestGorilla vs Codility
Technical hiring assessments · September 2026 Edition

TestGorilla vs Codility

Zero of twelve models named TestGorilla first on the direct prompt; one named Codility. Both were named by all twelve models and TestGorilla carries 31 labels and Codility 44, so the shares are not directly comparable.

TestGorilla

accepted challenger

Named in four categories this edition.

Codility

criticized challenger

Named in three categories this edition.

First-choice share5%5%Of first choices across the direct, paraphrase, budget and scale prompts, 0 to 100.
Negative rate16%32%Negative labels as a share of the product's labels, 0 to 100.
Rank in category#5#6A position in a field of 6; printed, not drawn.
Labels3144A count; the two differ.
The two percentage rows are drawn on one 0 to 100 track, TestGorilla reading right to left. Rank and label count are printed, not drawn.HackerRank was named alongside these two in ten of the twelve direct answers. HackerRank vs TestGorilla · HackerRank vs Codility · CoderPad vs TestGorilla

Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the technical hiring assessments page.

By framing

How many of the twelve models made each the first choice, per way of asking, and how many argued against it.
TestGorillaFirst choices, of twelve modelsCodility
Direct011 against Codility
Paraphrase002 against Codility
Comparative02
Budget-constrained201 against TestGorilla · 4 against Codility
Scale-constrained01
Negative004 against TestGorilla · 7 against Codility
Bars are first choices, 0 to 12 each sideModels that argued againstA model can name both, so the two sides of a row do not sum to twelve.

Across every category in the September 2026 Edition, TestGorilla and Codility were named in the same answer 100 times, of the 238 answers naming TestGorilla and the 186 naming Codility. In those answers Codility took the first choice nine times and TestGorilla twenty-five.

Every model, every framing

The seventy-two answers behind the chart above, one cell each: where TestGorilla and Codility stood in it.
ModelDirectParaphraseComparativeBudget-constrainedScale-constrainedNegative
Claude Haiku 4.5
GPT-5.4 mini
Gemini 3.5 Flash
Perplexity Sonar
Grok 4.1 Fast
Mistral Small
DeepSeek V4 Flash
Llama 4 Maverick
Qwen 3.7 Flash
Kimi K2
GLM 4.7 FlashX
MiniMax M2.5
TestGorilla Codility first choice named as an alternative argued againstblank: not namedEach cell is one answer, TestGorilla on the left and Codility on the right.

The direct prompt

The plain question, one answer per model, grouped by where TestGorilla and Codility stood in it.

Codility first, TestGorilla an alternative

1 of 12 modelsTestGorilla was named in the answer but not as the choice, or not at all.
Kimi K2Codility alternatives: CodeSignal, Coderbyte, HackerRank, TestGorilla

Neither was the first choice, one was named

4 of 12 modelsThe answer put something else first and named one of the two as an alternative.
GPT-5.4 miniHackerRank alternatives: CodeSignal, CoderPad, Codility
Gemini 3.5 FlashCoderPad alternatives: CodeSignal, HackerEarth, TestGorilla
DeepSeek V4 FlashCodeSignal alternatives: Codility, HackerRank, TestGorilla
MiniMax M2.5CodeSignal alternatives: Augment Code, HackerRank, TestGorilla

Neither was named

7 of 12 modelsThe answer made no first choice from these two in this category.
Claude Haiku 4.5no first choice
Perplexity SonarHackerRank alternatives: CodeSignal, CoderPad, Coderbyte, Goodfit
Grok 4.1 FastHackerRank alternatives: CodeSignal, Coderbyte
Mistral SmallHackerRank alternatives: Augment Code, Coderbyte
Llama 4 Maverickno first choice alternatives: CodeSignal, CoderPad, Coderbyte, Goodfit, HackerRank
Qwen 3.7 FlashHackerRank alternatives: CodeSignal, TestDome, Vervoe
GLM 4.7 FlashXHackerRank alternatives: CodeSignal

Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.

By buyer segment

The same question asked on behalf of a different buyer. Each standing is computed within its segment and they are never added together. The figures above are the mid-market standing, which is the one the category orders by.
Small business
TestGorilla leads by twenty-one points.
TestGorilla21%#2 of 7
Codility0%#7 of 7
The full small business standing →
Mid-marketThe figures above
Level: the same share of first choices.
TestGorilla5%#5 of 6
Codility5%#6 of 6
The full mid-market standing →
Enterprise
The order flips: Codility leads at enterprise.
Codility19%#2 of 6
TestGorilla0%#6 of 6
The full enterprise standing →

What the models said about TestGorilla

Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of four in this category shown.

“Billing issues: Multiple reviews mention unethical business practices, with charges continuing after cancellation” MiniMax M2.5 · negative prompt · hard negative
“It has a reputation for being buggy and confusing... making it a poor measure of senior engineering ability” DeepSeek V4 Flash · negative prompt · hard negative
“Frequently called one of the "worst online assessment tools" due to poor usability, confusing interfaces” Grok 4.1 Fast · negative prompt · hard negative

What the models said about Codility

Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Five of five in this category shown.

“it is often described as "quite pricey" ... Codility has had issues with "test integrity," where questions and answers are not sufficiently hidden from recruiters” GLM 4.7 FlashX · negative prompt · hard negative
“Opaque scoring: Candidates and reviewers criticize edge-case scoring and lack of transparency” MiniMax M2.5 · negative prompt · hard negative
“Why we recommend avoiding HackerRank or Codility (for now)” Gemini 3.5 Flash · paraphrase prompt · hard negative
“Codility offers the best balance of capability, reputation, and price” Kimi K2 · direct prompt · first choice
“Often ranked as the best overall, especially for enterprise use.” Mistral Small · comparative prompt · first choice
Also compared

Comparisons are drawn for the top eight products in each category, each against each. The output is the models' output; nothing here is a recommendation by the index.