Five of twelve models named Deputy first on the direct prompt; zero named Homebase. Deputy was named by eleven of the twelve models and Homebase by twelve and Deputy carries 26 labels and Homebase 25, so the shares are not directly comparable.
Named in eight categories this edition.
Named in eleven categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the workforce scheduling page.
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of three in this category shown.
“1. Deputy (Best Overall Balance) ... Deputy strikes the best balance between sophisticated feature sets (like AI forecasting) and user-friendly design.” Qwen 3.7 Flash · direct prompt · first choice
“I'd recommend Deputy if you have multiple locations or compliance needs, or Connecteam if you have deskless/field workers” Kimi K2 · paraphrase prompt · first choice
“Look at Homebase, Deputy, or When I Work\u2014these are entirely different tools from the booking apps above.” DeepSeek V4 Flash · comparative prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of three in this category shown.
“Sling or Homebase (note: these cap out quickly at mid-scale)” Qwen 3.7 Flash · paraphrase prompt · soft negative
“Recommendation: Start with Homebase if you have one location and want time tracking included for free—it's the most balanced for limited budgets” Grok 4.1 Fast · budget prompt · first choice
“When I Work or Homebase are excellent starting points due to their balance of features, ease of use, and affordability” Mistral Small · paraphrase prompt · first choice
Comparisons are drawn for the top three products in each category. The output is the models' output; nothing here is a recommendation by the index.