Models · scored, not promised

Pick the model by its score on your work.

Staffbox runs open-weight models on a Mac in your building. Every option is scored on the same test, and a cloud model is used only for the jobs you choose, on your own key.

Local first, scored on the same test

Local on the boxDefault · measuredqwen3:8b on a 16 GB Mac mini, or gpt-oss-20b on a 32 GB box.
Your own modelScored firstIf it runs in Ollama or speaks the OpenAI-compatible API, we run your test on it.
Cloud on your keyOff by defaultOnly for the jobs you name, in writing. You pay the provider.

Every model below ran the same graded tests: two fictional companies with 40 questions each (quotes, policy lookups, email routing, and questions it should decline), plus 15 messy quote requests written before our latest changes. Same brain, same quote tool, graded by machine. No customer has run these yet; your pilot runs the same test on your own past requests.

ModelWho makes itRuns onFieldstone ITPeachtree CabinetMessy 15
qwen3:8b (default)Alibaba (Qwen)16 GB Mac mini40 of 4039 of 4015 of 15
gpt-oss-20bOpenAI (open weights, Apache 2.0)32 GB Mac (measured on a GPU server)39 of 4040 of 4015 of 15
qwen3.8 27BAlibaba (Qwen)Mac Studio class (measured on a GPU server)40 of 4039 of 40—
Same 8B, no brain1 of 401 of 40

On the 16 GB Mac mini the 8B answered in about 1 to 3.2 s on average, one question at a time; several people asking at once still queue. Mac Studio and Apple-hardware speed for gpt-oss-20b are not measured yet. measured

6 Oct, quote action v2.4 on the 16 GB Mac mini: 351 of 360 generated messy requests right, and 0 wrong totals in 475 answers; of 11 misses, 10 declined and asked the customer and 1 was a routing miss. On 5 Oct the previous version scored 345 of 360 with 6 wrong totals, all on requests that also asked about an item we don't stock (it priced the nearest stocked item or skipped the question); v2.4 flags the unstocked item and never substitutes. A person checks every draft. measured

Every graded answer: the Staffbox scorecard on Hugging Face. Raw files: evals/results.

How to choose

Most sites: the 8B on a Mac mini

The default. Scores above, about 1 to 3.2 s average per answer, and 4 W idle according to Apple. Nothing leaves the building.

Can't use Chinese-made models: gpt-oss-20b

OpenAI's open-weight model scored 39, 40 and 15 of 15. It needs a 32 GB box. Some public buyers prohibit Alibaba products; this is the option for them. See Government.

Bigger jobs: a Mac Studio or a cloud model you choose

If your test shows a bigger model is needed, we size a Mac Studio for your site (planned, not yet measured), or point chosen jobs at a cloud provider on your own key. opinion

Cloud AI, only if you turn it on

Cloud is off by default. If you want a cloud model for some jobs, ask in writing, add your own key, and we point those jobs at it. You pay the provider at their rates; Staffbox does not bill for tokens. For those jobs, the data goes to that provider under its terms.

ProviderWhat you getStatus with Staffbox
Anthropic, Google Gemini, Azure OpenAIFrontier modelsSupported on your key. Gemini Pro with the same notes scored 39 of 40 on a clean set and 14 of 20 on messy requests. measured
OpenRouterMany models behind one keySupported on your key
SaladOpen models on a shared GPU networkSupported on your key; not yet scored planned
BasetenLarge open models such as gpt-oss-120bSupported on your key; not yet scored planned
CrusoeManaged open models, including Hermes Agent support (Crusoe guide)Supported on your key; not yet scored planned

Supported means it speaks the OpenAI-compatible API our worker and test harness use. For public buyers we only propose US-made models; catalogs that include Qwen or DeepSeek models are filtered for them.

Bring your own model

Already licensed a model, or running one on your own servers? If it speaks the OpenAI-compatible API or runs in Ollama, we run your test on it and show you the score next to ours. Fine-tuning on your past requests is experimental and not part of a pilot. planned

Questions

Which model does Staffbox use by default?

qwen3:8b on a 16 GB Mac mini: 40 of 40 and 39 of 40 on our two test companies and 15 of 15 messy quote requests, about 1 to 3.2 s average per answer, one at a time. Fictional companies; your pilot runs the same test on your requests.

Is there a model that is not made in China?

Yes. OpenAI's open-weight gpt-oss-20b scored 39 of 40, 40 of 40 and 15 of 15 with the same brain and quote tool. It needs a 32 GB box; its speed on Apple hardware is not measured yet.

Can we use ChatGPT, Claude or Gemini instead?

For the jobs you choose, on your own key, after you ask in writing. You pay the provider directly. Those jobs' data goes to that provider.

Why not just use a cloud model?

Gemini Pro scored 14 of 20 on messy requests; our local 27B with the quote tool scored 16 of 20 on the same set (14 of 17 quotes exactly right). Our 8B with the quote action scored 15 of 15 on an unseen messy set (19 of 20 on a set used while fixing). On clean questions Gemini Pro matched us, 39 of 40. Our quote tool does the arithmetic in code from your price list, and the box keeps your data in the building.

Call HadiTextFree test