Models · scored, not promised
Pick the model by its score on your work.
Staffbox runs open-weight models on a Mac in your building. Every option is scored on the same test, and a cloud model is used only for the jobs you choose, on your own key.
Local first, scored on the same test
Every model below ran the same graded tests: two fictional companies with 40 questions each (quotes, policy lookups, email routing, and questions it should decline), plus 15 messy quote requests written before our latest changes. Same brain, same quote tool, graded by machine. No customer has run these yet; your pilot runs the same test on your own past requests.
| Model | Who makes it | Runs on | Fieldstone IT | Peachtree Cabinet | Messy 15 |
|---|---|---|---|---|---|
| qwen3:8b (default) | Alibaba (Qwen) | 16 GB Mac mini | 40 of 40 | 39 of 40 | 15 of 15 |
| gpt-oss-20b | OpenAI (open weights, Apache 2.0) | 32 GB Mac (measured on a GPU server) | 39 of 40 | 40 of 40 | 15 of 15 |
| qwen3.8 27B | Alibaba (Qwen) | Mac Studio class (measured on a GPU server) | 40 of 40 | 39 of 40 | — |
| Same 8B, no brain | 1 of 40 | 1 of 40 |
On the 16 GB Mac mini the 8B answered in about 1 to 3.2 s on average, one question at a time; several people asking at once still queue. Mac Studio and Apple-hardware speed for gpt-oss-20b are not measured yet. measured
6 Oct, quote action v2.4 on the 16 GB Mac mini: 351 of 360 generated messy requests right, and 0 wrong totals in 475 answers; of 11 misses, 10 declined and asked the customer and 1 was a routing miss. On 5 Oct the previous version scored 345 of 360 with 6 wrong totals, all on requests that also asked about an item we don't stock (it priced the nearest stocked item or skipped the question); v2.4 flags the unstocked item and never substitutes. A person checks every draft. measured
Every graded answer: the Staffbox scorecard on Hugging Face. Raw files: evals/results.
How to choose
Most sites: the 8B on a Mac mini
The default. Scores above, about 1 to 3.2 s average per answer, and 4 W idle according to Apple. Nothing leaves the building.
Can't use Chinese-made models: gpt-oss-20b
OpenAI's open-weight model scored 39, 40 and 15 of 15. It needs a 32 GB box. Some public buyers prohibit Alibaba products; this is the option for them. See Government.
Bigger jobs: a Mac Studio or a cloud model you choose
If your test shows a bigger model is needed, we size a Mac Studio for your site (planned, not yet measured), or point chosen jobs at a cloud provider on your own key. opinion
Cloud AI, only if you turn it on
Cloud is off by default. If you want a cloud model for some jobs, ask in writing, add your own key, and we point those jobs at it. You pay the provider at their rates; Staffbox does not bill for tokens. For those jobs, the data goes to that provider under its terms.
| Provider | What you get | Status with Staffbox |
|---|---|---|
| Anthropic, Google Gemini, Azure OpenAI | Frontier models | Supported on your key. Gemini Pro with the same notes scored 39 of 40 on a clean set and 14 of 20 on messy requests. measured |
| OpenRouter | Many models behind one key | Supported on your key |
| Salad | Open models on a shared GPU network | Supported on your key; not yet scored planned |
| Baseten | Large open models such as gpt-oss-120b | Supported on your key; not yet scored planned |
| Crusoe | Managed open models, including Hermes Agent support (Crusoe guide) | Supported on your key; not yet scored planned |
Supported means it speaks the OpenAI-compatible API our worker and test harness use. For public buyers we only propose US-made models; catalogs that include Qwen or DeepSeek models are filtered for them.
Bring your own model
Already licensed a model, or running one on your own servers? If it speaks the OpenAI-compatible API or runs in Ollama, we run your test on it and show you the score next to ours. Fine-tuning on your past requests is experimental and not part of a pilot. planned
Questions
Which model does Staffbox use by default?
qwen3:8b on a 16 GB Mac mini: 40 of 40 and 39 of 40 on our two test companies and 15 of 15 messy quote requests, about 1 to 3.2 s average per answer, one at a time. Fictional companies; your pilot runs the same test on your requests.
Is there a model that is not made in China?
Yes. OpenAI's open-weight gpt-oss-20b scored 39 of 40, 40 of 40 and 15 of 15 with the same brain and quote tool. It needs a 32 GB box; its speed on Apple hardware is not measured yet.
Can we use ChatGPT, Claude or Gemini instead?
For the jobs you choose, on your own key, after you ask in writing. You pay the provider directly. Those jobs' data goes to that provider.
Why not just use a cloud model?
Gemini Pro scored 14 of 20 on messy requests; our local 27B with the quote tool scored 16 of 20 on the same set (14 of 17 quotes exactly right). Our 8B with the quote action scored 15 of 15 on an unseen messy set (19 of 20 on a set used while fixing). On clean questions Gemini Pro matched us, 39 of 40. Our quote tool does the arithmetic in code from your price list, and the box keeps your data in the building.