# How to choose an LLM for your company

Canonical: https://www.swfte.com/how-to-choose-an-llm-for-your-company
Last verified: 2026-10-06
Difficulty: Beginner
Time: About 2 to 5 working days, mostly collecting the test cases and reading terms. The test runs themselves take hours.
Cost: Low: API usage for a few hundred test calls, plus staff time. Local candidates cost nothing beyond your hardware.
Hardware: None for hosted candidates. A laptop or workstation with enough memory for any small open-weight model you want to try.

## Short answer

Choose an LLM by writing your requirements first, turning the non-negotiables into pass or fail gates, shortlisting three to five candidates across hosted and open-weight options, then testing each on 30 to 100 of your own real cases with the same prompts. Score the results on a weighted scorecard, check each licence and data-handling term, and keep an exit plan. Leaderboards only help you shortlist.

## Who this is for

- Technology and product leads choosing a model or a provider for a specific company use case.
- Procurement, security and data protection staff who need a record of how a model was chosen.
- Teams starting with one hosted API who want to know when an open-weight model is a real alternative.

Not for:
- Anyone looking for a single best model for everyone. There is no such model; the answer depends on your task, data and constraints, and it changes monthly.
- Teams that need the open-weight testing method in detail. See [how to evaluate an open-source LLM](https://www.swfte.com/how-to-evaluate-an-open-source-llm) for that.

## Prerequisites

- A named use case, not a general ambition: one task with an owner and an example of a good result.
- 30 to 100 real examples of that task, with the answer a qualified person would accept, collected with permission and with personal data removed.
- Someone from security or data protection available to read vendor terms.
- A couple of hours of API access to each hosted candidate, and a machine to try open-weight ones (see [how to run LLMs locally](https://www.swfte.com/how-to-run-llms-locally)).

## Hosted API or open-weight model? The deciding questions

Most shortlists contain both kinds of candidate, so understand the trade first. A hosted model is called over an API run by a provider. An open-weight model is a file you can run on infrastructure you choose, under the licence that ships with it. Neither is better in general.

| Question | Points to a hosted API | Points to an open-weight model |
| --- | --- | --- |
| Must prompts and outputs stay on your infrastructure? | No, and the provider terms are acceptable | Yes, or a regulator or customer requires it |
| How hard is the task? | Open-ended reasoning or broad knowledge where the best available model matters | Narrow and well specified, where a small model with a good prompt or an adapter is enough |
| Volume and cost shape | Low or spiky volume: pay per use, no servers to run | High, steady volume where owning capacity is cheaper than per-token pricing |
| Who will operate it? | You have no capacity to run servers | You have, or will hire, people who can run and patch an inference server |
| How much control over change do you need? | You can accept the provider changing or retiring models on their schedule | You need to pin a version and keep it as long as you choose |

> TIP: Plan for both on the shortlist. A hosted model is a good reference for what is possible; an open-weight model shows what you could own. The test in step 5 will tell you how big the gap is on your task.

## Steps

### Step 1: Write the requirements in one page

Outcome: A one-page requirements note with an owner, written before any model is named.

Write the workload in concrete terms. What does the input look like: length, language, format, whether it contains personal or confidential data? What must the output be: free text, a label, JSON, a tool call? How many requests a day, at what peak? How fast must a reply arrive, and what happens if it is wrong or late?

Add the constraints that are not about quality. Where may the data be processed and stored? Which regulations apply to this use case? Who must approve the choice, and what evidence will they ask for? What is the budget per month, and what is the cost of a wrong answer?

This note is the thing every candidate is judged against. Keep it short enough that people read it. If two stakeholders disagree about what a good answer is, resolve that now: it is the most common reason a model test produces a result nobody believes.

**Requirements template**

| Area | Write down |
| --- | --- |
| Task | One sentence, with who uses the result and what they do next |
| Inputs | Typical and worst-case length, languages, formats, data sensitivity |
| Outputs | Format and a description of a correct result; examples of unacceptable ones |
| Volume and latency | Requests per day and at peak; acceptable reply time |
| Data and region | Classification, where processing and storage may happen, retention limits |
| Risk | What a wrong answer costs; what needs a human check |
| Budget and owner | Monthly ceiling; the named person who owns the choice |

### Step 2: Turn constraints into gates and the rest into weights

Outcome: A list of pass or fail gates and a weighted scorecard with weights summing to 100.

A gate is a requirement a candidate must meet or be dropped, however good it is otherwise: data may not leave the EU, licence must allow commercial use, must handle your languages. Gates remove candidates cheaply and make the decision defensible.

Everything else is a weighted preference. Agree the weights before testing, with the people who will own and use the result, and write them down. The weights below are an example of the shape only; they are not a recommendation, and yours should reflect your task. Setting them after you have seen results is how a scorecard gets bent to justify a favourite.

Keep the number of criteria small, five to eight. Many criteria dilute the ones that matter. Include cost and operating effort alongside quality, because a slightly better model that is three times the cost or needs a team to run is often the wrong answer.

**Example scorecard shape (illustrative weights, not a recommendation)**

| Criterion | Weight | How you measure it |
| --- | --- | --- |
| Task accuracy on your test set | 40 | Score from step 5 |
| Cost per 1,000 tasks | 20 | Tokens used in the test x price, scaled to your volume |
| Latency at your typical input | 10 | Median and slowest reply time in the test |
| Data control and terms | 15 | Checklist from step 4 |
| Operating effort | 10 | Who runs it and how much work that is |
| Exit ease | 5 | How hard it is to switch away, from step 7 |

### Step 3: Build a shortlist of three to five

Outcome: A shortlist that mixes at least one hosted model and, where your gates allow, at least one open-weight model.

Use public rankings and our own pages to shortlist, not to decide. Benchmarks measure generic tasks, often with prompts and settings that are not yours, and the best score on a leaderboard may be the wrong model for your task. The [AI leaderboard](https://www.swfte.com/ai/leaderboard), [best self-hosted models for enterprises](https://www.swfte.com/best-self-hosted-ai-model-for-enterprises) and [model pricing](https://www.swfte.com/ai/pricing) are good places to find candidates and see rough costs.

Aim for variety: a strong hosted model as a reference, a cheaper hosted model, and one or two open-weight models in sizes you could run. If you are under a data-residency gate, drop any hosted option that cannot meet it and keep the rest. Write each candidate down with the exact model identifier and the date, because names are reused across versions.

For open-weight candidates, first look at the licence on the model card. As an example of what you will find, the Hugging Face model API reported these tags on 6 October 2026. A tag is a pointer to the licence, not the licence itself, so open the licence file before you rely on it.

**Licence tags read from Hugging Face on 2026-10-06 (examples, not a recommendation)**

| Model | Licence tag | Gated download |
| --- | --- | --- |
| Qwen/Qwen3-8B | apache-2.0 | No |
| google/gemma-4-E2B-it | apache-2.0 | No |
| microsoft/Phi-4-mini-instruct | mit | No |
| HuggingFaceTB/SmolLM3-3B | apache-2.0 | No |
| meta-llama/Llama-3.1-8B-Instruct | llama3.1 (custom licence) | Yes, manual approval |
| mistralai/Ministral-8B-Instruct-2410 | other (read the licence) | No |

### Step 4: Check licences and data terms before you test

Outcome: A completed terms checklist per candidate, and any that fail a gate removed.

Do this early because it is free and decisive. For open-weight models, read the licence file that ships with the exact checkpoint. Permissive licences such as Apache-2.0 and MIT are the simpler starting point for commercial work; custom or restricted licences can carry limits on use, scale or region. Save the licence text with the date you read it. This is orientation, not legal advice, and your legal team should confirm it for your case.

For hosted providers, read the data terms for the exact service and plan you would use, and note the date, because they change. Two examples read on 6 October 2026. OpenAI's data controls page says data sent to its API is not used to train or improve its models unless you opt in, that abuse-monitoring logs are retained for up to 30 days, that a zero-data-retention control exists, and that data residency is offered in several regions including Europe (EEA and Switzerland), with some regions needing approval and additional amendments. Anthropic's privacy centre says that by default it does not use inputs or outputs from its commercial products, including its API, to train its models; it also notes that feedback you submit can be stored for up to five years. Those are examples of the level of detail to look for, not a verdict on either provider.

Repeat that exercise for every hosted candidate and ask the questions in the table. If a provider cannot answer one in writing, treat it as a no for gated data.

**Terms checklist for each candidate**

| Question | Where to find it |
| --- | --- |
| Is our data used to train models by default? | Provider data-usage page and the contract |
| How long are prompts and outputs retained, and can retention be set to zero? | Data controls page, plan terms |
| Where is data processed and stored, and can we choose a region? | Residency documentation and the agreement |
| Who are the sub-processors, and will we be told of changes? | Sub-processor list, data processing agreement |
| What does the licence allow: commercial use, modification, redistribution, scale limits, regions? | The licence file shipped with the weights |
| What happens when the model is updated or retired? | Deprecation policy for hosted models; your own pinning for open weights |

### Step 5: Test every candidate on your own tasks

Outcome: A table of accuracy, token use and latency for each candidate on the same test set.

Use your 30 to 100 real cases, the same prompt and the same settings for every candidate. Hosted providers and local servers such as Ollama, LM Studio and llama.cpp all expose an OpenAI-style chat endpoint, so one short script can run every candidate by changing the base URL and model name. Ollama's endpoint is `http://localhost:11434/v1/`; hosted providers publish theirs.

The script below is our own. It uses exact match, which suits labels and fixed formats; swap in a rubric or a human review for free text, and read [how to validate your AI](https://www.swfte.com/how-to-validate-your-ai) for the pitfalls of using a model as a judge. Some hosted models reject a `temperature` setting; if a call errors, remove it for that candidate and note it. If a server does not return token usage, count tokens another way, because you need them for the cost line.

Look at the misses, not just the totals. Group them: wrong facts, wrong format, refusals, language problems, timeouts. A candidate with a slightly lower score whose misses are all easy to fix with a better prompt may beat one whose misses are random. Run each candidate twice to see whether results are stable.

candidates.json: one entry per candidate:

```json
[
  {"name": "local-small", "base_url": "http://localhost:11434/v1/", "model": "qwen3:4b", "key_env": "NONE"},
  {"name": "hosted-reference", "base_url": "https://<provider-base-url>/v1", "model": "<model-id-from-provider>", "key_env": "PROVIDER_API_KEY"}
]
```

Run all candidates on the test set (compare.py):

```python
import json, os, time
from openai import OpenAI

candidates = json.load(open("candidates.json"))
cases = [json.loads(line) for line in open("data/test.jsonl", encoding="utf-8") if line.strip()]

for c in candidates:
    client = OpenAI(base_url=c["base_url"], api_key=os.environ.get(c["key_env"], "none"))
    hits = tokens_in = tokens_out = 0
    seconds = []
    for case in cases:
        messages = case["messages"]
        prompt, gold = messages[:-1], messages[-1]["content"].strip()
        start = time.time()
        reply = client.chat.completions.create(model=c["model"], messages=prompt, temperature=0)
        seconds.append(time.time() - start)
        hits += int(reply.choices[0].message.content.strip() == gold)
        tokens_in += reply.usage.prompt_tokens
        tokens_out += reply.usage.completion_tokens
    seconds.sort()
    print(f'{c["name"]}: {hits}/{len(cases)} correct | tokens in {tokens_in} out {tokens_out} | median {seconds[len(seconds)//2]:.2f}s | slowest {seconds[-1]:.2f}s')
```

Run it:

```bash
pip install openai
python compare.py
```

> WARNING: Never put real API keys in candidates.json. The file names an environment variable; set the variable in your shell. Keep the test set out of any prompt you share with a provider unless the data terms you checked allow it.

### Step 6: Score, decide and write the decision down

Outcome: A filled-in scorecard and a one-page decision record.

Convert each criterion to a score from 0 to 10 using the rules you set in step 2, multiply by the weights and add up. For cost, work from tokens: take the tokens used in the test, divide by the number of test cases to get tokens per task, multiply by the provider's price per million tokens (read from their pricing page on the day) and scale to 1,000 tasks. For a self-hosted model, use the monthly cost of the server divided by the tasks it handles in a month. Our [token cost calculator](https://www.swfte.com/ai/token-cost-calculator) does the first part.

Be sceptical of small differences. With 50 test cases, one case is two points of accuracy. If the top two candidates are within the noise, pick on operating effort, data control and exit ease, and say so in the record.

Write the decision record: the requirements note, the candidates and their exact identifiers, the licence and terms you read with dates, the test set description, the scores, the weights, the decision and who approved it. A reviewer or regulator can follow that, and so can whoever inherits the system. For the evidence your organisation may need around this, see [how to validate your AI](https://www.swfte.com/how-to-validate-your-ai).

### Step 7: Plan the exit and the re-test

Outcome: A way to switch models without rewriting the application, and a date to re-test.

Assume you will change models. Keep the model name, base URL and prompt in configuration rather than code, and call models through an OpenAI-style interface so that swapping a provider or moving to your own server is a configuration change. A gateway helps once you have more than one: see [how to set up an LLM gateway](https://www.swfte.com/how-to-set-up-an-llm-gateway).

Keep your assets portable: the test set, the prompts, the scorecard and the decision record belong to you, and they are what let you evaluate the next candidate in a day instead of a month. Avoid features that exist only at one provider unless they earn their place on the scorecard.

Write down the exit triggers: price changes, a deprecation notice, a failed re-test, a change in your data or region requirements. Schedule the re-test, for example every quarter, and re-run it whenever a candidate or your task changes. If you later want a model of your own, [how to create your own local model](https://www.swfte.com/how-to-create-your-own-local-model) shows what that involves.

## When to stop, and what a good answer looks like

Stop when one candidate clears every gate and wins the scorecard by a margin you can explain to a sceptical colleague. Do not keep testing new releases forever: models change monthly and a stable, tested choice beats a perfect one you never ship. Re-run the same test when a candidate is updated, when your task changes, or on a fixed schedule, and treat that as routine maintenance rather than a new project.

If no candidate clears the gates, the usual causes are a task that is too broad, test cases that disagree with each other, or a missing capability such as a language. Fix the task or the test first, then re-run. If the gap is knowledge, add retrieval ([how to build a RAG system](https://www.swfte.com/how-to-build-a-rag-system)) before you conclude the model is the problem.

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| Every candidate scores about the same | The test set is too easy, too small, or the metric is too coarse. | Add hard and rare cases, use more cases, and switch to a metric that separates them, such as field-level accuracy or a rubric. |
| The hosted candidate returns an error about `temperature` | Some models do not accept the setting you passed. | Remove the parameter for that candidate and note it in the record so the comparison stays honest. |
| The script crashes on `usage` being missing | The server did not return token counts. | Count tokens with the provider's tokeniser, or estimate from characters and say so in the record. |
| Stakeholders reject the result | Requirements or pass rules were agreed after results were seen, or the test cases are disputed. | Go back to the requirements note and get sign-off on the cases and weights, then re-run. |
| The best model fails a data gate | Provider terms or region do not meet the requirement. | Drop it. Ask whether a different plan or region meets the gate in writing; if not, move to the next candidate or an open-weight model you host. |
| Results change between runs | Sampling randomness, a changed hosted model, or a different prompt. | Use temperature zero where allowed, pin the model identifier, run twice and report the range. |
| A small open-weight model is much worse than the hosted reference | The task needs more capability or knowledge than a small model has. | Try a larger open-weight model, improve the prompt with examples, add retrieval, or accept a hosted model if your gates allow it. |

## Verify it worked

- [ ] The requirements note was written and agreed before any model was tested.
- [ ] Gates and weights were fixed before the first test run.
- [ ] Every candidate was tested on the same cases with the same prompt, and each result was repeated.
- [ ] For every candidate you have the licence or data terms read, with the date and a saved copy.
- [ ] The decision record names the model identifiers, the scores, the approver and the date.
- [ ] A switch to another model is a configuration change, and a re-test date is in the calendar.

## Next steps

- [How to evaluate an open-source LLM](https://www.swfte.com/how-to-evaluate-an-open-source-llm): the deeper method for open-weight candidates, including licence and supply chain
- [How to run LLMs locally](https://www.swfte.com/how-to-run-llms-locally): get open-weight candidates running on your own machine for the test
- [How to validate your AI](https://www.swfte.com/how-to-validate-your-ai): turn the test set into regression gates and an evidence pack
- [Open-source model testing](https://www.swfte.com/open-source-model-testing): how we test open-weight models, including a published log

## FAQ

### Which LLM is best for a company?

There is no single best model. It depends on your task, data sensitivity, volume, languages and budget, and rankings change monthly. Shortlist from public rankings, then test three to five candidates on your own cases and score them against requirements you wrote down first.

### Should we use a hosted API or an open-weight model?

Use a hosted API when data may leave your infrastructure, volume is modest and you want the strongest general capability. Use an open-weight model when prompts must stay in your environment, volume is high and steady, or you need to pin a version. Many companies use both.

### How many test cases do I need to compare LLMs?

Start with 30 to 100 real cases that include hard and rare ones. With 50 cases a single case is two points of accuracy, so treat small gaps as noise. More cases make the comparison steadier. This is our practical guidance, not a statistical rule.

### Can I use an open-source LLM commercially?

It depends on the licence of the exact checkpoint. Apache-2.0 and MIT are permissive; custom licences can add limits on use, scale or region. Open the licence file shipped with the weights, save it with the date, and have your legal team confirm for your case.

### Is my data used to train the provider's models?

Read the data terms for the exact plan. OpenAI states that API data is not used for training unless you opt in, and Anthropic states that it does not use commercial product inputs or outputs for training by default. Both pages carry other conditions, so read them and confirm in writing.

### How do I avoid vendor lock-in with LLMs?

Call models through an OpenAI-style interface, keep model names, base URLs and prompts in configuration, and keep your test set and scorecard so you can evaluate a replacement quickly. Ollama, LM Studio and llama.cpp all expose OpenAI-compatible endpoints, which makes local fallbacks easy to test.

### How often should we re-test our choice?

Re-run the same test when the model is updated or retired, when your task or data changes, and on a fixed schedule such as quarterly. Because you kept the test set and scorecard, a re-test is a day of work, not a new project.

## How Swfte can help

You can run this process with a spreadsheet and a script. If you want one interface in front of several hosted and local models while you compare, and a place to see usage and cost, these pages describe what exists.

- [Connect](https://www.swfte.com/products/connect): one OpenAI-compatible gateway in front of several models, with routing and usage tracking
- [Model leaderboard](https://www.swfte.com/ai/leaderboard): a starting point for the shortlist
- [Best self-hosted models for enterprises](https://www.swfte.com/best-self-hosted-ai-model-for-enterprises): open-weight candidates and what to check
- [Open-source model testing](https://www.swfte.com/open-source-model-testing): our method and log for testing open-weight models

Swfte is optional here. Nothing in the process depends on it.

## Sources

- [Hugging Face model pages and API](https://huggingface.co/Qwen/Qwen3-8B): licence tags and gating of the example models, read on 2026-10-06
- [OpenAI data controls](https://developers.openai.com/api/docs/guides/your-data): API data not used for training unless opted in, 30-day abuse-monitoring logs, zero data retention, data residency regions
- [Anthropic: is my data used for model training?](https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training): commercial products not used for training by default; feedback storage up to 5 years
- [Ollama OpenAI compatibility](https://docs.ollama.com/openai): local endpoint http://localhost:11434/v1/ and supported chat completions endpoint
- [LM Studio OpenAI compatibility](https://lmstudio.ai/docs/developer/openai-compat): LM Studio exposes an OpenAI-style endpoint at http://localhost:1234/v1
- [llama.cpp server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md): llama-server exposes /v1/chat/completions

Last verified against these sources on 2026-10-06.
