Model sizing

SLM vs LLM: when a small language model is enough

Learn how primary sources define a small language model, when one is enough, when it is not, and how to route each task to the smallest model that passes.

A small language model (SLM) is, in the NVIDIA research paper “Small Language Models are the Future of Agentic AI”, a model that fits on a common consumer device and answers one user’s requests with practical latency. A large language model (LLM) is any model that is not an SLM. The paper gives a hedged figure of 10 billion parameters for 2025 and warns that such thresholds go stale, so decide by the task: narrow, repeated and bounded work often suits an SLM, and open-ended work often needs a larger model.

Last verified 2026-10-07. Sources are listed at the end of the page.

How do primary sources define a small language model?

The sources define “small” differently, which is the useful finding. Read on 2026-10-07.

SourceWhat it saysWhat to take from it
NVIDIA research paper, “Small Language Models are the Future of Agentic AI”Section 2.1 defines an SLM as a language model that can fit onto a common consumer electronic device and perform inference with latency low enough to be practical for one user’s agentic requests. An LLM is a model that is not an SLM. The authors say they would be comfortable with most models below 10 billion parameters being SLMs as of 2025, and chose definitions that avoid hardware-specific metrics such as parameter count, which they say quickly become obsolete.The definition is about a device and a latency, not a number. The figure is the authors’ own 2025 view.
Microsoft Phi-4-mini-instruct model card3.8 billion parameters and a 128K-token context. Released under the MIT licence. Intended for memory or compute constrained environments and latency-bound scenarios. The card also states the model does not have the capacity to store much factual knowledge, and that languages other than English perform worse.A vendor “small” model states its own limits: facts and non-English text.
Google Gemma 4 model cardOne family spans E2B (2.3 billion effective parameters), E4B (4.5 billion effective), 12B, a 26B mixture-of-experts model with 3.8 billion active, and a 31B dense model. The card describes the sizes as deployable from high-end phones to laptops and servers.One vendor ships “small” and “mid-size” in the same family, so size labels differ by vendor.

The NVIDIA paper is a position paper by NVIDIA researchers. It argues a case, and the vendor cards describe the vendors’ own models.

When is a small language model enough?

The NVIDIA paper’s argument is that many agentic systems have language models performing a small number of specialised tasks repetitively and with little variation, and that an SLM can be sufficient for many of those calls. The conditions it points to are a narrow task, a stable format, high volume, and a need for low latency or local execution.

Four signals favour a small model. The task has a fixed output shape, such as a label, a JSON object or a short rewrite. The volume is high enough that cost per call matters. The latency budget is tight, or the data may not leave a device or network. And you can write a test that says what a correct answer is. If the task fails any of the four, plan for a larger model, or for a router.

Fine-tuning can move a task from the second group to the first. A small model tuned on your reviewed examples can match a larger prompted one on that narrow task. That is a hypothesis to test on a held-out set, never an assumption.

When is a small language model not enough?

  • Facts and knowledge. Phi-4-mini’s card says it lacks capacity to store much factual knowledge. Give any model a retrieval step for facts instead of expecting it to remember them.
  • Open-ended conversation. The NVIDIA paper itself says that where general-purpose conversational ability is essential, heterogeneous systems that call several models are the natural choice.
  • Long, multi-step reasoning across broad context. A small model may lose the thread. Test it on your longest real case before trusting it.
  • Languages beyond the model’s strengths. Phi-4-mini’s card warns of worse performance outside English, and a language list on any card is not a quality measurement.
  • Unfamiliar tool use. The Phi-4-mini card says the model could sometimes hallucinate function names or URLs in function-calling scenarios. Validate every tool call.
  • High-stakes answers with no check behind them. Size is not the main risk there. The missing check is.

How do cost and memory compare, without made-up numbers?

You can reason about cost with arithmetic you can check, and measure the rest on your own hardware.

FactorWhat changes with model sizeHow to check it
Memory for weightsRoughly parameters multiplied by bytes per parameter. A 4-bit build needs about half the bytes of an 8-bit one. A mixture-of-experts model keeps every expert in memory, so active parameters do not cut memory.Use the model card’s own memory statement where it gives one, such as Google’s Gemma table, and add the context window.
Compute per tokenFalls with the parameters used per token. Active parameters, not total, set the compute of a mixture-of-experts model.Measure tokens per second on your hardware at your concurrency.
LatencyA smaller model on local hardware can avoid a network hop. Shared servers add queueing.Time to first token and inter-token latency, measured on your real prompts.
Quality on your taskNot a function of size alone. A tuned small model can beat a larger prompted one on a narrow task, and lose on a broad one.Score both on the same held-out set before choosing.
OperationsSeveral small models mean several artefacts to host, version and re-test.Count the models you would have to keep current, and who owns each.

How do I route each task to the smallest model that works?

Routing makes the answer to “SLM or LLM” a per-task setting, so you are never choosing once for everything.

  1. 1. List your task types

    Group real traffic into types such as classify, extract, summarise, draft, answer from documents and plan. Count the volume of each.

  2. 2. Build one test set per type

    Use real examples with a written pass rule. Keep the sets in your own storage.

  3. 3. Find the smallest passing model per type

    Run small and large candidates on each set. Pick the smallest that clears your bar with the margin you chose in advance.

  4. 4. Add an escalation rule

    If a small model’s output fails a schema check or a confidence test, send the request to a larger model. This is a fallback chain.

  5. 5. Re-test on a schedule

    Models and traffic change. Re-run the sets whenever you swap a model and every quarter.

Should I fine-tune a small model?

Fine-tune when a prompt cannot make the model behave consistently on a narrow, stable task and you hold reviewed examples of correct output. Adapter methods such as LoRA keep the cost of training low, which is why small models and adapters often go together. See LoRA LLM fine-tuning for what is trained, how adapters are merged or served, and when it is the wrong tool.

Do not fine-tune to add facts. Use retrieval for knowledge the model lacks, as in how to build a RAG system, and tune only for behaviour. Whatever you train, the model inherits the base model’s licence, so read it before you start.

Where Swfte fits

You do not need Swfte to choose between a small and a large model: the test set above does that. Swfte helps with routing and hosting once you have chosen. Status words are Built, In progress and Roadmap.

CapabilityStatusWhat it means here
Connect routingBuiltOne OpenAI-compatible API. Routing rules per workspace with a mandatory fallback chain, and named strategies that include cost-optimised and fallback-chain. You set the rules, and Connect does not judge task difficulty for you. Usage caps can downgrade to a cheaper model when a cap is hit. See Connect.
Model VaultBuiltHosts weights you bring, with versions, stage promotion and deploy to a dedicated endpoint, so a small model you trained or chose can sit behind Connect. See own your model and deploy and serve.
EvaluationBuiltStudio scores single-turn chat. See evaluation and safety.
Managed fine-tuning or distillationRoadmapNot offered. Swfte research on specialised models is exploratory. Train elsewhere with open tools and bring the weights.

Sources and last verified

Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.

Frequently asked questions

What is the difference between an SLM and an LLM?

An SLM is a language model small enough to run on a common consumer device with practical latency for one user, per the NVIDIA paper’s definition. An LLM is any model that is not an SLM. The paper adds that it would be comfortable calling most models below 10 billion parameters SLMs as of 2025, but warns that size thresholds go out of date.

How many parameters is a small language model?

No source fixes a firm number. The NVIDIA paper’s working view is that most models below 10 billion parameters count as small as of 2025. Microsoft’s Phi-4-mini has 3.8 billion, and Google’s Gemma 4 family runs from 2.3 billion effective parameters to a 31 billion dense model. Define small by your device and latency instead.

Are SLMs cheaper than LLMs?

Per call they usually need less memory and compute, which is arithmetic you can check, but total cost also depends on volume, quality and operations. A small model that fails often and needs retries, or a larger fleet to maintain, can cost more. Measure cost per correct answer on your own test set.

Can a small language model replace GPT-class models for agents?

For some calls, yes. The NVIDIA paper argues that many agent invocations are narrow and repetitive, and suit small models, while open-ended conversation favours systems that mix models. Treat that as a position to test. Move one task type at a time, behind a fallback to a larger model.

Can I fine-tune a small language model on Swfte?

No. Swfte does not offer managed fine-tuning, and it is on the Roadmap as research only. You can train a small model elsewhere with open tools and bring the weights to the Model Vault, which hosts and versions them. Evaluation and routing in Connect then work around it.

How do I decide between an SLM and an LLM for one task?

Write a test with a pass rule, then run both on real examples. Choose the smaller model if it clears your bar with a margin and the task has a fixed output shape, high volume or a tight latency or privacy limit. If it fails some cases, keep it and escalate those to the larger model.

Route each task to the smallest model that passes your tests.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.