Read · beginner
Choose the right LLM for the job
A practical guide to choosing fast, general, reasoning, long-context, multimodal and specialist models by workflow phase.
Choosing an LLM is easier when you stop asking “Which model is best?” and ask “What does this step need to do?”
A model that is excellent at a difficult coding decision may be wasteful for sorting incoming messages. A fast model may be perfect for extraction and frustrating for a task that needs several dependent decisions. The right choice depends on the task, the workflow phase, and the cost of being wrong.
This guide gives you a practical default for each situation, then shows how to test that default instead of choosing by reputation.
The short version
| Model shape | Start here when you need | Typical tasks |
|---|---|---|
| Fast and small | Low latency and low cost | Classification, extraction, routing, rewrites |
| General workhorse | Balanced quality and speed | Drafting, summarization, chat, tool calls, everyday coding |
| Reasoning | Several dependent decisions or hard constraints | Planning, debugging, trade-offs, difficult analysis |
| Long-context | More relevant input than a normal context window can hold | Large documents, codebases, transcripts, cross-file synthesis |
| Multimodal | Understanding images, audio or video is part of the task | Screenshot diagnosis, document vision, transcription, video review |
| Specialist | A narrow output is more important than open-ended conversation | Embeddings, speech, image generation, moderation, reranking |
These are starting points, not permanent labels. A smaller model with a clear prompt and a useful tool can beat a larger model with missing context. Measure the task you actually have.
Choose by workflow phase
Most useful AI systems do not need one model for everything. They use a small model for cheap, repeatable steps and reserve stronger models for decisions that justify their cost.
| Workflow phase | Good default | Escalate when |
|---|---|---|
| Intake and routing | Fast and small | The request is ambiguous or the route has a high cost of failure |
| Retrieval and extraction | Fast and small, with structured output | The source is long, messy, visual or spread across many files |
| Planning | General workhorse | The plan has many branches, dependencies or competing constraints |
| Execution with tools | General workhorse | The model must recover from unexpected tool results or make risky changes |
| Verification | Fast model for checks; stronger model for review | A missed error would be expensive, unsafe or hard to reverse |
| Final explanation | General workhorse or fast model | The explanation must reconcile conflicting evidence or teach a difficult concept |
The important pattern is simple: cheap steps filter and prepare; expensive steps decide and verify.
Intake and routing
Use a fast model when the output is a label, a route, or a small piece of normalized data:
- identify whether a support message is billing, technical or sales;
- extract a date, amount and customer ID;
- decide whether a request needs a human;
- choose which tool or workflow should handle the next step.
Give it a closed set of possible outputs and validate the result in code. If the model can return any sentence, it is doing more work than the router needs.
Retrieval and extraction
Start with a fast model when the source is short and the job is mechanical. Use a long-context or multimodal model only when the input itself creates the difficulty.
Do not use a larger context window as an excuse to send an entire database. Retrieve the relevant records, remove repeated boilerplate, and preserve the source location of each extracted fact. More input is not automatically more useful input.
Planning
A general model is usually enough for a short plan with known steps. Try a reasoning model when the plan must satisfy several constraints at once, compare alternatives, or respond to discoveries made during the work.
For example, “write a three-step checklist” is ordinary generation. “Plan a migration that keeps the service available, respects this dependency order, and includes rollback steps” is a reasoning task.
Execution with tools
Choose a model with reliable tool calling and enough instruction-following quality for the consequences of its actions. The model does not need to be the most intelligent available if the workflow gives it narrow tools, explicit permissions and a clear stop condition.
For risky actions, split the work:
- Ask one model to propose the action.
- Validate arguments and permissions in your application.
- Execute the tool.
- Ask the model to inspect the result and either continue or stop.
The application remains responsible for authorization, validation and irreversible operations. Model choice cannot replace those controls.
Verification
Verification deserves its own model call or, better, a deterministic check. Use tests, schemas, type checkers, database constraints and parsers whenever they can answer the question.
Use a model to review what deterministic checks cannot easily judge: whether an explanation answers the question, whether a draft follows a tone, or whether a proposed design missed an important trade-off. A reviewer should receive the output, the original requirements and a clear rubric—not only “is this good?”
Match model capability to failure cost
The strongest model is not always the correct model. Start by naming what failure means for your task:
- Low cost: a slightly awkward rewrite or an imperfect tag. Prefer speed and volume.
- Medium cost: a customer receives an unhelpful answer or a report needs manual cleanup. Use a balanced model and a fallback.
- High cost: code is changed incorrectly, money is moved, private data is exposed or a safety decision is wrong. Use stronger review, deterministic guardrails and human approval where needed.
This gives you a better rule than “always use the biggest model”: spend quality budget where an error is costly, not where the prompt looks impressive.
A practical routing policy
Keep routing logic explicit. You can start with a small decision function and replace the labels with the model IDs supported by your provider:
type Task = {
inputTokens: number;
needsVision: boolean;
needsReasoning: boolean;
highVolume: boolean;
highRisk: boolean;
};
const chooseModelFamily = (task: Task) => {
if (task.needsVision) return "multimodal";
if (task.inputTokens > 80_000) return "long-context";
if (task.highRisk || task.needsReasoning) return "reasoning";
if (task.highVolume) return "fast";
return "general";
};
Real routing needs more than token count. Add the capabilities your task requires: structured output, tool calling, streaming, regional availability, data controls and maximum response length. Keep those requirements in configuration so changing providers does not require rewriting your workflow.
For a multi-step agent, route each step separately. A single model setting at the top of the agent is convenient, but it hides cost and makes it harder to improve one phase without changing every other phase.
How to compare models in practice
Make a small evaluation set from real work. Twenty representative examples are more useful than a leaderboard score that does not match your prompts.
For each example, record:
- Task result: Did it produce a usable answer?
- Failure type: What exactly went wrong?
- Latency: How long did the user or next step wait?
- Cost: What did the request consume at your real input and output sizes?
- Operational fit: Did it support the context, tools, output format and region you need?
Compare models on the same prompt, context, tools and checks. Run the evaluation more than once when outputs are variable. A model that wins one dramatic example may lose on the ordinary cases that make up most of your bill.
Keep a simple routing table beside the evaluation:
| Route | Default family | Success check | Fallback |
|---|---|---|---|
| Ticket classification | Fast | Valid label from allowed set | General |
| Code change proposal | General | Tests and review rubric pass | Reasoning |
| Large policy comparison | Long-context | Required claims have source locations | Reasoning |
| Production action | Reasoning or general | Validated tool arguments and approval | Human |
Update the table when the evidence changes. Do not update it because a new model name is trending.
Common mistakes
Choosing by brand or benchmark alone
Benchmarks can tell you something about a model’s capabilities, but they do not measure your prompt, context, tools, language, latency target or failure cost. Test your route.
Using reasoning for simple work
Reasoning is useful when there is a real decision to make. It is unnecessary for a fixed classification, a short rewrite or a schema-validated extraction. Extra thinking can add latency and cost without improving the answer.
Sending a huge context instead of improving retrieval
Long context helps when the relevant evidence is genuinely large or cross-document. It does not fix duplicated, stale or irrelevant input. Better retrieval often beats a larger model.
Treating a model as a permission system
The model may suggest an action. Your application must decide whether that action is allowed. Validate arguments, scope tools narrowly, log decisions and require approval for irreversible work.
Having no fallback
Every route should define what happens when the preferred model times out, rejects the input, returns invalid structure or fails its quality check. A fallback can be another model, a retry with less context, a deterministic path or a human.
A five-question checklist
Before choosing a model, answer:
- What is the smallest acceptable output?
- Is the task classification, generation, reasoning, retrieval, transformation or verification?
- What input modality, context size, tool support and output format are required?
- How costly is a wrong answer, and what can verify it?
- What is the cheapest model that passes a representative evaluation set?
If you cannot answer question five, you are not choosing yet—you are guessing. Start with a baseline, collect failures and let the evaluation move the route.
The takeaway
Use fast models for volume, general models for everyday work, reasoning models for difficult decisions, long-context models for genuinely large inputs, multimodal models for non-text evidence and specialist models for narrow operations.
Then split your workflow by phase. Let inexpensive models prepare work. Spend quality on decisions and high-risk reviews. Keep deterministic checks and application permissions outside the model.
The correct LLM is not the one with the biggest reputation. It is the least expensive model that reliably passes your task’s real quality bar.