← Library

Read · beginner

Choose the right LLM for the job

A practical guide to choosing fast, general, reasoning, long-context, multimodal and specialist models by workflow phase.

Choosing an LLM is easier when you stop asking “Which model is best?” and ask “What does this step need to do?”

A model that is excellent at a difficult coding decision may be wasteful for sorting incoming messages. A fast model may be perfect for extraction and frustrating for a task that needs several dependent decisions. The right choice depends on the task, the workflow phase, and the cost of being wrong.

This guide gives you a practical default for each situation, then shows how to test that default instead of choosing by reputation.

The short version

LLM model families and their best starting use cases
Model shapeStart here when you needTypical tasks
Fast and smallLow latency and low costClassification, extraction, routing, rewrites
General workhorseBalanced quality and speedDrafting, summarization, chat, tool calls, everyday coding
ReasoningSeveral dependent decisions or hard constraintsPlanning, debugging, trade-offs, difficult analysis
Long-contextMore relevant input than a normal context window can holdLarge documents, codebases, transcripts, cross-file synthesis
MultimodalUnderstanding images, audio or video is part of the taskScreenshot diagnosis, document vision, transcription, video review
SpecialistA narrow output is more important than open-ended conversationEmbeddings, speech, image generation, moderation, reranking

These are starting points, not permanent labels. A smaller model with a clear prompt and a useful tool can beat a larger model with missing context. Measure the task you actually have.

Choose by workflow phase

Most useful AI systems do not need one model for everything. They use a small model for cheap, repeatable steps and reserve stronger models for decisions that justify their cost.

Workflow phase Good default Escalate when
Intake and routing Fast and small The request is ambiguous or the route has a high cost of failure
Retrieval and extraction Fast and small, with structured output The source is long, messy, visual or spread across many files
Planning General workhorse The plan has many branches, dependencies or competing constraints
Execution with tools General workhorse The model must recover from unexpected tool results or make risky changes
Verification Fast model for checks; stronger model for review A missed error would be expensive, unsafe or hard to reverse
Final explanation General workhorse or fast model The explanation must reconcile conflicting evidence or teach a difficult concept

The important pattern is simple: cheap steps filter and prepare; expensive steps decide and verify.

Intake and routing

Use a fast model when the output is a label, a route, or a small piece of normalized data:

  • identify whether a support message is billing, technical or sales;
  • extract a date, amount and customer ID;
  • decide whether a request needs a human;
  • choose which tool or workflow should handle the next step.

Give it a closed set of possible outputs and validate the result in code. If the model can return any sentence, it is doing more work than the router needs.

Retrieval and extraction

Start with a fast model when the source is short and the job is mechanical. Use a long-context or multimodal model only when the input itself creates the difficulty.

Do not use a larger context window as an excuse to send an entire database. Retrieve the relevant records, remove repeated boilerplate, and preserve the source location of each extracted fact. More input is not automatically more useful input.

Planning

A general model is usually enough for a short plan with known steps. Try a reasoning model when the plan must satisfy several constraints at once, compare alternatives, or respond to discoveries made during the work.

For example, “write a three-step checklist” is ordinary generation. “Plan a migration that keeps the service available, respects this dependency order, and includes rollback steps” is a reasoning task.

Execution with tools

Choose a model with reliable tool calling and enough instruction-following quality for the consequences of its actions. The model does not need to be the most intelligent available if the workflow gives it narrow tools, explicit permissions and a clear stop condition.

For risky actions, split the work:

  1. Ask one model to propose the action.
  2. Validate arguments and permissions in your application.
  3. Execute the tool.
  4. Ask the model to inspect the result and either continue or stop.

The application remains responsible for authorization, validation and irreversible operations. Model choice cannot replace those controls.

Verification

Verification deserves its own model call or, better, a deterministic check. Use tests, schemas, type checkers, database constraints and parsers whenever they can answer the question.

Use a model to review what deterministic checks cannot easily judge: whether an explanation answers the question, whether a draft follows a tone, or whether a proposed design missed an important trade-off. A reviewer should receive the output, the original requirements and a clear rubric—not only “is this good?”

Match model capability to failure cost

The strongest model is not always the correct model. Start by naming what failure means for your task:

  • Low cost: a slightly awkward rewrite or an imperfect tag. Prefer speed and volume.
  • Medium cost: a customer receives an unhelpful answer or a report needs manual cleanup. Use a balanced model and a fallback.
  • High cost: code is changed incorrectly, money is moved, private data is exposed or a safety decision is wrong. Use stronger review, deterministic guardrails and human approval where needed.

This gives you a better rule than “always use the biggest model”: spend quality budget where an error is costly, not where the prompt looks impressive.

A practical routing policy

Keep routing logic explicit. You can start with a small decision function and replace the labels with the model IDs supported by your provider:

type Task = {
  inputTokens: number;
  needsVision: boolean;
  needsReasoning: boolean;
  highVolume: boolean;
  highRisk: boolean;
};

const chooseModelFamily = (task: Task) => {
  if (task.needsVision) return "multimodal";
  if (task.inputTokens > 80_000) return "long-context";
  if (task.highRisk || task.needsReasoning) return "reasoning";
  if (task.highVolume) return "fast";
  return "general";
};

Real routing needs more than token count. Add the capabilities your task requires: structured output, tool calling, streaming, regional availability, data controls and maximum response length. Keep those requirements in configuration so changing providers does not require rewriting your workflow.

For a multi-step agent, route each step separately. A single model setting at the top of the agent is convenient, but it hides cost and makes it harder to improve one phase without changing every other phase.

How to compare models in practice

Make a small evaluation set from real work. Twenty representative examples are more useful than a leaderboard score that does not match your prompts.

For each example, record:

  1. Task result: Did it produce a usable answer?
  2. Failure type: What exactly went wrong?
  3. Latency: How long did the user or next step wait?
  4. Cost: What did the request consume at your real input and output sizes?
  5. Operational fit: Did it support the context, tools, output format and region you need?

Compare models on the same prompt, context, tools and checks. Run the evaluation more than once when outputs are variable. A model that wins one dramatic example may lose on the ordinary cases that make up most of your bill.

Keep a simple routing table beside the evaluation:

Route Default family Success check Fallback
Ticket classification Fast Valid label from allowed set General
Code change proposal General Tests and review rubric pass Reasoning
Large policy comparison Long-context Required claims have source locations Reasoning
Production action Reasoning or general Validated tool arguments and approval Human

Update the table when the evidence changes. Do not update it because a new model name is trending.

Common mistakes

Choosing by brand or benchmark alone

Benchmarks can tell you something about a model’s capabilities, but they do not measure your prompt, context, tools, language, latency target or failure cost. Test your route.

Using reasoning for simple work

Reasoning is useful when there is a real decision to make. It is unnecessary for a fixed classification, a short rewrite or a schema-validated extraction. Extra thinking can add latency and cost without improving the answer.

Sending a huge context instead of improving retrieval

Long context helps when the relevant evidence is genuinely large or cross-document. It does not fix duplicated, stale or irrelevant input. Better retrieval often beats a larger model.

Treating a model as a permission system

The model may suggest an action. Your application must decide whether that action is allowed. Validate arguments, scope tools narrowly, log decisions and require approval for irreversible work.

Having no fallback

Every route should define what happens when the preferred model times out, rejects the input, returns invalid structure or fails its quality check. A fallback can be another model, a retry with less context, a deterministic path or a human.

A five-question checklist

Before choosing a model, answer:

  1. What is the smallest acceptable output?
  2. Is the task classification, generation, reasoning, retrieval, transformation or verification?
  3. What input modality, context size, tool support and output format are required?
  4. How costly is a wrong answer, and what can verify it?
  5. What is the cheapest model that passes a representative evaluation set?

If you cannot answer question five, you are not choosing yet—you are guessing. Start with a baseline, collect failures and let the evaluation move the route.

The takeaway

Use fast models for volume, general models for everyday work, reasoning models for difficult decisions, long-context models for genuinely large inputs, multimodal models for non-text evidence and specialist models for narrow operations.

Then split your workflow by phase. Let inexpensive models prepare work. Spend quality on decisions and high-risk reviews. Keep deterministic checks and application permissions outside the model.

The correct LLM is not the one with the biggest reputation. It is the least expensive model that reliably passes your task’s real quality bar.

Knowledge retrieval

What are you working through?

Start typing to search every guide, video, tool and snippet.