The Frontier Model Gets the Attention, the Small Models Get the Volume
You know the pattern by now. A new frontier model drops, the leaderboard screenshots go around, someone posts a thread about how this changes everything, and two weeks later it happens again. Meanwhile, if you open up an actual trace from an agent running in production, the frontier model barely shows up. Most of the lines in that trace belong to small specialized models for AI agents: dense encoders, cross-encoder rerankers, extraction models, OCR.
The claim is simple enough to state and easy to verify on your own traffic. Generation fires once or twice per turn. Embedding, reranking, and extraction fire continuously, on every chunk, every retrieval, every document that lands in the pipeline. That asymmetry is the whole story, and it’s the part that shows up on the invoice.
This post is about cost and latency behaviour, not benchmark rankings. I’m not going to tell you which model tops MTEB this month, partly because it will have changed by the time you read this. What I want to do instead is describe how the work distributes across model classes and what that distribution does to your bill.
I come at this from hosting and cloud infrastructure rather than from ML research. Seven years of looking at cPanel servers, AWS accounts, and OCI tenancies has given me a fairly narrow but useful instinct: the thing that runs on every request is the thing that determines your monthly spend. Not the expensive thing. The frequent thing. A slow query that runs once a day is an annoyance. A slow query in the page render path is a capacity problem. Inference follows the same logic, and a lot of teams are still budgeting as if the frontier model were the main cost centre.
How a RAG Pipeline Splits Work Between Models
Take a single agent turn over a document-heavy workload and break it into tasks.
The user asks something. That query gets embedded, which is a dense encoder job (stella, bge-m3, that family). Retrieval pulls back candidates from the vector store, more than you want to put in context, so a cross-encoder reranker scores them and you keep the top few. If new documents entered the system, each one got chunked and embedded on the way in, and probably passed through an extraction model like GLiNER to pull out entities, dates, and amounts into structured fields. If any of those documents were scans or page images, an OCR model read them first. Only after all that does the frontier or mid-size LLM get called to actually generate an answer.
Count the calls. Generation: once, maybe twice if there’s a tool loop. Embedding: once per query plus once per new chunk, which on an ingest-heavy day could be thousands. Reranking: once per retrieval, and every turn does a retrieval. Extraction and OCR scale with documents, not with turns, which for most document workloads means they scale with whatever your users are uploading.
The expensive model sits on the least frequent path. That’s not an accident anyone stumbled into, it’s a deliberate architecture. You put the general-purpose reasoning model where you need general-purpose reasoning, which is once, at the end, and you put narrow models on the repetitive work because narrow models are cheap and the repetitive work is where volume lives.
Three terms worth pinning down before the next section:
- Self-hosted embedding models means running the encoder on your own GPU rather than calling an embeddings API. Same output shape, different cost structure.
- Reranker model for RAG means a cross-encoder that scores query-document pairs directly. Slower per pair than a vector similarity lookup, much more precise, and it only runs over your candidate set.
- Open weight models vs frontier LLMs is not a quality argument in the abstract. It’s a question of whether the specific job needs general reasoning or a narrow transformation. Most jobs in a RAG pipeline are the latter.
Are Small Specialized Models for AI Agents Actually Good Enough?
For embedding, reranking, extraction, and OCR, yes. For the hardest novel reasoning, no. That’s the honest short answer, and most of the disagreement I see online comes from people arguing about one while thinking about the other.
The numbers worth citing come from a Superlinked comparison written up in Small Specialized Models Are Eating the AI Stack (While Everyone Watches Frontier LLMs) (opens in new tab). Across eight standard MTEB retrieval tasks, self-hosted open models averaged 0.600 nDCG against 0.615 for hosted frontier embedding models from Voyage, OpenAI, and Cohere. The open models won outright on two of the eight. That is close enough that on most corpora the difference will be inside the noise of your own evaluation set.
The cost comparison on the same corpus is not close. Embedding a billion tokens ran roughly $10 to $17 self-hosted against roughly $120 to $130 through the hosted APIs. Call it an order of magnitude for a couple of points of retrieval quality.
On the generation side the picture is fuzzier and the caveats matter more. The same write-up cites an Artificial Analysis Intelligence Index snapshot putting Qwen3.6-27B around 46 against GPT-5.1’s 48. Treat that as a snapshot and nothing more. Index scores move with reasoning mode, with harness configuration, and with every model release, so a number quoted in a blog post is a starting point for your own check rather than a fact you can plan against. The more grounded data point is operational: Intercom reported cutting around $250,000 a month by swapping one hosted GPT call in a Fin AI pipeline task for a fine-tuned 14B Qwen model. One task. Not the whole product.
That’s the shape of the opportunity. Not “replace your LLM,” but “find the specific high-volume step where a narrow model does the same job and stop paying frontier rates for it.”
Choosing Models for a RAG Pipeline
Decide by job, not by model. Here’s how the split tends to fall.
Embeddings. Self-hosting wins on cost, speed, and control. This is the first workload most teams should pull in-house, because the volume is high, the output is a fixed-size vector, and swapping the implementation behind an interface is a contained change. Hosted APIs still make sense when your volume is low and spiky and you’d rather not own a GPU at all.
Reranking. Rarely worth an API call. The reranker fires on every retrieval, adds network round-trip latency to a step that sits directly in your response path, and cross-encoders in the 100M-parameter range are small enough to sit alongside other models on one GPU. Precision per dollar here is very good.
Extraction. Small models handle schema-shaped output well, and extraction is almost always schema-shaped. Where you have fuzzy, open-ended reads with no fixed target fields, the hosted model earns its keep.
OCR and document vision. Predictable batch cost favours local. Rare formats and one-off reads favour hosted, because tuning a local pipeline for a format you’ll see twice is not a good use of anyone’s week.
Hardest novel reasoning. Frontier model, genuinely. Don’t fight this one.
What self-hosting adds is operational work, and this is where I’ve seen projects quietly fall over. You need GPU capacity planning, request batching to get useful throughput, warm pools so cold starts don’t land on user requests, pinned model versions so retrieval quality doesn’t shift underneath you, and monitoring that distinguishes “model is slow” from “queue is deep.”
There’s also a trap the source describes well. Most inference tooling was built for one large model spread across many GPUs. Small-model serving is the inverse: several models on one GPU with fast switching between them. If you give each model its own deployment and its own pool, you provision five GPUs for peak, run them at low utilisation, and pay for four to sit idle. This is the same mistake as running five VPS instances because five services each got their own box, and any site owner who has consolidated a sprawl of underused instances onto one properly sized server already understands the arithmetic.
Watch p99 while you do this. The models that fire on every request set your tail latency, and a reranker that’s fine at p50 and terrible at p99 will show up as a slow product long before it shows up as a slow model.
What to Expect Next in AI Agent Inference Cost
Benchmark snapshots go stale fast, which is the main reason I’m reluctant to build an argument on any single score. Re-check them yourself against the live comparison, and more importantly, build a small evaluation set from your own documents and queries. Public retrieval benchmarks tell you about public corpora. Your corpus has its own vocabulary, its own chunk shapes, and its own failure modes.
The part that keeps changing the maths is that open weights can be adapted cheaply. Fine-tuning a small encoder or extractor on domain data is within reach of one engineer and a modest GPU budget, and for narrow repetitive tasks a tuned small model can beat a general large one on the specific thing you care about. That gap tends to widen as your task gets narrower.
My expectation, and I’ll flag this as expectation rather than fact: hybrid routing becomes the default architecture rather than an optimisation people bolt on later. The frontier model gets reserved for the reasoning step, and embedding, reranking, extraction, and OCR run locally on shared GPU capacity, behind an interface that lets you swap either direction. I think the tooling for multi-model serving on single GPUs improves noticeably over the next year or two, because the demand is obvious and the current tooling clearly wasn’t designed for it.
Some constraints won’t dissolve. GPU availability stays uneven by region and instance type. Inference serving is a genuine specialism, and hiring for it competes with everyone else hiring for it. Running your own stack means owning upgrades, security patches, driver versions, and the on-call rotation that comes with all of it. Self-hosting trades a variable API bill for a fixed infrastructure bill plus engineering time, and that trade is only good above a certain volume.
Audit One Pipeline Before You Change Anything
Pick one RAG pipeline you already run in production. Instrument it so that every model call is logged with the model name, token or item count, and latency. Run it for a week of normal traffic without changing anything else.
Then count calls per model per turn and multiply out to a month. You’ll get a cost per model class from your own traffic rather than from someone’s benchmark, and in my experience the ranking surprises people. Take the top line by call volume, usually the embedding or reranking step, and price the self-hosted equivalent against it, including the GPU hours you’d actually need at your peak.
Move that one step first. Put it behind an interface, keep the hosted path live as a fallback, and validate retrieval quality against your own evaluation set before you cut over. If quality holds, you have a number to justify the next step. If it doesn’t, you flip a config flag and you’ve lost a week instead of a quarter.