← Back to blog
LLM Serving·October 8, 2026·9 min read

vLLM, SGLang, or TensorRT-LLM: What Actually Differs When You Self-Host an LLM

vLLM, SGLang, and TensorRT-LLM all promise fast self-hosted inference, but they make different bets on KV cache management, structured output, and hardware lock-in — here's what actually changes depending on your traffic pattern.

Ask three teams why they picked their inference engine and you'll get three different, confidently-stated reasons that all boil down to "it was fast in our benchmark." That's not wrong, exactly — it's just incomplete. vLLM, SGLang, and TensorRT-LLM all do the same two jobs (batch requests efficiently, manage GPU memory for the KV cache), but they made different architectural bets about how, and those bets matter more than whatever tokens-per-second number shows up in a vendor's blog post. The engine that wins your benchmark on a synthetic workload of 500-token completions can lose badly on your actual traffic, which is probably bursty, prefix-heavy, and full of tool-calling agents re-sending the same system prompt eight times a minute.

This is about what each engine actually optimizes for, where that optimization breaks down, and how to read past the marketing numbers to the question that matters: does this engine's scheduling and memory model match the shape of your workload.

The Problem All Three Are Solving

Naive LLM serving — one request, one forward pass, wait for it to finish, start the next — wastes most of a GPU's compute. A single request rarely saturates a modern accelerator's matrix units, especially during the token-by-token decode phase, which is memory-bandwidth bound, not compute bound. The fix is batching: run many requests' forward passes together so the GPU's compute is shared across them.

The complication is that LLM requests don't arrive, or finish, at the same time. A naive "static batch" — group N requests, run them together, wait for all N to finish before starting the next batch — stalls on whichever request in the batch generates the most tokens. If one request wants 1,000 tokens and the other 31 want 50, you're holding 31 finished slots idle for the long tail.

All three engines solve this with some form of continuous batching (also called in-flight batching): as soon as a request finishes, a new one is slotted into the batch immediately, without waiting for the whole batch to drain. This alone was the single biggest serving-throughput improvement of the last few years — bigger than any single model optimization — because it keeps the GPU doing useful work instead of waiting on stragglers.

Static batching vs. continuous batching

Static batch (GPU idles on stragglers) next batch waits for all three

Continuous batching (slots refill as requests finish) A finished request's slot is reassigned to a new one on the very next scheduler step — no request waits on another request's tail. This is table stakes for all three engines; the real differences are in how they manage the KV cache memory behind each slot.

Once you accept continuous batching as the baseline, the interesting differences are in KV cache memory management — how each engine stores and reuses the key/value tensors that attention needs for every token already generated — and that's where vLLM, SGLang, and TensorRT-LLM genuinely diverge.

vLLM: PagedAttention and Breadth

vLLM, out of UC Berkeley's Sky Computing Lab, is the reason continuous batching with efficient memory management became the default expectation rather than a nice-to-have. Its core contribution, PagedAttention, treats the KV cache the way an operating system treats virtual memory: instead of pre-allocating one large contiguous buffer per request (which forces you to reserve for the worst-case sequence length and wastes most of it on shorter requests), it splits the cache into fixed-size blocks and allocates them on demand, non-contiguously, the way an OS pages physical memory. That alone recovered a large share of GPU memory previously lost to fragmentation and over-reservation, which translates directly into being able to hold more concurrent requests in memory at once.

On top of that, vLLM added automatic prefix caching: when two requests share an identical token prefix (the same system prompt, the same few-shot examples, the same tool schema block), the KV blocks for that shared prefix are computed once and reused, keyed by content hash rather than by request ID. It also supports multi-LoRA serving (swapping adapter weights per request without reloading the base model), speculative decoding, and an OpenAI-compatible server that runs on Nvidia, AMD, Intel Gaudi, and Google TPU hardware — the broadest hardware story of the three. The tradeoff is that breadth comes with somewhat more general-purpose, less hardware-specific tuning than an engine built for exactly one vendor's silicon.

SGLang: Built Around Shared-Prefix and Structured Workloads

SGLang, from the same research lineage as the LMSYS Chatbot Arena project, takes the prefix-sharing idea further with RadixAttention: KV cache blocks are organized in a radix tree shared across all requests the server has handled recently, not just within a single conversation. If a hundred concurrent agent sessions all start with the same 2,000-token tool-calling system prompt, SGLang's tree structure means that prefix exists in GPU memory once, and every new request that matches it skips recomputing those tokens entirely — eviction is managed with a cache-aware LRU policy across the whole tree, not per-request.

That design pays off specifically on workloads with heavy prefix reuse: RAG pipelines re-sending the same retrieved context, agent loops re-sending the same tool schema and growing history, few-shot prompting at scale. SGLang also bakes structured output deep into its execution model — constrained JSON generation, regex-constrained decoding, and grammar-based decoding run with less overhead than bolting a logit-mask layer onto a generic server, because the constraint-aware decoding integrates with its own scheduler rather than wrapping it. It's newer and has had a smaller hardware support matrix than vLLM historically (Nvidia and AMD are the primary targets), though that gap has been closing. Several frontier labs have talked publicly about adapting SGLang-style radix caching internally for exactly this reason: agentic and RAG traffic is disproportionately prefix-shared, and caching that well is worth more than raw peak throughput on isolated, unrelated requests.

TensorRT-LLM: Maximum Throughput, One Vendor

TensorRT-LLM is Nvidia's own inference library, built on top of the general-purpose TensorRT compiler. Where vLLM and SGLang run your Hugging Face checkpoint close to as-is, TensorRT-LLM compiles the model into a hardware-specific optimized engine ahead of time — fusing kernels, selecting tuned kernel implementations for your exact GPU, and applying quantization (down to FP8 and, on Blackwell-class hardware, FP4) at the compilation step rather than at load time. That compile step is the main practical cost: swapping to a new checkpoint or a new quantization scheme means rebuilding the engine, which is slower iteration than vLLM's or SGLang's "point at a Hugging Face repo and go" workflow.

What you get for that cost is, in most published and independently reproduced benchmarks, the best raw throughput and lowest latency achievable on Nvidia hardware when the engine is well-tuned for your exact model and batch shape — which makes sense, since it's compiled specifically for that GPU generation rather than running a general-purpose kernel path. It supports in-flight batching and a paged KV cache conceptually similar to vLLM's, plus deep integration with Nvidia Triton Inference Server for production deployment, multi-GPU/multi-node serving, and the full quantization ladder Nvidia ships for each new architecture. The tradeoff is explicit vendor lock-in and a materially steeper ops learning curve — you're not just running a Python server, you're managing a model-compilation pipeline as part of your deployment process.

The Comparison Table, and What It Doesn't Show

vLLMSGLangTensorRT-LLM
LicenseApache 2.0Apache 2.0Apache 2.0 (Nvidia-maintained)
HardwareNvidia, AMD, Intel Gaudi, TPUPrimarily Nvidia, AMDNvidia only
KV cache strategyPagedAttention + prefix cachingRadixAttention (tree-shared cache)Paged KV cache, compiled per-engine
Deployment modelLoad checkpoint, runLoad checkpoint, runCompile model into engine, then run
Structured outputSupported via integrationNative, scheduler-levelSupported via Triton + grammar backends
Best raw throughput (tuned, Nvidia)StrongStrong, often best on shared-prefix trafficTypically highest for isolated, well-tuned workloads
Iteration speed on new checkpointsFastFastSlow (recompile step)
Multi-LoRA servingYesYesLimited

The table tells you what each engine supports. It doesn't tell you that the "best throughput" column is almost always measured on a synthetic benchmark of independent, non-overlapping requests — which is exactly the traffic pattern least like a production agent or RAG system. If your real traffic is a swarm of tool-calling agents re-sending a 3,000-token system prompt on every turn, SGLang's or vLLM's prefix/radix caching can out-perform a raw-throughput-optimized TensorRT-LLM deployment that doesn't exploit that overlap, even though the latter wins the vendor's own microbenchmark. Conversely, if you're running offline batch scoring of millions of independent, non-overlapping documents on a fixed Nvidia fleet with a stable model version, TensorRT-LLM's compiled-engine overhead pays for itself and the iteration-speed cost barely matters because you compile once and run for weeks.

The other thing the table hides is operational gravity. TensorRT-LLM commits you to Nvidia for the life of the deployment; that's fine if you've already standardized on Nvidia and want every last percentage point of throughput, and a real cost if you want the option to move workloads to AMD MI300-class hardware or a cloud TPU later. vLLM's broader hardware support is a hedge against exactly that lock-in, at the cost of not being the single fastest option on any one vendor's silicon. SGLang sits in between: Nvidia-first but not Nvidia-only, and its engineering effort has gone disproportionately into the caching and structured-decoding problems that matter most for agentic and RAG workloads specifically, rather than general-purpose peak throughput.

Spinning One Up

All three ship (or integrate with) an OpenAI-compatible HTTP API, which matters because it means the serving layer is swappable later without rewriting your client code. A minimal vLLM deployment looks like this:

pip install vllm

vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-prefix-caching \
  --max-model-len 8192 \
  --port 8000

SGLang's equivalent is structurally similar — point it at a checkpoint, get back an OpenAI-compatible endpoint — with RadixAttention's prefix sharing on by default rather than an opt-in flag. TensorRT-LLM instead requires a trtllm-build step against your checkpoint and target GPU before you ever serve a request, which is the concrete form the "compile vs. interpret" tradeoff takes in practice.

If you're serving a model behind a tool-calling agent or an MCP server — where the client sends requests through an OpenAI-compatible or Anthropic-compatible API regardless of what's actually running behind it — this choice is invisible to the agent itself. It only shows up as latency, cost per token, and how gracefully the system handles the bursty, prefix-heavy traffic that agent loops actually produce.

The Takeaway

Don't pick an inference engine off a leaderboard number measured on a traffic pattern you don't have. Profile your actual request shape first — how much prefix overlap exists across concurrent requests, how bursty arrivals are, how often the underlying model or quantization changes — and then match it: vLLM for the broadest hardware support and fastest iteration, SGLang when prefix-heavy agent or RAG traffic dominates, TensorRT-LLM when you've already committed to Nvidia and want to trade build-step friction for the last mile of throughput on a stable, high-volume workload.

#llm-inference#vllm#sglang#tensorrt-llm#kv-cache#self-hosted-ai

Related reading

LLM Serving
KV Cache Reuse and the Hidden Latency Budget of Agent Loops
LLM Serving
How Constrained Decoding Actually Forces an LLM to Emit Valid JSON
AI Hardware
The Memory Wall: Why On-Device LLM Inference Is Bottlenecked by Bandwidth, Not Compute