vLLM vs SGLang on a Single A10G: What 16 Models Actually Showed
May 15, 2026 · 15 min read
The Same Chart, Every Few Weeks
Every few weeks, I see the same chart. A few inference engines. A few colorful bars. Tokens per second on the Y-axis. One bar is taller. And just like that, we have a winner. vLLM wins. Or SGLang wins. Or TensorRT-LLM wins. Everyone argues in the comments for a day, bookmarks the chart, and moves on. I could never quite do that. Because every time I looked at one of those charts, I had the same four questions: > Which model? Which workload? What concurrency? Which metric? Change any one of those, and I had a feeling the answer might change with it. So instead of finding another benchmark, I decided to build the one I wanted to read. 16 models. 5 scenarios. Both engines. One 24 GB GPU. By the time I was done, I had 533 result files. And the conclusion was much less satisfying than a clean bar chart: There is no winner. vLLM wins some cells. SGLang wins others. At 7B and above, they sometimes become almost indistinguishable. And in the two places where one engine absolutely destroys the other, the reason is not some magical scheduler design. It is basically one execution flag. That turned out to be the most useful part of the entire experiment.
The Setup: 16 Models, 5 Scenarios, One 24 GB Card
Most inference benchmarks I come across are run on H100s. That makes complete sense if you have H100s. I did not. What I cared about was a much more ordinary question: > If I am putting a RAG endpoint behind a g5.2xlarge, what should I actually run? That instance gives me one NVIDIA A10G with 24 GB of VRAM, and it costs around $1.21/hr. That is a very different world from benchmarking on an 80 GB accelerator. And I could not find a comparison that answered the question for this hardware. So I ran it myself. | Component | Detail | |-----------|--------| | GPU | NVIDIA A10G, 24 GB | | Instance | AWS g5.2xlarge, 8 vCPU, 32 GB RAM | | vLLM | v0.18.0-cu130 | | SGLang | nightly-dev-cu13-20260321 | | Precision | bfloat16 | | Execution | Sequential, one engine resident at a time | The models range from 2B to 9B. Every model goes through every scenario on both engines. That gives me 160 baseline cells, and all 160 completed. Then I kept going. I added a five-iteration variance phase. A concurrency sweep all the way to 64. A decode-length sweep. Speculative decoding. Every run writes its own JSON result. By the end, the folder looked like this: 162 baseline files, 18 speculative decoding runs, 201 variance runs, 8 concurrency-64 runs, and 144 decode-length runs. A slightly unreasonable amount of benchmarking for one A10G. But at least now I had the matrix I wanted. > [!NOTE] Anything larger than 9B is out of scope here. Qwen3-30B-A3B needs roughly 60 GB at bf16 and Gemma 3 12B leaves no room for a KV cache on 24 GB. This post is about what fits on the card you can actually afford.
The Scoreboard
Before getting into what happened, here is the short version. | Metric | vLLM | SGLang | |--------|------|--------| | Lower TTFT, single request | 15 / 16 | 1 / 16 | | Higher throughput, 4B and below | 6 / 7 | 1 / 7 | | Higher throughput, 7B to 9B | tied within 3% | tied within 3% | | Structured generation | 14 / 16 | 2 / 16 | | Prefix-sharing TTFT | 6 / 16 | 10 / 16 | | Best single-request TTFT | 20 ms, Gemma 2 2B | 30 ms | | Peak throughput | 265 tok/s, Gemma 2 2B | 258 tok/s | If I stopped the blog here, the recommendation would look obvious. Use vLLM by default. Switch to SGLang if you have lots of shared prefixes. That is not a terrible conclusion. It is just not the part of the benchmark I would actually use to make a production decision. The interesting stuff showed up when I stopped looking at averages.
Scenario 1: Single Request Latency (vLLM Wins 15 of 16)
I started with the simplest possible request. One model. One request. No queue. No concurrency battle. Just: > How quickly can you give me the first token? vLLM won on 15 of the 16 models. Some examples: - 20 ms vs 30 ms on Gemma 2 2B - 24 ms vs 57 ms on SmolLM3 3B - 41 ms vs 66 ms on Qwen 2.5 7B - 43 ms vs 69 ms on Llama 3.1 8B Across the models, the advantage ranges from roughly 18% to 58%. And I expected the explanation to show up in decode speed. It did not. Once the model starts generating tokens, both engines are basically tied. On almost every model, decode speed lands within about 1 tok/s. So the first-token gap is happening before steady-state generation. That points toward scheduling overhead, kernel launches, and the work the engine has to do to get the request moving. vLLM's CUDA graph capture helps a lot here. Which is important, because the only model where vLLM loses this test is also the model where CUDA graphs disappear. And that one exception ended up explaining several of the strangest results in the benchmark.
The Eager Mode Tax
Gemma 3 4B was the first result that made me stop and check the logs. SGLang hit first token in 78 ms. vLLM took 87 ms. Not a huge difference by itself. But it was the only model where SGLang beat vLLM in the single-request test. At first glance, I could have written: > SGLang schedules Gemma 3 better. That would have been wrong. Gemma 3 mixes sliding-window attention with global attention layers. vLLM cannot capture CUDA graphs across that pattern, so the model has to run with --enforce-eager. That flag changes everything. Without graph capture, vLLM loses the optimization path that helped it on the other 15 models. And once I noticed it here, I started seeing the same penalty everywhere else. > [!WARNING] One flag, set for one architecture, inverts four separate results in this benchmark. If you are comparing engines and one of them is running eager, you are not comparing engines. You are measuring a compatibility gap and calling it a benchmark. I started calling this the Eager Mode Tax. You will see it several more times. Every time Gemma 3 produces a ridiculous-looking gap, it is basically the same bill arriving again.
Scenario 2: Sustained Throughput (vLLM Wins Small, Everyone Ties at 7B)
Next I started adding concurrency. 1 request. Then more. All the way to 32. This is where I expected the engines to start separating clearly. Instead, I got a pretty clean size-dependent pattern: - Gemma 2 2B: 265 vs 258 tok/s, vLLM by 3% - SmolLM3 3B: 230 vs 205, vLLM by 12% - Phi-4 mini 4B: 189 vs 176, vLLM by 7% - Gemma 4 E4B: 83.8 vs 81.3, vLLM by 3% - Qwen 2.5 7B: 105 vs 106, inside 1% - Mistral 7B: 107 vs 107, tie - Llama 3.1 8B: 102 vs 102, tie Below roughly 4B, vLLM has room to pull ahead. The advantage ranges from about 3% to 12%. Then I hit the 7B models. And the race basically stopped. Qwen 2.5 7B: tie. Mistral 7B: tie. Llama 3.1 8B: tie. Once you get into the 7B-to-9B range, the difference stays under roughly 3%. The GPU has become the bottleneck. That is the easiest way I found to think about it. With a small model, the engine still has enough breathing room for scheduling, batching, and memory-management differences to matter. vLLM's continuous batching and block allocator can squeeze out a few more tokens. Once the model is big enough to saturate the A10G, both engines are standing in the same traffic jam. There is only so much GPU left to optimize. And then Gemma 3 walks into the room. SGLang: 149 tok/s. vLLM: 84 tok/s. A 77% difference. That is not a typo. vLLM spends 2,137 seconds generating the same 179K tokens that SGLang finishes in 1,200 seconds. At this point I already knew what I was looking at. The Eager Mode Tax. > [!TIP] Gemma 4 does not inherit the problem. It uses the same hybrid attention and adds QK-norm, but vLLM ships a native loader that keeps CUDA graphs intact, and the E4B numbers land right back in normal territory. Same family, opposite recommendation. There were two other results I did not want to smooth over just because they made the story messier. SmolLM3 3B is about 12% slower on SGLang. SGLang often gets described as the engine that shines as workloads scale, but on this architecture, vLLM's kernel selection simply worked better in the versions I tested. And Gemma 2 9B had a different problem. Its medians looked fine. Its tail latency absolutely did not. I will come back to that one.
Scenario 3: Long Context at 8K (SGLang Takes the Lead Back)
Then I changed the shape of the request. Instead of small prompts, I pushed twenty 8,192-token prompts through at roughly concurrency 4. And suddenly the scoreboard started moving toward SGLang. SGLang wins TTFT on 8 of the 16 models. More importantly, those wins cluster around the larger models: - Qwen 2.5 7B: 59 ms vs 91 ms - Llama 3.1 8B: 63 ms vs 90 ms - Qwen3 8B: 70 ms vs 102 ms This is where RadixAttention starts earning its keep. With an 8K prompt, prefill is no longer a small setup cost. It is a major part of the request. SGLang's radix-tree KV cache can reuse overlapping prefix segments between concurrent requests automatically. vLLM also has prefix caching, but it works at block granularity, so it gets less reuse once the prompts start diverging in the middle. The practical effect is simple: The longer the shared context becomes, the more expensive it is to recompute. And SGLang is better at avoiding some of that work. Decode throughput still follows the model-size pattern. vLLM tends to stay ahead around 3B and below. SGLang starts looking stronger around 7B to 9B. And yes, Gemma 3 4B shows up with 180 vs 98 tok/s. By this point, you already know what that is.
Scenario 4: Prefix Sharing at 60% Overlap (SGLang Wins 10 of 16)
The next test was almost designed for SGLang. I gave requests 60% shared prefixes and measured what happened. This is basically the request shape you get in a lot of RAG and multi-turn chat systems. Same system prompt. Similar retrieved context. Different user question near the end. SGLang wins TTFT on 10 of the 16 models. And the biggest wins again show up around 7B to 9B: - Llama 3.1 8B: 65 ms vs 93 ms - Qwen3 8B: 59 ms vs 95 ms - Granite 3.3 8B: 66 ms vs 110 ms - DeepSeek-R1-Distill Llama 8B: 55 ms vs 94 ms That is roughly 20 to 45 ms removed from every shared-prefix request. Twenty milliseconds does not sound dramatic when it is written in a benchmark table. But if your application serves the same system prompt and overlapping retrieved context thousands of times a day, you are paying that cost on almost every interaction. This test also exposed a useful split between the two engines. SGLang usually gets the first token out faster. vLLM can still finish the response slightly faster once the prefix cost has been amortized. So which one matters? For a batch pipeline, maybe total throughput. For an interactive RAG application, the user is sitting there waiting for the response to begin. They feel TTFT. > [!NOTE] The only small models where vLLM still wins prefix TTFT are Gemma 4 E2B and E4B. The trie advantage does not really take over until around 7B, so if you are serving 3B models this whole scenario matters less than you would expect.
Scenario 5: Structured JSON (vLLM Wins 14 of 16)
This result genuinely surprised me. SGLang has a compressed finite-state machine for constrained decoding, and structured generation is one of the workloads it is explicitly built to handle well. So I expected SGLang to do very well here. On the A10G, vLLM won 14 out of 16 models. A few of the numbers: - Gemma 2 2B: 1,225 vs 957 tok/s, vLLM by 28% - SmolLM3 3B: 930 vs 774, by 20% - Qwen 2.5 7B: 456 vs 384, by 19% - Phi-4 mini 4B: 736 vs 669, by 10% - Llama 3.1 8B: 426 vs 423, tie The biggest gaps show up on the smaller models. That actually makes sense after staring at it for a while. When normal decoding is already very fast, constraint checking becomes a larger fraction of the work. The GPU is not hiding that overhead anymore. SGLang wins only two models. Granite 3.3 8B, by about 3%. And Gemma 3 4B, by 81%. You can probably finish that sentence for me now. Eager Mode Tax. Again.
Does Output Length Change Anything? (No)
At this point I had another suspicion. Maybe the benchmark was flattering one engine because the outputs were too short. That happens. So I changed maxoutputtokens. 64. 256. 1024. 4096. Concurrency stayed at 8, and I ran three iterations for every cell. I expected the lines to cross all over the place. They mostly did not. The ranking stays remarkably stable. Gemma 3 4B remains around 1.8x faster on SGLang whether I ask for 64 tokens or 4096. Gemma 2 2B and Phi-4 mini remain within a few percent of each other. vLLM has a small advantage at short outputs, and the lead crosses over somewhere around 1024. The more interesting result was not which engine won. It was what happened when I simply asked too much from the card. Llama 3.1 8B. 4096 output tokens. Concurrency 8. p99 end-to-end latency climbs to roughly five minutes. SGLang tail TTFT reaches 99 seconds. vLLM reaches 36 seconds. At that point, I am not tuning an inference engine anymore. I am asking a 24 GB A10G to serve a workload it cannot serve gracefully. No clever scheduler flag fixes physics.
The Median Mirage: Where the Engines Actually Differ
This is where the benchmark stopped being a scorecard for me. It became something I could actually use. I pushed the 7B-to-9B models to concurrency 64. Each engine got 900 requests. I expected the 9B models to OOM. Or at least start throwing errors once the queue became deep enough. Instead: zero errors across all 8 cells. That was the first surprise. A single 24 GB A10G can sustain 7B-to-9B-class models end to end at 128-token prompts and 256-token outputs, even at this concurrency level. That is a useful number to have before someone tells you the solution automatically requires an A100. Then I looked at the median. SGLang looked great. Median TTFT stayed around 69 to 74 ms. vLLM sat around 93 to 131 ms. The radix tree was clearly doing its job under load. If I stopped there, SGLang would have won this section. Then I looked at p99. Gemma 2 9B changed the entire story. vLLM p99 TTFT: 1,859 ms. SGLang: 29,427 ms. I checked that number more than once. It is roughly a 16x gap. The chart literally needs a logarithmic axis just to make both bars visible. Gemma 2 alternates local and global attention, and under a deep queue that pattern appears to stall SGLang's continuous-batch scheduler. The same kind of tail problem shows up in inter-token latency. SGLang versus vLLM at TPOT p99: - Qwen3 8B: 256.2 ms vs 54.3 ms, which is 4.7x worse - Llama 3.1 8B: roughly 2.4x worse - SmolLM3 3B: around 2.0x worse Below 4B, the gap mostly disappears and stays under 5%. This was probably the moment that changed how I looked at the entire benchmark. The medians had told me the engines were basically interchangeable at 7B and above. Technically, the medians were correct. Operationally, they were almost useless. > [!IMPORTANT] The medians said these engines were interchangeable at 7B and up. The medians were right and completely useless. Interchangeable at p50 and 16x apart at p99 is not a small print detail, it is the difference between a service that works and a service that is on fire twice an hour.
Throughput Is Not Goodput
This is also where tokens per second stopped being the metric I cared about most. Throughput asks: > How many tokens can this engine generate per second? Goodput asks something much closer to the question I actually care about in production: > How many requests can this engine serve without breaking my latency promise? That means a request has to satisfy the TTFT budget and the TPOT budget. Not one. Both. And only one of those metrics has ever caused me to look at a dashboard at 3 AM. For this test, I set the target to TTFT under 100 ms and TPOT under 35 ms. Then I recalculated the ranking. Three models completely rearranged the story. First: Gemma 3 4B on vLLM. The model generates plenty of tokens. Its normal throughput number does not look catastrophic. But it satisfies the latency SLO zero percent of the time. That means 0 rps of goodput. SGLang reaches 0.31 rps with a 42% pass rate. The Eager Mode Tax pushes vLLM's TPOT beyond 35 ms on essentially every request. Second: Llama 3.1 8B. SGLang reaches 0.21 rps at a 35.5% pass rate. vLLM reaches only 0.07 rps at 14.5%. About a 3x goodput advantage for SGLang. Remember the single-request benchmark? vLLM had better TTFT there. And it still loses this workload. Because once requests arrive concurrently, SGLang keeps more of them inside the complete latency window. Third: Gemma 2 9B. This one reaches 0 rps on both engines. Its native TPOT is already around 44 ms. So a 35 ms TPOT budget is simply impossible for this model on this card. That is not an engine problem. That is an SLO problem. The lesson for me was simple: Do not choose the model and then invent an SLO it cannot hit. Measure the model first. Then write the SLO around reality. None of these conclusions are visible in a normal tokens-per-second chart.
Speculative Decoding: The Draft Model That Ate the KV Cache
This was the most expensive lesson in the project. I expected speculative decoding to be one of the easy wins. Use a cheaper draft mechanism to predict tokens. Verify them with the main model. Generate faster. That is the pitch. On a 24 GB A10G with 8B-class models, the accounting goes the other way. Compared with baseline throughput: - Ngram costs vLLM about 6% on Llama 3.1 8B and 7% on Qwen3 8B - On SGLang, the same idea costs about 28% and 36% - Eagle3 costs vLLM about 20% - SGLang Eagle3 never completed At first this looks like speculative decoding is broken. It is not. It is running out of room. The main model plus the Eagle3 draft model need roughly 16.8 GiB just for weights. The KV cache has not even entered the conversation yet. To make the setup fit, I had to use --gpu-memory-utilization 0.95, --enforce-eager, and --max-model-len 2048. Now I have squeezed the system into such a small memory envelope that there is not enough batch width left to amortize the draft proposal overhead. The optimization has eaten the resource it needed in order to become an optimization. On an A100 40 GB or H100, I would expect this tradeoff to look very different. On this A10G, baseline wins. > [!TIP] There is one exception worth keeping. On Gemma 4 E4B, Ngram cuts TTFT by roughly 45% on both engines, from 84 ms to 47 ms on vLLM and 87 ms to 48 ms on SGLang, at a cost of 8% and 29% peak throughput respectively. If first token latency is what you are selling and you have throughput to spend, that trade is on the table.
What This Run Does Not Tell You
After 533 result files, it becomes very tempting to treat the spreadsheet like truth. It is not. There are four caveats I would want to know before using these numbers for anything important. First, Gemma 4 E2B. It logs valid TTFT in singlerequestlatency and throughputramp, but reports zero output tokens. So I cannot publish its per-request tok/s numbers. The other three scenarios are clean, and a rerun is queued. Second, the medians are much more stable than the tails. Across five iterations per cell, the coefficient of variation for p50 TTFT and tokens/sec stays under 1% almost everywhere. p95 is noisier. Several cells cross 5% CV, and Phi-4 mini under SGLang's throughput ramp reaches 61%. So I trust the median trends much more strongly than any individual p95 or p99 value. Treat the tail numbers as directional. Third, SGLang Eagle3 on Llama 3.1 8B never completed because the nightly image I pinned has since been retired. So when I say SGLang Eagle3 will not fit comfortably on this A10G, that is based on memory arithmetic and the vLLM Eagle3 footprint. It is not a completed SGLang run. And finally, this is one GPU. One precision. One set of engine versions. Both projects move quickly. A benchmark like this is a photograph, not a law of physics.
Quick Reference Guide
If you just want the answer I would use when choosing an engine for this exact hardware: | Workload | Pick | Why | |----------|------|-----| | Latency-sensitive single request | vLLM | Wins TTFT on 15/16 | | Structured or JSON output | vLLM | Wins 14/16 | | Prefix-heavy RAG and multi-turn chat | SGLang | Saves 20 to 45 ms at 7B and up | | Long context, 8K and above, at 7B to 9B | SGLang | Wins TTFT on 8/16, concentrated at that size | | High-throughput batch at 7B to 9B | Either | Tied within 3% | | Tight p99 at high concurrency | vLLM | 16x better tail TTFT on Gemma 2 9B | | Anything Gemma 3 | SGLang | The Eager Mode Tax is worth 77% throughput | | Anything Gemma 4 | vLLM | Native loader, lower TTFT, 3 to 6% throughput edge | | Speculative decoding on an A10G | Neither | Baseline beats Ngram and Eagle3 | The more runs I collected, the less useful the question "which engine is faster?" became. The better question was: faster for what? It is a matrix, not a ranking. The cell you care about is a combination of the engine, model, workload, concurrency, and metric. Move to another cell and the winner can change completely.
Closing Thoughts: Run It On Your Own Model
If I had to compress all 533 result files into one sentence, it would be this: H100 benchmarks do not predict A10G behavior. SGLang's throughput advantage shrinks into a tie at 7B to 9B on this card. vLLM's structured-generation advantage is wider here than the H100 numbers would have made me expect. And Gemma 3's CUDA graph compatibility issue is large enough to invert four separate benchmark results. But the biggest lesson was not actually about either engine. It was about what I was measuring. The two results I would be most likely to make a production decision around were the p99 blowup and the goodput inversion. Neither one appears in the metric almost everyone publishes. Tokens per second is useful. It just does not tell me whether users are waiting 30 seconds for one unlucky request. It does not tell me whether 90% of my traffic violates the latency target. And it definitely does not tell me whether a benchmark advantage came from an architectural improvement or because the other engine quietly fell back to eager mode. That is why I am publishing the harness too. The prompt packs, Terraform module, benchmark scripts, and raw JSON from all 533 runs are open source. Reproducing the setup costs roughly a dollar an hour. So do not take my winner. Run your model. On your hardware. With your prompts. At your concurrency. Against your SLO. Then trust those numbers. Because the marketing chart is always downstream of somebody else's hardware budget. > [!TIP] Harness and raw data: github.com/varad-more/inference-engine-benchmark-system
Method and References
Method: sequential execution, one engine resident at a time on a single A10G. 160/160 baseline cells across 16 models and 5 scenarios, plus variance at 5 iterations per cell with 95% confidence intervals, a concurrency-64 ramp at 0% error rate, a 4-length decode sweep, and speculative decoding. Reproducible via scripts/runallbenchmarks.sh and scripts/runnewbenchmarks.sh. Figures regenerate via python -m analysis.generatefigure. > [!NOTE] References > > 1. vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention > 2. SGLang: Efficient Execution of Structured Language Model Programs (RadixAttention) > 3. EAGLE-3: Scaling up Inference Acceleration of Large Language Models > 4. Amazon EC2 G5 Instances