← Engineering notes

Note 02 · Inference · Deployment

Llama 3.1 70B on a single L40S — what we measured.

A single NVIDIA L40S (48 GB), Llama 3.1 70B Instruct, vLLM as the serving layer. We benchmarked quantization variants, batch sizes and prompt lengths against each other. This note documents real latency and throughput numbers from a production customer-support agent.

~22 min May 2026 From a live deployment
TL;DR

Llama 3.1 70B in Q4_K_M passt mit 4k Kontext komfortabel auf eine L40S. Time-to-First-Token: P50 ~480 ms, P95 ~780 ms. Die Decode-Rate liegt bei ~15–20 Token/s pro Stream — eine 200-Token-Antwort ist damit nach ~11–13 s komplett, dank Token-Streaming sieht der Nutzer die ersten Wörter aber praktisch sofort. Bei batch_size=4 und 8k Kontext steigt die P95-TTFT Richtung 1.6 s und die Decode-Rate pro Stream sinkt spürbar. Q8 liefert ~6% bessere Faithfulness gegen 35% mehr VRAM und höhere Latenz — für die meisten Mittelstands-Use-Cases nicht das Geld wert. FP16 ist auf einer einzelnen L40S nicht praktikabel: 4-bit-Quantisierung ist hier kein Kompromiss, sondern die korrekte Wahl.

1 · Hardware und Modell

The configuration we're talking about:

  • GPU — NVIDIA L40S, 48 GB GDDR6 ECC, 350 W TDP. One card, no NVLink interconnect.
  • Host — AMD EPYC 9354P, 128 GB RAM, local NVMe. No network storage on the inference path.
  • Modell — meta-llama/Llama-3.1-70B-Instruct, quantized to Q4_K_M (~40 GB on disk, ~38 GB in VRAM).
  • Stack — vLLM 0.6.x, CUDA 12.4, Python 3.11. No TensorRT-LLM, no Triton — vLLM is simple enough and fast enough for our workloads.

Use-Case: ein Customer-Support-Agent mit RAG-Kontext (Top-5 Chunks, je ~400 Token), durchschnittliche Antwortlänge 180–240 Token. Die Anforderung: Time-to-First-Token P95 ≤ 1 s — die Antwort wird per Token-Streaming ausgeliefert, entscheidend ist also, wie schnell die ersten Wörter erscheinen, nicht wann das letzte Token fertig ist.

2 · Quantization — what we compared

We ran three quantization levels against our ground-truth set (see Note 01). Identical prompts, identical decoding parameters (temperature=0.2, top-p=0.9), 200 tokens max.

Variant VRAM TTFT P50 / P95 Faithfulness
Q4_K_M~38 GB~480 / ~780 ms0.891
Q8_0~52 GB ⚠——
FP16~140 GB ⚠——

Q8 and FP16 exceed the capacity of one L40S: Q8 would require paged offloading, while FP16 exceeds it considerably. We measured Q8 on two L40S GPUs with tensor parallelism: faithfulness 0.945 and TTFT about 590/950 ms. Faithfulness improves by 6%, with roughly twice the hardware required. Convincing for an internal knowledge system, usually less so for a customer chatbot.

3 · Batching — where the knee sits

vLLMs Continuous-Batching ist beeindruckend, aber bei großem Kontext (Top-5 RAG-Chunks plus System-Prompt) verschiebt sich das Optimum schnell. Wir haben batch_size 1, 4, 8 gegen 4k und 8k Kontext gemessen — die Werte sind gerundete Bereiche aus mehreren Läufen, keine Punktmessungen:

Kontext Batch TTFT P95 Decode pro Stream
4k1~0.8 s~18–20 Token/s
4k4~1.1 s~15–17 Token/s
4k8~1.5 s~12–15 Token/s
8k1~1.3 s~17–19 Token/s
8k4~1.6 s~13–16 Token/s
8k8~3 s ⚠~10–12 Token/s

Zur Einordnung: Bei ~15–20 Token/s pro Stream ist eine 200-Token-Antwort nach ~11–13 s vollständig ausgeliefert. Weil wir per Token-Streaming ausliefern, ist das für Nutzer unkritisch — die ersten Wörter erscheinen nach der TTFT, also unter einer Sekunde. Für unseren Use-Case (TTFT P95 ≤ 1 s, ~3–4 parallele Streams in der Spitze) ist 4k Kontext mit batch_size=4 der sweet spot. Bei 8k-Kontext-Anfragen routen wir bewusst auf einen separaten Pool mit niedrigerem Batch, um die TTFT nicht zu sprengen.

4 · vLLM configuration that saved us pain

# vllm-server.yaml (Auszug)

model: /models/llama-3.1-70b-q4km
quantization: gguf
dtype: auto

# Kontext und KV-Cache
max-model-len: 8192
kv-cache-dtype: fp8        # spart ~30% VRAM, Qualitätsverlust nicht messbar
enable-prefix-caching: true   # für RAG-Workloads kritisch

# Throughput
max-num-seqs: 8
max-num-batched-tokens: 32768
gpu-memory-utilization: 0.92  # nicht 0.95 — lässt Headroom für CUDA graphs

# Stabilität
swap-space: 4  # GB
disable-log-stats: false

# Networking
host: 0.0.0.0
port: 8000
api-key: "${VLLM_API_KEY}"

Three settings we learned the hard way:

  • enable-prefix-caching: true — in RAG workloads the system prompt and often even large parts of the context don't change per request. Prefix caching halves time-to-first-token. One-line change, huge impact.
  • kv-cache-dtype: fp8 — one third less VRAM, no measurable quality loss on the eval set. Lets us run batch_size=8 instead of 6.
  • gpu-memory-utilization: 0.92 — the vLLM default 0.9 is conservative, 0.95 caused us OOM crashes as soon as CUDA graphs kicked in. 0.92 is the value that runs robustly.

5 · When a second GPU is worth it

We recommend a second L40S to a customer in exactly three scenarios:

  • Q8-Quantisierung ist für den Use-Case nötig (regulatorisch oder messbar bessere Faithfulness > 5%).
  • Nachhaltig mehr als ~6–8 parallele Streams, nicht nur in Peaks.
  • Active-active HA requirement (two servers in two fire zones).

In the other cases, a single L40S is sufficient, without the additional hardware and complexity of tensor parallelism. We prefer a second L40S system for failover to combining two into one logical unit.

6 · What didn't fit in the TL;DR

Three things we'd do differently on a next deployment:

  • Wir würden früher mit AWQ-Quantisierung experimentieren. Q4_K_M ist gut, aber AWQ liefert in unseren ersten Tests vergleichbare Latenz mit messbar besserer Faithfulness. Die Produktionsstabilität müssen wir noch verifizieren.
  • We'd use system-prompt caching more aggressively. In the current deployment the system prompt changes monthly — that's wasted opportunity.
  • We'd evaluate speculative decoding. vLLM supports it stably by now; we simply haven't tested it yet.