← Note tecniche

Note 02 · Inference · Deployment

Llama 3.1 70B su una L40S — cosa abbiamo misurato.

Una sola NVIDIA L40S (48 GB), Llama 3.1 70B Instruct, vLLM come serving layer. Abbiamo misurato varianti di quantizzazione, batch size e lunghezze di prompt l'una contro l'altra. Questa nota documenta i valori reali di latenza e throughput da un agente di customer support in produzione.

~22 min Maggio 2026 Da un deployment live
TL;DR

Llama 3.1 70B in Q4_K_M passt mit 4k Kontext komfortabel auf eine L40S. Time-to-First-Token: P50 ~480 ms, P95 ~780 ms. Die Decode-Rate liegt bei ~15–20 Token/s pro Stream — eine 200-Token-Antwort ist damit nach ~11–13 s komplett, dank Token-Streaming sieht der Nutzer die ersten Wörter aber praktisch sofort. Bei batch_size=4 und 8k Kontext steigt die P95-TTFT Richtung 1.6 s und die Decode-Rate pro Stream sinkt spürbar. Q8 liefert ~6% bessere Faithfulness gegen 35% mehr VRAM und höhere Latenz — für die meisten Mittelstands-Use-Cases nicht das Geld wert. FP16 ist auf einer einzelnen L40S nicht praktikabel: 4-bit-Quantisierung ist hier kein Kompromiss, sondern die korrekte Wahl.

1 · Hardware und Modell

La configurazione di cui stiamo parlando:

  • GPU — NVIDIA L40S, 48 GB GDDR6 ECC, 350 W TDP. Una scheda, niente connessione NVLink.
  • Host — AMD EPYC 9354P, 128 GB RAM, NVMe locale. Niente network storage sul percorso di inferenza.
  • Modell — meta-llama/Llama-3.1-70B-Instruct, quantizzato in Q4_K_M (~40 GB su disco, ~38 GB in VRAM).
  • Stack — vLLM 0.6.x, CUDA 12.4, Python 3.11. Niente TensorRT-LLM, niente Triton — vLLM è abbastanza semplice e abbastanza veloce per i nostri carichi.

Use-Case: ein Customer-Support-Agent mit RAG-Kontext (Top-5 Chunks, je ~400 Token), durchschnittliche Antwortlänge 180–240 Token. Die Anforderung: Time-to-First-Token P95 ≤ 1 s — die Antwort wird per Token-Streaming ausgeliefert, entscheidend ist also, wie schnell die ersten Wörter erscheinen, nicht wann das letzte Token fertig ist.

2 · Quantizzazione — cosa abbiamo confrontato

Abbiamo fatto girare tre livelli di quantizzazione contro il nostro ground-truth set (vedi Nota 01). Prompt identici, parametri di decoding identici (temperature=0.2, top-p=0.9), 200 token max.

Variante VRAM TTFT P50 / P95 Faithfulness
Q4_K_M~38 GB~480 / ~780 ms0.891
Q8_0~52 GB ⚠——
FP16~140 GB ⚠——

Q8 e FP16 superano la capacità di una L40S: Q8 richiederebbe offloading, FP16 la supera nettamente. Abbiamo misurato Q8 su due GPU L40S con parallelismo tensoriale: fedeltà 0,945 e TTFT circa 590/950 ms. La fedeltà aumenta del 6%, con un fabbisogno di hardware circa doppio. Convincente per un sistema interno, di solito meno per un chatbot clienti.

3 · Batching — dove sta il punto di rottura

vLLMs Continuous-Batching ist beeindruckend, aber bei großem Kontext (Top-5 RAG-Chunks plus System-Prompt) verschiebt sich das Optimum schnell. Wir haben batch_size 1, 4, 8 gegen 4k und 8k Kontext gemessen — die Werte sind gerundete Bereiche aus mehreren Läufen, keine Punktmessungen:

Kontext Batch TTFT P95 Decode pro Stream
4k1~0.8 s~18–20 Token/s
4k4~1.1 s~15–17 Token/s
4k8~1.5 s~12–15 Token/s
8k1~1.3 s~17–19 Token/s
8k4~1.6 s~13–16 Token/s
8k8~3 s ⚠~10–12 Token/s

Zur Einordnung: Bei ~15–20 Token/s pro Stream ist eine 200-Token-Antwort nach ~11–13 s vollständig ausgeliefert. Weil wir per Token-Streaming ausliefern, ist das für Nutzer unkritisch — die ersten Wörter erscheinen nach der TTFT, also unter einer Sekunde. Für unseren Use-Case (TTFT P95 ≤ 1 s, ~3–4 parallele Streams in der Spitze) ist 4k Kontext mit batch_size=4 der sweet spot. Bei 8k-Kontext-Anfragen routen wir bewusst auf einen separaten Pool mit niedrigerem Batch, um die TTFT nicht zu sprengen.

4 · Configurazione vLLM che ci ha risparmiato dolori

# vllm-server.yaml (Auszug)

model: /models/llama-3.1-70b-q4km
quantization: gguf
dtype: auto

# Kontext und KV-Cache
max-model-len: 8192
kv-cache-dtype: fp8        # spart ~30% VRAM, Qualitätsverlust nicht messbar
enable-prefix-caching: true   # für RAG-Workloads kritisch

# Throughput
max-num-seqs: 8
max-num-batched-tokens: 32768
gpu-memory-utilization: 0.92  # nicht 0.95 — lässt Headroom für CUDA graphs

# Stabilität
swap-space: 4  # GB
disable-log-stats: false

# Networking
host: 0.0.0.0
port: 8000
api-key: "${VLLM_API_KEY}"

Tre impostazioni che abbiamo imparato a caro prezzo:

  • enable-prefix-caching: true — nei carichi RAG il system prompt e spesso anche grandi porzioni di contesto non cambiano per richiesta. Il prefix caching dimezza il time-to-first-token. Modifica da una riga con impatto enorme.
  • kv-cache-dtype: fp8 — un terzo di VRAM in meno, nessuna perdita di qualità misurabile sull'eval-set. Ci permette batch_size=8 invece di 6.
  • gpu-memory-utilization: 0.92 — il default vLLM 0.9 è conservativo, 0.95 ci ha causato crash OOM appena attivati i CUDA graph. 0.92 è il valore che gira robusto.

5 · Quando vale la seconda GPU

Raccomandiamo a un cliente una seconda L40S esattamente in tre scenari:

  • Q8-Quantisierung ist für den Use-Case nötig (regulatorisch oder messbar bessere Faithfulness > 5%).
  • Nachhaltig mehr als ~6–8 parallele Streams, nicht nur in Peaks.
  • Requisito di HA active-active (due server in due compartimenti antincendio).

Negli altri casi basta una singola L40S, senza l’hardware aggiuntivo e la complessità del parallelismo tensoriale. Preferiamo un secondo sistema L40S per il failover a due schede collegate in un’unica unità logica.

6 · Cosa non entrava nel TL;DR

Tre cose che faremmo diversamente al prossimo deployment:

  • Wir würden früher mit AWQ-Quantisierung experimentieren. Q4_K_M ist gut, aber AWQ liefert in unseren ersten Tests vergleichbare Latenz mit messbar besserer Faithfulness. Die Produktionsstabilität müssen wir noch verifizieren.
  • Useremmo il caching del system prompt in modo più aggressivo. Nel deployment attuale il system prompt cambia mensilmente — è uno spreco.
  • Valuteremmo lo speculative decoding. vLLM ormai lo supporta stabilmente, semplicemente non l'abbiamo ancora testato.