Stack & Architecture

Production-ready components — documented and interchangeable.

Our technology stack is a consolidated overview of all systems we deploy in bespoke AI solutions. All components are open-source or commercially available, long-term maintainable, and not locked to a single model or tool. Complete architecture decisions and benchmarks are in our Engineering Notes.

01 · Models

Open-source LLMs, deployable on-premise.

Wir arbeiten mit bewährten Open-Source-Modellen, die sich zu 100 % auf Kundenhardware deployen lassen. Die Modellwahl erfolgt nach Anforderung, nicht nach Marketing-Budget.

Large Language Model Llama 3.1 Meta model in three sizes: 8B for latency, 70B for intelligence. Quantized in Q4_K_M or Q8. Open weights, no restrictions. Open Weights · 8B bis 70B · Multi-Sprache
Large Language Model Mistral · DeepSeek Mistral 7B for speed, Mistral Large 2 for longer contexts. DeepSeek-V3 for domain-specific tasks. All deployable on customer hardware. Mistral 7B/Large · DeepSeek-V3 · Effizient
Why this model mix: Llama leads in quality and community support, Mistral 7B in throughput, DeepSeek for domain-specific work. We choose based on eval-set results and latency requirements — no vendor lock-in, models stay replaceable.

02 · Inference & Serving

Production-grade inference engine — batching & KV-cache optimization.

vLLM is our core inference server for L40S and H100 GPUs. Ollama for simple setups or local development. Both enable efficient batching and persistent cache patterns.

Inference server vLLM High-performance server with prefix caching, dynamic KV cache, batched inference. Llama 3.1 70B in Q4_K_M on L40S: time-to-first-token P50 ~480 ms, P95 ~780 ms (batch_size=1). Prefix Caching · HNSW-Support · Batched · NVIDIA-native
Alternative / Local Ollama Simplified locally-run alternative for development or proof-of-concept. Straightforward for demos, but fewer tuning options than vLLM. Einfach · Cross-Platform · PoC-freundlich
Why vLLM in production: Nativ auf NVIDIA GPUs optimiert, Caching-Features reduzieren Latenz um 30–50 %, Batching ermöglicht höheren Durchsatz. Ollama genügt für Piloten, aber vLLM ist der Standard für SLA-Workloads.

03 · Retrieval & Vectors

pgvector — vectors alongside your business data.

Not a separate vector DB (Pinecone, Weaviate). Instead: pgvector in PostgreSQL with HNSW index. One database system, one backup regime, no sync complexity.

Vector search pgvector + HNSW PostgreSQL 16+ mit pgvector Extension. HNSW-Index für schnelle Nearest-Neighbor-Suche, IVFFlat für größere Korpora. Recall@5 > 95 % auf echten Kundendaten. HNSW · IVFFlat · PostgreSQL 16 · Sub-100ms
Fallback (only >50M embeddings) Qdrant · Pinecone Separate Vektor-DB nur wenn Corpus >50M Embeddings oder Strict P50 <50ms Latenz. Für 99 % der Mittelstands-Use-Cases overkill. Enterprise · Optional · Specialist
Why pgvector standard: Operational simplicity (one DB backup, one monitoring system), tenant-scoped filtering via SQL, alongside business data. Benchmark shows: faster and cheaper than external storage for 500k–2M embeddings.

04 · App integration & agents

AI lives in the code, not beside it.

The AI agent is a Django app or plugin. Tool-calling via native Anthropic/OpenAI APIs, audit logs for every agent call in the database.

Web framework Django 5.x All AI agents are Django apps or plugins. Direct access to ORM models (patient, appointment, invoice), permissions, authentication. Deployment via standard Django process. Django 5.0+ · ORM-native · Batteries-included
Agent pattern Tool-Calling + Audit Agents via Claude/OpenAI/Llama tool-use. Every call logged in DB with user, timestamp, input, output. Refusal logic and fallback behavior defined per agent. Tool-Use · Reflexion · Deterministic Fallback
Why embedded in Django: No Zapier workflows, no external API synchronisation. Agents have direct access to your ORM models, permissions and business logic. One codebase, one deployment, one permissions audit trail.

05 · Quality metrics

7-metric evaluation — before every deployment.

Unser internes Eval-Harness läuft bei jedem Modell- oder Index-Wechsel. Drei Ground-Truth-Sets (Realität, Adversarial, Regression), automatisierte Reports, Regression-Blocking (>5 % blockt den Merge).

The 7 metrics:
Retrieval
recall@5
Ranking
MRR
Hit-rate
@1
Generation
faithfulness
Relevance
answer-relevance
Safety
refusal-rate
Evaluation process: Three sets per project: real user questions, adversarial edge cases and hallucinations, and regression queries from 6–12 months. Every model or index change generates an automated regression report. Regression above 5% blocks release.
Evaluation in detail — Engineering Note 01 →

06 · Compliance & Data Protection

Regulated industries. No backdoors.

All components can run fully on-premise. GDPR compliance by architecture (not by paperwork), EU-AI-Act classification for every use case.

Data sovereignty On-Premise LLM, embedding engine, and vector DB run on customer hardware in their own network. No data transfer to third parties, GDPR-compliant by design. For authorities, law firms, practices, industry. Vollständig dezentralisiert · Keine Cloud · Kein Datentransfer
Regulation EU AI Act · GDPR Risk classification (minimal · limited · high risk), mandatory documentation and bias audit. Pragmatic approach for mid-market — no 200-page law firm opinions, just what's legally required. Risk-Klassifizierung · Transparenz-Dokumentation · Audit-Trail
Why on-premise architecture: We plan hardware and operations around your requirements. We prepare an individual proposal. Please contact us.
GDPR-compliant by architecture
EU-AI-Act risk classification
On-premise deployable · no vendor lock-in

Stack overview

All components together.

This stack is proven in several customer projects in production. Long-term maintainable, documented, and not tied to any single application.

Layer Component Rationale
Modelle Llama 3.1, Mistral, DeepSeek Open Weights, on-premise deploybar, quantisierbar, austauschbar
Inferenz vLLM (Prod) oder Ollama (PoC) vLLM: Batching, KV-Cache, Prefix-Caching; Ollama: einfach & schnell
Retrieval pgvector in PostgreSQL + HNSW Vektoren neben Geschäftsdaten, ein Backup-System, tenant-scoped filtering
App-Layer Django 5.x + Tool-Calling KI sitzt im Code, direkter ORM-Zugriff, Audit-Logs pro Call
Evaluation 7-Metriken-Harness recall@5, MRR, hit-rate@1, faithfulness, answer-relevance, refusal-rate — vor jedem Deployment
Compliance On-Premise, DSGVO, EU AI Act Datensouveränität durch Design, Risikoklassifizierung pro Use-Case

Architecture philosophy

Why this stack.

Proven components over trend tools. Operational simplicity over complexity. Fewer moving parts, less to debug, less to maintain.

01 Real benchmarks, no guesses All components are deployed in multiple customer projects. vLLM latency, pgvector recall, eval-metric calibration — all measured on real workloads, not theoretical.
02 Operational simplicity One database system (PostgreSQL), no separate vector DB. One deployment process (Django), no external API dependencies for the core setup. Less to secure, less to monitor.
03 Long-term maintainable All components are open-source or established commercial services, actively maintained and without esoteric dependencies. Still deployed in 3 years without asking vendors.
04 No backdoors On-premise is optional but standard by default. All components can run on own hardware. No forced cloud dependency for data residency or compliance.
Architecture decisions in detail — engineering notes →

Deeper insights

Engineering Notes — benchmarks & decisions.

This page is an overview. For the details — concrete latency numbers, index tuning, eval methodology — see our Engineering Notes.

Note 01 RAG evaluation Eval harness, ground-truth construction, regression blocking, calibration for consistent measurements across model changes. To the note →
Note 02 Llama 3.1 on L40S Quantization, vLLM batching, KV-cache tuning, latency profiles (P50/P95), where the knee lies and when a second GPU pays off. To the note →
Note 03 pgvector & Django ORM Why pgvector over Pinecone, index strategy, tuning parameters, recall comparison, limits of the approach, when external vector DB is needed. To the note →

Technology & stack · documented · on-premise-capable

Schauen wir uns gemeinsam an, wie dieser Stack zu Ihrem Projekt passt.

30-minute call. We analyze your requirements and honestly say which components make sense for you — and which are overkill. No upsell, no hype, just engineering reality.

+49 941 20 90 28 62

Reply within 24 hours on business days · Mon–Fri 9–18 · CODLAB · St.-Jakob-Str. 6, 93161 Sinzing

Documented development GDPR · servers in the EU, in Germany on request Direct developer · no call center