Stack & Architecture
Production-ready components — documented and interchangeable.
Our technology stack is a consolidated overview of all systems we deploy in bespoke AI solutions. All components are open-source or commercially available, long-term maintainable, and not locked to a single model or tool. Complete architecture decisions and benchmarks are in our Engineering Notes.
01 · Models
Open-source LLMs, deployable on-premise.
Wir arbeiten mit bewährten Open-Source-Modellen, die sich zu 100 % auf Kundenhardware deployen lassen. Die Modellwahl erfolgt nach Anforderung, nicht nach Marketing-Budget.
02 · Inference & Serving
Production-grade inference engine — batching & KV-cache optimization.
vLLM is our core inference server for L40S and H100 GPUs. Ollama for simple setups or local development. Both enable efficient batching and persistent cache patterns.
03 · Retrieval & Vectors
pgvector — vectors alongside your business data.
Not a separate vector DB (Pinecone, Weaviate). Instead: pgvector in PostgreSQL with HNSW index. One database system, one backup regime, no sync complexity.
04 · App integration & agents
AI lives in the code, not beside it.
The AI agent is a Django app or plugin. Tool-calling via native Anthropic/OpenAI APIs, audit logs for every agent call in the database.
05 · Quality metrics
7-metric evaluation — before every deployment.
Unser internes Eval-Harness läuft bei jedem Modell- oder Index-Wechsel. Drei Ground-Truth-Sets (Realität, Adversarial, Regression), automatisierte Reports, Regression-Blocking (>5 % blockt den Merge).
06 · Compliance & Data Protection
Regulated industries. No backdoors.
All components can run fully on-premise. GDPR compliance by architecture (not by paperwork), EU-AI-Act classification for every use case.
Stack overview
All components together.
This stack is proven in several customer projects in production. Long-term maintainable, documented, and not tied to any single application.
| Layer | Component | Rationale |
|---|---|---|
| Modelle | Llama 3.1, Mistral, DeepSeek | Open Weights, on-premise deploybar, quantisierbar, austauschbar |
| Inferenz | vLLM (Prod) oder Ollama (PoC) | vLLM: Batching, KV-Cache, Prefix-Caching; Ollama: einfach & schnell |
| Retrieval | pgvector in PostgreSQL + HNSW | Vektoren neben Geschäftsdaten, ein Backup-System, tenant-scoped filtering |
| App-Layer | Django 5.x + Tool-Calling | KI sitzt im Code, direkter ORM-Zugriff, Audit-Logs pro Call |
| Evaluation | 7-Metriken-Harness | recall@5, MRR, hit-rate@1, faithfulness, answer-relevance, refusal-rate — vor jedem Deployment |
| Compliance | On-Premise, DSGVO, EU AI Act | Datensouveränität durch Design, Risikoklassifizierung pro Use-Case |
Architecture philosophy
Why this stack.
Proven components over trend tools. Operational simplicity over complexity. Fewer moving parts, less to debug, less to maintain.
Deeper insights
Engineering Notes — benchmarks & decisions.
This page is an overview. For the details — concrete latency numbers, index tuning, eval methodology — see our Engineering Notes.
Technology & stack · documented · on-premise-capable
Schauen wir uns gemeinsam an, wie dieser Stack zu Ihrem Projekt passt.
30-minute call. We analyze your requirements and honestly say which components make sense for you — and which are overkill. No upsell, no hype, just engineering reality.