We develop AI systems within your Django application and run models on your own hardware. The same team handles development and operation.
Customized AI systems
for regulated companies.
AI engineering for SMEs, law firms, medical practices and industry. On-premise LLM (Llama, Mistral) on your own hardware, RAG with data sovereignty, agents integrated directly into your applications. No mere API wrapper, no cloud lock-in.
Engineering notes
How we build — notes from the engineering bench.
Documented architecture decisions, benchmarks from production deployments and engineering patterns we reuse. No generic tutorials, no marketing whitepapers — only the questions we had to answer on real customer projects.
How we evaluate a RAG system before it goes live
Our internal eval harness: 7 metrics, three ground-truth sets from real customer questions, automated regression on every model swap. What we measure — and why “works in the demo” isn't enough.
Lesen → Note 02 · InferenceLlama 3.1 70B on a single L40S — what we measured
Quantization (Q4/Q8), batched inference, vLLM configuration. Where the latency knee sits, when a second GPU is worth it, and which token sizes halve throughput. Real numbers from a production deployment.
Lesen → Note 03 · Architekturpgvector next to the Django ORM — pattern and benchmark
Why we put vector search directly in Postgres instead of running a separate Pinecone or Weaviate instance. Indexing strategy, HNSW tuning, recall comparison. With the ORM mixin we reuse in every project.
Lesen →Our AI stack
Open-source LLMs, integrated in Django, on your hardware.
We rely on production-ready, open components. Models and tools interchangeable, no vendor lock-in, no hidden cloud dependencies.
Engagement models
Four ways to work with us.
From AI strategy audit to fully operated on-premise cluster. Each step with a fixed price after the strategy conversation, without hourly trap.
01
AI Strategy & Audit
Workflow analysis, use case assessment, architecture recommendation. Including EU AI Act classification of your existing systems and a written roadmap.
02
Pilot Implementation
A production-ready AI agent or RAG prototype, embedded in your Django application. Weekly demos, iterative until go-live.
03
Production rollout
Complete on-premise setup with own hardware, vLLM/Ollama stack, monitoring, GDPR documentation, and maintenance contract from day one.
04
Managed AI
Ongoing operation of your AI systems: model updates, GPU maintenance, vLLM tuning, compliance reviews and extended availability on request.
FAQ
The most common questions before the call.
That is exactly what we are here for. We analyse your workflow, build, deploy and maintain the system — including model updates and ongoing operation. You need no in-house AI expertise: you get documentation, a clear maintenance path and a direct contact person, no call centre.
After the strategy audit, a pilot delivers a production-ready use case in 8–12 weeks — with weekly demos. No big bang after a year: you see early whether it pays off and decide after every 2-week sprint whether we continue.
Then we tell you — in the strategy audit, before you invest heavily. We honestly assess whether a use case brings real ROI. Better a clear “not worth it here” than an expensive project with no impact. The audit fixed fee is credited against a subsequent pilot.
Let's outline your AI solution.
30 minutes on the phone. We analyze your workflow and honestly tell you whether AI brings ROI for you — and if so, with which architecture (agent, RAG, on-premise or cloud). No AI hype, no sales pitch.
+49 941 20 90 28 62 [email protected]