Pillar 01 · AI Engineering

Tailored AI systems — securely integrated into your applications.

We build AI agents, RAG systems and automation directly within your Django app. Llama or Mistral on your own GPU and pgvector for RAG connect AI to your existing systems.

Methodology

Measure first, then optimize.

Most AI projects don't fail on the technology, but on measures nobody checked for impact beforehand. We work the other way around: we measure first where the problem really lies — and recommend a costly optimization only when it demonstrably helps on the eval set.

Before you invest We also say no. Reranker, fine-tuning, heavier prompts — every measure costs money and latency. Before we recommend it, we check on the eval set whether it improves quality at all. Often the honest answer is: this optimization isn't worth it here. Then you save the cost.
Reliable measurement Quality that stays stable even after a model switch. If you simply let an AI model judge answer quality, the number changes with the model — the more expensive model isn't automatically the more reliable one. We calibrate the evaluation per use case and with a human sample, so your metrics stay comparable.
How we measure quality — Engineering Notes →

Model selection

The right model — not the most well-known brand.

For secretarial agents, a local Llama 3.1 is often sufficient. For complex RAG workflows with long contexts, we use Claude. We choose based on requirements, not marketing budget.

Open-source · On-Premise Llama 3.1 · Mistral · DeepSeek Fully deployable on customer hardware. Quantization Q4/Q8 for a lean VRAM footprint. Models and weights interchangeable, no vendor lock-in. vLLM · Ollama · pgvector
Commercial APIs Claude · GPT-4 · Gemini Where latency, context length, or multimodality are decisive — and no special data protection requirements oppose it. EU region if available. Anthropic · OpenAI · Google

Deployment options

Cloud, on-premise, or hybrid — depending on your compliance requirements.

01 · Cloud API Fastest start Integration through Anthropic, OpenAI or Google APIs. For use cases without special data protection requirements: API connectivity, current model versions, elastic scaling.
02 · On-Premise Full data sovereignty Models run on customer hardware. No data transfer to third parties, GDPR-compliant by architecture. For regulated industries — healthcare, law, authorities, industry.
03 · Hybrid Sensitive data local, complex in the cloud PII anonymization and sensitive classification locally, large generation steps via API. Gradual migration to full on-premise possible.

Customized Solutions

Six application areas — embedded in your application, not alongside.

Every solution is integrated into your Django application and talks directly to your database. No white-label wrapper, no Make/Zapier kludge in between.

Healthcare AI agent for secretariats Telephone appointment scheduling, caller triage and patient enquiries — connected directly to your Django or PostgreSQL database. In production in customer projects.
Industry On-Premise IT monitoring Llama or Mistral on local server: log analysis, incident triage, monitoring evaluation. No logs leave the corporate network.
Law · Contracts · Knowledge RAG on corporate data Chatbot that searches internal PDFs, wikis, contracts, and databases. Answers with source citation — pgvector and Django.
Backoffice · Automation Workflow agents Email classification, invoice extraction, report generation. Agents reside in your Django app with direct ORM access.
Domain language Fine-tuning on your own data Train Llama or Mistral on your industry language — medical, legal, technical. LoRA and QLoRA, deployed on your GPU.
Compliance EU AI Act · GDPR Risk classification, compliance audit, documentation roadmap. Pragmatic for SMEs — no 200-page law firm opinions.
All industry solutions →

Stack

Production-ready components — no experiments.

Open-source stack, documented, interchangeable. We choose tools that have been running in our live systems for months — not what is trending on Hacker News.

GPU compute module for on-premise LLM inference on customer hardware
On-premise inference on your own GPU (RTX 4090 / L40S / H100) — your models stay on your hardware.
Models Llama · Mistral Open-source LLMs, deployable on-premise, quantifiable.
Inference vLLM · Ollama Production serving on your own GPU. Batched, persistent.
Retrieval pgvector Vector search directly next to your business data in Postgres.
App layer Django · Tool-Calling Agents live in the application, audit logs included.
Free initial consultation →

Why CODLAB

We build AI solutions — no slide decks.

We develop software for Bavarian SMEs. We understand the realities: limited budgets, no in-house IT team and strict GDPR requirements. We deliver concrete engineering work.

01 Already delivered, not promised AI reception and IT monitoring on local hardware.
02 Integrated into the application We build production Django applications. AI lives in your codebase, with access to your ORM models and business logic, instead of being external in a Zapier workflow.
03 On-premise really on-premise We deploy Llama, Mistral, and DeepSeek on local GPU servers at our clients — including maintenance, updates, and vLLM tuning. No OpenAI backdoor.
04 EU AI Act · pragmatic We classify your use cases (minimal · limited · high risk) and honestly tell you what you need to document. No 200-page law firm reports.
Free initial consultation →

Compliance

How we translate regulations into processes.

The EU AI Act classifies AI systems by risk. Documentation is not optional. CODLAB leads you systematically through these four steps — reliably, verifiable, pragmatic for SMEs.

01 · Classification minimal, limited or high risk? We systematically document your use case: input data (what and from whom), system logic, outputs, scope of liability. From this follows the risk category per EU AI Act Annex III. Example: a RAG chatbot over legal documents typically lands in limited risk (decision support). A secretary agent with notification function likewise.
02 · Data Origin Which data flows where? Registry: All training data, all eval sets, all inference input (chat logs, embeddings). For on-premise: no data egress. For cloud API: data processing agreement (DPA), data residency, deletion deadlines. We provide this overview to your data protection officer in writing.
03 · Transparency & Proof Bias? Hallucinations? Proven. Before going live: bias audit via eval set (minority representation, gender bias). Faithfulness measurement (how often does the model make unsupported statements?). Refusal rate (when does it rightly say no?). We document these metrics per model version and record them in the compliance dossier.
04 · Ongoing Operation Monitoring instead of surprises After go-live: automated quality regression on every model update. Chat-log sampling every 500 requests (capturing bias, hallucinations, unexpected outputs). Monthly compliance reporting to you and, where applicable, the supervisory authority.
Required Documentation for Limited Risk You get this in writing. Transparency statement (1–2 pages): What does your AI system do, who uses it, what errors are realistic? Datasheet: training sources, size, eval metrics. Bias audit: genders, age groups, regional disparities (if relevant). Boundaries of use: where the agent is NOT deployed (e.g., final medical diagnosis).
Contrary to common fears No 200-page law firm opinion. For limited-risk systems you do not need an external conformity assessment (unlike high-risk). You need clean documentation — which CODLAB prepares for you — and an internal four-eyes principle for release decisions. That is the rulebook, nothing more.
Request compliance audit at fixed price →

Frequently Asked Questions about AI Integration

No standard chatbots or white-label wrapper around ChatGPT. We analyse your workflow — appointments, calls, documents, logs — and build an AI agent embedded directly in your Django app and connected to your database.

Yes — and we have already built it. Llama 3.1 or Mistral on a local GPU server (RTX 4090, L40S), Ollama or vLLM as inference engine, pgvector for RAG. No data leaves your network. Especially relevant for practices, law firms, authorities, and all regulated sectors with strict DSGVO requirements.

The AI agent is a Django app or plugin inside your existing codebase. It has direct access to your ORM models (patient, appointment, invoice, ticket), uses your permissions and authentication, and runs in the same deployment process. No Make/Zapier workflow in between, no external API synchronisation.

We prepare an individual proposal. Please contact us.

We classify your use case according to the AI Act risk levels (minimal/limited/high risk), create the mandatory documentation, and advise on transparency requirements. Most medium-sized business use cases (chatbots, RAG, internal tools) fall into "limited risk" with manageable obligations — no 200-page law firm opinion needed.

Tailored AI · Django-integrated · On-premise

Let's outline your AI solution.

30 Minuten kostenloses Erstgespräch. Wir analysieren Ihren Workflow, sagen ehrlich, ob KI hier Sinn ergibt — und wenn ja, in welcher Form (Agent, RAG, Automatisierung, on-premise oder Cloud). Kein KI-Hype, keine Stundenfalle, keine PowerPoint.

+49 941 20 90 28 62

Reply within 24 hours on business days · Mon–Fri 9–18 · CODLAB · St.-Jakob-Str. 6, 93161 Sinzing

Documented development GDPR · servers in the EU, in Germany on request Direct developer · no call center