Flagship Projects
Three production-grade AI systems — each designed to publish six benchmark numbers and run on real infrastructure.
| # | Project | Domain | Infra | Key differentiator |
|---|---|---|---|---|
| P1 | Enterprise Knowledge Platform | Any / cloud | Managed cloud | 10k+ docs, hybrid RAG, eval CI, full observability |
| P2 | Workflow Automation Platform | Business workflows | Cloud / hybrid | MCP servers, agent orchestration, human-in-the-loop, reliability SLOs |
| P3 | Insurance Compliance Copilot | Regulated / sovereign | minicloud k8s | AI Act compliance, PII pipeline, self-hosted models, Cosign, audit lineage |
The six benchmark numbers
Every project publishes these with methodology:
- Groundedness rate (%)
- Citation accuracy (%)
- Refusal rate on out-of-scope (%)
- p50 / p95 / p99 latency (ms)
- Cost per resolved query (€)
- Hallucination rate on adversarial set (%)
Progression logic
P1 → P2 → P3 is a deliberate ladder. P3 requires everything from P1 and P2, plus compliance-governance and full AI Act controls. Do not skip.