CDAAugmented document knowledge
Innovation · Data & analytics
A graph of 10,000 entities extracted from 11,000 pages of dossiers.
The innovation division held free-text dossiers that analytics could barely use. A generative extraction POC produced a structured, explorable, auditable graph.
Key metrics
- Client
- Bpifrance
- Sector
- Innovation
- Function
- Data & analytics
- Status
- POC delivered — industrialisation underway
- Deployment
- Managed cloud (client tenant)
The context
Hundreds of dossiers, more than 11,000 pages: qualitative natural-language data largely unused for cohort analysis or public-policy indicators.
The problem
- Links between startups, labs, patents and funding were unstructured.
- Manual analyses were long and non-reproducible. No continuously usable overview.
The solution
Co-designed ontology (node and relation types), generative extraction, relational storage, visualisation, expert correction loop. The Evidence Panel cites source sentences back to the original document.
NEXA components in scope
| Component | Role |
|---|---|
| Arbitration algorithm | Orchestration of extraction models |
| Evidence Panel | Source sentences back to the document |
| HITL | Validation / correction of nodes and relations |
| AI Knowledge Vault | Co-designed, reusable ontology |
| Cockpit | Precision, coverage, footprint |
The impact
| Axis | Concrete impact |
|---|---|
| Data | Quantitative ecosystem indicators inaccessible before structuring. |
| Precision | 80% precision; no hallucinations observed in control. |
| Scalability | Architecture extensible to other corpora. |
Your business ritual has the same potential.
Evidence is not our constraint. It is our product.