InsureIntel Zimbabwe
Turning insurance documents into structured evidence for a broker’s first-pass review.
- Python
- FastAPI
- spaCy
- scikit-learn
- PostgreSQL
- ChromaDB
The problem & my contribution
Insurance brokers must read policy wordings, claims forms, treaty agreements, and correspondence to find terms, identify missing clauses, and assess insurer exposure. InsureIntel Zimbabwe explored how document intelligence could bring those tasks into one review workflow. I designed, implemented, and evaluated the academic prototype.
- Run locally on an 8th-generation Intel i5 with 16 GB RAM, without a GPU or paid model APIs.
- Handle digital PDFs and scanned documents with limited authentic training material.
- Keep findings reviewable and preserve human responsibility for insurance decisions.
From input to useful output
The pipeline tries pdfplumber text extraction first and routes low-text documents to Tesseract OCR. TF-IDF with Logistic Regression classifies four document types; a custom spaCy model extracts six entity types, including premiums, coverage limits, deductibles, and exclusions. Clause checks and a document-risk score support triage. A separate Corporate Solvency Profiling module combines four financial indicators and an XGBoost anomaly model. Regulatory retrieval uses MiniLM embeddings and ChromaDB to surface supporting passages with citations.
How the pieces fit
FastAPI exposes the processing services to a server-rendered Jinja2 workspace. JWT authentication and role checks distinguish administrators, brokers, and analysts. Three stores separate structured records, raw extraction output, and regulatory retrieval; this flexibility also increases local deployment complexity.
- PostgreSQL
- Users, document records, compliance findings, insurer profiles, and audit records.
- MongoDB
- Raw OCR output and variable page structures.
- ChromaDB
- Embedded regulatory passages drawn from the Insurance Act, 12 IPEC circulars, and an FSR-1 template.
Engineering trade-offs
Fit the model to the constraint
TF-IDF + Logistic Regression rather than a heavier classifier.
A small, CPU-friendly model suited the vocabulary differences between the four document classes. Training used 847 labelled samples.
Trade-off: Overlapping vocabulary made treaty agreements and correspondence harder to distinguish.
Fail closed on missing rules
Raise a configuration error when required clause files are missing.
During development, missing knowledge files could silently produce a 100% compliance score. A startup guard and regression test replaced that misleading success path.
Trade-off: Processing stops until configuration is repaired; availability does not take priority over trustworthy findings.
Results, with boundaries
These are dissertation-reported proof-of-concept results, not independently reproduced portfolio benchmarks. The document corpus contained 45 documents: 11 authentic seeds and 34 synthetic documents. The report describes 20% held-out splits for classification and NER; compliance thresholds and profiling weights were calibrated against the evaluation/reference data, so those results are not independent validation.
- Classification F1
- 0.87
- Weighted average across four document types.
- Entity extraction F1
- 0.76
- Micro-average; below the 0.85 research target.
- Clause-check precision
- 0.91
- Mandatory clause detection; recall was 0.78.
- Digital / scanned PDF
- 12.4 / 34.7 s
- Mean processing time over ten runs per type on the local CPU setup.
- The report records 68 passing development tests, including authentication, role restrictions, extraction, and the missing-rules guard.
- Deductibles were the weakest entity type (F1 0.69): conditional wording led to incomplete extraction spans.
- Solvency rankings achieved reported Kendall’s τ = 0.74 against a regulatory reference, after weight optimisation—not a prospective insurer-risk test.
Source: University of Zimbabwe capstone dissertation, June 2026, implementation and results chapters. This condensed account anonymises client details and omits the full dissertation and source documents.