Skip to main content
← All work
CASE STUDY / 01

InsureIntel Zimbabwe

Turning insurance documents into structured evidence for a broker’s first-pass review.

  • Python
  • FastAPI
  • spaCy
  • scikit-learn
  • PostgreSQL
  • ChromaDB
01 / CONTEXT

The problem & my contribution

Insurance brokers must read policy wordings, claims forms, treaty agreements, and correspondence to find terms, identify missing clauses, and assess insurer exposure. InsureIntel Zimbabwe explored how document intelligence could bring those tasks into one review workflow. I designed, implemented, and evaluated the academic prototype.

  • Run locally on an 8th-generation Intel i5 with 16 GB RAM, without a GPU or paid model APIs.
  • Handle digital PDFs and scanned documents with limited authentic training material.
  • Keep findings reviewable and preserve human responsibility for insurance decisions.
02 / APPROACH

From input to useful output

The pipeline tries pdfplumber text extraction first and routes low-text documents to Tesseract OCR. TF-IDF with Logistic Regression classifies four document types; a custom spaCy model extracts six entity types, including premiums, coverage limits, deductibles, and exclusions. Clause checks and a document-risk score support triage. A separate Corporate Solvency Profiling module combines four financial indicators and an XGBoost anomaly model. Regulatory retrieval uses MiniLM embeddings and ChromaDB to surface supporting passages with citations.

Digital PDFs or scans enter text extraction, then document classification and entity extraction. Document checks and insurer profiling provide evidence for broker review; regulatory retrieval supports questions.
Original conceptual workflow. Outputs support broker judgment; they do not certify compliance. Open diagram ↗
03 / ARCHITECTURE

How the pieces fit

FastAPI exposes the processing services to a server-rendered Jinja2 workspace. JWT authentication and role checks distinguish administrators, brokers, and analysts. Three stores separate structured records, raw extraction output, and regulatory retrieval; this flexibility also increases local deployment complexity.

A Jinja2 broker workspace connects to a FastAPI processing layer backed by PostgreSQL for structured records, MongoDB for OCR text, and ChromaDB for regulatory retrieval.
Original architecture summary, redrawn from the implementation description—not a screenshot of the application. Open diagram ↗
PostgreSQL
Users, document records, compliance findings, insurer profiles, and audit records.
MongoDB
Raw OCR output and variable page structures.
ChromaDB
Embedded regulatory passages drawn from the Insurance Act, 12 IPEC circulars, and an FSR-1 template.
04 / DECISIONS

Engineering trade-offs

Fit the model to the constraint

TF-IDF + Logistic Regression rather than a heavier classifier.

A small, CPU-friendly model suited the vocabulary differences between the four document classes. Training used 847 labelled samples.

Trade-off: Overlapping vocabulary made treaty agreements and correspondence harder to distinguish.

Fail closed on missing rules

Raise a configuration error when required clause files are missing.

During development, missing knowledge files could silently produce a 100% compliance score. A startup guard and regression test replaced that misleading success path.

Trade-off: Processing stops until configuration is repaired; availability does not take priority over trustworthy findings.

05 / EVALUATION

Results, with boundaries

These are dissertation-reported proof-of-concept results, not independently reproduced portfolio benchmarks. The document corpus contained 45 documents: 11 authentic seeds and 34 synthetic documents. The report describes 20% held-out splits for classification and NER; compliance thresholds and profiling weights were calibrated against the evaluation/reference data, so those results are not independent validation.

Classification F1
0.87
Weighted average across four document types.
Entity extraction F1
0.76
Micro-average; below the 0.85 research target.
Clause-check precision
0.91
Mandatory clause detection; recall was 0.78.
Digital / scanned PDF
12.4 / 34.7 s
Mean processing time over ten runs per type on the local CPU setup.
  • The report records 68 passing development tests, including authentication, role restrictions, extraction, and the missing-rules guard.
  • Deductibles were the weakest entity type (F1 0.69): conditional wording led to incomplete extraction spans.
  • Solvency rankings achieved reported Kendall’s τ = 0.74 against a regulatory reference, after weight optimisation—not a prospective insurer-risk test.

Source: University of Zimbabwe capstone dissertation, June 2026, implementation and results chapters. This condensed account anonymises client details and omits the full dissertation and source documents.

Discuss this work →