Digitization Accuracy Report
Designed to document chart-element detection, coordinate mapping, and measurement extraction against known ground-truth values.
Research and validation
The June 2026 SRS defines an academic prototype, planned studies, and performance targets. It does not report completed model or clinical validation findings.
Formal deliverables
Designed to document chart-element detection, coordinate mapping, and measurement extraction against known ground-truth values.
Designed to evaluate clinical relevance, review workflows, and the safe interpretation of digitized trends and anomaly flags.
Parent-facing guidance written at or below the specified fifth-grade comprehension level with accessible, active instructions.
Technical guidance for clinic partners and potential future exchange with EMR systems through the API.
Dataset creation
The proposed training dataset uses printed Kenya MOH Road-to-Health chart pages filled by volunteers. It does not use real patient charts in this academic prototype.
Use separate weight-for-age and height-for-age pages for boys and girls.
Give volunteers pre-defined age and measurement values to plot.
Use blue or black ballpoint dots, crosses, or short lines.
Photograph each chart by two people, on two devices, under at least two lighting conditions.
Intentionally crease, fold, or lightly stain a subset.
The team lead checks images before annotation.
Annotation and model training
Roboflow is specified for annotation. The final test images must be photographed by people who did not fill the training charts.
Known plot values and varied physical conditions.
→Axes, percentile curves, doctor plot points, and related noise classes.
→Cloud GPU training is planned because the team has no local GPU workstation.
→Used during model development and hyperparameter selection.
Held out from training and photographed by separate participants.
→Compare detections and extracted values with the original assigned coordinates.
→Document method, errors, limits, and performance.
→Assess meaning and workflow beyond technical accuracy.
Specified performance targets
The SRS defines thresholds the implementation is expected to meet. Evidence must come from the held-out test set and clinical comparison work.
Minimum target for bounding boxes and standard orientation markers across varied lighting.
Target stated for the custom YOLOv11 chart-element pipeline.
Target extraction error tolerance for weight readings compared with manual assessment.
Target extraction error tolerance for infant height readings.
Target for the future FastAPI and PostgreSQL cloud deployment.
WHO LMS calculations
The analytics engine is specified to select sex-specific WHO reference tables by exact age in days and calculate a Z-score using the LMS formula.
Y is the observed measurement; M is the median; L is the Box-Cox power; S is the coefficient of variation.
Does the system choose the correct table for sex, measure, and exact age?
Do outputs match a trusted independent WHO calculation?
Are values beyond ±3 SD handled with the specified extended formula?
Are Z-scores translated into the right labelled percentile bands?
Research limitations
The model is constrained to the Kenya MOH Road-to-Health chart geometry. Results cannot be generalized to other chart formats without recalibration and testing.
Volunteer-filled charts can model handwriting and physical wear, but they do not capture every condition present in real clinical records.
The SRS proposes a Clinical Validation Study but provides no outcomes. Clinical safety and workflow value remain to be demonstrated.
The project is constrained to a semester prototype. Full clinical deployment is explicitly outside the stated scope.
Model training depends on external cloud GPU resources because the team does not have a local GPU workstation.
The project scope calls immunization tracking out of scope, while Module 8 defines it as an MVP requirement. This needs formal requirements reconciliation before validation.
Future reports should publish the dataset protocol, split logic, evaluation definitions, error distribution, device and lighting coverage, failure cases, manual override rate, clinical review method, and limitations—not only one headline metric.
Research collaboration