AI & Data Analytics

Applied AI for biological and health data — built on your data, validated honestly, and deployed only where it earns its place.

Artificial intelligence in health research has a credibility problem, and most of it is self-inflicted: models validated on the data they were trained on, performance reported without a baseline, and systems built on populations that look nothing like the ones they are eventually used on.

We approach it the other way round. Every project starts with the simplest defensible baseline and an honest feasibility assessment. If a complex model does not beat that baseline under external validation, we say so — and that is a useful result, not a failed project.

What We Build

Each engagement uses whichever of these the problem actually needs.

01

Clinical Prediction Models

Risk and outcome models for diagnosis, prognosis and triage, developed and reported to TRIPOD+AI standards with calibration, decision curve analysis and subgroup performance included as standard.

02

Genomic Machine Learning

Variant effect prediction, expression-based classification, cell type annotation and biomarker discovery, with feature selection kept inside the cross-validation loop to prevent leakage.

03

Medical Imaging Analysis

Classification and segmentation on radiology, pathology and microscopy images, with transfer learning where sample sizes are realistic and an explicit account of where the model fails.

04

Natural Language Processing

Extracting structured data from clinical notes, pathology reports and literature at scale, including handling of the multilingual and code-switched text common in African clinical records.

05

Large Language Models for Research

Literature triage and evidence synthesis, structured extraction from unstructured documents, and retrieval-augmented pipelines over your own document collections — with the outputs verified rather than trusted.

06

Forecasting and Early Warning

Time-series and spatio-temporal models for outbreak detection, disease burden forecasting and resource planning from routine surveillance data.

Analytics and Decision Support

Turning analysis into something a decision-maker can actually use.

Interactive Dashboards

Shiny, Dash, Streamlit and Power BI dashboards over your surveillance, trial or programme data — hosted by you, with the source code handed over.

Automated Reporting

Scheduled pipelines that turn a weekly data drop into a formatted report or a repository submission package without anyone touching a spreadsheet.

Data Warehousing

Consolidating fragmented sources — laboratory systems, REDCap, DHIS2, spreadsheets — into a single queryable store with a documented schema.

Genomic Data Portals

Searchable interfaces over variant, sample and surveillance datasets, with role-based access so collaborators see only what they should.

Real-Time Surveillance Views

Nextstrain builds and live dashboards that update as each sequencing batch completes, hosted on your own infrastructure.

Metrics and Programme Evaluation

Indicator frameworks, coverage estimates and evaluation analytics for public health programmes and funders.

How We Keep AI Honest

The baseline comes first

Before any model is trained we establish what a simple approach achieves — logistic regression, an existing clinical score, or the current standard of care. Every subsequent result is reported against that number. A deep learning model that matches logistic regression is a logistic regression problem.

Validation that means something

Internal cross-validation tells you very little on its own. Where the data allows we hold out a site, a time period or an entire cohort, because that is closer to how the model would actually be deployed. We report:

  • Discrimination and calibration, not accuracy alone
  • Confidence intervals on every performance metric
  • Performance broken down by age, sex, site and any subgroup where failure would matter
  • Decision curve analysis showing whether the model beats treating everyone or no one
  • Explicit statement of the population the model should not be used on

The transferability problem

Models trained overwhelmingly on European and North American data routinely lose most of their performance when applied to African populations — different disease prevalence, different case mix, different measurement practice, different genetic background. Polygenic risk scores are the starkest example, losing the majority of their predictive power across ancestry groups.

Where you want to use a published model, we validate it on your population before you rely on it. Frequently the honest answer is that it needs recalibration, or that it should not be used at all.

Interpretability is not optional

For anything that could touch a clinical decision, a reviewer or a clinician has to be able to see what the model is using. We deliver SHAP values, feature importance and partial dependence alongside the model, and we look specifically for models that have learned a shortcut — the scanner rather than the disease, the ward rather than the diagnosis.

Data governance for AI

We do not train models on client data for any purpose other than that client's project without written permission. We do not pool data across clients. Where a model is built on identifiable data, membership inference risk is assessed before anything is shared or published.

Responsible AI for African Health Data

Why this needs saying

Africa is being offered a great deal of AI built elsewhere, validated elsewhere, and sold as though geography were irrelevant. Some of it is genuinely useful. A significant amount of it will underperform badly, and the failures will fall on patients who were never in the training data.

We think the counterweight is capacity: African institutions that can independently validate what they are being sold, and build what is missing. That is a training problem as much as a technical one, which is why every AI engagement we run includes skills transfer.

What we will not do

  • Deploy a clinical model without external validation on the target population
  • Report performance without a baseline comparison
  • Build a diagnostic tool intended for clinical use without the regulatory pathway being clear
  • Train on client data for other purposes without explicit written permission
  • Present a model as production-ready when it is a research prototype

Research use only

Our AI work is delivered for research purposes. It is not validated or approved for clinical diagnosis, treatment decisions or individual patient management. Where a client intends to pursue clinical deployment we can advise on what additional validation and regulatory work that would require.

Contact DataCore Analytics

Tell us about your data and we will scope it — free, within one working day.

+233 558 017 827

Free scoping call · reply within 1 working day Get a Quote