How DataCore Analytics Keeps Data Confidential
12 May 2026
Applied AI for biological and health data — built on your data, validated honestly, and deployed only where it earns its place.
Artificial intelligence in health research has a credibility problem, and most of it is self-inflicted: models validated on the data they were trained on, performance reported without a baseline, and systems built on populations that look nothing like the ones they are eventually used on.
We approach it the other way round. Every project starts with the simplest defensible baseline and an honest feasibility assessment. If a complex model does not beat that baseline under external validation, we say so — and that is a useful result, not a failed project.
Each engagement uses whichever of these the problem actually needs.
Risk and outcome models for diagnosis, prognosis and triage, developed and reported to TRIPOD+AI standards with calibration, decision curve analysis and subgroup performance included as standard.
Variant effect prediction, expression-based classification, cell type annotation and biomarker discovery, with feature selection kept inside the cross-validation loop to prevent leakage.
Classification and segmentation on radiology, pathology and microscopy images, with transfer learning where sample sizes are realistic and an explicit account of where the model fails.
Extracting structured data from clinical notes, pathology reports and literature at scale, including handling of the multilingual and code-switched text common in African clinical records.
Literature triage and evidence synthesis, structured extraction from unstructured documents, and retrieval-augmented pipelines over your own document collections — with the outputs verified rather than trusted.
Time-series and spatio-temporal models for outbreak detection, disease burden forecasting and resource planning from routine surveillance data.
Turning analysis into something a decision-maker can actually use.
Shiny, Dash, Streamlit and Power BI dashboards over your surveillance, trial or programme data — hosted by you, with the source code handed over.
Scheduled pipelines that turn a weekly data drop into a formatted report or a repository submission package without anyone touching a spreadsheet.
Consolidating fragmented sources — laboratory systems, REDCap, DHIS2, spreadsheets — into a single queryable store with a documented schema.
Searchable interfaces over variant, sample and surveillance datasets, with role-based access so collaborators see only what they should.
Nextstrain builds and live dashboards that update as each sequencing batch completes, hosted on your own infrastructure.
Indicator frameworks, coverage estimates and evaluation analytics for public health programmes and funders.
Before any model is trained we establish what a simple approach achieves — logistic regression, an existing clinical score, or the current standard of care. Every subsequent result is reported against that number. A deep learning model that matches logistic regression is a logistic regression problem.
Internal cross-validation tells you very little on its own. Where the data allows we hold out a site, a time period or an entire cohort, because that is closer to how the model would actually be deployed. We report:
Models trained overwhelmingly on European and North American data routinely lose most of their performance when applied to African populations — different disease prevalence, different case mix, different measurement practice, different genetic background. Polygenic risk scores are the starkest example, losing the majority of their predictive power across ancestry groups.
Where you want to use a published model, we validate it on your population before you rely on it. Frequently the honest answer is that it needs recalibration, or that it should not be used at all.
For anything that could touch a clinical decision, a reviewer or a clinician has to be able to see what the model is using. We deliver SHAP values, feature importance and partial dependence alongside the model, and we look specifically for models that have learned a shortcut — the scanner rather than the disease, the ward rather than the diagnosis.
We do not train models on client data for any purpose other than that client's project without written permission. We do not pool data across clients. Where a model is built on identifiable data, membership inference risk is assessed before anything is shared or published.
Africa is being offered a great deal of AI built elsewhere, validated elsewhere, and sold as though geography were irrelevant. Some of it is genuinely useful. A significant amount of it will underperform badly, and the failures will fall on patients who were never in the training data.
We think the counterweight is capacity: African institutions that can independently validate what they are being sold, and build what is missing. That is a training problem as much as a technical one, which is why every AI engagement we run includes skills transfer.
Our AI work is delivered for research purposes. It is not validated or approved for clinical diagnosis, treatment decisions or individual patient management. Where a client intends to pursue clinical deployment we can advise on what additional validation and regulatory work that would require.
