Data Submission Guidelines
- Home
- Working With Us
- Data Submission
How to send us your data: accepted formats, transfer methods, the metadata we need, and what we check on arrival.
We are an analysis-only organisation and never handle physical samples. Everything below concerns digital data transfer.
Read this before your first transfer. Most delays at the start of a project are caused by missing metadata or incomplete file sets, not by the analysis itself.
The Transfer Process
Data is only accepted after an agreement is signed.
Agreement First
No data is accepted before a written agreement is in place covering scope, storage location, retention period, access, authorship and intellectual property. Where the study involves human subjects we ask to see ethical approval and any applicable data transfer agreement.
You Receive a Transfer Link
We send an encrypted, time-limited upload link to a Google Cloud bucket provisioned for your project alone. Do not email data, and do not send it through consumer file-sharing services.
Upload Data and Metadata Together
Upload your files along with the completed metadata sheet we send you. Incomplete metadata is the single most common cause of delay.
We Verify on Arrival
Within two working days we confirm receipt, verify checksums, check that every expected file is present and readable, and run initial quality control. If something is missing or corrupted we tell you immediately.
Analysis Begins
Once verification passes, your named analyst starts work to the approved plan and confirms the delivery date.
Retention and Deletion
At the end of the project data is returned, retained for the agreed period, or securely deleted — whichever your agreement specifies. We confirm deletion in writing.
Accepted File Formats
| Data type | Formats we accept | Notes |
|---|---|---|
| Sequence reads | FASTQ (.fastq.gz, .fq.gz) — gzip compressed | Keep read pairs as separate R1/R2 files with consistent naming |
| Aligned reads | BAM, CRAM | Include the index (.bai / .crai) and tell us the reference genome build used |
| Nanopore signal | POD5, FAST5 | Only needed if you want rebasecalling; otherwise send FASTQ |
| Variants | VCF, gVCF, BCF (bgzip compressed) | Include the .tbi index and the reference build |
| Expression matrices | CSV, TSV, MTX, RDS, H5AD | State whether values are raw counts, TPM, FPKM or normalised |
| Single-cell | CellRanger outs directory, H5AD, RDS, Loom | Send the filtered and raw matrices where both exist |
| Methylation arrays | IDAT (both Grn and Red per sample) | Include the sample sheet with sentrix ID and position |
| Genotype arrays | PLINK (.bed/.bim/.fam), VCF, Illumina final report | State the array and genome build |
| Clinical / tabular | CSV, XLSX, Stata (.dta), SPSS (.sav), REDCap export | Send the data dictionary with it |
| Reference files | FASTA, GTF, GFF3 | Only if you need a non-standard reference or annotation |
Metadata We Need
The sample sheet
Every transfer must include a sample sheet as CSV or XLSX with one row per sample. Without it we cannot begin, because file names alone rarely tell us which sample belongs to which group.
- Sample ID — exactly matching the file names, no spaces or special characters
- File names — the exact files belonging to that sample, including R1 and R2
- Group or condition — the comparison you want to make
- Batch information — sequencing run, extraction date, plate, chip position
- Covariates — age, sex, site, timepoint, or anything else that may confound
- Collection date and location — required for surveillance and phylogenetic work
Study information
Alongside the sample sheet, tell us:
- The organism and reference genome build you expect us to use
- The library preparation kit and sequencing platform
- Whether the data is stranded, paired-end, and the read length
- Any samples you already know are problematic, and why
- The research question, restated in one or two sentences
Naming conventions
Use only letters, numbers, hyphens and underscores in file names. Avoid spaces, brackets, ampersands and non-ASCII characters — they break pipelines silently. Keep the sample identifier at the start of the file name and consistent across all files belonging to that sample.
Checksums
Generate MD5 or SHA-256 checksums before upload and send them with the data. Large transfers do occasionally corrupt, and a checksum mismatch caught on day one is far cheaper than a result you cannot reproduce three weeks later.
Data Security and Human Subjects
Where your data is stored
Data is held on Google Cloud with encryption in transit and at rest, in a project-specific bucket. Access is limited to your named analyst and the reviewing consultant, and every access is logged. We do not store client research data on personal devices.
De-identification
Please de-identify data before transfer wherever the analysis allows it. Replace names, hospital numbers and national identifiers with study codes, and keep the linking key at your own institution — we do not need it and prefer not to hold it.
Where identifiable data is genuinely unavoidable for the analysis, the handling terms are set out explicitly in the project agreement before transfer.
Ethical approval
For projects involving human subjects we ask for the approval reference and approving committee, and a copy of the approval letter. This is not bureaucracy for its own sake: journals ask for it, and it is easier to produce at the start than during revision.
What we will not do
We do not reuse client data for other projects, for method development or for training machine learning models without written permission. We do not share data between clients. We do not deposit data in public repositories on your behalf unless you ask us to and the consent framework allows it.
Common Problems and How to Avoid Them
These are the issues that most often delay a project by a week or more:
- Missing R2 files. Paired-end datasets arriving with only forward reads — check file counts before uploading
- Sample sheet does not match file names. Trailing spaces and inconsistent capitalisation are the usual culprits
- Unknown genome build. Data aligned to an unstated reference cannot be safely combined with anything else
- Normalised values sent as raw counts. Differential expression tools require raw counts; TPM will produce wrong results silently
- No batch information. If we cannot model batch, we cannot rule it out as the explanation for your finding
- Zipped archives of zipped archives. Send files as they are; nested archives slow verification and sometimes corrupt
