Donor genetics, age, sex, anatomical location, disease history, agonal state, post-mortem interval, dissociation chemistry, capture platform, and sequencing depth can all modify the observed transcriptome before computational analysis begins.
The practical consequence is direct: two cells assigned to the same retinal population may not occupy the same molecular state, while cells separated by a clustering algorithm may differ because of handling history rather than biology. In ocular tissue, this distinction is especially unstable. The retina, retinal pigment epithelium, choroid, and adjacent posterior-segment structures do not degrade at the same rate, do not tolerate dissociation equally, and do not contribute comparable amounts of RNA or nuclear material.
A human retinal atlas can therefore become more than a map of cell types. It can become a map of procurement conditions unless the upstream variance stack is recorded and modeled.
The dual-stack variance model: biology versus technical noise
The first analytical failure is usually classification. A batch is treated as a technical container, while the donor is treated as a biological replicate. In ocular multiomics, those categories often overlap.
A donor may contribute one eye to one processing batch and another eye to a different batch. Diseased tissue may be available only after a longer post-mortem interval. A specific retinal subregion may be dissected by one operator, while another operator handles a different anatomical region. A capture platform may be used for one cohort and a different platform for another. Once these variables align, the statistical model cannot cleanly determine whether an expression shift reflects disease, ancestry, cell state, tissue location, or processing history.
The variance stack has two main layers:
| Variance layer | Typical components | Primary effect on the dataset |
|---|---|---|
| Biological variance | Age, sex, donor genetics, disease, anatomical subregion, agonal history | Changes genuine cell-state distributions and gene expression |
| Technical variance | Post-mortem interval, dissociation, capture chemistry, ambient RNA, sequencing depth | Alters cell recovery, transcript abundance, and apparent cell identity |
| Confounded variance | Biological and technical variables aligned within the same donor or batch | Makes causal interpretation unstable |
| Computational variance | Normalization, feature selection, integration, clustering resolution | Changes the representation and boundaries of cell populations |
This is not a semantic distinction. It determines whether a donor-specific transcriptomic signature is retained as signal or removed as batch effect.
For example, if all control eyes are processed rapidly while all AMD eyes experience longer procurement latency, the resulting expression difference may contain disease-associated biology and post-mortem degradation at the same time. If the same disease group is also enriched for older donors, age becomes a third aligned factor. A correction algorithm may compress the combined signal, but it cannot reconstruct the counterfactual dataset that was never collected.
Donor is not an interchangeable replicate
In many tissues, a large number of cells can create the impression of strong statistical power. Ocular scRNA-seq makes this impression dangerous. Tens of thousands of cells from one donor do not replace independent donor-level replication.
Cells from the same retina share:
- Genetic background.
- Systemic disease context.
- Medication exposure and agonal history.
- Procurement latency.
- Temperature trajectory.
- Dissection and dissociation conditions.
- Anatomical sampling decisions.
The effective replication unit is therefore closer to the donor than to the individual cell. A dataset with approximately 290,000 viable cells from 18 fresh living and post-mortem retinal specimens can provide substantial cellular resolution, but its biological interpretation still depends on how those specimens are distributed across age, sex, population background, disease status, retinal region, and processing conditions.
The same principle applies to high-throughput single-nucleus work. A posterior-segment study identifying 37 transcriptomically distinct cell types from approximately 151,000 nuclei across six extraretinal structures provides broad cellular coverage. It does not eliminate donor variability. It gives the field more opportunities to observe where that variability enters.
Cell count increases resolution. It does not erase donor structure.
The confounding problem in retinal clustering
Retinal cell type clustering is often described as if the central task were selecting the right algorithm. In practice, the more difficult task is deciding which variation the algorithm is allowed to preserve.
A cluster may be driven by:
1. A stable lineage program.
2. A transient physiological state.
3. A disease-associated response.
4. Regional specialization.
5. Differential survival during dissociation.
6. Ambient RNA contamination.
7. Post-mortem transcript decay.
8. Unequal sequencing depth.
9. Donor-specific expression.
10. A combination of several of these factors.
The resulting cluster boundaries are operational. They describe the structure visible in the measured matrix, not necessarily the full biological structure of the tissue.
This is particularly relevant for rare retinal interneuron sub-clusters. Population-specific expression differences may reflect genetic ancestry, but they may also reflect local microenvironmental history, disease exposure, handling, or unmeasured donor characteristics. Current evidence does not support collapsing those possibilities into a single explanation.
Post-mortem interval and the kinetics of transcriptomic decay
Post-mortem interval is often reduced to a single timestamp: death to preservation. That timestamp is necessary but not sufficient. The tissue experiences a trajectory, not a binary condition.
Temperature, ischemia, residual enzymatic activity, tissue exposure, dissection order, and preservation chemistry all interact with the interval. Their effects are also cell-type-specific. A transcript that remains stable in one population may decline rapidly in another. A membrane-bound cell may be lost during dissociation before its RNA quality is measured. A nucleus may remain recoverable while the corresponding whole-cell transcriptome is no longer representative.
The known post-mortem-dependent downregulation of MALAT1 in rod photoreceptors illustrates the problem. A reduction in a commonly detected transcript may be interpreted as a state shift, lower cellular activity, or a photoreceptor-specific disease signature. If the reduction tracks post-mortem interval, the expression change is instead part of the degradation kinetics.
The target death-to-preservation interval used in procurement planning is six hours. This should be treated as a quality objective, not as a universal boundary between valid and invalid material. The exact interval at which cell-state identification becomes irreversibly corrupted is not established for every ocular cell subtype.
PMI affects cell populations unevenly
A retina is not a homogeneous RNA reservoir. Photoreceptors, ganglion cells, Müller glia, vascular cells, RPE, and choroidal populations have distinct structural and biochemical properties. Their recovery after death and during processing can diverge.
This creates several interpretive hazards:
- A population may appear depleted because its cells are more fragile.
- A transcriptomic state may appear enriched because resistant cells survive preferentially.
- A stress-response program may reflect procurement latency rather than pathology.
- A lower gene count may be mistaken for a low-complexity biological state.
- A donor may show altered cell proportions because tissue layers were not sampled equivalently.
The issue is not simply that longer PMI produces “lower quality.” That description is too coarse for multiomics. The relevant question is which transcripts, nuclei, cells, and tissue compartments are preferentially affected at each stage of the interval.
For bulk RNA-seq, degradation can alter the aggregate expression profile. For scRNA-seq, it can additionally change which cells enter the library. For snRNA-seq, nuclei may improve recovery from difficult or partially degraded tissue, but nuclear RNA is not a direct substitute for the whole-cell transcriptome. It emphasizes a different molecular compartment and therefore requires a different interpretation of expression programs.
Procurement metadata must be treated as analytical input
PMI belongs in the data model beside donor age, disease status, and anatomical region. It should not remain in a free-text field inside a biobank record.
A usable ocular procurement record should distinguish at least:
- Time of death or the best available estimate.
- Time of enucleation or tissue recovery.
- Time tissue reached controlled preservation.
- Temperature conditions during transport.
- Time of dissection.
- Retinal or posterior-segment region sampled.
- Preservation method.
- Whole-cell versus nuclei workflow.
- Operator and processing batch.
- Library preparation chemistry.
- Sequencing depth and platform.
These fields allow downstream analysts to model degradation as a continuous process rather than as an afterthought. They also make it possible to exclude or stratify samples without discarding the entire cohort.
Dissociation sensitivity and tissue-specific survival bias
The dissociation protocol is not a neutral conversion step. It is a selection process.
Enzymatic exposure, mechanical force, incubation time, temperature, and tissue fragmentation determine which cells remain intact, which cells release ambient RNA, and which populations fail before capture. The effect is not uniform across the posterior segment. Choroid and RPE are particularly sensitive to protocol-dependent changes in cell survival and gene profiles.
This produces a form of selection bias that can be difficult to detect from a final matrix. A low-quality cell may be filtered out. A missing population leaves no equivalent record. If a fragile cell type disappears consistently from one protocol, the analysis may report a clean but incomplete cellular landscape.
Whole-cell scRNA-seq and snRNA-seq capture different failures
Whole-cell scRNA-seq is often favored when cytoplasmic transcripts and detailed cellular states are required. It is also more exposed to tissue integrity, dissociation stress, and the availability of intact cells.
Single-nucleus RNA-seq can extend the usable window for certain specimens and improve access to structurally difficult tissue. However, it changes the measurement target. Nuclear transcripts, retained intronic reads, and reduced cytoplasmic representation alter the feature space. A state detected in whole-cell data may be attenuated in nuclear data, while another program may be more visible.
The two modalities should therefore be integrated with explicit modality-aware expectations. They should not be treated as identical assays with different packaging.
| Workflow | Main strength | Main vulnerability | Interpretation constraint |
|---|---|---|---|
| Whole-cell scRNA-seq | Stronger access to cytoplasmic transcript programs and intact cell states | Sensitive to dissociation, cell fragility, and post-mortem tissue integrity | Cell recovery may be biased toward robust populations |
| snRNA-seq | Useful for difficult, archived, or partially degraded tissue; broad structural access | Nuclear transcriptome differs from whole-cell transcriptome | Expression programs require modality-specific annotation |
| Bulk RNA-seq | High aggregate material and broad molecular coverage | Loses cellular resolution and masks composition changes | Expression shifts may reflect both regulation and cell proportion |
| Multiomic integration | Connects transcriptomic, epigenetic, and proteomic dimensions | Amplifies cross-platform and donor-level confounding | Alignment must preserve known biological differences |
Ambient RNA is not a minor contamination term
The retina contains abundant transcripts from highly represented cell populations. When cells are damaged during dissociation, released RNA can enter the surrounding suspension and be captured by other droplets. A low-abundance cell may then acquire transcripts that do not originate from its own cytoplasm.
This can blur boundaries between neighboring cell types. RPE-associated, photoreceptor-associated, or glial transcripts may appear in unexpected populations, especially when tissue disruption is substantial. The effect is not always obvious in dimensionality reduction. It may present as a subtle shift in marker intensity rather than as a separate contamination cluster.
Ambient RNA correction can reduce the problem, but it cannot restore the original cellular context. The stronger control remains upstream: rapid handling, tissue-specific dissociation optimization, and explicit measurement of suspension quality.
The protocol should be validated by tissue, not only by sample
A single dissociation recipe applied to retina, RPE, choroid, and optic-nerve-adjacent structures assumes equivalent mechanical and enzymatic behavior. That assumption is operationally convenient and biologically weak.
Protocol validation should examine:
- Viability by tissue layer.
- Median genes detected per cell or nucleus.
- Mitochondrial and ribosomal transcript burden.
- Recovery of fragile and rare populations.
- Marker leakage across cell types.
- Fraction of doublets.
- Cell-cycle and stress-response activation.
- Agreement between technical replicates.
- Stability of cell proportions across donors.
The objective is not to maximize the total number of captured cells. It is to maximize the fidelity of the cellular distribution that enters the dataset.
Population-specific signatures: ancestry, environment, or processing?
Human retinal reference atlases are beginning to expose transcriptomic diversity across populations. One Chinese human retina atlas generated from approximately 290,000 viable cells across 18 specimens identified population-specific cellular diversity. That finding matters because reference atlases are frequently used as annotation scaffolds for new disease cohorts.
The risk is straightforward. A population-associated expression program may be retained as meaningful biology in one analysis and removed as a batch effect in another. The correct decision cannot be made from the expression matrix alone if population background is aligned with collection site, tissue quality, age distribution, disease prevalence, or processing workflow.
A donor-specific transcriptomic signature has at least three possible sources:
1. Genetic background. Cis-regulatory variation and other inherited factors alter baseline expression or response thresholds.
2. Microenvironmental history. Systemic disease, medication, nutrition, inflammation, and local ocular conditions influence cell state.
3. Collection structure. Procurement latency, dissection, preservation, and library preparation create technical differences correlated with the donor group.
These sources can coexist. The problem is not solved by selecting one label.
Reference atlases are not neutral coordinate systems
An atlas is a reference only within the limits of its sampling frame. If it overrepresents one population, age band, anatomical region, or preservation workflow, its cell-state definitions may reflect those conditions.
This does not invalidate the atlas. It defines the scope of its use.
When a disease dataset is mapped to a reference, analysts should ask:
- Were the same tissue regions sampled?
- Were whole cells or nuclei profiled?
- Are the donor populations comparable?
- Is the reference built from fresh living tissue, post-mortem tissue, or both?
- Are disease and control groups balanced for PMI?
- Do rare cell states replicate across independent donors?
- Does the state remain visible when donor identity is modeled explicitly?
Annotation confidence should decrease when a cell state appears only in one donor, one population, one platform, or one procurement condition. The state may be real. It may also be a local artifact.
Inter-donor variance is not automatically unwanted variance
A common analytical objective is to minimize inter-donor gene expression variance. That objective is appropriate only when the variance is known to be technical.
Some donor-level differences are the biology under investigation. This is especially true in glaucoma, inherited retinal dystrophy, age-related macular degeneration, and optic neuropathy, where susceptibility, compensatory response, and progression may vary between individuals.
Removing all donor structure can produce a visually coherent embedding with a biologically flattened result. The integrated manifold may look cleaner while clinically relevant heterogeneity has been compressed.
A corrected embedding is not evidence that the correction was biologically correct.
Beyond computational integration: what batch correction can and cannot recover
Integration tools are useful for representing datasets in a common space. They are not time machines. If post-mortem interval, disease status, and processing chemistry were fully confounded before sequencing, no algorithm can observe the missing combinations.
Harmony, scVI, and related methods can reduce technical separation when the design contains enough overlap between batches and biological groups. They can also help align shared cell populations across platforms or cohorts. But they cannot determine whether a feature is technical or biological without support from the experimental design and metadata.
A robust integration workflow therefore begins before normalization.
The minimum design logic
A stronger ocular scRNA-seq study distributes major variables across processing conditions wherever feasible:
1. Balance disease and control material across batches. Do not assign one disease state to one library-preparation run if the comparison can be avoided.
2. Record PMI continuously. A categorical label such as short or long loses useful degradation information.
3. Preserve anatomical identity. Retina, RPE, choroid, and neighboring structures should not be collapsed into a single tissue field.
4. Retain donor-level identifiers through all analyses. Cells should never become anonymous observations after quality control.
5. Use technical replicates strategically. Replicate processing can separate protocol effects from donor effects.
6. Compare modalities explicitly. Whole-cell and nuclear data should be integrated with known differences in transcript capture.
7. Model cell composition and expression separately. A change in average expression may result from altered cell proportions rather than altered regulation within a cell type.
8. Validate high-value states orthogonally. Imaging, targeted assays, bulk measurements, proteomics, or independent donors can test whether a computationally defined state persists.
This is not an argument against computational correction. It is an argument against assigning it responsibility for upstream design failures.
Pseudobulk remains useful
Cell-level resolution is essential for identifying states, but donor-level summaries remain valuable for inference. Pseudobulk aggregation within a defined cell type can reduce the false precision created by treating thousands of cells from one donor as independent biological replicates.
For each donor and cell population, analysts can summarize counts or expression and then compare donors across conditions. This approach does not solve cell-recovery bias or confounding. It does, however, align the statistical unit more closely with the experimental unit.
The choice of aggregation strategy matters. If a cell type is rare or differentially recovered across donors, its pseudobulk profile may reflect both true expression and sampling instability. Cell number, recovery fraction, and quality metrics should remain visible alongside the aggregated expression matrix.
Quality control should describe failure modes
A single pass-fail threshold is insufficient for ocular tissue. Filtering out cells with low gene counts may remove damaged cells, but it may also remove biologically small or fragile populations. Filtering by mitochondrial content may eliminate degraded cells while also disproportionately affecting certain tissues.
Quality control should therefore be layered:
- Cell or nucleus viability and integrity.
- Gene and transcript complexity.
- Mitochondrial and ribosomal burden.
- Doublet probability.
- Ambient RNA burden.
- Donor and batch representation.
- Tissue-specific recovery.
- Marker consistency.
- Sensitivity of clusters to threshold changes.
A cluster that disappears when the mitochondrial threshold moves slightly is not necessarily false. It is unstable. That instability should be reported rather than hidden behind a single selected parameter.
A procurement-to-analysis data pipeline
Ocular multiomics is often described as a sequencing problem. The more accurate description is a chain-of-custody problem with a sequencing endpoint.
Every transition can add variance:
- Donor identification.
- Consent and eligibility screening.
- Death notification.
- Eye recovery.
- Transport.
- Temperature control.
- Dissection.
- Tissue partitioning.
- Dissociation or nuclei isolation.
- Library preparation.
- Sequencing.
- Cell filtering.
- Integration.
- Biological interpretation.
The pipeline should preserve provenance across all transitions. A molecular profile without a procurement record has limited interpretive value. Conversely, a well-documented specimen can remain useful even when one modality underperforms, because the failure mode is known and can be incorporated into study design.
A practical data model should connect:
- Donor metadata.
- Eye-level metadata.
- Tissue-region metadata.
- Aliquot and preservation records.
- Processing events.
- Library identifiers.
- Sequencing runs.
- Quality-control outputs.
- Cell-level and donor-level analytical objects.
The key is not the number of fields. It is the ability to trace a signal backward. If a photoreceptor program is reduced, the system should allow the analyst to determine whether the effect is associated with disease, donor group, PMI, tissue region, library depth, or cell recovery.
Without that traceability, the dataset becomes a static matrix. With it, the dataset becomes an auditable biological process.
The operational standard for interpreting divergent cell states
When cell states diverge between donors, the first question should not be whether one donor is an outlier. The first question should be which variance stack explains the divergence.
A disciplined interpretation sequence is:
1. Verify tissue comparability. Confirm that the same anatomical compartment and sampling depth were used.
2. Inspect procurement latency. Compare PMI, preservation timing, and transport conditions.
3. Check modality and chemistry. Whole-cell and nuclear libraries can produce different state visibility.
4. Review cell recovery. A shift in state frequency may result from selective survival.
5. Separate donor from cell effects. Reanalyze with donor-aware models or pseudobulk summaries.
6. Test integration sensitivity. Examine whether the state persists under alternative normalization and correction parameters.
7. Assess marker specificity. Confirm that the signature is not driven by ambient RNA or a small number of unstable genes.
8. Seek orthogonal support. Use another molecular layer or an independent specimen set where available.
9. Report unresolved ambiguity. Do not convert an underdetermined result into a definitive biological claim.
This sequence is slower than producing a single integrated embedding. It is also more compatible with translational use.
What the field can conclude now
The evidence supports a clear but bounded assessment.
Donor variability in ocular single-cell RNA-seq is generated by both biological and technical factors. Post-mortem interval contributes cell-type-specific degradation kinetics, including measurable effects on rod photoreceptor transcript profiles. Dissociation sensitivity can change survival and gene expression in tissues such as the choroid and RPE. Population-specific transcriptomic diversity is detectable in human retinal reference data, but its relationship to genetic ancestry and donor microenvironment remains incompletely resolved.
Large cell counts and sophisticated integration improve resolution. They do not substitute for balanced procurement, explicit metadata, tissue-specific protocol validation, or donor-aware statistics. Computational correction can reduce unwanted structure when the study design contains the overlap needed to identify it. It cannot reliably recover biological signals that were confounded with handling conditions before sequencing.
The strict operational conclusion is therefore simple: cell-state divergence should be treated as a property of the entire ocular biospecimen pipeline, not as a defect discovered at the clustering stage. A trustworthy retinal atlas begins with procurement latency, tissue identity, and chain-of-custody metadata. Sequencing reveals the consequences. It does not retroactively control them.
