EV Biomarker discovery
A biomarker program arrives with forty plasma samples, a panel of surface protein targets, and a plan to find the markers separating cases from controls. The assay works. The samples are good. And the study cannot produce a result that will survive a second cohort. Not because anything went wrong at the bench, but because the sample size was fixed before anyone counted how many hypotheses the panel would test. That ordering is backwards, and it is the most common reason promising candidates evaporate on replication.
How large does an EV biomarker discovery cohort need to be?
It depends on how many markers you measure. Cohort size and panel breadth are one decision, not two. A twelve-marker panel needs more samples than a three-marker panel to clear the same false-discovery threshold at equal power. Fix the panel first, then size the cohort to it.
Key takeaways
- Panel breadth and cohort size are a single joint design decision. Adding markers after the sample size is fixed lowers your effective power without changing anything you can see at the bench.
- Median statistical power in neuroscience research has been estimated at roughly 8–31%, and low-powered studies systematically overstate effect size [9].
- Effect sizes reported in highly cited biomarker papers are consistently larger than the same associations measured in subsequent meta-analyses [4]. Powering a study on a published effect size therefore under-powers it.
- Retrospective biobank cohorts carry fixed aliquot volumes and undocumented storage histories. Long-term frozen storage measurably reduces recoverable vesicle counts and plasma proteome coverage [6].
- Split the cohort into discovery and replication sets before running a single plate, not after a candidate list appears.
The forty-sample pilot and what it can honestly tell you
Small discovery cohorts are not simply weaker versions of large ones. They fail in a specific and misleading way: they inflate the effect sizes of whatever they do detect.
The mechanism is selection. With a small sample, only large apparent differences clear the significance threshold, so the survivors are systematically those exaggerated by sampling variation. An analysis of statistical power across the neurosciences estimated median power between 8% and 31%, and identified overestimated effect size as a direct consequence rather than a side effect [9]. The result looks like a finding. It behaves like noise.
This is measurable in the biomarker literature specifically. A comparison of effect sizes reported in highly cited individual biomarker papers against the same associations in later meta-analyses found the initial reports were consistently larger [4]. That has a direct and uncomfortable consequence for study design: if you power a new study on an effect size taken from a published paper, you are almost certainly powering it on an inflated number, and your study will be under-powered by construction.
Fluid biomarker work in neurodegeneration has examined its own reproducibility problem in some depth. A review of why findings fail to replicate identified cohort composition, sample handling, and analytical variation as interacting sources, and argued that reproducibility has to be designed in at recruitment rather than assessed afterwards [5]. Cohort design is not a preliminary step before the interesting part. It is the part that determines whether the interesting part means anything.
None of this argues against pilot studies. A forty-sample run is a reasonable way to confirm that an assay behaves on your matrix and that dilution ranges are sensible. It is not a way to identify biomarkers. Conflating the two is how a technical validation exercise gets written up as a discovery result.
Every marker you add raises the bar
Here is where panel breadth stops being free.
Testing one marker at a conventional significance level of 0.05 means accepting a one-in-twenty chance of a false positive. Test twenty markers and, absent correction, you should expect roughly one spurious hit even when nothing real is present. The standard remedy is to control the false discovery rate — the expected proportion of declared discoveries that are actually false — rather than the probability of any single error [8]. False discovery rate control is less punishing than Bonferroni correction and is now the default in panel-based work.
It is not free either. Controlling the false discovery rate across a panel means each individual marker faces a stricter effective threshold as the panel grows. To keep the same power to detect a real effect, the sample size has to rise. This is the arithmetic that panel expansion runs into: a marker added to the panel late in planning does not just add a measurement, it raises the bar for every other marker in the run.
Large-scale plasma proteomics shows what the corrected end of this spectrum looks like in practice. The UK Biobank Pharma Proteomics Project profiled 2,923 proteins across 54,219 participants and reported 14,287 primary genetic associations [10]. That scale is what makes a several-thousand-target panel statistically tractable. A targeted panel of eight to twenty EV surface markers sits at a very different point — but the same arithmetic governs it, just with smaller numbers on both sides.
The practical move is to decide the panel first and treat it as fixed. Then size the cohort against it. A design in which the panel grows through the planning phase while the sample size stays where it started is a design losing power at every step, invisibly, with no signal at the bench that anything is wrong. The corollary is that there is no statistical saving in running markers one at a time instead — sequential single-analyte testing carries its own costs in sample volume and batch structure.
Figure 1. How required cohort size scales with panel breadth
| Markers in panel | Effective per-marker threshold (BH, FDR 0.05) | Relative sample size needed for 80% power |
|---|---|---|
| 1 | 0.050 | 1.0× (reference) |
| 5 | 0.010 | ~1.4× |
| 12 | 0.004 | ~1.6× |
| 25 | 0.002 | ~1.8× |
| 50 | 0.001 | ~2.0× |
Illustrative. Thresholds shown are the most stringent rung of the Benjamini-Hochberg step-up procedure [8], which is the relevant case when only one or two markers are truly associated. Relative sample sizes are order-of-magnitude illustrations for a two-group comparison at fixed effect size, not values from a specific dataset. The point is the shape of the curve: cost rises steeply at first, then flattens.
That flattening matters, and it cuts in a direction people do not expect. Going from one marker to five costs you roughly 40% more samples. Going from twelve to fifty costs you another 25%. If you are already committed to a multiplex panel, adding markers is comparatively cheap in statistical terms — the expensive step was leaving single-marker territory in the first place. Panels should be designed generously at the outset rather than expanded incrementally.
The samples you actually have, not the ones you would design
Most large-population discovery work in this field does not run on prospectively collected material. It runs on biobank plasma: aliquots drawn years ago, for a different question, under protocols that were reasonable at the time and are documented to varying degrees.
Three constraints follow, and none of them can be assayed away.
Volume is fixed and often small. A biobank aliquot is what it is. This is the constraint that most directly limits panel breadth: with 200 µL and an assay requiring 100 µL per run, you get one attempt and no repeats. Methods that consume material in an upfront isolation step spend part of that volume before any measurement happens. A workflow that measures multiple surface markers from a single small-volume input changes what is possible from a fixed archive — this is the practical case for panel-based approaches like the LuminEV Research Kit in retrospective study designs, rather than any claim about detection limits.
Storage history is a variable, whether or not it was recorded. A systematic study of pre-analytical factors in plasma EV proteomics found that long-term frozen storage progressively reduced proteome coverage, with lower recoverable vesicle counts and altered size distributions, and explicitly cautioned against treating long-archived biobank material as equivalent to fresh [6]. The same work found that high-speed centrifugation of plasma at 8,000 × g before analysis cut both vesicle and protein numbers to around 30% of the original — a processing decision, made once, years ago, that no downstream assay can reverse.
Freeze-thaw is less catastrophic than commonly assumed, but not neutral. A systematic comparison of isolation methods on freshly frozen versus freeze-thawed plasma found that a single freeze-thaw cycle had minimal impact on vesicle integrity, morphology, or protein composition [11]. That is genuinely reassuring for biobank work. It is also specifically about one cycle. Multiple cycles and multi-year storage are different questions with less comfortable answers.
The same applies to the mechanics of measuring surface proteins in plasma, and the pattern generalizes beyond vesicles. When the Standardization of Alzheimer’s Blood Biomarkers working group empirically tested variations in blood collection and handling across a panel of blood-based markers, collection tube type alone produced different values for every marker assessed, and delayed centrifugation affected a subset [7]. Pre-analytical variation is not a small correction term. It is frequently larger than the biological difference you are hunting, and it sits upstream of everything you can fix later through assay precision across plates and sites.
Reporting infrastructure exists for exactly this. The ISEV Blood EV Task Force developed a minimal information framework for blood EV research, noting that hundreds of pre-analytical protocols and more than forty distinct variables are in circulation, and providing a structure for recording which were used [2]. The broader field guidance covers pre-processing variables, separation, and characterization requirements [1]. Neither will repair a cohort assembled without them. Both will tell you, before you start, whether the archive you have been offered can answer the question you are asking.
Figure 2. Biobank cohort constraints and their design consequences
| Constraint | Typical situation | What it forces in the design |
|---|---|---|
| Aliquot volume | Fixed, often 100–500 µL, no repeat draws | Caps panel breadth and eliminates re-runs; favours low-input, multi-marker methods |
| Storage duration | Variable across the cohort; years in some cases | Include storage time as a covariate; avoid confounding it with case status |
| Freeze-thaw count | Often undocumented | Verify records exist before committing; single cycle is tolerable, unknown counts are not |
| Collection tube and processing | May differ across contributing sites | Block or stratify by site; never let site correlate with group |
| Metadata completeness | Partial, especially for older material | Determines which covariates can be modelled at all |
Compiled from published pre-analytical guidance and study findings [1][2][6][7][11].
Note the recurring theme in the right-hand column: the danger is not variability itself, it is variability that lines up with your comparison. A cohort where cases came from one site and controls from another will produce real, reproducible differences. They will be about the sites.
Plan the replication cohort before you run the discovery cohort
The last structural decision is the one most often deferred: where the confirmation comes from.
Deciding this after a candidate list exists is a weaker position than deciding it beforehand, because by then the candidates have already shaped what counts as confirmation. The pivotal-evaluation framework for biomarker studies addresses this directly, proposing a design in which specimens are collected prospectively and stored before outcomes are known, then retrieved and assayed under blinding [3]. Two features of that framework transfer usefully to discovery-stage work even when a full pivotal design is out of scope. Cases and controls should be drawn randomly from a defined cohort rather than hand-selected, because selecting well-characterized cases and unusually healthy controls produces performance estimates that will not hold in a real population. And performance criteria should be specified in advance, so the study has a stated bar rather than whatever the data eventually clears.
Blinding deserves a note of its own, because in retrospective work it is easy to skip. If the person running the plates knows which samples are cases, that knowledge reaches the result through unconscious routes: which outliers look worth re-running, which plate gets repeated. Randomising sample position and withholding group labels until the data are locked costs nothing and removes the whole category of problem.
In practice, for a retrospective archive, that means partitioning the samples at the start. Assign a discovery set and hold back a replication set, and do not look at the second one. It costs statistical power in the discovery phase, which feels expensive and is the reason people resist it. What it buys is a candidate list with an independent confirmation attached, which is what determines whether the program advances.
What to do next
If you are designing a large-population discovery study on plasma vesicle surface markers, the sequence matters more than any individual choice:
- Fix the panel. Decide what you are measuring and stop adding to it, informed by what vesicle-level measurement adds for specific analytes. Design it generously — the statistical cost of breadth flattens quickly.
- Establish the effect size you are powering against, and discount published values rather than adopting them directly [4].
- Size the cohort against the fixed panel at your chosen false discovery rate and power, not against convenience or budget.
- Audit the archive before committing. Volume per aliquot, storage duration, freeze-thaw records, collection site, tube type. Check whether any of these correlate with case status [2].
- Partition discovery and replication sets up front, and keep the replication set closed [3].
- Confirm your assay input volume fits the smallest aliquot in the cohort, with margin. A workflow that needs less material per marker gives you more panel for the same archive; this is the specific constraint LuminEV was designed around.
The uncomfortable implication is that a well-designed discovery study is often larger and narrower than the one people want to run. That trade is worth making. The field’s replication problem is not primarily a measurement problem — assay performance has improved considerably while replication rates have not followed. What has changed less is study design, and design is the cheaper thing to fix — which is consistent with where the field’s binding constraints now sit. As pre-analytical reporting frameworks mature and biobanks begin capturing the variables that matter [2], the constraint is shifting from what we can measure to whether we asked a question the samples could answer.
References
- Welsh JA, Goberdhan DCI, O’Driscoll L, et al. Minimal information for studies of extracellular vesicles (MISEV2023): From basic to advanced approaches. Journal of Extracellular Vesicles. 2024;13(2):e12404. doi:10.1002/jev2.12404
- Lucien F, Gustafson D, Lenassi M, et al. MIBlood-EV: Minimal information to enhance the quality and reproducibility of blood extracellular vesicle research. Journal of Extracellular Vesicles. 2023;12(12):e12385. doi:10.1002/jev2.12385
- Pepe MS, Feng Z, Janes H, Bossuyt PM, Potter JD. Pivotal evaluation of the accuracy of a biomarker used for classification or prediction: standards for study design. Journal of the National Cancer Institute. 2008;100(20):1432–1438. doi:10.1093/jnci/djn326
- Ioannidis JPA, Panagiotou OA. Comparison of effect sizes associated with biomarkers reported in highly cited individual articles and in subsequent meta-analyses. JAMA. 2011;305(21):2200–2210. doi:10.1001/jama.2011.713
- Mattsson-Carlgren N, Palmqvist S, Blennow K, Hansson O. Increasing the reproducibility of fluid biomarker studies in neurodegenerative studies. Nature Communications. 2020;11(1):6252. doi:10.1038/s41467-020-19957-6
- Suresh PS, Zhang Q. Impact of preanalytical factors on plasma extracellular vesicles and human plasma proteome. Clinical Proteomics. 2026;23(1). doi:10.1186/s12014-026-09590-8
- Verberk IMW, Misdorp EO, Koelewijn J, et al. Characterization of pre-analytical sample handling effects on a panel of Alzheimer’s disease-related blood-based biomarkers: Results from the Standardization of Alzheimer’s Blood Biomarkers (SABB) working group. Alzheimer’s & Dementia. 2022;18(8):1484–1497. doi:10.1002/alz.12510
- Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B. 1995;57(1):289–300. doi:10.1111/j.2517-6161.1995.tb02031.x
- Button KS, Ioannidis JPA, Mokrysz C, et al. Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience. 2013;14(5):365–376. doi:10.1038/nrn3475
- Sun BB, Chiou J, Traylor M, et al. Plasma proteomic associations with genetics and health in the UK Biobank. Nature. 2023;622(7982):329–338. doi:10.1038/s41586-023-06592-6
- Li X, Li X, Tong L, et al. Systematic evaluation of isolation techniques and freeze-thaw effects on plasma extracellular vesicle heterogeneity and subpopulation profiling. Journal of Extracellular Biology. 2025;4(6):e70058. doi:10.1002/jex2.70058



