A lab identifies an extracellular vesicle surface protein that separates patients from controls. The effect size is large, the p-value is convincing, the figure is clean. Three years later, no one has replicated it in a second cohort, and the assay has never been run at a second site. This is the ordinary outcome, not the exception — and the reason usually has nothing to do with whether the biology was real. It has to do with how the measurement was made.
Key takeaways
- The bench-to-clinic gap in EV biomarkers is primarily analytical. Candidates discovered under conditions that cannot be reproduced across labs, sites, or sample volumes do not survive validation, regardless of the strength of the original signal.
- Reported inter-laboratory variability in EV measurements has been described as reaching 94%, which makes cross-study comparison — the foundation of validation — unreliable by default [1].
- Standards exist. MISEV2023 covers production, separation, and characterization; MIBlood-EV covers preanalytical reporting for blood-derived EVs. Adherence, not availability, is the limiting factor [2,3].
- A candidate’s context of use should be defined before discovery finishes, because context determines which analytical characteristics matter.
- Four practical gates decide whether an EV assay can enter a trial: serial sampling stability, low-volume compatibility, multi-site throughput, and matrix tolerance.
The gap is a measurement problem wearing a biology costume
EV research has grown for two decades on a genuinely strong premise: vesicles carry cargo that reflects the state of the cell that released them, which makes them a plausible route to reading tissue biology from blood [2]. The premise has held up. The translational record has not kept pace with it.
The usual explanation — that EV biology is too heterogeneous to yield clean biomarkers — is only partly right, and it obscures the more actionable problem. EVs occupy an awkward analytical space: smaller than cells, larger than individual proteins, and present in a background of highly abundant non-vesicular material. That combination means most established analytical methods were not designed for them, and the workarounds labs adopt during discovery tend to be labor-intensive, operator-dependent, and difficult to specify tightly enough for another lab to reproduce [4].
The consequence is measurable. A survey referenced in a recent liquid biopsy review put inter-laboratory variability in EV measurements as high as 94% [1]. Take that number seriously for a moment. If two labs measuring the same analyte in the same matrix can differ by that margin, then a discovery cohort and a validation cohort run at different sites are not really testing the same hypothesis. The validation study is testing the assay and the biology at once, and when it fails, the field usually blames the biology.
Where the signal actually gets lost
Three failure points recur, and none of them are biological:
Preanalytical drift. Blood collection tube type, time to processing, centrifugation protocol, freeze-thaw history, and platelet contamination all shift EV measurements. When these variables go unreported — which is common — a result cannot be reproduced even by a lab that wants to. The MIBlood-EV initiative was created specifically to standardize reporting of these variables for plasma and serum, on the reasoning that reproducibility in liquid biopsy work depends on the steps that happen before the assay ever runs [3].
Ambiguous attribution. If a workflow does not adequately separate vesicles from co-isolated non-EV material, a measured protein cannot be confidently assigned to EVs at all. This is the problem MISEV2023 addresses most directly, and incomplete adherence to its expectations continues to weaken confidence in reported EV specificity — which in turn creates obstacles at analytical validation and regulatory review [1,2].
Nomenclature slippage. “Exosome,” “small EV,” and “microvesicle” are still used inconsistently across the literature. Two papers reporting the same marker in the same disease may not be describing the same particle population, which makes meta-analysis and evidence-building unreliable [5].
Figure 1 makes the mismatch concrete.
Figure 1. Discovery-phase conditions versus trial-phase requirements. The columns rarely match, and every mismatch is a point where a candidate biomarker fails validation for non-biological reasons. Illustrative, based on constraints described in [1,3,4].
| Variable | Typical discovery setting | What a clinical study requires |
|---|---|---|
| Sample volume | 2–5 mL plasma, sometimes more | 0.5–1 mL, often shared across multiple assays |
| Cohort size | 20–60 samples, single site | Hundreds to thousands, multi-site |
| Isolation | Ultracentrifugation or gradient, hands-on, operator-sensitive | Specified, transferable, documented workflow |
| Preanalytical control | Convenience sampling, variable handling | Fixed protocol, reported per MIBlood-EV-style framework |
| Timepoints | Cross-sectional, single draw | Serial draws over 6–18 months per participant |
| Readout | One analyte per run | Multiple analytes from one aliquot |
| Success criterion | Group separation | Predefined performance against a stated context of use |
Standards are not the bottleneck — adoption is
It is tempting to describe EV translation as a field waiting for guidelines. It is not. MISEV has been revised three times since 2014, and MISEV2023 was compiled from ISEV task force input and feedback from more than a thousand researchers, covering production, separation, and characterization across cell culture, biofluids, and solid tissue, plus newer sections on EV release, uptake, and in vivo work [2]. Its stated purpose is rigor, reproducibility, and transparency, and its authors were explicit that it functions as informed guidance rather than a rigid checklist [2,6].
Complementary frameworks have followed. MIBlood-EV targets preanalytical reporting for blood-derived EVs [3]. Reviews in adjacent subfields have argued that following MISEV is not a compliance exercise but a practical route to results other labs can build on [7]. Work on single-vesicle methods has reached the same conclusion from a different direction: without agreed protocols for isolation, characterization, and interpretation, cross-study comparison and regulatory acceptance both stall [8].
So the guidance exists and the field broadly agrees with it. What remains is the harder problem: many discovery-stage workflows physically cannot meet the standard at trial scale. A protocol requiring 4 mL of plasma and a day of hands-on ultracentrifugation per sample is not going to be run across 600 participants at eight sites, no matter how well it is documented. Standardization and practicality have to be solved together, which is why method choice at discovery is a translational decision, not just a technical one.
Define context of use before discovery finishes
The single change with the largest effect on translational odds is also the least technical: state what the biomarker is for, in specific terms, before the discovery phase ends.
Regulatory frameworks have made this vocabulary precise. The FDA–NIH BEST resource distinguishes diagnostic, prognostic, predictive, monitoring, pharmacodynamic/response, and safety biomarkers, and pairs each with a context of use — the specific stated purpose and setting in which the measurement is qualified [9]. These categories are not bureaucratic labels. They determine which analytical properties matter.
Consider two candidates built on the same marker:
- As a diagnostic, the assay needs a defined cutoff, established specificity against relevant confounders, and performance in the population where it would be used at presentation.
- As a pharmacodynamic marker, the cutoff matters less than within-subject stability over time. What the assay must demonstrate is that a change measured across serial draws reflects a real change in biology rather than assay drift, handling variation, or seasonal shifts in sample processing.
A team that never chooses between these ends up optimizing for neither. They report group separation — the discovery-stage default — and then discover at validation that no one asked whether the assay could detect a within-subject change of the size a drug would plausibly produce.
Four gates a translational EV assay has to pass
Context of use sets the target. These four constraints decide whether an assay can reach it. All four are knowable at the start of a program, and all four are cheaper to design for than to retrofit.
Serial sampling stability. Monitoring and pharmacodynamic applications require repeated measurement in the same participant over months. The relevant performance figure is within-subject precision across runs, batches, and operators — not the between-group difference from the discovery paper. If assay-to-assay variation is the same size as the biological change being tracked, the biomarker cannot function as an endpoint.
Low sample volume. Clinical protocols allocate plasma across many assays, and neurology cohorts in particular are volume-constrained. An assay needing several milliliters competes with every other planned measurement and usually loses. Sub-milliliter compatibility is close to a precondition for inclusion in a trial panel.
Multi-site throughput. Validation requires large, prospective, multi-center cohorts, and that requirement is only satisfiable when the assay runs consistently in more than one pair of hands [1]. Plate-based formats with defined hands-on time transfer between sites; extended manual isolation workflows do not, which is why isolation-free approaches have become a translational consideration rather than a convenience. This is the constraint the LuminEV Research Kit was built around — multiplexed measurement of EV surface proteins directly in plasma or culture media, in a 96-well format, without a prior isolation step.
Matrix tolerance. Plasma is a difficult background. Lipoproteins, protein aggregates, and platelet-derived material all co-purify or interfere, and the degree of interference varies with collection and handling. An assay that performs well on purified vesicles from culture supernatant may behave differently on clinical plasma, and that difference has to be characterized rather than assumed [4].
Figure 3. The four gates, with the question each one answers. Illustrative framework.
| Gate | The question | Fails when |
|---|---|---|
| Serial stability | Can it detect change within a participant? | Within-subject CV approaches the expected effect size |
| Volume | Can it run on what the protocol allocates? | Requires > 1 mL of a shared aliquot |
| Throughput | Can a second site produce the same number? | Hands-on isolation dominates the workflow |
| Matrix | Does it hold up in clinical plasma? | Only characterized in purified or cultured material |
What crossing the gap looks like in practice
The abstract argument is easier to accept with a concrete instance, so here is one from ALS.
TDP-43 is a defining pathological feature of ALS, present in the large majority of cases, and it has been a difficult target to demonstrate engagement against in living patients. A blood-based measure of neuronal TDP-43 is therefore useful in a specific, stated way: not as a diagnostic, but as a pharmacodynamic marker capable of showing whether a treatment changes the underlying biology.
The PARADIGM trial (NCT05357950) was a randomized, double-blind, placebo-controlled Phase 2b study whose prespecified primary biomarker outcome was plasma neuron-derived exosomal TDP-43 or prostaglandin J2 — with results published in JAMA Neurology in 2026 and the prespecified EV analyses reported separately following completion of assay development [10]. That structure is what distinguishes translational work from discovery work. The biomarker was named in advance, tied to a specific role, and measured on serial samples across every scheduled visit in an interventional trial, rather than being retrofitted onto data after the fact.
Two details are worth extracting for anyone designing an EV biomarker program:
The context of use came first. The marker was positioned as a treatment-response measure, which set the analytical requirement — within-subject change detection across longitudinal draws — before any samples were run.
The assay was fixed before the trial, not during it. Serial samples from all study visits were processed on a single defined EV isolation and measurement workflow. That is the difference between a biomarker result that can be interpreted and one that is confounded with method drift.
It is also worth being precise about what this kind of result does and does not establish. A prespecified biomarker outcome in one Phase 2b study is evidence that an EV-based measurement can function as a trial endpoint. It is not qualification, and it does not substitute for confirmation in additional studies. But it demonstrates that the translational path is passable when the analytical work is done in the right order.
What to do next
If you are running an EV biomarker program, four decisions carry most of the translational weight, and all four are available at the start:
- Write the context of use down. One sentence naming the biomarker category, the population, and the decision it would inform. Revisit it whenever the assay changes.
- Report preanalytical variables as though someone will try to reproduce your result. Use an established reporting framework rather than a bespoke methods paragraph [2,3].
- Choose the discovery assay with the validation cohort in mind. Ask what the assay would require at 600 samples across eight sites with 0.8 mL of plasma each, and let the answer influence the method you use at n = 40.
- Characterize within-subject precision early. If the biomarker will ever be used to track change, that figure matters more than the group difference you plan to publish.
None of this is a substitute for good biology. A candidate with no mechanistic connection to disease will not be rescued by analytical rigor. But the reverse holds far more often than the field acknowledges: real biology is routinely lost to measurement that was never designed to travel.
The direction of the field is toward closing that gap structurally rather than case by case — standardized preanalytical reporting, method-agnostic characterization requirements, and assay formats that can move from a discovery bench to a multi-site protocol without being rebuilt. As those become the default rather than the exception, the interesting question stops being whether EV biomarkers can reach the clinic and becomes which ones deserve to.
References
[1] Extracellular vesicle biomarkers: current status and future perspectives as novel tools in liquid biopsy. Frontiers in Immunology. 2026. doi:10.3389/fimmu.2026.1853671
[2] Welsh JA, Goberdhan DCI, O’Driscoll L, et al. Minimal information for studies of extracellular vesicles (MISEV2023): from basic to advanced approaches. Journal of Extracellular Vesicles. 2024;13(2):e12404. doi:10.1002/jev2.12404
[3] Gustafson D, Nieuwland R, Lucien F. MIBlood-EV: an online reporting tool to facilitate the standardized reporting of preanalytical variables and quality control of plasma and serum to enhance rigor and reproducibility in liquid biopsy research. Biopreservation and Biobanking. 2025. doi:10.1089/bio.2024.0083
[4] Bridging laboratory innovation to translational research and commercialization of extracellular vesicle isolation and detection. Biosensors and Bioelectronics. 2025.
[5] Minimal Information for Studies of Extracellular Vesicles (MISEV): ten-year evolution (2014–2023). 2024.
[6] Théry C, Witwer KW, Aikawa E, et al. Minimal information for studies of extracellular vesicles 2018 (MISEV2018). Journal of Extracellular Vesicles. 2018;7(1):1535750. doi:10.1080/20013078.2018.1535750
[7] Saint-Pol J. Reproducibility and transparency: why following MISEV guidelines is beneficial for the studies on EVs and brain barriers. Extracellular Vesicles and Circulating Nucleic Acids. 2025. doi:10.20517/evcna.2024.63
[8] Unveiling heterogeneity: innovations and challenges in single-vesicle analysis for clinical translation. 2025.
[9] FDA-NIH Biomarker Working Group. BEST (Biomarkers, EndpointS, and other Tools) Resource. Silver Spring, MD: US Food and Drug Administration; 2016 (updated periodically). https://www.ncbi.nlm.nih.gov/books/NBK326791/
[10] Cudkowicz ME, et al. Safety and efficacy of PrimeC in amyotrophic lateral sclerosis: the PARADIGM randomized clinical trial. JAMA Neurology. 2026;83(5):471–480. doi:10.1001/jamaneurol.2026.0230



