Layers of the research pipeline
A living, annotated collection of multiverse-style methods literature
This is a curated, annotated reference collection on multiverse-style analysis — specification curves, vibration of effects, many-analysts studies, and their relatives — organized by where in the research pipeline the forking happens. It accompanies my blog series on arbitrary-yet-reasonable decisions in science.
How it is structured. Layer 0 collects the debate about whether we should multiverse at all — the critiques deserve top billing, and the strongest ones come first. A foundations section follows, opening with a chronological cartography of the guides and field maps the method has produced about itself. Then the pipeline: Layer 1 (psychometrics — how we measure what we measure), Layer 2 (primary studies — how we gather evidence), Layer 3 (meta-analysis — how we combine knowledge), and Layer 4 (meta-meta — how we combine the combinations). A software and tutorials section closes it out. Empirical entries carry a Paths varied note stating which specifications were explored and, where the paper reports it, which decisions proved consequential.
A living document. I intend to review and extend this list about once a year. It is certainly incomplete — the field moves fast, and search favors the canon over last quarter’s preprints. I am genuinely happy about suggestions and pointers, especially when a new paper or preprint appears: ideally via an issue or pull request on GitHub, or by email. If you think something belongs here, you are probably right.
⭐ marks the one entry per section to read first. A companion BibTeX file (multiverse_resources.bib, starter set) lives alongside this document for one-click import into your reference manager. On freezing each version I archive the collection on Zenodo for a citable DOI. Open territory notes in each layer flag where the map is still blank — these gaps are invitations.
Compiled 2026-08-20 (v1.0).
Layer 0 — Rationale & critiques: should we multiverse at all?
Before any layer of the pipeline, there is a prior question: whether running all reasonable analyses is the honest response to analytic flexibility, or an abdication of the researcher’s duty to identify the single best-justified analysis. This section collects both sides — the foundational critical frameworks, the strongest “there is only one correct analysis” position, the manufacturing-of-doubt worry, the garbage-in-garbage-out problem of ill-chosen specifications, and evidence that multiverse results themselves can fail to replicate — along with the statistical work on how to draw valid inferences across a multiverse at all.
- ⭐ Del Giudice, M., & Gangestad, S. W. (2021). A traveler’s guide to the multiverse: Promises, pitfalls, and a framework for the evaluation of analytic decisions. AMPPS, 4(1). https://doi.org/10.1177/2515245920954925 (the foundational critical framework — equivalence vs. non-equivalence of specifications)
- Lakens, D., Rasti, S., & Tunç, M. N. (2026). There is only one correct analysis. OSF Preprints (June 2026). (the strongest “turtles bottom out” position: epistemic uncertainty should be reduced, not embraced; argues for integrative experiments over multiverse/many-analyst approaches). https://osf.io/preprints/psyarxiv/hvjxf_v1
- Modrák, M. (2026). Multiverse analysis, abdication of responsibility and manufacturing of doubt. arXiv:2607.14623. https://doi.org/10.48550/arXiv.2607.14623
- Rohrer, J. M. (2021, March 7). Mülltiverse analysis. The 100% CI (blog). https://www.the100.ci/2021/03/07/mulltiverse-analysis/ (garbage specifications in, garbage multiverse out — and the “false sense of certainty” worry; also argues the multiverse may be most powerful as pedagogy)
- Rohrer, J. M., Hullman, J., & Gelman, A. (2026). What’s a multiverse good for anyway? [preprint at https://osf.io/preprints/psyarxiv/37g29_v1; discussed on Statistical Modeling blog Feb 2026, https://statmodeling.stat.columbia.edu/2026/02/12/what-a-multiverse-good-for-anyway/ reflection/critique vs. persuasion vs. serious inference; fails as persuasion when researchers disagree on included variations, fails as inference without a coherent estimand]
- Saltelli, A., Aleksankina, K., Becker, W., et al. (2019). Why so many published sensitivity analyses are false: A systematic review of sensitivity analysis practices. Environmental Modelling & Software, 114, 29–39. https://doi.org/10.1016/j.envsoft.2019.01.012 (the older sibling method, done badly at scale — a warning for multiverse practice)
- Rubin, M. (2024). Preregistration does not improve the transparent evaluation of severity in Popper’s philosophy of science or when deviations are allowed. arXiv. https://doi.org/10.48550/arXiv.2408.12347
- Boulesteix, A.-L., & Hoffmann, S. (2024). To adjust or not to adjust: it is not the tests performed that count, but how they are reported and interpreted. BMJ Medicine, 3(1). https://doi.org/10.1136/bmjmed-2023-000783
- Olsson-Collentine, A., van Aert, R. C. M., Bakker, M., & Wicherts, J. (2025). Meta-analyzing the multiverse: A peek under the hood of selective reporting. Psychological Methods, 30(3), 441–461. https://doi.org/10.1037/met0000559 (multiverse effects often fail to replicate across samples; proposes underlying multiverse variability index — critique from within) — Paths varied: counterfactual multiverses over 211 Registered Replication Report samples; only ~30% of studies contained any significant hypothesized effect, and a given researcher DF’s influence was inconsistent across samples — multiverse results themselves often fail to replicate.
- Heyman, T., & Vanpaemel, W. (2022). Multiverse analyses in the classroom. Meta-Psychology. https://doi.org/10.15626/MP.2020.2718 (the pedagogy angle: teaching forks by making students walk them)
- Heyman, T., Pronizius, E., Lewis, S. C., Acar, O. A., Adamkovič, M., et al. (2025, in press). Crowdsourcing multiverse analyses to explore the impact of different data-processing and analysis decisions: A tutorial. Psychological Methods. https://doi.org/10.1037/met0000770
Foundational papers (cross-layer)
The papers that defined the method and its vocabulary. These are not tied to any one layer: they introduce the multiverse and specification-curve idea, demonstrate empirically that independent analysts diverge on identical data, extend the logic to data-collection decisions, and map how the family of multiverse-style methods has spread across fields. Start here if the concept is new.
Cartography of the multiverse: guides and field maps
Attempts to map the territory itself — first field-specific guides, then landscape surveys, now a consolidated multidisciplinary standard. Read chronologically, they show a method growing self-aware:
2016 — the founding map: Steegen, Tuerlinckx, Gelman, & Vanpaemel introduce the multiverse and walk through the first worked example (full citation below).
2019 — first landscape, meta-analytic territory: Voracek, Kossmeier, & Tran chart which-data × how-to-analyze as the coordinate system for multiverse meta-analysis (full citation in Layer 3).
2020 — the visualization landscape: Kossmeier, M., Tran, U. S., & Voracek, M. (2020). Charting the landscape of graphical displays for meta-analysis and systematic reviews: A comprehensive review, taxonomy, and feature analysis. BMC Medical Research Methodology, 20, 26. https://doi.org/10.1186/s12874-020-0911-9 (the Vienna group’s second map: every meta-analytic graph in existence — 150+ textbooks checked cover to cover — taxonomized; the display vocabulary that multiverse meta-analyses draw on)
2021 — the traveler’s guide: Del Giudice & Gangestad supply the critical map legend — equivalence classes (Type E/N/U) that determine which paths belong on the map at all (full citation in Layer 0).
2024 — the practical tutorial: Götz, Sarma, & O’Boyle, hands-on with the
multiverseR package (full citation in Software).2025 — second landscape, the whole territory: Voracek, Tran, & Kern (2025). Mapping the landscape of multiverse-style data-analytic methods: A scoping review. PsyArXiv. https://doi.org/10.31234/osf.io/g4zck_v1 (424 sources; the first comprehensive survey of everything the method family has touched — still preprint, check for the journal version before citing)
2025 — the book: Young & Cumberworth, Multiverse Analysis: Computational Methods for Robust Results (Cambridge UP; full citation below).
⭐ 2026 — the consolidated standard: Short, C. A., Breznau, N., Bruntsch, M., Burkhardt, M., Busch, N. A., Cesnaite, E., Frank, M., Gießing, C., Krähmer, D., Kristanto, D., Lonsdorf, T., Neuendorf, C., Nguyen, H. H. V., Rausch, M., Schmalz, X., Schneck, A., Tabakci, C., & Hildebrandt, A. (2026). Multicurious: A multidisciplinary guide to multiverse analysis. Advances in Methods and Practices in Psychological Science. https://doi.org/10.1177/25152459261434881 — The current state-of-the-art procedural guide (META-REP consortium): formalizes defensibility vs. equivalence (defensible vs. principled multiverse), harmonizes the terminology zoo (multiverse, SCA, VoE, multimodel, manyverse…), covers preregistration (SMART), pipeline similarity/pseudoreplication, and the covariates-as-estimand-shift debate. Open access; the section-by-section companion to this whole page.
⭐ Steegen, S., Tuerlinckx, F., Gelman, A., & Vanpaemel, W. (2016). Increasing transparency through a multiverse analysis. Perspectives on Psychological Science, 11(5), 702–712. https://doi.org/10.1177/1745691616658637 — Paths varied: data-processing choices (exclusions, operationalizations of the outcome and moderators) on the fertility/religiosity dataset; showed the original finding survived only a minority of reasonable processing pipelines.
Simonsohn, U., Simmons, J. P., & Nelson, L. D. (2020). Specification curve analysis. Nature Human Behaviour, 4, 1208–1214. https://doi.org/10.1038/s41562-020-0912-z — Paths varied: re-analyzed three published findings (e.g., Black-sounding names discrimination, female-named hurricanes) across all justified specifications; one finding proved robust, one weak, one not robust — inference via joint permutation test across the curve.
- Earlier working-paper version: Simonsohn, Simmons, & Nelson (2015). Specification curve: Descriptive and inferential statistics on all reasonable specifications. SSRN. https://doi.org/10.2139/ssrn.2694998
Haaf, J. M., Hoogeveen, S., Berkhout, S., Gronau, Q. F., & Wagenmakers, E.-J. (2020). A Bayesian multiverse analysis of Many Labs 4: Quantifying the evidence against mortality salience. PsyArXiv. https://doi.org/10.31234/osf.io/cb9er — Paths varied: Bayesian model specifications (priors, exclusion criteria) applied to Many Labs 4 mortality-salience data; evidence pointed against the effect across the multiverse.
Shunsen, H., Haojie, C., Xiaoxiong, L., Xinran, D., & Yun, W. (2023). Multiverse-style analysis: Introduction and application. Advances in Psychological Science, 31(2), 196. https://doi.org/10.3724/SP.J.1042.2023.00196 (Chinese-language introduction)
⭐ Silberzahn, R., Uhlmann, E. L., Martin, D. P., et al. (2018). Many analysts, one data set. AMPPS, 1, 337–356. https://doi.org/10.1177/2515245917747646 (the 29-teams red-card study) — Paths varied: 29 teams freely chose models, covariates, and operationalizations on identical data; effect estimates ranged from strong to null, with no single decision explaining the spread — the point is that defensible model choice itself is the fork.
Hoffmann, S., Schönbrodt, F., Elsas, R., Wilson, R., Strasser, U., & Boulesteix, A.-L. (2021). The multiplicity of analysis strategies jeopardizes replicability: lessons learned across disciplines. Royal Society Open Science, 8(4), 201925. https://doi.org/10.1098/rsos.201925 — Cross-disciplinary synthesis (psychology, finance, hydrology, biometrics): analysis-strategy multiplicity threatens replicability everywhere, arguing for multiverse-style reporting as a general remedy.
Harder, J. A. (2020). The multiverse of methods: Extending the multiverse analysis to address data-collection decisions. Perspectives on Psychological Science, 15(5), 1158–1177. https://doi.org/10.1177/1745691620917678
Aczel, B., Szaszi, B., Nilsonne, G., et al. (2021). Consensus-based guidance for conducting and reporting multi-analyst studies. eLife, 10, e72185.
Many-analysts studies (the empirical backbone across disciplines):
- Botvinik-Nezer, R., et al. (2020). Variability in the analysis of a single neuroimaging dataset by many teams. Nature, 582, 84–88. https://doi.org/10.1038/s41586-020-2314-9 (NARPS: 70 fMRI teams, same data, divergent conclusions)
- Breznau, N., Rinke, E. M., Wuttke, A., Nguyen, H. H., et al. (2022). Observing many researchers using the same data and hypothesis reveals a hidden universe of uncertainty. PNAS, 119(44), e2203150119. https://doi.org/10.1073/pnas.2203150119 (sociology’s many-analysts study; researcher characteristics barely explain the spread)
- Muñoz, J., & Young, C. (2018). We ran 9 billion regressions: Eliminating false positives through computational model robustness. Sociological Methodology, 48(1), 1–33. https://doi.org/10.1177/0081175018777988
- Schweinsberg, M., Feldman, M., Staub, N., van den Akker, O. R., van Aert, R. C. M., van Assen, M. A. L. M., et al. (2021). Same data, different conclusions: Radical dispersion in empirical results when independent analysts operationalize and test the same hypothesis. Organizational Behavior and Human Decision Processes. https://doi.org/10.1016/j.obhdp.2021.02.003 (the second-generation many-analysts study: dispersion arises already at operationalization, before any model is fit)
- Trübutschek, D., Yang, Y. F., Gianelli, C., Cesnaite, E., et al. (2024). EEGManyPipelines: A large-scale, grassroots multi-analyst study of electroencephalography analysis practices in the wild. Journal of Cognitive Neuroscience, 36(2), 217–224. https://doi.org/10.1162/jocn_a_02087
Forking-paths canon & precursors:
- Leamer, E. E. (1985). Sensitivity analyses would help. American Economic Review, 75(3), 308–313. (the 1985 ancestor of the whole idea)
- Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22, 1359–1366. https://doi.org/10.1177/0956797611417632 (researcher degrees of freedom, the founding demonstration)
- Gelman, A., & Loken, E. (2013). The garden of forking paths. Columbia University working paper. https://sites.stat.columbia.edu/gelman/research/unpublished/forking.pdf
- Gelman, A., & Loken, E. (2014). The statistical crisis in science. American Scientist, 102(6), 460–465. https://doi.org/10.1511/2014.111.460
- Bryan, C. J., Yeager, D. S., & O’Brien, J. M. (2019). Replicator degrees of freedom allow publication of misleading failures to replicate. PNAS, 116(51), 25535–25545. https://doi.org/10.1073/pnas.1910951116 (the forks cut both ways — replicators have them too)
- Rubin, M. (2017). An evaluation of four solutions to the forking paths problem: Adjusted alpha, preregistration, sensitivity analyses, and abandoning the Neyman–Pearson approach. Review of General Psychology, 21, 321–329. https://doi.org/10.1037/gpr0000135
- Young, C., & Holsteen, K. (2017). Model uncertainty and robustness: A computational framework for multimodel analysis. Sociological Methods & Research, 46(1), 3–40. https://doi.org/10.1177/0049124115610347
- Wacker, J. (2017). Increasing the reproducibility of science through close cooperation and forking path analysis. Frontiers in Psychology, 8, 1332. https://doi.org/10.3389/fpsyg.2017.01332
- Young, C., & Cumberworth, E. (2025). Multiverse analysis: Computational methods for robust results. Cambridge University Press. (the book-length treatment, from the “9 billion regressions” lineage)
- Ramsey, J. B. (1969). Tests for specification errors in classical linear least-squares regression analysis. JRSS B, 31(2), 350–371. https://doi.org/10.1111/j.2517-6161.1969.tb00796.x (1969: the deepest root)
- Neumayer, E., & Plümper, T. (2017). Robustness tests for quantitative research. Cambridge University Press.
- Auspurg, K., & Brüderl, J. (2021). Has the credibility of the social sciences been credibly destroyed? Reanalyzing the “Many Analysts, One Data Set” project. Socius, 7. https://doi.org/10.1177/23780231211024421 (the estimand critique of Silberzahn: much of the spread came from teams answering different questions)
Robustness-reproduction frameworks:
- Dreber, A., & Johannesson, M. (2025). A framework for evaluating reproducibility and replicability in economics. Economic Inquiry, 63(2), 338–356. https://doi.org/10.1111/ecin.13244
- Brodeur, A., Cook, N., Hartley, J., & Heyes, A. (2024). Do pre-registration and pre-analysis plans reduce p-hacking and publication bias? Journal of Political Economy Microeconomics, 2(3), 527–561. https://doi.org/10.1086/730455
- Ankel-Peters, J., Brodeur, A., Dreber, A., Johannesson, M., Neubauer, F., & Rose, J. (2024). A protocol for structured robustness reproductions and replicability assessments. Institute for Replication, No. 143.
- Chen, G., Cai, Z., & Taylor, P. A. (2024). Through the lens of causal inference: Decisions and pitfalls of covariate selection. Aperture Neuro, 4. https://doi.org/10.52294/001c.124817 (the causal backbone of the covariates-as-decision-node debate)
Nepomuceno, A., Ghosal, A., Lentisco, A., & Ioannidis, J. (2026). Uptake and implementation of multiverse-style analyses across 613 studies. https://doi.org/10.64898/2026.07.15.738584
Inference across specifications
- Girardi, P., Vesely, A., Lakens, D., Altoè, G., Pastore, M., Calcagnì, A., & Finos, L. (2024). Post-selection inference in multiverse analysis (PIMA): An inferential framework based on the sign flipping score test. Psychometrika, 89(2), 542–568. https://doi.org/10.1007/s11336-024-09973-6
- Mandl, M. M., Becker-Pennrich, A. S., Hinske, L. C., Hoffmann, S., & Boulesteix, A.-L. (2024). Addressing researcher degrees of freedom through minP adjustment. BMC Medical Research Methodology, 24(1), 152. https://doi.org/10.1186/s12874-024-02279-2
- Bartoš, F., Hoogeveen, S., Sarafoglou, A., & Pawel, S. (2025). Single-dataset meta-analysis for many-analysts and multiverse studies. arXiv:2511.17064. https://doi.org/10.48550/arXiv.2511.17064
- Vesely, A., & Andreella, A., et al. (2026). PIMAX: Post-selection inference for multiverse analysis in mixed-effects models. arXiv:2607.03225
- Ozenne, B., Nørgaard, M., Pernet, C., & Ganz, M. (2025). A sensitivity analysis of preprocessing pipelines: Toward a solution for multiverse analyses. Imaging Neuroscience, 3, imag_a_00523. https://doi.org/10.1162/imag_a_00523
Layer 1 — Psychometrics: how we measure what we measure
The forks begin before any data are analyzed, inside the instruments themselves. This layer collects multiverse-style work on measurement: differential item functioning assessed across the full space of defensible detection choices, test-validation coefficients that shift with design decisions, reliability that moves with data-processing rules, and the field’s own ambivalence about whether advanced psychometric models add value over simple sum scores. Open territory: systematic multiverses over scale choice itself (which of 280 depression scales) remain rare; empirical equivalence assessment for measurement pipelines is underdeveloped.
- ⭐ Flake, J. K., & Fried, E. I. (2020). Measurement schmeasurement: Questionable measurement practices and how to avoid them. AMPPS, 3(4), 456–465. https://doi.org/10.1177/2515245920952393 — QMPs as a “stunning source of researcher degrees of freedom”: the measurement layer’s manifesto, arguing the replication debate obsesses over statistics while the deeper forks sit in how constructs are measured at all.
- Plessen, C. Y., Fischer, F., Hartmann, C., Liegl, G., Schalet, B., Kaat, A. J., Pesantez, R., Joeris, A., Heng, M., Rose, M., & the AOBERT Consortium (2024). Differential item functioning between English, German, and Spanish PROMIS® physical function ceiling items. Quality of Life Research, 34(5). https://doi.org/10.1007/s11136-024-03866-y (DIF assessed via multiverse across the full plausible space of logistic ordinal regression choices) — Paths varied: the full plausible space of logistic ordinal regression DIF choices (criteria, thresholds, anchor decisions) across English/German/Spanish samples; DIF-item identification proved robust across specifications — multiverse as an answer to arbitrary DIF cutoffs.
- Companion software:
lordifMultiverseR package — github.com/cyplessen
- Companion software:
- Webb, S. S., & Demeyere, N. (2022). Using multiverse analysis to highlight differences in convergent correlation outcomes due to data analytical and study design choices. Assessment. https://doi.org/10.1177/10731911221127904 (2,220 analyses on psychometric test validation) — Paths varied: 2,220 analyses crossing sample group (healthy vs. stroke), sample size, test metrics, and covariate inclusion for convergent-validity correlations of an executive-function test — validation coefficients themselves vibrate with design choices.
- Byrne, et al. (2026). ‘We’re not going to start lifting stones now…’: Stakeholder perspectives on the role of psychometric methods in outcome measurement. British Journal of Clinical Psychology, 1–16. https://doi.org/10.1111/bjc.70067 — The “added benefit” question put to the field itself: 21 stakeholder interviews across psychometrics, clinical practice, applied research, and statistics on whether IRT/SEM add value over sum scores in applied work — the perceived-benefit side of the fork this layer maps. [Complete author list from the PDF]
- Harrison, C. J., Plessen, C. Y., Liegl, G., Rodrigues, J. N., Sabah, S. A., Cook, J. A., Beard, D. J., & Fischer, F. (2023). Item response theory may account for unequal item weighting and individual-level measurement error in trials that use PROMs: A psychometric sensitivity analysis of the TOPKAT trial. Journal of Clinical Epidemiology, 158, 62–69. https://doi.org/10.1016/j.jclinepi.2023.03.013 — Paths varied: re-scoring a real RCT’s PROM outcomes under IRT vs. sum scores — the “does the fancy model change trial conclusions?” question tested on actual trial data rather than argued in the abstract; the empirical companion to the Byrne perception study.
- Hanel, P. H. P., & Zarzeczna, N. (2023). From multiverse analysis to multiverse operationalisations: 262,143 ways of measuring well-being. Religion, Brain & Behavior, 13(3), 309–313. https://doi.org/10.1080/2153599X.2022.2070259 — Paths varied: which items/subscales operationalize “well-being” at all — the construct’s operationalization as its own combinatorial multiverse, upstream of any model choice.
- Parsons, S. (2020). Exploring reliability heterogeneity with multiverse analyses: Data processing decisions unpredictably influence measurement reliability. PsyArXiv. https://doi.org/10.31234/osf.io/y6tcz (reliability itself as a multiverse outcome — perfect Layer 1 fit) — Paths varied: data-processing decisions (trial-level exclusions, outlier rules, scoring) for cognitive-behavioral task measures; reliability estimates shifted unpredictably — the instrument’s own precision is specification-dependent.
- Bauer, D. J., Belzak, W. C., & Cole, V. T. (2020). Regularized MNLFA to detect DIF. Structural Equation Modeling, 27(1), 43–55. https://doi.org/10.1080/10705511.2019.1642754
- Preprocessing multiverses in biophysiological measurement (EEG/ERP/fMRI/SCR):
- Clayson, P. E., Baldwin, S. A., Rocha, H. A., & Larson, M. J. (2021). The data-processing multiverse of event-related potentials (ERPs): A roadmap. NeuroImage, 245, 118712. https://doi.org/10.1016/j.neuroimage.2021.118712
- Feuerriegel, D., & Bode, S. (2022). Bring a map when exploring the ERP data processing multiverse. NeuroImage, 259, 119443. https://doi.org/10.1016/j.neuroimage.2022.119443 (why effect size must never be the pipeline-ranking criterion)
- Clayson, P. E. (2024). Beyond single paradigms, pipelines, and outcomes: Embracing multiverse analyses in psychophysiology. International Journal of Psychophysiology, 197, 112311. https://doi.org/10.1016/j.ijpsycho.2024.112311
- Kuhn, M., Gerlicher, A. M., & Lonsdorf, T. B. (2022). Navigating the manyverse of skin conductance response quantification. Psychophysiology, 59(9), e14058. https://doi.org/10.1111/psyp.14058
- Jacobsen, N. S., Kristanto, D., Welp, S., Inceler, Y. C., & Debener, S. (2025). Preprocessing choices for P3 analyses with mobile EEG: Systematic review and interactive exploration. Psychophysiology, 62(1), e14743. https://doi.org/10.1111/psyp.14743
- Kristanto, D., Burkhardt, M., Thiel, C. M., Debener, S., Gießing, C., & Hildebrandt, A. (2024). The multiverse of data preprocessing and analysis in graph-based fMRI: Systematic review + decision support tool. Neuroscience & Biobehavioral Reviews, 165, 105846. https://doi.org/10.1016/j.neubiorev.2024.105846
- Dafflon, J., et al. (2022). A guided multiverse study of neuroimaging analyses. Nature Communications, 13, 3758. https://doi.org/10.1038/s41467-022-31347-8 (ML sampling when the multiverse is too big to compute)
- Paul, K., Short, C. A., Beauducel, A., et al. (2022). The methodology and dataset of the CoScience EEG-personality project — a large-scale, multi-laboratory project grounded in cooperative forking paths analysis. Personality Science, 3(1), e7177. https://doi.org/10.5964/ps.7177
- Paul, K., Beauducel, A., et al. (2025). Frontal alpha asymmetry as a marker of approach motivation? Insights from a cooperative forking path analysis. JPSP, 128(1), 196–210. https://doi.org/10.1037/pspp0000503 (preregistered cooperative multiverse settling a classic EEG-marker question)
- Fried E. I. (2017). The 52 symptoms of major depression: Lack of content overlap among seven common depression scales. Journal of affective disorders, 208, 191–197. https://doi.org/10.1016/j.jad.2016.10.019
Layer 2 — Primary studies: how we gather evidence
One dataset, one research question, thousands of defensible analyses. This layer covers multiverse and specification-curve applications to primary studies — observational and experimental — including the screen-time debate fought entirely in specification curves, and the vibration-of-effects tradition from epidemiology, which quantifies how much estimates move under different covariate-adjustment, measurement, and sampling choices. Open territory: multiverse analyses of full clinical-trial analysis pipelines (model × missing-data strategy) are scarce relative to observational work.
- Orben, A., & Przybylski, A. K. (2019). The association between adolescent well-being and digital technology use. Nature Human Behaviour. https://doi.org/10.1038/s41562-018-0506-1 — Paths varied: thousands of specifications across three large adolescent datasets (which well-being measures, which technology measures, which covariates); the tiny negative association shrank or vanished depending on mostly arbitrary choices — comparable in size to eating potatoes.
- Twenge, J. M., Haidt, J., Lozano, J., & Cummins, K. M. (2022). Specification curve analysis shows that social media use is linked to poor mental health, especially among girls. Acta Psychologica, 224, 103512. https://doi.org/10.1016/j.actpsy.2022.103512 (the counter-SCA) — Paths varied: re-ran Orben & Przybylski’s SCA with four alterations; the consequential decisions were separating social media from TV within “screen time” and how one multi-subscale mental-health measure was weighted — with those changed, associations for girls emerged more clearly.
- Rengasamy, M., Moriarity, D., Kraynak, T., Tervo-Clemmens, B., & Price, R. (2023). Exploring the multiverse: the impact of researchers’ analytic decisions on relationships between depression and inflammatory markers. Neuropsychopharmacology : official publication of the American College of Neuropsychopharmacology, 48(10), 1465–1474. https://doi.org/10.1038/s41386-023-01621-4 (58,000+ specifications) — Paths varied: 9 common analytic decisions (log-transformation, covariate count, outlier handling, etc.) yielding 58,000+ combinations on NHANES inflammatory-marker/depression data, testing robustness of the association across all of them.
- Lonsdorf, T. B., Gerlicher, A., Klingelhöfer-Jens, M., & Krypotos, A.-M. (2022). Multiverse analyses in fear conditioning research (+
multifear). Behaviour Research and Therapy, 153, 104072. https://doi.org/10.1016/j.brat.2022.104072 — Paths varied: statistical model family (Bayesian vs. frequentist ANOVA/t-tests/mixed models) crossed with data-reduction approaches for SCR fear-conditioning data; both the size and direction of effects shifted with model and reduction choice. - Stern, J., Arslan, R. C., Gerlach, T. M., & Penke, L. (2019). No robust evidence for cycle shifts in preferences for men’s bodies in a multiverse analysis. Evolution and Human Behavior, 40(6), 517–525. https://doi.org/10.1016/j.evolhumbehav.2019.08.005 — Paths varied: specifications for cycle-phase estimation and preference outcomes; the claimed cycle shifts did not survive the multiverse.
- Engzell, P., & Mood, C. (2023). Understanding patterns and trends in income mobility through multiverse analysis. American Sociological Review, 88(4), 600–626. https://doi.org/10.1177/00031224231180607 (flagship sociology application)
- Lefort-Besnard, J., Nichols, T. E., & Maumet, C. (2025). Statistical inference for same-data meta-analysis in neuroimaging multiverse analyses. Imaging Neuroscience. https://doi.org/10.1162/imag_a_00513
- Vibration of effects family (observational designs):
- ⭐ Patel, C. J., Burford, B., & Ioannidis, J. P. A. (2015). Assessment of vibration of effects due to model specification. J Clin Epidemiol, 68(9), 1046–1058. https://doi.org/10.1016/j.jclinepi.2015.05.029 — Paths varied: all 8,192 combinations of 13 adjustment covariates in Cox models for 417 variables’ associations with mortality; adjustment-set choice alone flipped significance and even sign for a nontrivial share of associations (their “Janus effect”).
- Klau, S., Hoffmann, S., Patel, C. J., Ioannidis, J. P., & Boulesteix, A.-L. (2021). Model, measurement and sampling uncertainty in the VoE framework. IJE, 50(1), 266–278. https://doi.org/10.1093/ije/dyaa164 — Paths varied: extends VoE to decompose result variability into model choice, measurement error, and sampling uncertainty in one framework — showing how much of the “vibration” each source contributes.
- Chu, L., Ioannidis, J. P. A., et al. (2020). VoE in epidemiologic studies of alcohol and breast cancer. IJE, 49(2), 608–618. https://doi.org/10.1093/ije/dyz271 — Paths varied: analytic approaches within and across published observational studies of alcohol and breast cancer, quantifying how much reported estimates vibrate with model choices.
- Tierney, B. T., et al. (2021). Leveraging VoE analysis for robust discovery. PLOS Biology, 19, e3001398. https://doi.org/10.1371/journal.pbio.3001398 — Paths varied: covariate-adjustment vibration deployed prospectively as a discovery filter in large biomedical datasets — VoE turned from critique into a robustness screen.
- Vinatier, C., Hoffmann, S., Patel, C., DeVito, N. J., Cristea, I. A., Tierney, B., Ioannidis, J. P. A., & Naudet, F. (2024). What is the vibration of effects? BMJ Evidence-Based Medicine. https://doi.org/10.1136/bmjebm-2023-112747 (accessible explainer) — Explainer, not an empirical multiverse: defines VoE and situates it among multi-analyst and multiverse approaches.
Layer 3 — Meta-analysis: how we combine knowledge
Every fork from the layers below arrives here, plus new ones: which studies to include, which effect size to compute, which model to pool with, which bias corrections to apply, which outliers to exclude. This is currently the most colonized layer — three independent groups (Vienna, Rennes, Amsterdam) have converged on multiverse meta-analysis from psychology, epidemiology, and clinical trials — alongside the Bayesian model-averaging alternative, which resolves the forks by weighting rather than mapping them. Open territory: no dedicated multiverse over risk-of-bias assessment methods exists yet.
⭐ Voracek, M., Kossmeier, M., & Tran, U. S. (2019). Which data to meta-analyze, and how? Zeitschrift für Psychologie, 227(1), 64–82. https://doi.org/10.1027/2151-2604/a000357 (the founding multiverse-MA paper) — Framework paper: crosses which data (combinatorial study inclusion) with how (model and estimator choice) to define the meta-analytic multiverse; the template your own MAs and Pietschnig’s implement.
Plessen multiverse meta-analyses:
- Plessen, C. Y., Karyotaki, E., & Cuijpers, P. (2022). Protocol. BMJ Open, 12(1). https://doi.org/10.1136/bmjopen-2021-050197
- ⭐ Plessen, C. Y., Karyotaki, E., Miguel, C., Ciharova, M., & Cuijpers, P. (2023). Exploring the efficacy of psychotherapies for depression: a multiverse meta-analysis. BMJ Mental Health, 26, e300626. https://doi.org/10.1136/bmjment-2022-300626 — Paths varied: all plausible meta-analyses from the Metapsy depression-trial pool; consequential forks were inclusion of high risk-of-bias studies, waitlist comparators, and (not) adjusting for publication bias — each inflating the summary effect.
- Plessen, C. Y., Panagiotopoulou, O. M., Tong, L., Cuijpers, P., & Karyotaki, E. (2024). Digital mental health interventions for depression: a multiverse meta-analysis. Journal of Affective Disorders, 369, 1031–1044. https://doi.org/10.1016/j.jad.2024.10.018 — Paths varied: 3,638 meta-analyses over 125 RCTs (populations, intervention characteristics, designs); g averaged 0.43 and stayed positive from the 10th to 90th percentile — larger with adults, LMICs, guidance, waitlist controls; smaller with publication-bias adjustment and ≥24-week outcomes.
Pietschnig group (Vienna) multiverse meta-analyses:
- Pietschnig, J., Gerdesmann, D., Zeiler, M., & Voracek, M. (2022). The meta-analytical multiverse of brain volume and IQ associations. Royal Society Open Science, 9(5), 211621. https://doi.org/10.1098/rsos.211621 (432 specifications) — Paths varied: 432 specifications from 27 which-data choices (healthy vs. clinical samples, correction types, g-ness of measures) × 16 how-to-analyze choices (Hedges–Olkin, Hunter–Schmidt, unweighted, RVE); the brain-volume/IQ association held across most, but its magnitude varied enough to explain prior meta-analytic disagreement.
- Dürlinger, F., & Pietschnig, J. (2022). Meta-analyzing intelligence and religiosity associations: Evidence from the multiverse. PLOS ONE. https://doi.org/10.1371/journal.pone.0262699 — Paths varied: 192 specifications; 70.4% of summary effects significant and all of those negative — direction robust; strength moved with intelligence measure type (psychometric tests vs. GPA proxies) and sample type (pre-college vs. adult).
- Oberleiter, S., & Pietschnig, J. (2023). The Mozart effect myth: a multiverse meta-analysis. Scientific Reports, 13, 3175. https://doi.org/10.1038/s41598-023-30206-w — Paths varied: three independent analytic approaches on k=8 usable studies of Mozart KV448 and epilepsy; trivial-to-small nonsignificant effects throughout — the multiverse as myth-busting, with reporting opacity blocking even inclusion.
- Patzl, S., Oberleiter, S., & Pietschnig, J. (2024). Self-assessed intelligence through the lens of the multiverse. Journal of Intelligence, 12(9), 81. https://doi.org/10.3390/jintelligence12090081 — Paths varied: 278 effect sizes across specification-curve and combinatorial analyses of the self-assessed/psychometric intelligence link, probing generality across SAI measurement methods and samples.
- Fries, J., Oberleiter, S., Bodensteiner, F.A. et al. Multilevel multiverse meta-analysis indicates lower IQ as a risk factor for physical and mental illness. Commun Psychol 3, 74 (2025). https://doi.org/10.1038/s44271-025-00245-2
- Gerdesmann, D., & Pietschnig, J. (2025). Brain volume–IQ multiverse update. OSF: osf.io/y6msp
Naudet group (Rennes) — VoE applied to meta-analysis and trial data:
- Palpacuer, C., Hammas, K., Duprez, R., Laviolle, B., Ioannidis, J. P. A., & Naudet, F. (2019). Vibration of effects from diverse inclusion/exclusion criteria and analytical choices: 9216 different ways to perform an indirect comparison meta-analysis. BMC Medicine, 17(1), 174. https://doi.org/10.1186/s12916-019-1409-3 — Paths varied: 9,216 indirect-comparison meta-analyses of nalmefene vs. naltrexone from crossed inclusion/exclusion criteria and analytic choices; conclusions about comparative efficacy depended on these review-level decisions.
- El Bahri, M., Wang, X., Biaggi, T., Falissard, B., Naudet, F., & Barry, C. (2022). A multiverse analysis of meta-analyses assessing acupuncture efficacy for smoking cessation evidenced vibration of effects. J Clin Epidemiol, 152, 140–150. https://doi.org/10.1016/j.jclinepi.2022.09.001 — Paths varied: review-level choices across meta-analyses of acupuncture for smoking cessation; documented VoE at the synthesis level.
- Gouraud, H., Wallach, J. D., Boussageon, R., Ross, J. S., & Naudet, F. (2022). Vibration of effect in more than 16,000 pooled analyses of individual participant data from 12 RCTs comparing canagliflozin and placebo: multiverse analysis. BMJ Medicine, 1(1). https://doi.org/10.1136/bmjmed-2022-000154 (IPD-level multiverse — bridges Layers 2 and 3) — Paths varied: >16,000 pooled analyses of IPD from 12 canagliflozin RCTs (endpoint definitions, populations, analytic choices); shows vibration exists even with gold-standard trial data pooled at the individual level.
Other multiverse-MA applications:
- Kang, H., Sartorius, A. M., Deilhaug, E., Walle, K. M., & Quintana, D. S. (2025). A multiverse meta-analysis of intranasal oxytocin studies. Psychoneuroendocrinology, 172, 107312.
- Sanabria, D., Ciria, L., Holgado, D., Bartoš, F., Luque-Casado, A., Fernández-Del-Olmo, M., Perakakis, P., & Román-Caballero, R. (2026). Running ahead of the evidence? Rethinking the exercise-cognition consensus. Journal of Science and Medicine in Sport. https://doi.org/10.1016/j.jsams.2026.07.003
Robust Bayesian / model-averaging family (Bartoš–Maier–Wagenmakers):
- Bartoš F, Maier M, Wagenmakers E-J, Doucouliagos H, Stanley TD. Robust Bayesian meta-analysis: Model-averaging across complementary publication bias adjustment methods. Res Syn Meth. 2023; 14(1): 99-116. https://doi.org/10.1002/jrsm.1594
- Maier, M., Bartoš, F., & Wagenmakers, E.-J. (2023). Robust Bayesian meta-analysis: Addressing publication bias with model-averaging. Psychological Methods, 28(1), 107–122. https://doi.org/10.1037/met0000405
- Bartoš, F., Maier, M., Stanley, T. D., & Wagenmakers, E.-J. (2025). Robust Bayesian meta-regression. Psychological Methods. [psycnet 2025-81744-001]
- RoBMA vignettes: https://fbartos.github.io/RoBMA/articles/ (the former CustomEnsembles vignette was renamed upstream; see v20-bayesian-model-averaging)
- Critique: Román-Caballero, R., & Vadillo, M. A. (2024). A meta-analyst should make informed decisions: Issues with Bayesian model-averaging meta-analyses. OSF preprint https://osf.io/tm7dv
Harrer, M., Miguel, C., et al. (2025). Standardized effect sizes are far from “standardized”. PLOS Mental Health. https://doi.org/10.1371/journal.pmen.0000347
Harrer, M., Miguel, C., Hussey, I., Cristea, I. A., van Ballegooijen, W., Basic, D., Wang, Y., Pfund, R. A., Quero, S., van Spreckelsen, P., Schnurr, P. P., van Straten, A., Furukawa, T. A., Papola, D., & Cuijpers, P. (2025). Implausible effects of psychological interventions: Meta-epidemiological study and development of a simple flagging tool. medRxiv. https://doi.org/10.1101/2025.11.12.25340062 — The outlier fork made principled: instead of ad-hoc exclusion rules, a meta-epidemiological flagging tool for effects too large to be plausible — turning “which studies do we drop?” from an arbitrary decision into a criterion. Paths varied: ways to compute the “same” standardized mean difference in depression MAs (SD pooling, change vs. endpoint, correlation assumptions); the effect size itself is a fork before any pooling happens.
Network meta-analysis:
- Vinatier, C., Palpacuer, C., Scanff, A., & Naudet, F. (2024). Vibration of effects resulting from treatment selection in mixed-treatment comparisons: a multiverse analysis on network meta-analyses of antidepressants in major depressive disorder. BMJ evidence-based medicine, 29(5), 324–332. https://doi.org/10.1136/bmjebm-2024-112848
- PRISMA-NMA: Hutton, B., et al. (2015). Annals of Internal Medicine, 162(11), 777–784. https://doi.org/10.7326/M14-2385
- PRISMA-NMA update: Veroniki, A. A., et al. (2025). Scoping review protocol. JBI Evidence Synthesis, 23(3), 517–526. https://doi.org/10.11124/JBIES-24-00308; and medRxiv https://doi.org/10.1101/2025.02.06.25321746
- PRISMA-IPD: Stewart, L. A., et al. (2015). JAMA, 313(16), 1657–1665. https://doi.org/10.1001/jama.2015.3656
Layer 4 — Meta-meta: how we combine combinations of evidence
The thinnest layer, and possibly the next frontier: umbrella reviews and overviews of meta-analyses, where the units of analysis are themselves syntheses stacked on syntheses. Multiverse-style analysis at this level barely exists — the entries here are early examples and infrastructure, and the gap itself is an argument. Open territory: essentially everything — umbrella-level multiverse analysis barely exists.
- Gougeon, A., Aribi, I., Guernouche, S., Lega, J. C., Wright, J. M., Verstuyft, C., Lajoinie, A., Gueyffier, F., & Grenet, G. (2025). Publication bias in pharmacogenetics of statin-associated muscle symptoms: A meta-epidemiological study. Atherosclerosis, 400, 118624. https://doi.org/10.1016/j.atherosclerosis.2024.118624
- Cuijpers, P., Miguel, C., Harrer, M., Plessen, C. Y., Ciharova, M., Papola, D., Ebert, D., & Karyotaki, E. (2023). Psychological treatment of depression: A systematic overview of a ‘Meta-Analytic Research Domain’. Journal of Affective Disorders, 335, 141–151. https://doi.org/10.1016/j.jad.2023.05.011
- I am currently working on an umbrella review with multiverse meta-analysis on digital interventions for anxiety disorders
Applications beyond psychology & medicine
The method has escaped its home field. A non-exhaustive sample of where multiverse thinking has landed — and where this list will grow fastest:
- Fairness in machine learning: One model, many scores: Using multiverse analysis to prevent fairness hacking and evaluate the influence of model design decisions. arXiv:2308.16681
- Bibliometrics: Specification uncertainty: What the disruption index tells us about the (hidden) multiverse of bibliometric indicators. arXiv:2406.13367
- Computational social science: Making uncertainty visible: Multiverse analysis for robust computational social science. arXiv:2605.19745
- Sociology: Engzell & Mood on income mobility (see Layer 2); Breznau et al. and Muñoz & Young (see Foundations)
- Neuroimaging & psychophysiology: see the preprocessing-multiverse subsection in Layer 1
- AI benchmark evaluation: Plessen — multiverse/specification-curve analysis of LLM leaderboard scoring pipelines (~123,000 specifications) [will add preprint link when public]; companion IRT/DIF/measurement-invariance analysis of Open LLM Leaderboard data [will add preprint link]
Software / tutorials
The tooling: R packages for building, running, and visualizing multiverses at each layer, hands-on tutorials, and reporting guidance.
specr— https://masurp.github.io/specr/multiverseR package — Götz, M., Sarma, A., & O’Boyle, E. H. (2024). Tutorial. International Journal of Psychology, 59(6), 1003–1014. https://doi.org/10.1002/ijop.13229metaMultiverse(Plessen) — github.com/cyplessen + tutorial postlordifMultiverse(Plessen) — DIF multiverse detectionmultifear— fear conditioning multiversesvoe(Patel group) — https://www.chiragjpgroup.org/voe/RoBMA— robust Bayesian model-averaged meta-analysis- Visualization: Hall, B. D., Liu, Y., Jansen, Y., Dragicevic, P., Chevalier, F., & Kay, M. (2022). A survey of tasks and visualizations in multiverse analysis reports. Computer Graphics Forum, 41(1), 402–426. https://doi.org/10.1111/cgf.14443
metaviz(Kossmeier, Tran & Voracek) — rainforest/thick-forest plots, sunset funnel plots, visual funnel-plot inference: https://CRAN.R-project.org/package=metaviz- multiverse-tools (Vesely) — generalized R/Python code + interactive dashboard for specification-curve and multiverse meta-analysis, adapting Voracek et al. (2019): https://github.com/d-vesely/multiverse-tools
- Visualization: Krähmer, D., & Young, C. (2026). Visualizing vastness: Graphical methods for multiverse analysis. PLOS ONE, 21(2), e0339452. https://doi.org/10.1371/journal.pone.0339452 (multiverse plots for very large specification spaces)
- SMART — Systematic Multiverse Analysis Registration Tool (Short, Inceler, et al., 2025): guided identification, documentation, and preregistration of pipelines
- Boba (Liu et al., 2020, IEEE TVCG) — language-agnostic multiverse authoring/visualization: https://github.com/uwdata/boba
- Milliways: Sarma, A., Hwang, K., Hullman, J., & Kay, M. (2024). Taming multiverses through principled evaluation of data analysis paths. CHI ’24. https://doi.org/10.1145/3613904.3642375 — interactive exploration; probabilistic vs. possibilistic uncertainty
- multiverse R package (formal citation): Sarma, A., Kale, A., Moon, M. J., Taback, N., Chevalier, F., Hullman, J., & Kay, M. (2023). Multiverse: Multiplexing alternative data analyses in R notebooks. CHI ’23. https://doi.org/10.1145/3544548.3580726
- SMART (formal citation): Short, C. A., Inceler, Y. C., Frank, M., & Hildebrandt, A. (2025). The Systematic Multiverse Analysis Registration Tool (SMART) for defining multiverse analyses. Royal Society Open Science, 12(1), 250800. https://doi.org/10.1098/rsos.250800
- Sampling (formal citation): Short, C. A., Hildebrandt, A., et al. (2025). Lost in a large EEG multiverse? Comparing sampling approaches for representative pipeline selection. Journal of Neuroscience Methods, 424, 110564. https://doi.org/10.1016/j.jneumeth.2025.110564
- Worked preregistration example: Sabey, H. (2023). A multiverse analysis on Wu et al. (2020). OSF. https://doi.org/10.17605/OSF.IO/BEVHG
- multitool (Young & Vermeent, 2024): https://ethan-young.github.io/multitool/
- comet (Burkhardt & Gießing, 2026) — network-neuroscience multiverses. Imaging Neuroscience. https://doi.org/10.1162/IMAG.a.1122
- mverse — student-friendly extension of the multiverse package: https://github.com/mverseanalysis/mverse
- Multiverse Sampling Tool (Short, Hildebrandt, et al., 2025) — pipeline sampling for infeasibly large multiverses: https://github.com/cassiesh/MultiverseSamplingTool
- Explorable multiverse documents: Dragicevic, P., Jansen, Y., Sarma, A., Kay, M., & Chevalier, F. (2019). CHI ’19. https://doi.org/10.1145/3290605.3300295
- Preregistration template (neuroimaging): Flournoy et al., https://johnflournoy.science/multiverse-preregistration/
- FORRT glossary — https://forrt.org/glossary/english/multiverse_analysis/
AI disclaimer: This resource collection was compiled with Claude. Claude ran the literature searches, pulled and cross-checked DOIs against PubMed and publisher pages, merged in my Zotero exports, and drafted the section summaries and per-reference notes, which I reviewed and edited. The selection of layers, the framing of the series, and anything written in my own voice elsewhere on this blog remain LLM-free — this page is the exception, and the reason is simple: verifying forty DOIs by hand is exactly the kind of work I am happy to delegate.