Layers of the research pipeline

A living, annotated collection of multiverse-style methods literature

Curated references on multiverse analysis, specification curves, vibration of effects, and many-analysts studies — organised by where in the research pipeline the forking happens.
Author

Constantin Yves Plessen

Modified

August 20, 2026

This is a curated, annotated reference collection on multiverse-style analysis — specification curves, vibration of effects, many-analysts studies, and their relatives — organized by where in the research pipeline the forking happens. It accompanies my blog series on arbitrary-yet-reasonable decisions in science.

How it is structured. Layer 0 collects the debate about whether we should multiverse at all — the critiques deserve top billing, and the strongest ones come first. A foundations section follows, opening with a chronological cartography of the guides and field maps the method has produced about itself. Then the pipeline: Layer 1 (psychometrics — how we measure what we measure), Layer 2 (primary studies — how we gather evidence), Layer 3 (meta-analysis — how we combine knowledge), and Layer 4 (meta-meta — how we combine the combinations). A software and tutorials section closes it out. Empirical entries carry a Paths varied note stating which specifications were explored and, where the paper reports it, which decisions proved consequential.

A living document. I intend to review and extend this list about once a year. It is certainly incomplete — the field moves fast, and search favors the canon over last quarter’s preprints. I am genuinely happy about suggestions and pointers, especially when a new paper or preprint appears: ideally via an issue or pull request on GitHub, or by email. If you think something belongs here, you are probably right.

⭐ marks the one entry per section to read first. A companion BibTeX file (multiverse_resources.bib, starter set) lives alongside this document for one-click import into your reference manager. On freezing each version I archive the collection on Zenodo for a citable DOI. Open territory notes in each layer flag where the map is still blank — these gaps are invitations.

Compiled 2026-08-20 (v1.0).


Layer 0 — Rationale & critiques: should we multiverse at all?

Before any layer of the pipeline, there is a prior question: whether running all reasonable analyses is the honest response to analytic flexibility, or an abdication of the researcher’s duty to identify the single best-justified analysis. This section collects both sides — the foundational critical frameworks, the strongest “there is only one correct analysis” position, the manufacturing-of-doubt worry, the garbage-in-garbage-out problem of ill-chosen specifications, and evidence that multiverse results themselves can fail to replicate — along with the statistical work on how to draw valid inferences across a multiverse at all.

  • ⭐ Del Giudice, M., & Gangestad, S. W. (2021). A traveler’s guide to the multiverse: Promises, pitfalls, and a framework for the evaluation of analytic decisions. AMPPS, 4(1). https://doi.org/10.1177/2515245920954925 (the foundational critical framework — equivalence vs. non-equivalence of specifications)
  • Lakens, D., Rasti, S., & Tunç, M. N. (2026). There is only one correct analysis. OSF Preprints (June 2026). (the strongest “turtles bottom out” position: epistemic uncertainty should be reduced, not embraced; argues for integrative experiments over multiverse/many-analyst approaches). https://osf.io/preprints/psyarxiv/hvjxf_v1
  • Modrák, M. (2026). Multiverse analysis, abdication of responsibility and manufacturing of doubt. arXiv:2607.14623. https://doi.org/10.48550/arXiv.2607.14623
  • Rohrer, J. M. (2021, March 7). Mülltiverse analysis. The 100% CI (blog). https://www.the100.ci/2021/03/07/mulltiverse-analysis/ (garbage specifications in, garbage multiverse out — and the “false sense of certainty” worry; also argues the multiverse may be most powerful as pedagogy)
  • Rohrer, J. M., Hullman, J., & Gelman, A. (2026). What’s a multiverse good for anyway? [preprint at https://osf.io/preprints/psyarxiv/37g29_v1; discussed on Statistical Modeling blog Feb 2026, https://statmodeling.stat.columbia.edu/2026/02/12/what-a-multiverse-good-for-anyway/ reflection/critique vs. persuasion vs. serious inference; fails as persuasion when researchers disagree on included variations, fails as inference without a coherent estimand]
  • Saltelli, A., Aleksankina, K., Becker, W., et al. (2019). Why so many published sensitivity analyses are false: A systematic review of sensitivity analysis practices. Environmental Modelling & Software, 114, 29–39. https://doi.org/10.1016/j.envsoft.2019.01.012 (the older sibling method, done badly at scale — a warning for multiverse practice)
  • Rubin, M. (2024). Preregistration does not improve the transparent evaluation of severity in Popper’s philosophy of science or when deviations are allowed. arXiv. https://doi.org/10.48550/arXiv.2408.12347
  • Boulesteix, A.-L., & Hoffmann, S. (2024). To adjust or not to adjust: it is not the tests performed that count, but how they are reported and interpreted. BMJ Medicine, 3(1). https://doi.org/10.1136/bmjmed-2023-000783
  • Olsson-Collentine, A., van Aert, R. C. M., Bakker, M., & Wicherts, J. (2025). Meta-analyzing the multiverse: A peek under the hood of selective reporting. Psychological Methods, 30(3), 441–461. https://doi.org/10.1037/met0000559 (multiverse effects often fail to replicate across samples; proposes underlying multiverse variability index — critique from within)Paths varied: counterfactual multiverses over 211 Registered Replication Report samples; only ~30% of studies contained any significant hypothesized effect, and a given researcher DF’s influence was inconsistent across samples — multiverse results themselves often fail to replicate.
  • Heyman, T., & Vanpaemel, W. (2022). Multiverse analyses in the classroom. Meta-Psychology. https://doi.org/10.15626/MP.2020.2718 (the pedagogy angle: teaching forks by making students walk them)
  • Heyman, T., Pronizius, E., Lewis, S. C., Acar, O. A., Adamkovič, M., et al. (2025, in press). Crowdsourcing multiverse analyses to explore the impact of different data-processing and analysis decisions: A tutorial. Psychological Methods. https://doi.org/10.1037/met0000770

Foundational papers (cross-layer)

The papers that defined the method and its vocabulary. These are not tied to any one layer: they introduce the multiverse and specification-curve idea, demonstrate empirically that independent analysts diverge on identical data, extend the logic to data-collection decisions, and map how the family of multiverse-style methods has spread across fields. Start here if the concept is new.

Cartography of the multiverse: guides and field maps

Attempts to map the territory itself — first field-specific guides, then landscape surveys, now a consolidated multidisciplinary standard. Read chronologically, they show a method growing self-aware:

  • 2016 — the founding map: Steegen, Tuerlinckx, Gelman, & Vanpaemel introduce the multiverse and walk through the first worked example (full citation below).

  • 2019 — first landscape, meta-analytic territory: Voracek, Kossmeier, & Tran chart which-data × how-to-analyze as the coordinate system for multiverse meta-analysis (full citation in Layer 3).

  • 2020 — the visualization landscape: Kossmeier, M., Tran, U. S., & Voracek, M. (2020). Charting the landscape of graphical displays for meta-analysis and systematic reviews: A comprehensive review, taxonomy, and feature analysis. BMC Medical Research Methodology, 20, 26. https://doi.org/10.1186/s12874-020-0911-9 (the Vienna group’s second map: every meta-analytic graph in existence — 150+ textbooks checked cover to cover — taxonomized; the display vocabulary that multiverse meta-analyses draw on)

  • 2021 — the traveler’s guide: Del Giudice & Gangestad supply the critical map legend — equivalence classes (Type E/N/U) that determine which paths belong on the map at all (full citation in Layer 0).

  • 2024 — the practical tutorial: Götz, Sarma, & O’Boyle, hands-on with the multiverse R package (full citation in Software).

  • 2025 — second landscape, the whole territory: Voracek, Tran, & Kern (2025). Mapping the landscape of multiverse-style data-analytic methods: A scoping review. PsyArXiv. https://doi.org/10.31234/osf.io/g4zck_v1 (424 sources; the first comprehensive survey of everything the method family has touched — still preprint, check for the journal version before citing)

  • 2025 — the book: Young & Cumberworth, Multiverse Analysis: Computational Methods for Robust Results (Cambridge UP; full citation below).

  • 2026 — the consolidated standard: Short, C. A., Breznau, N., Bruntsch, M., Burkhardt, M., Busch, N. A., Cesnaite, E., Frank, M., Gießing, C., Krähmer, D., Kristanto, D., Lonsdorf, T., Neuendorf, C., Nguyen, H. H. V., Rausch, M., Schmalz, X., Schneck, A., Tabakci, C., & Hildebrandt, A. (2026). Multicurious: A multidisciplinary guide to multiverse analysis. Advances in Methods and Practices in Psychological Science. https://doi.org/10.1177/25152459261434881 — The current state-of-the-art procedural guide (META-REP consortium): formalizes defensibility vs. equivalence (defensible vs. principled multiverse), harmonizes the terminology zoo (multiverse, SCA, VoE, multimodel, manyverse…), covers preregistration (SMART), pipeline similarity/pseudoreplication, and the covariates-as-estimand-shift debate. Open access; the section-by-section companion to this whole page.

  • ⭐ Steegen, S., Tuerlinckx, F., Gelman, A., & Vanpaemel, W. (2016). Increasing transparency through a multiverse analysis. Perspectives on Psychological Science, 11(5), 702–712. https://doi.org/10.1177/1745691616658637Paths varied: data-processing choices (exclusions, operationalizations of the outcome and moderators) on the fertility/religiosity dataset; showed the original finding survived only a minority of reasonable processing pipelines.

  • Simonsohn, U., Simmons, J. P., & Nelson, L. D. (2020). Specification curve analysis. Nature Human Behaviour, 4, 1208–1214. https://doi.org/10.1038/s41562-020-0912-zPaths varied: re-analyzed three published findings (e.g., Black-sounding names discrimination, female-named hurricanes) across all justified specifications; one finding proved robust, one weak, one not robust — inference via joint permutation test across the curve.

    • Earlier working-paper version: Simonsohn, Simmons, & Nelson (2015). Specification curve: Descriptive and inferential statistics on all reasonable specifications. SSRN. https://doi.org/10.2139/ssrn.2694998
  • Haaf, J. M., Hoogeveen, S., Berkhout, S., Gronau, Q. F., & Wagenmakers, E.-J. (2020). A Bayesian multiverse analysis of Many Labs 4: Quantifying the evidence against mortality salience. PsyArXiv. https://doi.org/10.31234/osf.io/cb9erPaths varied: Bayesian model specifications (priors, exclusion criteria) applied to Many Labs 4 mortality-salience data; evidence pointed against the effect across the multiverse.

  • Shunsen, H., Haojie, C., Xiaoxiong, L., Xinran, D., & Yun, W. (2023). Multiverse-style analysis: Introduction and application. Advances in Psychological Science, 31(2), 196. https://doi.org/10.3724/SP.J.1042.2023.00196 (Chinese-language introduction)

  • ⭐ Silberzahn, R., Uhlmann, E. L., Martin, D. P., et al. (2018). Many analysts, one data set. AMPPS, 1, 337–356. https://doi.org/10.1177/2515245917747646 (the 29-teams red-card study)Paths varied: 29 teams freely chose models, covariates, and operationalizations on identical data; effect estimates ranged from strong to null, with no single decision explaining the spread — the point is that defensible model choice itself is the fork.

  • Hoffmann, S., Schönbrodt, F., Elsas, R., Wilson, R., Strasser, U., & Boulesteix, A.-L. (2021). The multiplicity of analysis strategies jeopardizes replicability: lessons learned across disciplines. Royal Society Open Science, 8(4), 201925. https://doi.org/10.1098/rsos.201925 — Cross-disciplinary synthesis (psychology, finance, hydrology, biometrics): analysis-strategy multiplicity threatens replicability everywhere, arguing for multiverse-style reporting as a general remedy.

  • Harder, J. A. (2020). The multiverse of methods: Extending the multiverse analysis to address data-collection decisions. Perspectives on Psychological Science, 15(5), 1158–1177. https://doi.org/10.1177/1745691620917678

  • Aczel, B., Szaszi, B., Nilsonne, G., et al. (2021). Consensus-based guidance for conducting and reporting multi-analyst studies. eLife, 10, e72185.

  • Many-analysts studies (the empirical backbone across disciplines):

    • Botvinik-Nezer, R., et al. (2020). Variability in the analysis of a single neuroimaging dataset by many teams. Nature, 582, 84–88. https://doi.org/10.1038/s41586-020-2314-9 (NARPS: 70 fMRI teams, same data, divergent conclusions)
    • Breznau, N., Rinke, E. M., Wuttke, A., Nguyen, H. H., et al. (2022). Observing many researchers using the same data and hypothesis reveals a hidden universe of uncertainty. PNAS, 119(44), e2203150119. https://doi.org/10.1073/pnas.2203150119 (sociology’s many-analysts study; researcher characteristics barely explain the spread)
    • Muñoz, J., & Young, C. (2018). We ran 9 billion regressions: Eliminating false positives through computational model robustness. Sociological Methodology, 48(1), 1–33. https://doi.org/10.1177/0081175018777988
    • Schweinsberg, M., Feldman, M., Staub, N., van den Akker, O. R., van Aert, R. C. M., van Assen, M. A. L. M., et al. (2021). Same data, different conclusions: Radical dispersion in empirical results when independent analysts operationalize and test the same hypothesis. Organizational Behavior and Human Decision Processes. https://doi.org/10.1016/j.obhdp.2021.02.003 (the second-generation many-analysts study: dispersion arises already at operationalization, before any model is fit)
    • Trübutschek, D., Yang, Y. F., Gianelli, C., Cesnaite, E., et al. (2024). EEGManyPipelines: A large-scale, grassroots multi-analyst study of electroencephalography analysis practices in the wild. Journal of Cognitive Neuroscience, 36(2), 217–224. https://doi.org/10.1162/jocn_a_02087
  • Forking-paths canon & precursors:

    • Leamer, E. E. (1985). Sensitivity analyses would help. American Economic Review, 75(3), 308–313. (the 1985 ancestor of the whole idea)
    • Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22, 1359–1366. https://doi.org/10.1177/0956797611417632 (researcher degrees of freedom, the founding demonstration)
    • Gelman, A., & Loken, E. (2013). The garden of forking paths. Columbia University working paper. https://sites.stat.columbia.edu/gelman/research/unpublished/forking.pdf
    • Gelman, A., & Loken, E. (2014). The statistical crisis in science. American Scientist, 102(6), 460–465. https://doi.org/10.1511/2014.111.460
    • Bryan, C. J., Yeager, D. S., & O’Brien, J. M. (2019). Replicator degrees of freedom allow publication of misleading failures to replicate. PNAS, 116(51), 25535–25545. https://doi.org/10.1073/pnas.1910951116 (the forks cut both ways — replicators have them too)
    • Rubin, M. (2017). An evaluation of four solutions to the forking paths problem: Adjusted alpha, preregistration, sensitivity analyses, and abandoning the Neyman–Pearson approach. Review of General Psychology, 21, 321–329. https://doi.org/10.1037/gpr0000135
    • Young, C., & Holsteen, K. (2017). Model uncertainty and robustness: A computational framework for multimodel analysis. Sociological Methods & Research, 46(1), 3–40. https://doi.org/10.1177/0049124115610347
    • Wacker, J. (2017). Increasing the reproducibility of science through close cooperation and forking path analysis. Frontiers in Psychology, 8, 1332. https://doi.org/10.3389/fpsyg.2017.01332
    • Young, C., & Cumberworth, E. (2025). Multiverse analysis: Computational methods for robust results. Cambridge University Press. (the book-length treatment, from the “9 billion regressions” lineage)
    • Ramsey, J. B. (1969). Tests for specification errors in classical linear least-squares regression analysis. JRSS B, 31(2), 350–371. https://doi.org/10.1111/j.2517-6161.1969.tb00796.x (1969: the deepest root)
    • Neumayer, E., & Plümper, T. (2017). Robustness tests for quantitative research. Cambridge University Press.
    • Auspurg, K., & Brüderl, J. (2021). Has the credibility of the social sciences been credibly destroyed? Reanalyzing the “Many Analysts, One Data Set” project. Socius, 7. https://doi.org/10.1177/23780231211024421 (the estimand critique of Silberzahn: much of the spread came from teams answering different questions)
  • Robustness-reproduction frameworks:

    • Dreber, A., & Johannesson, M. (2025). A framework for evaluating reproducibility and replicability in economics. Economic Inquiry, 63(2), 338–356. https://doi.org/10.1111/ecin.13244
    • Brodeur, A., Cook, N., Hartley, J., & Heyes, A. (2024). Do pre-registration and pre-analysis plans reduce p-hacking and publication bias? Journal of Political Economy Microeconomics, 2(3), 527–561. https://doi.org/10.1086/730455
    • Ankel-Peters, J., Brodeur, A., Dreber, A., Johannesson, M., Neubauer, F., & Rose, J. (2024). A protocol for structured robustness reproductions and replicability assessments. Institute for Replication, No. 143.
    • Chen, G., Cai, Z., & Taylor, P. A. (2024). Through the lens of causal inference: Decisions and pitfalls of covariate selection. Aperture Neuro, 4. https://doi.org/10.52294/001c.124817 (the causal backbone of the covariates-as-decision-node debate)
  • Nepomuceno, A., Ghosal, A., Lentisco, A., & Ioannidis, J. (2026). Uptake and implementation of multiverse-style analyses across 613 studies. https://doi.org/10.64898/2026.07.15.738584

Inference across specifications

  • Girardi, P., Vesely, A., Lakens, D., Altoè, G., Pastore, M., Calcagnì, A., & Finos, L. (2024). Post-selection inference in multiverse analysis (PIMA): An inferential framework based on the sign flipping score test. Psychometrika, 89(2), 542–568. https://doi.org/10.1007/s11336-024-09973-6
  • Mandl, M. M., Becker-Pennrich, A. S., Hinske, L. C., Hoffmann, S., & Boulesteix, A.-L. (2024). Addressing researcher degrees of freedom through minP adjustment. BMC Medical Research Methodology, 24(1), 152. https://doi.org/10.1186/s12874-024-02279-2
  • Bartoš, F., Hoogeveen, S., Sarafoglou, A., & Pawel, S. (2025). Single-dataset meta-analysis for many-analysts and multiverse studies. arXiv:2511.17064. https://doi.org/10.48550/arXiv.2511.17064
  • Vesely, A., & Andreella, A., et al. (2026). PIMAX: Post-selection inference for multiverse analysis in mixed-effects models. arXiv:2607.03225
  • Ozenne, B., Nørgaard, M., Pernet, C., & Ganz, M. (2025). A sensitivity analysis of preprocessing pipelines: Toward a solution for multiverse analyses. Imaging Neuroscience, 3, imag_a_00523. https://doi.org/10.1162/imag_a_00523

Layer 1 — Psychometrics: how we measure what we measure

The forks begin before any data are analyzed, inside the instruments themselves. This layer collects multiverse-style work on measurement: differential item functioning assessed across the full space of defensible detection choices, test-validation coefficients that shift with design decisions, reliability that moves with data-processing rules, and the field’s own ambivalence about whether advanced psychometric models add value over simple sum scores. Open territory: systematic multiverses over scale choice itself (which of 280 depression scales) remain rare; empirical equivalence assessment for measurement pipelines is underdeveloped.

  • ⭐ Flake, J. K., & Fried, E. I. (2020). Measurement schmeasurement: Questionable measurement practices and how to avoid them. AMPPS, 3(4), 456–465. https://doi.org/10.1177/2515245920952393 — QMPs as a “stunning source of researcher degrees of freedom”: the measurement layer’s manifesto, arguing the replication debate obsesses over statistics while the deeper forks sit in how constructs are measured at all.
  • Plessen, C. Y., Fischer, F., Hartmann, C., Liegl, G., Schalet, B., Kaat, A. J., Pesantez, R., Joeris, A., Heng, M., Rose, M., & the AOBERT Consortium (2024). Differential item functioning between English, German, and Spanish PROMIS® physical function ceiling items. Quality of Life Research, 34(5). https://doi.org/10.1007/s11136-024-03866-y (DIF assessed via multiverse across the full plausible space of logistic ordinal regression choices)Paths varied: the full plausible space of logistic ordinal regression DIF choices (criteria, thresholds, anchor decisions) across English/German/Spanish samples; DIF-item identification proved robust across specifications — multiverse as an answer to arbitrary DIF cutoffs.
    • Companion software: lordifMultiverse R package — github.com/cyplessen
  • Webb, S. S., & Demeyere, N. (2022). Using multiverse analysis to highlight differences in convergent correlation outcomes due to data analytical and study design choices. Assessment. https://doi.org/10.1177/10731911221127904 (2,220 analyses on psychometric test validation)Paths varied: 2,220 analyses crossing sample group (healthy vs. stroke), sample size, test metrics, and covariate inclusion for convergent-validity correlations of an executive-function test — validation coefficients themselves vibrate with design choices.
  • Byrne, et al. (2026). ‘We’re not going to start lifting stones now…’: Stakeholder perspectives on the role of psychometric methods in outcome measurement. British Journal of Clinical Psychology, 1–16. https://doi.org/10.1111/bjc.70067 — The “added benefit” question put to the field itself: 21 stakeholder interviews across psychometrics, clinical practice, applied research, and statistics on whether IRT/SEM add value over sum scores in applied work — the perceived-benefit side of the fork this layer maps. [Complete author list from the PDF]
  • Harrison, C. J., Plessen, C. Y., Liegl, G., Rodrigues, J. N., Sabah, S. A., Cook, J. A., Beard, D. J., & Fischer, F. (2023). Item response theory may account for unequal item weighting and individual-level measurement error in trials that use PROMs: A psychometric sensitivity analysis of the TOPKAT trial. Journal of Clinical Epidemiology, 158, 62–69. https://doi.org/10.1016/j.jclinepi.2023.03.013Paths varied: re-scoring a real RCT’s PROM outcomes under IRT vs. sum scores — the “does the fancy model change trial conclusions?” question tested on actual trial data rather than argued in the abstract; the empirical companion to the Byrne perception study.
  • Hanel, P. H. P., & Zarzeczna, N. (2023). From multiverse analysis to multiverse operationalisations: 262,143 ways of measuring well-being. Religion, Brain & Behavior, 13(3), 309–313. https://doi.org/10.1080/2153599X.2022.2070259Paths varied: which items/subscales operationalize “well-being” at all — the construct’s operationalization as its own combinatorial multiverse, upstream of any model choice.
  • Parsons, S. (2020). Exploring reliability heterogeneity with multiverse analyses: Data processing decisions unpredictably influence measurement reliability. PsyArXiv. https://doi.org/10.31234/osf.io/y6tcz (reliability itself as a multiverse outcome — perfect Layer 1 fit)Paths varied: data-processing decisions (trial-level exclusions, outlier rules, scoring) for cognitive-behavioral task measures; reliability estimates shifted unpredictably — the instrument’s own precision is specification-dependent.
  • Bauer, D. J., Belzak, W. C., & Cole, V. T. (2020). Regularized MNLFA to detect DIF. Structural Equation Modeling, 27(1), 43–55. https://doi.org/10.1080/10705511.2019.1642754
  • Preprocessing multiverses in biophysiological measurement (EEG/ERP/fMRI/SCR):
    • Clayson, P. E., Baldwin, S. A., Rocha, H. A., & Larson, M. J. (2021). The data-processing multiverse of event-related potentials (ERPs): A roadmap. NeuroImage, 245, 118712. https://doi.org/10.1016/j.neuroimage.2021.118712
    • Feuerriegel, D., & Bode, S. (2022). Bring a map when exploring the ERP data processing multiverse. NeuroImage, 259, 119443. https://doi.org/10.1016/j.neuroimage.2022.119443 (why effect size must never be the pipeline-ranking criterion)
    • Clayson, P. E. (2024). Beyond single paradigms, pipelines, and outcomes: Embracing multiverse analyses in psychophysiology. International Journal of Psychophysiology, 197, 112311. https://doi.org/10.1016/j.ijpsycho.2024.112311
    • Kuhn, M., Gerlicher, A. M., & Lonsdorf, T. B. (2022). Navigating the manyverse of skin conductance response quantification. Psychophysiology, 59(9), e14058. https://doi.org/10.1111/psyp.14058
    • Jacobsen, N. S., Kristanto, D., Welp, S., Inceler, Y. C., & Debener, S. (2025). Preprocessing choices for P3 analyses with mobile EEG: Systematic review and interactive exploration. Psychophysiology, 62(1), e14743. https://doi.org/10.1111/psyp.14743
    • Kristanto, D., Burkhardt, M., Thiel, C. M., Debener, S., Gießing, C., & Hildebrandt, A. (2024). The multiverse of data preprocessing and analysis in graph-based fMRI: Systematic review + decision support tool. Neuroscience & Biobehavioral Reviews, 165, 105846. https://doi.org/10.1016/j.neubiorev.2024.105846
    • Dafflon, J., et al. (2022). A guided multiverse study of neuroimaging analyses. Nature Communications, 13, 3758. https://doi.org/10.1038/s41467-022-31347-8 (ML sampling when the multiverse is too big to compute)
    • Paul, K., Short, C. A., Beauducel, A., et al. (2022). The methodology and dataset of the CoScience EEG-personality project — a large-scale, multi-laboratory project grounded in cooperative forking paths analysis. Personality Science, 3(1), e7177. https://doi.org/10.5964/ps.7177
    • Paul, K., Beauducel, A., et al. (2025). Frontal alpha asymmetry as a marker of approach motivation? Insights from a cooperative forking path analysis. JPSP, 128(1), 196–210. https://doi.org/10.1037/pspp0000503 (preregistered cooperative multiverse settling a classic EEG-marker question)
  • Fried E. I. (2017). The 52 symptoms of major depression: Lack of content overlap among seven common depression scales. Journal of affective disorders, 208, 191–197. https://doi.org/10.1016/j.jad.2016.10.019

Layer 2 — Primary studies: how we gather evidence

One dataset, one research question, thousands of defensible analyses. This layer covers multiverse and specification-curve applications to primary studies — observational and experimental — including the screen-time debate fought entirely in specification curves, and the vibration-of-effects tradition from epidemiology, which quantifies how much estimates move under different covariate-adjustment, measurement, and sampling choices. Open territory: multiverse analyses of full clinical-trial analysis pipelines (model × missing-data strategy) are scarce relative to observational work.

  • Orben, A., & Przybylski, A. K. (2019). The association between adolescent well-being and digital technology use. Nature Human Behaviour. https://doi.org/10.1038/s41562-018-0506-1Paths varied: thousands of specifications across three large adolescent datasets (which well-being measures, which technology measures, which covariates); the tiny negative association shrank or vanished depending on mostly arbitrary choices — comparable in size to eating potatoes.
  • Twenge, J. M., Haidt, J., Lozano, J., & Cummins, K. M. (2022). Specification curve analysis shows that social media use is linked to poor mental health, especially among girls. Acta Psychologica, 224, 103512. https://doi.org/10.1016/j.actpsy.2022.103512 (the counter-SCA)Paths varied: re-ran Orben & Przybylski’s SCA with four alterations; the consequential decisions were separating social media from TV within “screen time” and how one multi-subscale mental-health measure was weighted — with those changed, associations for girls emerged more clearly.
  • Rengasamy, M., Moriarity, D., Kraynak, T., Tervo-Clemmens, B., & Price, R. (2023). Exploring the multiverse: the impact of researchers’ analytic decisions on relationships between depression and inflammatory markers. Neuropsychopharmacology : official publication of the American College of Neuropsychopharmacology, 48(10), 1465–1474. https://doi.org/10.1038/s41386-023-01621-4 (58,000+ specifications)Paths varied: 9 common analytic decisions (log-transformation, covariate count, outlier handling, etc.) yielding 58,000+ combinations on NHANES inflammatory-marker/depression data, testing robustness of the association across all of them.
  • Lonsdorf, T. B., Gerlicher, A., Klingelhöfer-Jens, M., & Krypotos, A.-M. (2022). Multiverse analyses in fear conditioning research (+ multifear). Behaviour Research and Therapy, 153, 104072. https://doi.org/10.1016/j.brat.2022.104072Paths varied: statistical model family (Bayesian vs. frequentist ANOVA/t-tests/mixed models) crossed with data-reduction approaches for SCR fear-conditioning data; both the size and direction of effects shifted with model and reduction choice.
  • Stern, J., Arslan, R. C., Gerlach, T. M., & Penke, L. (2019). No robust evidence for cycle shifts in preferences for men’s bodies in a multiverse analysis. Evolution and Human Behavior, 40(6), 517–525. https://doi.org/10.1016/j.evolhumbehav.2019.08.005Paths varied: specifications for cycle-phase estimation and preference outcomes; the claimed cycle shifts did not survive the multiverse.
  • Engzell, P., & Mood, C. (2023). Understanding patterns and trends in income mobility through multiverse analysis. American Sociological Review, 88(4), 600–626. https://doi.org/10.1177/00031224231180607 (flagship sociology application)
  • Lefort-Besnard, J., Nichols, T. E., & Maumet, C. (2025). Statistical inference for same-data meta-analysis in neuroimaging multiverse analyses. Imaging Neuroscience. https://doi.org/10.1162/imag_a_00513
  • Vibration of effects family (observational designs):
    • ⭐ Patel, C. J., Burford, B., & Ioannidis, J. P. A. (2015). Assessment of vibration of effects due to model specification. J Clin Epidemiol, 68(9), 1046–1058. https://doi.org/10.1016/j.jclinepi.2015.05.029Paths varied: all 8,192 combinations of 13 adjustment covariates in Cox models for 417 variables’ associations with mortality; adjustment-set choice alone flipped significance and even sign for a nontrivial share of associations (their “Janus effect”).
    • Klau, S., Hoffmann, S., Patel, C. J., Ioannidis, J. P., & Boulesteix, A.-L. (2021). Model, measurement and sampling uncertainty in the VoE framework. IJE, 50(1), 266–278. https://doi.org/10.1093/ije/dyaa164Paths varied: extends VoE to decompose result variability into model choice, measurement error, and sampling uncertainty in one framework — showing how much of the “vibration” each source contributes.
    • Chu, L., Ioannidis, J. P. A., et al. (2020). VoE in epidemiologic studies of alcohol and breast cancer. IJE, 49(2), 608–618. https://doi.org/10.1093/ije/dyz271Paths varied: analytic approaches within and across published observational studies of alcohol and breast cancer, quantifying how much reported estimates vibrate with model choices.
    • Tierney, B. T., et al. (2021). Leveraging VoE analysis for robust discovery. PLOS Biology, 19, e3001398. https://doi.org/10.1371/journal.pbio.3001398Paths varied: covariate-adjustment vibration deployed prospectively as a discovery filter in large biomedical datasets — VoE turned from critique into a robustness screen.
    • Vinatier, C., Hoffmann, S., Patel, C., DeVito, N. J., Cristea, I. A., Tierney, B., Ioannidis, J. P. A., & Naudet, F. (2024). What is the vibration of effects? BMJ Evidence-Based Medicine. https://doi.org/10.1136/bmjebm-2023-112747 (accessible explainer) — Explainer, not an empirical multiverse: defines VoE and situates it among multi-analyst and multiverse approaches.

Layer 3 — Meta-analysis: how we combine knowledge

Every fork from the layers below arrives here, plus new ones: which studies to include, which effect size to compute, which model to pool with, which bias corrections to apply, which outliers to exclude. This is currently the most colonized layer — three independent groups (Vienna, Rennes, Amsterdam) have converged on multiverse meta-analysis from psychology, epidemiology, and clinical trials — alongside the Bayesian model-averaging alternative, which resolves the forks by weighting rather than mapping them. Open territory: no dedicated multiverse over risk-of-bias assessment methods exists yet.

  • ⭐ Voracek, M., Kossmeier, M., & Tran, U. S. (2019). Which data to meta-analyze, and how? Zeitschrift für Psychologie, 227(1), 64–82. https://doi.org/10.1027/2151-2604/a000357 (the founding multiverse-MA paper) — Framework paper: crosses which data (combinatorial study inclusion) with how (model and estimator choice) to define the meta-analytic multiverse; the template your own MAs and Pietschnig’s implement.

  • Plessen multiverse meta-analyses:

    • Plessen, C. Y., Karyotaki, E., & Cuijpers, P. (2022). Protocol. BMJ Open, 12(1). https://doi.org/10.1136/bmjopen-2021-050197
    • ⭐ Plessen, C. Y., Karyotaki, E., Miguel, C., Ciharova, M., & Cuijpers, P. (2023). Exploring the efficacy of psychotherapies for depression: a multiverse meta-analysis. BMJ Mental Health, 26, e300626. https://doi.org/10.1136/bmjment-2022-300626Paths varied: all plausible meta-analyses from the Metapsy depression-trial pool; consequential forks were inclusion of high risk-of-bias studies, waitlist comparators, and (not) adjusting for publication bias — each inflating the summary effect.
    • Plessen, C. Y., Panagiotopoulou, O. M., Tong, L., Cuijpers, P., & Karyotaki, E. (2024). Digital mental health interventions for depression: a multiverse meta-analysis. Journal of Affective Disorders, 369, 1031–1044. https://doi.org/10.1016/j.jad.2024.10.018Paths varied: 3,638 meta-analyses over 125 RCTs (populations, intervention characteristics, designs); g averaged 0.43 and stayed positive from the 10th to 90th percentile — larger with adults, LMICs, guidance, waitlist controls; smaller with publication-bias adjustment and ≥24-week outcomes.
  • Pietschnig group (Vienna) multiverse meta-analyses:

    • Pietschnig, J., Gerdesmann, D., Zeiler, M., & Voracek, M. (2022). The meta-analytical multiverse of brain volume and IQ associations. Royal Society Open Science, 9(5), 211621. https://doi.org/10.1098/rsos.211621 (432 specifications)Paths varied: 432 specifications from 27 which-data choices (healthy vs. clinical samples, correction types, g-ness of measures) × 16 how-to-analyze choices (Hedges–Olkin, Hunter–Schmidt, unweighted, RVE); the brain-volume/IQ association held across most, but its magnitude varied enough to explain prior meta-analytic disagreement.
    • Dürlinger, F., & Pietschnig, J. (2022). Meta-analyzing intelligence and religiosity associations: Evidence from the multiverse. PLOS ONE. https://doi.org/10.1371/journal.pone.0262699Paths varied: 192 specifications; 70.4% of summary effects significant and all of those negative — direction robust; strength moved with intelligence measure type (psychometric tests vs. GPA proxies) and sample type (pre-college vs. adult).
    • Oberleiter, S., & Pietschnig, J. (2023). The Mozart effect myth: a multiverse meta-analysis. Scientific Reports, 13, 3175. https://doi.org/10.1038/s41598-023-30206-wPaths varied: three independent analytic approaches on k=8 usable studies of Mozart KV448 and epilepsy; trivial-to-small nonsignificant effects throughout — the multiverse as myth-busting, with reporting opacity blocking even inclusion.
    • Patzl, S., Oberleiter, S., & Pietschnig, J. (2024). Self-assessed intelligence through the lens of the multiverse. Journal of Intelligence, 12(9), 81. https://doi.org/10.3390/jintelligence12090081Paths varied: 278 effect sizes across specification-curve and combinatorial analyses of the self-assessed/psychometric intelligence link, probing generality across SAI measurement methods and samples.
    • Fries, J., Oberleiter, S., Bodensteiner, F.A. et al. Multilevel multiverse meta-analysis indicates lower IQ as a risk factor for physical and mental illness. Commun Psychol 3, 74 (2025). https://doi.org/10.1038/s44271-025-00245-2
    • Gerdesmann, D., & Pietschnig, J. (2025). Brain volume–IQ multiverse update. OSF: osf.io/y6msp
  • Naudet group (Rennes) — VoE applied to meta-analysis and trial data:

    • Palpacuer, C., Hammas, K., Duprez, R., Laviolle, B., Ioannidis, J. P. A., & Naudet, F. (2019). Vibration of effects from diverse inclusion/exclusion criteria and analytical choices: 9216 different ways to perform an indirect comparison meta-analysis. BMC Medicine, 17(1), 174. https://doi.org/10.1186/s12916-019-1409-3Paths varied: 9,216 indirect-comparison meta-analyses of nalmefene vs. naltrexone from crossed inclusion/exclusion criteria and analytic choices; conclusions about comparative efficacy depended on these review-level decisions.
    • El Bahri, M., Wang, X., Biaggi, T., Falissard, B., Naudet, F., & Barry, C. (2022). A multiverse analysis of meta-analyses assessing acupuncture efficacy for smoking cessation evidenced vibration of effects. J Clin Epidemiol, 152, 140–150. https://doi.org/10.1016/j.jclinepi.2022.09.001Paths varied: review-level choices across meta-analyses of acupuncture for smoking cessation; documented VoE at the synthesis level.
    • Gouraud, H., Wallach, J. D., Boussageon, R., Ross, J. S., & Naudet, F. (2022). Vibration of effect in more than 16,000 pooled analyses of individual participant data from 12 RCTs comparing canagliflozin and placebo: multiverse analysis. BMJ Medicine, 1(1). https://doi.org/10.1136/bmjmed-2022-000154 (IPD-level multiverse — bridges Layers 2 and 3)Paths varied: >16,000 pooled analyses of IPD from 12 canagliflozin RCTs (endpoint definitions, populations, analytic choices); shows vibration exists even with gold-standard trial data pooled at the individual level.
  • Other multiverse-MA applications:

    • Kang, H., Sartorius, A. M., Deilhaug, E., Walle, K. M., & Quintana, D. S. (2025). A multiverse meta-analysis of intranasal oxytocin studies. Psychoneuroendocrinology, 172, 107312.
    • Sanabria, D., Ciria, L., Holgado, D., Bartoš, F., Luque-Casado, A., Fernández-Del-Olmo, M., Perakakis, P., & Román-Caballero, R. (2026). Running ahead of the evidence? Rethinking the exercise-cognition consensus. Journal of Science and Medicine in Sport. https://doi.org/10.1016/j.jsams.2026.07.003
  • Robust Bayesian / model-averaging family (Bartoš–Maier–Wagenmakers):

    • Bartoš F, Maier M, Wagenmakers E-J, Doucouliagos H, Stanley TD. Robust Bayesian meta-analysis: Model-averaging across complementary publication bias adjustment methods. Res Syn Meth. 2023; 14(1): 99-116. https://doi.org/10.1002/jrsm.1594
    • Maier, M., Bartoš, F., & Wagenmakers, E.-J. (2023). Robust Bayesian meta-analysis: Addressing publication bias with model-averaging. Psychological Methods, 28(1), 107–122. https://doi.org/10.1037/met0000405
    • Bartoš, F., Maier, M., Stanley, T. D., & Wagenmakers, E.-J. (2025). Robust Bayesian meta-regression. Psychological Methods. [psycnet 2025-81744-001]
    • RoBMA vignettes: https://fbartos.github.io/RoBMA/articles/ (the former CustomEnsembles vignette was renamed upstream; see v20-bayesian-model-averaging)
    • Critique: Román-Caballero, R., & Vadillo, M. A. (2024). A meta-analyst should make informed decisions: Issues with Bayesian model-averaging meta-analyses. OSF preprint https://osf.io/tm7dv
  • Harrer, M., Miguel, C., et al. (2025). Standardized effect sizes are far from “standardized”. PLOS Mental Health. https://doi.org/10.1371/journal.pmen.0000347

  • Harrer, M., Miguel, C., Hussey, I., Cristea, I. A., van Ballegooijen, W., Basic, D., Wang, Y., Pfund, R. A., Quero, S., van Spreckelsen, P., Schnurr, P. P., van Straten, A., Furukawa, T. A., Papola, D., & Cuijpers, P. (2025). Implausible effects of psychological interventions: Meta-epidemiological study and development of a simple flagging tool. medRxiv. https://doi.org/10.1101/2025.11.12.25340062 — The outlier fork made principled: instead of ad-hoc exclusion rules, a meta-epidemiological flagging tool for effects too large to be plausible — turning “which studies do we drop?” from an arbitrary decision into a criterion. Paths varied: ways to compute the “same” standardized mean difference in depression MAs (SD pooling, change vs. endpoint, correlation assumptions); the effect size itself is a fork before any pooling happens.

  • Network meta-analysis:

Layer 4 — Meta-meta: how we combine combinations of evidence

The thinnest layer, and possibly the next frontier: umbrella reviews and overviews of meta-analyses, where the units of analysis are themselves syntheses stacked on syntheses. Multiverse-style analysis at this level barely exists — the entries here are early examples and infrastructure, and the gap itself is an argument. Open territory: essentially everything — umbrella-level multiverse analysis barely exists.

  • Gougeon, A., Aribi, I., Guernouche, S., Lega, J. C., Wright, J. M., Verstuyft, C., Lajoinie, A., Gueyffier, F., & Grenet, G. (2025). Publication bias in pharmacogenetics of statin-associated muscle symptoms: A meta-epidemiological study. Atherosclerosis, 400, 118624. https://doi.org/10.1016/j.atherosclerosis.2024.118624
  • Cuijpers, P., Miguel, C., Harrer, M., Plessen, C. Y., Ciharova, M., Papola, D., Ebert, D., & Karyotaki, E. (2023). Psychological treatment of depression: A systematic overview of a ‘Meta-Analytic Research Domain’. Journal of Affective Disorders, 335, 141–151. https://doi.org/10.1016/j.jad.2023.05.011
  • I am currently working on an umbrella review with multiverse meta-analysis on digital interventions for anxiety disorders

Applications beyond psychology & medicine

The method has escaped its home field. A non-exhaustive sample of where multiverse thinking has landed — and where this list will grow fastest:

  • Fairness in machine learning: One model, many scores: Using multiverse analysis to prevent fairness hacking and evaluate the influence of model design decisions. arXiv:2308.16681
  • Bibliometrics: Specification uncertainty: What the disruption index tells us about the (hidden) multiverse of bibliometric indicators. arXiv:2406.13367
  • Computational social science: Making uncertainty visible: Multiverse analysis for robust computational social science. arXiv:2605.19745
  • Sociology: Engzell & Mood on income mobility (see Layer 2); Breznau et al. and Muñoz & Young (see Foundations)
  • Neuroimaging & psychophysiology: see the preprocessing-multiverse subsection in Layer 1
  • AI benchmark evaluation: Plessen — multiverse/specification-curve analysis of LLM leaderboard scoring pipelines (~123,000 specifications) [will add preprint link when public]; companion IRT/DIF/measurement-invariance analysis of Open LLM Leaderboard data [will add preprint link]

Software / tutorials

The tooling: R packages for building, running, and visualizing multiverses at each layer, hands-on tutorials, and reporting guidance.


AI disclaimer: This resource collection was compiled with Claude. Claude ran the literature searches, pulled and cross-checked DOIs against PubMed and publisher pages, merged in my Zotero exports, and drafted the section summaries and per-reference notes, which I reviewed and edited. The selection of layers, the framing of the series, and anything written in my own voice elsewhere on this blog remain LLM-free — this page is the exception, and the reason is simple: verifying forty DOIs by hand is exactly the kind of work I am happy to delegate.