Research

Publications, preprints, and conference presentations.

Plain-language summaries in English and German are available on selected first-author papers.

Google Scholar · ORCID · CV

First Author Publications

Plessen, C. Y. (2024). Research Synthesis of Psychological Interventions for Mental Health (Doctoral dissertation). Vrije Universiteit Amsterdam.

Plessen, C. Y., Fischer, F., Hartmann, C., Liegl, G., Rose, M., Garrido, C. J., Terrill, A. L., Kravitz, R. L., Zumbrunn, T., Latimer, S., Amtmann, D., & Morgan, E. M. (2025). Differential item functioning between English, German, and Spanish PROMIS® physical function ceiling items. Quality of Life Research, 34, 1377-1391.

What did we do? We tested whether health questionnaires work the same way in English, German, and Spanish. We looked at questions about physical activities that healthy people can usually do easily.

Why does it matter? Doctors and researchers need to know if questionnaires measure the same thing across different languages. If they don’t, we can’t fairly compare health between countries or language groups.

What did we find? Most questions worked similarly across all three languages. However, 9 out of 35 questions worked slightly differently - especially questions about running and jumping. This means we need to be careful when comparing these specific questions across languages.

Was haben wir gemacht? Wir haben getestet, ob Gesundheitsfragebögen in Englisch, Deutsch und Spanisch gleich funktionieren. Wir haben Fragen zu körperlichen Aktivitäten untersucht, die gesunde Menschen normalerweise leicht machen können.

Warum ist das wichtig? Ärzte und Forscher müssen wissen, ob Fragebögen in verschiedenen Sprachen dasselbe messen. Wenn nicht, können wir die Gesundheit zwischen Ländern oder Sprachgruppen nicht fair vergleichen.

Was haben wir herausgefunden? Die meisten Fragen funktionierten in allen drei Sprachen ähnlich. Aber 9 von 35 Fragen funktionierten etwas anders - besonders Fragen über Rennen und Springen. Das bedeutet, wir müssen vorsichtig sein, wenn wir diese speziellen Fragen zwischen Sprachen vergleichen.

Plessen, C. Y., Panagiotopoulou, O. M., Tong, L., Ciharova, M., & Cuijpers, P. (2024). Digital mental health interventions for the treatment of depression: A multiverse meta-analysis. Journal of Affective Disorders, 369, 1031-1044.

What did we do? We combined results from 125 studies (with over 32,000 people) to see if apps and websites can help people with depression. We tested thousands of different ways to analyze the data to see if the results hold up.

Why does it matter? Depression affects millions of people worldwide. Digital tools like apps could help many more people than traditional therapy alone, but we need to know if they really work.

What did we find? Digital interventions help reduce depression symptoms with a moderate effect. The results were consistent across thousands of different analyses. They work best when: someone guides you through the program, you’re an adult, you have diagnosed depression, and the program compares to no treatment. The effects last at least 24 weeks.

Was haben wir gemacht? Wir haben Ergebnisse von 125 Studien (mit über 32.000 Menschen) kombiniert, um zu sehen, ob Apps und Websites Menschen mit Depressionen helfen können. Wir haben tausende verschiedene Wege getestet, die Daten zu analysieren, um zu sehen, ob die Ergebnisse stabil sind.

Warum ist das wichtig? Depressionen betreffen Millionen von Menschen weltweit. Digitale Werkzeuge wie Apps könnten viel mehr Menschen helfen als traditionelle Therapie allein - aber wir müssen wissen, ob sie wirklich funktionieren.

Was haben wir herausgefunden? Digitale Interventionen helfen, Depressionssymptome mit einem mittleren Effekt zu reduzieren. Die Ergebnisse waren über tausende verschiedene Analysen hinweg konsistent. Sie funktionieren am besten wenn: jemand Sie durch das Programm begleitet, Sie erwachsen sind, Sie eine diagnostizierte Depression haben, und das Programm mit keiner Behandlung verglichen wird. Die Effekte halten mindestens 24 Wochen an.

Plessen, C. Y., Karyotaki, E., Miguel, C., Ciharova, M., & Cuijpers, P. (2023). Exploring the efficacy of psychotherapies for depression: a multiverse meta-analysis. BMJ Mental Health, 26(1), e300626.

What did we do? We analyzed 92 studies comparing psychotherapy to no treatment for depression. Instead of making one analysis, we ran 4,281 different analyses using different reasonable choices researchers could make.

Why does it matter? Research results can change depending on how researchers analyze their data. We wanted to see if psychotherapy for depression works no matter how you analyze the data, or if results depend on researcher choices.

What did we find? Good news: Psychotherapy works for depression across all 4,281 analyses. However, the size of the effect varied quite a bit (from small to large) depending on analysis choices. Studies with weaker methods, waitlist control groups, and no correction for publication bias showed bigger effects. This means we need to be transparent about our choices when doing research.

Was haben wir gemacht? Wir haben 92 Studien analysiert, die Psychotherapie mit keiner Behandlung für Depression verglichen. Anstatt eine Analyse zu machen, haben wir 4.281 verschiedene Analysen mit unterschiedlichen vernünftigen Entscheidungen durchgeführt, die Forscher treffen könnten.

Warum ist das wichtig? Forschungsergebnisse können sich ändern, je nachdem wie Forscher ihre Daten analysieren. Wir wollten sehen, ob Psychotherapie für Depression funktioniert, egal wie man die Daten analysiert, oder ob Ergebnisse von Forscherentscheidungen abhängen.

Was haben wir herausgefunden? Gute Nachricht: Psychotherapie funktioniert für Depression über alle 4.281 Analysen hinweg. Aber die Größe des Effekts variierte ziemlich stark (von klein bis groß), abhängig von Analyseentscheidungen. Studien mit schwächeren Methoden, Wartelisten-Kontrollgruppen und ohne Korrektur für Publikationsbias zeigten größere Effekte. Das bedeutet, wir müssen transparent über unsere Entscheidungen sein, wenn wir forschen.

Plessen, C. Y., Hartmann, C., Heng, M., Liegl, G., Fischer, F., & Rose, M. (2023). How are age, gender, and country differences associated with PROMIS physical function, upper extremity, and pain interference scores? Clinical Orthopaedics and Related Research, 482(2), 244-256.

What did we do? We created reference values for health questionnaires about physical function and pain. We studied how age, gender, and country (USA, UK, Germany) affect people’s scores, using data from over 15,000 people.

Why does it matter? When a patient gets a score on a health questionnaire, doctors need to know: “Is this normal for someone their age and gender?” Reference values help doctors interpret scores and make better treatment decisions.

What did we find? Physical function decreases with age, especially for upper body activities. Women report slightly more difficulty with physical tasks than men. Country differences were small. We built a website tool where doctors can enter a patient’s age, gender, and country to see how their score compares to similar people. This helps make scores meaningful and actionable.

Was haben wir gemacht? Wir haben Referenzwerte für Gesundheitsfragebögen über körperliche Funktion und Schmerzen erstellt. Wir haben untersucht, wie Alter, Geschlecht und Land (USA, UK, Deutschland) die Punktzahlen von Menschen beeinflussen, mit Daten von über 15.000 Menschen.

Warum ist das wichtig? Wenn ein Patient eine Punktzahl auf einem Gesundheitsfragebogen bekommt, müssen Ärzte wissen: “Ist das normal für jemanden in diesem Alter und Geschlecht?” Referenzwerte helfen Ärzten, Punktzahlen zu interpretieren und bessere Behandlungsentscheidungen zu treffen.

Was haben wir herausgefunden? Die körperliche Funktion nimmt mit dem Alter ab, besonders für Oberkörperaktivitäten. Frauen berichten etwas mehr Schwierigkeiten mit körperlichen Aufgaben als Männer. Länderunterschiede waren klein. Wir haben ein Website-Tool gebaut, wo Ärzte Alter, Geschlecht und Land eines Patienten eingeben können, um zu sehen, wie ihre Punktzahl im Vergleich zu ähnlichen Menschen ist. Dies hilft, Punktzahlen bedeutungsvoll und umsetzbar zu machen.

Plessen, C. Y., Karyotaki, E., & Cuijpers, P. (2022). Exploring the efficacy of psychological treatments for depression: a multiverse meta-analysis protocol. BMJ Open, 12(1), e050197.

What did we do? We published our plan (called a protocol) for how we would study psychotherapy for depression using multiverse analysis. We described our methods before seeing the results, to be transparent.

Why does it matter? When researchers decide how to analyze data after seeing results, they might (consciously or not) choose methods that make results look better. By publishing our plan first, we commit to our methods and make our research more trustworthy.

What did we find? This is a protocol paper - it describes what we planned to do. The actual results were published in 2023 (see that paper above). Publishing protocols like this is good research practice and helps prevent bias in science.

Was haben wir gemacht? Wir haben unseren Plan (genannt Protokoll) veröffentlicht, wie wir Psychotherapie für Depression mit Multiversum-Analyse studieren würden. Wir haben unsere Methoden beschrieben, bevor wir die Ergebnisse gesehen haben, um transparent zu sein.

Warum ist das wichtig? Wenn Forscher entscheiden, wie sie Daten analysieren, nachdem sie Ergebnisse gesehen haben, könnten sie (bewusst oder unbewusst) Methoden wählen, die Ergebnisse besser aussehen lassen. Indem wir unseren Plan zuerst veröffentlichen, verpflichten wir uns zu unseren Methoden und machen unsere Forschung vertrauenswürdiger.

Was haben wir herausgefunden? Dies ist ein Protokoll-Papier - es beschreibt, was wir geplant haben zu tun. Die tatsächlichen Ergebnisse wurden 2023 veröffentlicht (siehe das Papier oben). Protokolle wie dieses zu veröffentlichen ist gute Forschungspraxis und hilft, Verzerrungen in der Wissenschaft zu verhindern.

Plessen, C. Y., Franken, F., Kathofer, M., Kotlyar, E., Maierwieser, R. J., Mayer, A., Rattner, K., Schmid, R. R., Sobisch, M., Ster, C., Wolfmayer, C., & Tran, U. S. (2020). Humor-styles and personality traits: An updated meta-analysis. Personality and Individual Differences, 154, 109676.

What did we do? We combined results from many studies (a meta-analysis) to see how different humor styles relate to personality traits. We looked at four humor styles: affiliative (bringing people together), self-enhancing (staying positive), aggressive (putting others down), and self-defeating (making fun of yourself).

Why does it matter? Humor is an important part of social life and mental health. Understanding how humor relates to personality helps us understand individual differences and potentially identify people who might benefit from interventions.

What did we find? Positive humor styles (affiliative, self-enhancing) were linked to extraversion and openness. Negative humor styles (aggressive, self-defeating) were linked to neuroticism. Self-defeating humor was negatively related to conscientiousness. These patterns help us understand how humor fits into broader personality structures.

Was haben wir gemacht? Wir haben Ergebnisse aus vielen Studien (eine Meta-Analyse) kombiniert, um zu sehen, wie verschiedene Humor-Stile mit Persönlichkeitsmerkmalen zusammenhängen. Wir haben vier Humor-Stile betrachtet: affiliativ (Menschen zusammenbringen), selbst-verstärkend (positiv bleiben), aggressiv (andere herabsetzen) und selbst-erniedrigend (sich selbst lustig machen).

Warum ist das wichtig? Humor ist ein wichtiger Teil des sozialen Lebens und der psychischen Gesundheit. Zu verstehen, wie Humor mit Persönlichkeit zusammenhängt, hilft uns, individuelle Unterschiede zu verstehen und möglicherweise Menschen zu identifizieren, die von Interventionen profitieren könnten.

Was haben wir herausgefunden? Positive Humor-Stile (affiliativ, selbst-verstärkend) waren mit Extraversion und Offenheit verbunden. Negative Humor-Stile (aggressiv, selbst-erniedrigend) waren mit Neurotizismus verbunden. Selbst-erniedrigender Humor war negativ mit Gewissenhaftigkeit verbunden. Diese Muster helfen uns zu verstehen, wie Humor in breitere Persönlichkeitsstrukturen passt.

Plessen, C. Y., Boeckle, M., Liegl, G., Leitner, A., Schneider, A., Preining, B., & Pieh, C. (2016). Bedarfsanalyse für ambulante Psychotherapie in Österreich [Supply and demand analysis for psychotherapy in Austria]. Psychologische Medizin, 3, 4-9.

What did we do? We analyzed the supply and demand for outpatient psychotherapy in Austria. We looked at how many therapists are available, how many people need therapy, and whether supply meets demand across different regions of Austria.

Why does it matter? Many people with mental health problems need psychotherapy, but if there aren’t enough therapists or they’re in the wrong places, people can’t get help. Understanding gaps in care helps policymakers allocate resources better.

What did we find? There were significant gaps between supply and demand for psychotherapy in Austria. Some regions had much better access than others. The findings highlighted the need for better planning of psychotherapy services to ensure everyone who needs help can access it.

Was haben wir gemacht? Wir haben Angebot und Nachfrage für ambulante Psychotherapie in Österreich analysiert. Wir haben untersucht, wie viele Therapeuten verfügbar sind, wie viele Menschen Therapie brauchen, und ob Angebot und Nachfrage in verschiedenen Regionen Österreichs übereinstimmen.

Warum ist das wichtig? Viele Menschen mit psychischen Gesundheitsproblemen brauchen Psychotherapie, aber wenn es nicht genug Therapeuten gibt oder sie an den falschen Orten sind, können Menschen keine Hilfe bekommen. Versorgungslücken zu verstehen hilft politischen Entscheidungsträgern, Ressourcen besser zu verteilen.

Was haben wir herausgefunden? Es gab signifikante Lücken zwischen Angebot und Nachfrage für Psychotherapie in Österreich. Manche Regionen hatten viel besseren Zugang als andere. Die Ergebnisse zeigten die Notwendigkeit besserer Planung von Psychotherapie-Diensten, um sicherzustellen, dass jeder, der Hilfe braucht, diese auch bekommen kann.

Co-Authored Publications

Harrer, M., Miguel, C., van Ballegooijen, W., Ciharova, M., Plessen, C. Y., Kuper, P., Sprenger, A. A., Buntrock, C., Papola, D., Cristea, I. A., de Ponti, N., Bašić, Đ., Pauley, D., Driessen, E., Quero, S., Grimaldos, J., Fernández Buendía, S., Botella, C., Hamblen, J. L., … Cuijpers, P. (2025). Effectiveness of psychotherapy: Synthesis of a “Meta-Analytic Research Domain” across world regions and 12 mental health problems. Psychological Bulletin, 151(5), 600.

Cuijpers, P., Harrer, M., Miguel, C., Ciharova, M., Papola, D., Basic, D., Botella, C., Cowpertwait, L., Cristea, I. A., de Ponti, N., Driessen, E., Ebert, D. D., Furukawa, T. A., Hamblen, J. L., Harrer, M., Miguel, C., Plessen, C. Y., Quero, S., Schnurr, P. P., & van Ballegooijen, W. (2025). Cognitive behavior therapy for mental disorders in adults: A unified series of meta-analyses. JAMA Psychiatry, 82(6), 563-571. Metapsy Project

Plessen, C. Y., Fischer, F., Hartmann, C., Liegl, G., Schalet, B., Kaat, A., Terrill, A. L., Garrido, C. J., Kravitz, R. L., Zumbrunn, T., Amtmann, D., Morgan, E. M., & Rose, M. (2025). Multiverse analysis of differential item function of PROMIS Physical Function items. Advances in Patient-Reported Outcomes, 1(1).

Miguel, C., Harrer, M., Karyotaki, E., Plessen, C. Y., Ciharova, M., Furukawa, T. A., Cristea, I. A., & Cuijpers, P. (2025). Self-reports vs clinician ratings of efficacies of psychotherapies for depression: a meta-analysis of randomized trials. Epidemiology and Psychiatric Sciences, 34, e15.

Scholz, C., Schmigalle, P., Plessen, C. Y., Liegl, G., Vajkoczy, P., Prasser, F., Rose, M., & Obbarius, A. (2024). The effect of self-management techniques on relevant outcomes in chronic back pain: A systematic review and meta-analysis. European Journal of Pain, 28(4), 532-550.

Cuijpers, P., Miguel, C., Ciharova, M., Harrer, M., Basic, D., Cristea, I. A., de Ponti, N., Driessen, E., Hamblen, J., Larsen, S. E., Matbouriahi, M., Papola, D., Pauley, D., Plessen, C. Y., Pfund, R. A., Setkowski, K., Schnurr, P. P., van Ballegooijen, W., Wang, Y., … Karyotaki, E. (2023). Absolute and relative outcomes of psychotherapies for eight mental disorders: a systematic review and meta-analysis. World Psychiatry, 23(2), 267-275.

Cuijpers, P., Miguel, C., Harrer, M., Plessen, C. Y., Ciharova, M., Papola, D., Ebert, D., & Karyotaki, E. (2023). Psychological treatment of depression: A systematic overview of a ‘Meta-Analytic Research Domain’. Journal of Affective Disorders, 335, 141-151.

Harrison, C. J., Plessen, C. Y., Liegl, G., Rodrigues, J. N., Sabah, S. A., Beard, D. J., & Fischer, F. (2023). Item response theory may account for unequal item weighting and individual-level measurement error in trials that use PROMs: a psychometric sensitivity analysis of the TOPKAT trial. Journal of Clinical Epidemiology, 158, 62-69.

Harrison, C. J., Plessen, C. Y., Liegl, G., Rodrigues, J. N., Sabah, S. A., Beard, D. J., & Fischer, F. (2023). Overcoming floor and ceiling effects in knee arthroplasty outcome measurement: mapping the Oxford Knee Score and High Activity Arthroplasty Score onto a common scale. Bone & Joint Research, 12(10), 624-635.

Harrison, C. J., Plessen, C. Y., Liegl, G., Rodrigues, J. N., Sabah, S. A., Beard, D. J., & Fischer, F. (2023). Item response theory assumptions were adequately met by the Oxford hip and knee scores. Journal of Clinical Epidemiology, 158, 166-176.

Cuijpers, P., Miguel, C., Ciharova, M., Quero, S., Plessen, C. Y., Ebert, D., Harrer, M., van Straten, A., & Karyotaki, E. (2023). Psychological treatment of depression with other comorbid mental disorders: systematic review and meta-analysis. Cognitive Behaviour Therapy.

Cuijpers, P., Miguel, C., Harrer, M., Plessen, C. Y., Ciharova, M., Ebert, D., & Karyotaki, E. (2023). Cognitive behavior therapy vs. control conditions, other psychotherapies, pharmacotherapies and combined treatment for depression: a comprehensive meta‐analysis including 409 trials with 52,702 patients. World Psychiatry, 22(1), 105-115. OSF Data

Kossmeier, K., Vilsmeier, J., Dittrich, R., Fritz, T., Kolmanz, C., Plessen, C. Y., Slowik, A., Tran, U. S., & Voracek, M. (2019). Long-term trends (1980-2017) in the N-pact factor: Comparative meta-research of journals in personality psychology and individual differences research. Zeitschrift für Psychologie, 227, 293-302.

Liegl, G., Plessen, C. Y., Leitner, A., Boeckle, M., & Pieh, C. (2015). Guided self-help interventions for irritable bowel syndrome: a systematic review and meta-analysis. European Journal of Gastroenterology & Hepatology, 27, 1209-1221.

Conference Abstracts & Preprints

Riazy, L., Plessen, C. Y., & Rose, M. (2024). European general population reference data for the PROMIS depression metric and major depression scales. Quality of Life Research, 33, S27-S28. [Conference abstract]

Plessen, C. Y., Gyimesi, M. L., Kern, B. M. J., Fritz, T., Lorca, M. V. C., Voracek, M., & Tran, U. S. (2020). Associations between academic dishonesty and personality: A pre-registered multilevel meta-analysis. PsyArXiv. [Preprint]

Talks & Presentations

2022

DGPPN 2022

The robustness of the efficacy of digital interventions for anxiety: an umbrella review and multiverse meta-analysis November 23, 2022

Background: Over the last decade, several meta-analyses and meta-reviews synthesized the evidence on the efficacy of digital mental health interventions for anxiety disorders—unfortunately with diverging conclusions. This leaves clinicians, researchers, and funding agencies with inconsistent recommendations.

Methods: To provide a birds-eye-perspective of the entire field we conducted an umbrella review and a multiverse meta-analysis. We investigated whether the meta-analytical method or the inclusion criteria were responsible for these differences, or whether most potential meta-analyses would reach similar conclusions.

Results: Our umbrella review included six meta-analyses with 84 primary studies. The included meta-analyses differed substantially in their AMSTAR-2 ratings, indicating heterogeneous quality. Our multiverse meta-analysis produced 1193 meta-analyses resulting from all possible analytical decisions. We identified several analytical decisions that consistently led to inflated effect size estimates. Larger effect sizes were found for the comparisons with wait-list control groups than for the comparisons with active control groups (mean Hedges g = 0.58 95% CI [0.40, 0.76] vs g = 0.26 95% CI [0.20, 0.42]). Larger effect sizes were found for digital interventions combining smartphone with internet interventions compared to standalone smartphone or internet applications. Meta-analyses that focused exclusively on guided interventions produced twice the effect sizes than meta-analyses on unguided interventions (g = 0.71, 95% CI [0.51, 0.90] vs g = 0.30, 95% CI [0.13, 0.47]).

Conclusion: We identified several analytical decisions that consistently led to inflated effect size estimates. However, we also found that most decisions did not disproportionately influence the resulting summary effect size estimates, which suggests that meta-analytical findings on digital interventions for anxiety are robust.

Conference Program


EACLIPT 2022

Using multiverse meta-analyses to investigate the robustness of mental health research on psychological treatments for depression and digital interventions for anxiety November 12, 2022

Background: At several stages in any meta-analysis, researchers must decide between multiple equally defensible choices (e.g., different study inclusion criteria, different ways of dealing with low-quality studies, different choices of methods, etc.). These different analytical decisions frequently result in various meta-analyses with overlapping research questions reaching different conclusions—resulting in ambiguous recommendations for clinicians, researchers, and funding agencies.

Methods: In a multiverse meta-analysis, researchers identify all these possible stages for analytical decisions, determine alternative analysis steps at each stage, and implement them simultaneously. As a result, a multiverse meta-analysis reports the outcomes of all possible meta-analyses resulting from all of these possible combinations. Therefore, this method is a promising tool to help answer why some of these meta-analyses diverged, whether the meta-analytical method and exclusion criteria were decisive for these differences, or whether we would reach similar results with most analytical strategies.

Results: We present the preliminary results of multiverse meta-analyses to evaluate the influence different analytical decisions might have had on two research questions, namely 1) the efficacy of psychological treatments for depression and 2) the efficacy of digital interventions for anxiety disorders.

Conclusion: We could identify several analytical decisions that consistently lead to inflated effect size estimates (e.g., the comparison with wait-list control groups, the inclusion of high risk of bias studies, and sometimes ignoring effect size dependency). However, we also identified many decisions that did not disproportionately influence the resulting summary effect size estimates, suggesting the overall robustness of meta-analytical findings on psychological treatments for depression and digital mental health research for anxiety.

Conference Program


SIPS 2022

Workshop: Multiverse Analyses - Introduction and Applications June 23, 2022

This session will tackle the multiverse approach to psychological science and data analysis: In each research project, researchers need to make a multitude of decisions, including decisions about which measurements to choose, how to pre-process data, or which analysis to run. Each of these decisions is potentially impactful and may create a unique universe out of a multiverse of possible outcomes. After a general introduction to the concept of a multiverse, we will host several guest speakers highlighting different facets and implementations of multiverse analysis. Two talks will hereby discuss fundamentals of multiverse analysis, focusing on visualization (Matthew Kay) and available software. Four talks will discuss various implementations of multiverse analysis: Jessica Dafflon will discuss multiverse analysis in developmental neuroscience, Constantin Yves Plessen will talk about multiverse meta-analyses and Justin Landy will talk about his project on crowdsourcing hypothesis tests (fourth speaker is tbd).

Workshop Program


ESMR 2022

What if…? A very short primer on conducting multiverse meta-analyses in R April 23, 2022

Even though conventional meta-analyses provide an overview of the published literature on a given research question, they do not consider different paths that could have been taken in selecting or analyzing the data. Most importantly, multiple meta-analyses with overlapping research questions can reach different conclusions due to differences in inclusion and exclusion criteria, or data analytical decisions. It is therefore crucial to evaluate the influence such choices might have on the result of each meta-analysis. Was the meta-analytical method and exclusion criteria decisive, or is the same result reached via multiple analytical strategies? What if a meta-analysts would have decided to go a different path—would the same outcome occur? Ensuring that the conclusions of a meta-analysis are not disproportionately influenced by data analytical decisions, a multiverse meta-analysis can provide the entire picture and underpin the robustness of the findings—or lack thereof—by conducting multiple, namely all possible and reasonable meta-analyses at once.

Hereby, multiverse meta-analyses provide a research integration like umbrella reviews yet additionally investigate the influence flexibility in data analysis could have on the resulting summary effect size. Importantly, in contrast to umbrella reviews, a multiverse analysis also quantitatively summarizes the results and includes not yet conducted meta-analyses. During the talk I will give a more detailed insight into this potent method, and run through the multiverse of meta-analyses on the efficacy of psychological treatments for depression as an empirical example.

Watch Presentation


ISOQOL 2022

Modeling general populations’ reference data for PROMIS item banks (Physical Functioning, Upper Extremities, and Pain interference) in multiple countries using quantile regression January 1, 2022

Aims: The use of PROs in research and clinical practice requires availability of appropriate and relevant comparison data. We aim to model age, and sex-specific reference values for the PROMIS Physical Function (PF), Upper Extremity (UE), and Pain Interference (PI) scales in populations age 50 and older in Germany, the UK, and the US.

Methods: We collected PROMIS PF, UE, and PI data via telephone interviews from the general population in Germany (N = 921), the UK (N = 905), and the US (N = 900). We investigated differential item functioning (DIF) between countries using iterative hybrid ordinal logistic regression. To account for the measurement error of latent estimates and to obtain a continuous distribution of the latent variable, we imputed 25 data sets with plausible values. Each latent estimate was replaced by a random value drawn from the individual latent variable posterior distribution approximated by a normal distribution. We then utilized quantile regressions to model the 1st, 5th, 10th–90th, 95th, and 99th percentiles and their respective standard errors in each dataset based on different combinations of predictors (age, sex, country). According to Rubin’s rules, the estimated percentiles and corresponding standard errors were pooled across the imputed datasets to provide the respective reference values.

Results: Three items from the PROMIS PF scale showed negligible DIF by country, indicating that all scales are valid for inter-country comparisons. Median regressions revealed significant effects of age (PF: b𝜏=.50= -0.33; PI: b𝜏=50= 0.08), sex (PF: b𝜏=50= -3.22; PI: b𝜏=50= 1.62) and country, indicating that stratification of reference data is warranted. The PF and UE scales showed considerable ceiling effects in all countries aged 50-69. For PI, this applied to all ages.

Conclusions: This paper illustrates a novel approach to model reference values for PROMIS measures based on individual patients’ characteristics. Substantial differences in the PROMIS scores between sex, age, and countries highlight the importance of such patient-specific reference values, enabling clinicians to utilize personalized reference values. Due to the use of plausible value imputation, the obtained population reference values can be compared to data collected with other PROMIS short forms or computer-adaptive tests.

Conference Program


2021

PHO 2021

Differential item functioning of PROMIS physical functioning ceiling items across Argentina, Germany and the US November 23, 2021

Objective: We investigate the validity of comparisons across general populations from Argentina, Germany, and the US by evaluating differential item functioning (DIF) for the PROMIS physical function (PROMIS-PF 2.0) ceiling items that were implemented to measure high physical ability. DIF is investigated as it could introduce biases to inter-country comparisons by potentially leading to systematically different physical function scores. If DIF is detected, individuals with the same ‘true’ underlying physical ability would score systematically different due specific cultural contexts or language differences.

Methods: General population samples completed the 35 ceiling items of the PROMIS PF 2.0. DIF was assessed with hybrid logistic ordinal regression models and Nagelkerkes’ pseudo R2-change of > 0.02 as the critical cutoff value indicating the presence of DIF. The impact of DIF on item scores and the T-scores was additionally examined by inspecting both the item characteristic curves (ICCs) and test characteristic curves (TCCs).

Results: Overall, 3601 persons participated—1001 from Argentina (mean age of 35.6 years, ranging from 18 to 69 years; 51% were female), 1000 from Germany (mean age of 44.9 years, ranging from 18 to 69; 52% were female), and 1600 from the US (mean age of 44.3 years, ranging from 18 to 88; 58% were female). For the comparison between Germany and the USA, 2 out of 35 items were flagged for DIF. For the comparison between Argentina and the US, 4 items were flagged and for the comparison between German and Argentinian items, 5 out of 35 items were flagged. Most DIF items had R2 values just above the critical value of 0.02 and all showed uniform DIF. The ICCs and TCCs showed that the magnitude and impact of DIF on the item and T-scores were negligible.

Conclusions: Our study supports the universal applicability of PROMIS across general populations from Argentina, Germany, and the US. Comparisons across persons from the general population are valid, when applying the PROMIS-PF 2.0 ceiling items.

Conference Abstracts