Resource

Choosing the Right Effect Size for Meta-Analysis: Means, Binary Outcomes and Correlations

Learn how to choose an appropriate effect-size metric for meta-analysis when studies report continuous outcomes, binary outcomes, or correlations. The Resource explains how the target estimand, measurement scale, substantive comparability, effect direction, and sampling variance affect the choice and interpretation of a common effect-size scale.

Meta-analysis requires more than collecting whatever statistics each study reports and feeding them into the same model. The quantities being pooled must represent effects on a scale that has a coherent interpretation across the included studies.

That decision starts with the research question and the outcome structure. Studies reporting continuous outcomes may support a raw mean difference or standardized mean difference. Binary outcomes may support a risk ratio, odds ratio, or risk difference. Correlational studies may use the correlation coefficient. These measures are not interchangeable simply because each can be called an “effect size” (Borenstein et al., 2021).

Practical rule: Choose the effect-size metric that represents the substantive comparison you want to synthesize, make the direction and scale consistent across studies, and ensure that every study contributes an estimate and sampling variance for that same interpretable quantity.

Start With the Estimand, Not the Statistic Reported in the Paper

Before choosing an effect size for meta-analysis, define what the synthesis is intended to estimate.

Ask:

  1. Is the outcome continuous or binary?
  2. Are continuous outcomes measured on the same scale across studies?
  3. Is the scientific question about an absolute difference or a relative effect?
  4. For binary outcomes, is the target a difference in risks, a ratio of risks, or a ratio of odds?
  5. Are the studies estimating an association between continuous variables?
  6. Do differently reported statistics actually represent sufficiently comparable underlying relationships to justify conversion?
  7. What direction will consistently represent benefit, harm, or a positive association?

This matters because an effect-size scale is part of the scientific interpretation, not merely a computational choice. Borenstein et al. (2021) distinguish raw mean differences, standardized mean differences, risk ratios, odds ratios, risk differences, and correlations precisely because each represents an effect differently. Epidemiologic effect and association measures likewise distinguish absolute contrasts from relative contrasts and require attention to what populations and conditions are being compared (Lash et al., 2021).

Quick Decision Table: Which Effect Size Should You Consider?

Outcome or data reported Possible effect-size metric Interpretation Major caution
Continuous outcome; every study uses the same meaningful scale Raw mean difference Difference between group means in the original outcome units Do not pool raw mean differences when studies measure the outcome on different scales
Continuous outcome; studies use different instruments intended to represent a comparable outcome Standardized mean difference, typically bias-corrected as Hedges' g Mean difference expressed relative to within-study variability Standardization changes the scale and interpretation; it does not make substantively different constructs automatically comparable
Binary events in two groups Risk ratio (RR) Ratio of event risks between groups Relative measure; interpretation depends on which group and event are placed in the numerator
Binary events in two groups Odds ratio (OR) Ratio of event odds between groups Odds are not risks; an OR should not automatically be interpreted as an RR
Binary events in two groups Risk difference (RD) Absolute difference in event risks between groups Sensitive to baseline risk, so it may vary across populations even when a relative effect is similar
Association between two continuous variables Correlation coefficient, r Direction and strength of the linear association on the correlation scale Correlations must represent comparable relationships; association should not automatically be interpreted causally
Studies report different effect-size families but address a genuinely comparable relationship Converted common metric, where supported All studies are expressed on one selected effect-size scale Conversion requires assumptions, and statistical convertibility does not establish substantive comparability

The central decision is therefore not “Which statistic is most commonly reported?” It is which effect measure best matches the target question and can be estimated consistently across the studies (Borenstein et al., 2021).

1. Effect Sizes Based on Means

Raw mean difference: prefer the original units when they are genuinely common

When all studies measure an outcome using the same meaningful scale, the raw mean difference is often the most directly interpretable choice. It preserves the original units: a blood-pressure difference remains a difference in blood-pressure units rather than being converted into standard deviations.

Borenstein et al. (2021) recommend the raw mean difference when the outcome scale is inherently meaningful or familiar and all studies use the same scale. The crucial restriction is that raw mean differences cannot be meaningfully combined when the underlying scales differ.

Decision rule: Same outcome and same meaningful scale → consider the raw mean difference.

The advantage is interpretability. Readers can evaluate the magnitude directly in units relevant to the outcome.

Standardized mean difference: when instruments differ

Suppose studies assess a sufficiently comparable construct but use different measurement instruments. A five-point difference on one instrument is not numerically comparable with a five-point difference on another.

In this situation, the standardized mean difference expresses the between-group mean difference relative to the study's variability, placing effects onto a common standardized scale. Borenstein et al. describe d and its bias-corrected form, Hedges' g, as standardized mean-difference measures that permit synthesis across different outcome measures when such standardization is substantively defensible (Borenstein et al., 2021).

Decision rule: Comparable construct but different measurement scales → consider the standardized mean difference.

Important distinction: Do not interpret standardization as proof that the studies measure the same thing. Dividing effects by standard deviations solves a scale problem; it does not solve a construct-comparability problem. Studies still need to be comparable enough that a common summary effect has scientific meaning (Borenstein et al., 2021).

Direction must be harmonized before pooling

Whether using raw or standardized differences, the analyst must define what positive and negative values mean and enforce that convention across studies.

For example, if positive values are defined to favor treatment, an outcome where lower scores represent improvement may require reversing the effect direction. Borenstein et al. emphasize that although the subtraction order is arbitrary, the chosen convention must be applied consistently across studies and designs (Borenstein et al., 2021).

A pooled effect becomes difficult to interpret if positive effects mean benefit in some studies and harm in others.

2. Effect Sizes for Binary Outcomes

For two-group binary data, three major choices are the risk ratio, odds ratio, and risk difference. They describe different aspects of the same 2 × 2 data and should not be treated as interchangeable labels (Borenstein et al., 2021).

Risk ratio meta-analysis

The risk ratio compares two event risks:

RR = risk in one group / risk in the comparison group.

A risk ratio of 1 represents equal risks. Values above or below 1 indicate the direction of the relative association according to the numerator convention.

For meta-analysis, Borenstein et al. perform calculations for the risk ratio on the logarithmic scale and transform the summary results back to the ratio scale for presentation (Borenstein et al., 2021).

A risk ratio meta-analysis is therefore appropriate when the target question is naturally relative: how does risk under one condition compare proportionally with risk under another?

Odds ratio meta-analysis

The odds ratio compares odds rather than probabilities. For meta-analysis, the odds ratio is likewise analyzed using its logarithm and corresponding sampling variance, with the summary and confidence limits subsequently exponentiated back to the OR scale (Borenstein et al., 2021).

An odds ratio meta-analysis should retain that interpretation. An OR of 2 refers to twice the odds under the defined comparison; it should not automatically be rewritten as twice the risk.

Lash et al. (2021) similarly distinguish risk ratios and odds ratios and caution against using “relative risk” indiscriminately for different ratio measures. The exact measure being estimated matters for epidemiologic interpretation.

Risk difference: the absolute alternative

The risk difference subtracts one risk from another rather than dividing them. Its null value is therefore 0 rather than 1.

Unlike RR and OR analyses, which use logarithmic scales, Borenstein et al. describe the risk difference as being analyzed in its raw units (Borenstein et al., 2021).

The distinction between relative and absolute effects is consequential. Risk ratios and odds ratios tend to be less sensitive than risk differences to variation in baseline event risk. A risk difference, however, can communicate absolute clinical impact more directly. Thus, two measures computed from exactly the same data can both be correct while answering different questions (Borenstein et al., 2021).

Choosing among RR, OR, and RD

Do not choose between these measures solely because one produces a more impressive numerical value.

Risk ratio

Target: Relative comparison of risks.

Decision: RR may be natural.

Odds ratio

Target: Relative comparison of odds.

Decision: OR.

Risk difference

Target: Absolute change in risk.

Decision: RD.

Lash et al. (2021) place the same distinction in epidemiologic terms: difference measures express absolute contrasts, whereas ratio measures express relative contrasts and are dimensionless. They also distinguish an observed measure of association between different populations from a causal effect comparing outcomes under alternative conditions for the same target population.

Consequently, pooling adjusted or unadjusted observational measures does not make them causal merely because they are called “effects” in a meta-analysis (Lash et al., 2021).

3. Correlation-Based Effect Sizes

When studies report the association between two continuous variables, the sample correlation coefficient r can serve as the effect-size index. The correlation is standardized and therefore does not depend on the original measurement units in the same way that an unstandardized covariance would (Borenstein et al., 2021).

A correlation meta-analysis nevertheless requires a coherent substantive relationship. A positive correlation must represent the same directional relationship across studies, and the variables being correlated need to be sufficiently comparable for a pooled correlation to have meaning.

Correlation coefficients range from −1 to +1:

  • positive values indicate that higher values of one variable tend to accompany higher values of the other;
  • negative values indicate an inverse relationship;
  • zero represents the null value for the correlation.

For meta-analytic computation, Borenstein et al. use Fisher's z transformation of correlations in the analysis and convert results back to the correlation scale for interpretation. Their worked correlation analyses distinguish the computational scale from the final correlation scale used for reporting (Borenstein et al., 2021).

Interpretation boundary: A correlation is an association measure. Its presence does not by itself establish that changing one variable would cause a change in the other. Lash et al.'s distinction between measures of association and measures of causal effect is directly relevant when interpreting observational meta-analyses (Lash et al., 2021).

4. Can Different Effect Sizes Be Converted?

Sometimes the included literature reports fundamentally different statistics even though the studies address a common research question. Borenstein et al. provide conversions among standardized mean differences, odds ratios, and correlations, including conversion of the associated variance (Borenstein et al., 2021).

This can be useful, but it requires two separate judgments.

Statistical question: can the measures be converted?

Supported conversion formulas can sometimes put effects reported as an odds ratio, standardized mean difference, or correlation onto a common metric.

Scientific question: should they be converted?

This is the more important question. Statistical convertibility does not by itself establish that the studies are substantively suitable for combination.

Statistical question: can the measures be converted?

Supported conversion formulas can sometimes put effects reported as an odds ratio, standardized mean difference, or correlation onto a common metric.

For example, Borenstein et al. describe conversion between the log odds ratio and standardized mean difference and between standardized mean difference and correlation. The conversion from a log odds ratio to d relies on an assumption about an underlying continuous trait and its distribution (Borenstein et al., 2021).

Scientific question: should they be converted?

This is the more important question.

Suppose randomized trials all measure essentially the same underlying outcome, but some investigators retain it as continuous while others dichotomize it into success/failure. Conversion to a common scale may be reasonable.

Now consider observational studies in which one set reports correlations for one type of relationship and another reports odds ratios from substantively different designs or constructs. A formula may exist, but the resulting synthesis can still be scientifically inappropriate.

Borenstein et al. explicitly argue that conversion decisions must be made case by case: the studies should be ones that the analyst would have considered suitable for combination had they originally reported the same metric. They also recommend sensitivity analysis when converted studies are incorporated (Borenstein et al., 2021).

Convert because the studies are substantively comparable—not merely because algebra permits conversion.

5. Why You Cannot Simply Mix Effect-Size Scales

A pooled number is interpretable only when its components refer to a common effect-size meaning.

Consider what would happen if an analyst placed the following directly into one column:

  • standardized mean difference = 0.40;
  • odds ratio = 1.60;
  • correlation = 0.30.

A numerical average of those values has no coherent scientific interpretation. Zero is the null for a mean difference and correlation, whereas 1 is the null for a ratio. Their units, ranges, sampling distributions, and interpretations differ.

Borenstein et al. therefore require studies using different effect-size indices to be converted to a common index before a joint synthesis is attempted, and even then only when the studies themselves are substantively comparable (Borenstein et al., 2021).

The problem is not solved by calling every entry an “effect size.” Effect size is a category of quantities, not a single universal measurement scale.

6. Sampling Variance Determines How Much Information an Estimate Carries

Meta-analysis needs more than an effect estimate from each study. It also needs information about the estimate's sampling variability.

Because study estimates vary through sampling error, each effect has an associated variance and standard error. A more precise estimate has less sampling uncertainty than an imprecise estimate. Borenstein et al. use inverse-variance weighting in the fixed-effect model: studies with smaller within-study variances receive larger weights (Borenstein et al., 2021).

This is why extracting only an OR, g, RR, or r is insufficient. The meta-analyst must also obtain or derive the appropriate variance or standard error for the chosen effect-size metric.

The variance must correspond to the scale actually analyzed.

  • RR and OR analyses are conducted on logarithmic scales;
  • risk differences are analyzed on their difference scale;
  • converted effect sizes require corresponding converted variances;
  • correlation analyses use the appropriate transformed scale for meta-analytic computation.

Borenstein et al. explicitly pair effect-size conversions with variance conversions because transforming only the point estimate would leave the weighting and uncertainty calculation on the wrong scale (Borenstein et al., 2021).

7. Effect Magnitude Is Not the Same as Statistical Significance

Meta-analysis is fundamentally concerned with estimating effects, not merely counting which studies crossed a p-value threshold.

Borenstein et al. distinguish significance testing from effect-size estimation. A p-value addresses evidence relative to a null hypothesis, whereas an effect-size estimate addresses the magnitude of the effect (Borenstein et al., 2021).

Altman (1991) makes the same distinction in medical research. Statistical significance and clinical importance are not equivalent: small studies can yield nonsignificant results while remaining compatible with clinically important effects because their estimates are imprecise. Confidence intervals make this uncertainty visible in a way that a significance label does not (Altman, 1991).

Accordingly, do not choose or interpret a meta-analytic effect measure according to which metric generates more statistically significant individual studies.

Instead report and interpret: effect magnitude + direction + confidence interval + effect-size scale + substantive meaning.

For difference measures, the conventional null is 0. For ratio measures, it is 1. A confidence interval conveys the precision of the estimated effect and should be interpreted together with the point estimate rather than reduced solely to whether it crosses its null value (Borenstein et al., 2021; Altman, 1991).

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry