Correlation vs Agreement: Why a High Correlation Does Not Prove Two Measurement Methods Agree
A high correlation between two measurement methods does not establish that they agree. This Resource explains how paired differences, average bias, Bland-Altman limits of agreement, graphical assessment, repeatability, and diagnostic accuracy address distinct methodological questions.
A high correlation between two measurement methods does not establish that they agree.
This is one of the most persistent errors in method-comparison research. A researcher measures the same subjects with two instruments, obtains a large correlation coefficient, and concludes that the methods are interchangeable. The problem is that correlation and agreement answer fundamentally different questions.
Correlation describes the strength of a linear association between two variables. Agreement asks whether two methods give measurements sufficiently close to one another, in the measurement units that matter scientifically or clinically. Altman explicitly identifies correlation as an inappropriate primary analysis for method-comparison studies because strong linear association can coexist with clinically poor agreement (Altman, 1991).
Practical rule: If the research question is whether two methods can be used interchangeably, analyze their differences—not merely their correlation.
Start With the Research Question
Before calculating any statistic, decide what is actually being evaluated.
| Research question | Statistical concept | Appropriate interpretation |
|---|---|---|
| Do larger values from one quantitative variable tend to accompany larger values from another in a linear relationship? | Correlation | Strength and direction of linear association |
| Do two methods measuring the same quantity give sufficiently similar numerical results? | Agreement | Magnitude and pattern of within-subject differences, including average bias and variation in those differences |
| Does a medical test distinguish people with a target condition from those without it? | Diagnostic accuracy | Classification performance relative to disease or reference status, using measures such as sensitivity, specificity, true- and false-positive fractions, or ROC-based summaries where appropriate |
These are not alternative statistical techniques for the same estimand. They are different research questions.
Rosner describes correlation as a way to quantify association between continuous variables and specifically characterizes it as appropriate when the goal is to describe their linear relationship rather than predict one from another (Rosner, 2016). Agreement requires a different target.
What Correlation Actually Measures
Pearson's correlation coefficient, r, summarizes linear association. Positive correlation means that larger values of one variable tend to accompany larger values of the other; negative correlation indicates an inverse linear relationship. The coefficient ranges from −1 to +1 and is unaffected by changes in location or scale (Rosner, 2016).
That last property is important in correlation vs agreement.
If one method consistently produces measurements that are shifted or rescaled relative to another method, the two sets of measurements can remain very strongly correlated. Correlation is designed to capture whether the observations track one another linearly; it is not designed to quantify how close their numerical values are.
Altman therefore distinguishes the questions directly: correlation measures the strength of linear association, whereas agreement must be evaluated in terms related to the measurements themselves (Altman, 1991).
Correlation asks about tracking, not closeness
Imagine the conceptual situation in which subjects with low measurements under Method A also tend to have low measurements under Method B, and subjects with high measurements under Method A tend to have high measurements under Method B.
That ordering can produce strong correlation.
But the methods could still differ systematically. One method might tend to produce higher readings than the other, or the discrepancy might become larger as the magnitude of the measurement increases.
The observations can therefore follow a strong linear pattern without lying close to the line of equality, where the two measurements would be identical.
Key distinction: Correlation describes whether measurements track one another. Agreement asks whether the measurements themselves are sufficiently close.
Why High Correlation Can Coexist With Poor Agreement
A correlation coefficient depends strongly on how much subjects vary from one another.
If a study includes subjects covering a broad range of the measured quantity, the between-subject variation can be large. Both methods may successfully distinguish relatively low-valued subjects from relatively high-valued subjects, producing a strong correlation, while their measurements for the same individual remain too far apart for the intended use.
Altman emphasizes precisely this problem: a high correlation can result from large variation between subjects even when agreement between measurements is clinically poor. Consequently, correlation as an assessment of agreement can be highly sensitive to the particular sample of subjects selected for the study (Altman, 1991).
Fundamental mismatch: Correlation is influenced by between-subject variation. Agreement concerns within-subject discrepancies between methods.
A method-comparison analysis should therefore focus on the second quantity.
Method Comparison Is a Paired-Measurement Problem
When two methods measure the same subject, the resulting observations are paired.
The measurement from Method A for participant i belongs with the measurement from Method B for that same participant. Treating the two sets of values as though they came from unrelated samples discards this structure.
Altman describes paired observations as measurements linked within individuals and emphasizes that paired analyses focus on within-subject differences rather than between-subject variability (Altman, 1991).
For two measurement methods, define the within-subject difference consistently, for example:
di = Ai − Bi
where Ai and Bi are the measurements obtained from the two methods for subject i.
The sign convention is arbitrary, but it must be stated and retained because it determines the direction of estimated bias.
Central question: How large are these differences, in which direction do they occur, and how much do they vary across subjects?
Analyze the Differences Between Methods
Altman's method-comparison framework begins by calculating the difference between the two measurements for every subject. The analysis then examines both the average difference and the variability of individual differences (Altman, 1991).
Two components are particularly important.
1. Mean difference: average bias
The mean of the paired differences estimates the average bias of one method relative to the other:
d̄ = (1 / n) Σ di
If the difference is defined as Method A minus Method B, a positive mean difference indicates that Method A tends to give higher measurements on average. A negative value indicates the reverse.
Altman explicitly interprets the mean difference in a method-comparison study as an estimate of average bias (Altman, 1991).
2. Standard deviation of the differences
The standard deviation of the paired differences describes how much individual method discrepancies vary around their mean.
This is crucial because interchangeability is usually an individual-level question: If I measure a new subject with either method, how different might the two results be?
A small average bias does not guarantee small individual disagreement. Agreement analysis therefore needs both the center and the spread of the differences (Altman, 1991).
Average bias alone is not sufficient. Two methods could have almost no mean difference because positive and negative discrepancies cancel, while individual differences remain unacceptably large.
The Bland Altman Method and Limits of Agreement
The Bland Altman method summarizes agreement using the mean difference together with limits based on the variability of the paired differences.
For reasonably symmetric differences, Altman describes approximate 95% limits of agreement as:
d̄ ± 2sd
where:
- d̄ is the mean difference between methods; and
- sd is the standard deviation of the differences.
These limits describe a range expected to contain the differences between the two methods for most subjects under the assumptions of the analysis (Altman, 1991).
The modern normal-theory expression is often written using 1.96 rather than the approximate 2, but Altman's text presents the practical approximation as mean difference ± 2 SD.
Interpretation: Limits of agreement quantify disagreement in the original measurement units.
That makes them fundamentally different from a dimensionless correlation coefficient.
Statistical limits do not define acceptable agreement
A limits-of-agreement calculation does not decide whether agreement is scientifically or clinically adequate.
Altman is explicit that interpretation of the mean and standard deviation of the differences depends on the clinical circumstances; statistics cannot define what degree of agreement is acceptable (Altman, 1991).
Substantive judgment is still required: Would differences of the magnitude represented by these limits be acceptable for the intended use of the measurement?
That decision should not be replaced by asking whether r exceeds an arbitrary threshold.
Graphical Assessment Is Central to Method Comparison Statistics
A numerical summary should be accompanied by graphical assessment.
Plot the raw measurements against one another
A scatterplot of Method A against Method B can show the overall relationship between measurements. For agreement, however, the key reference is the line of equality, not merely a fitted regression line.
If both methods gave exactly the same measurement for every subject, the points would lie on that equality line.
Altman recommends that a raw-data method-comparison plot be square and display the line of equality (Altman, 1991).
Plot difference against average
The more informative agreement plot places the difference between methods on the vertical axis and their average on the horizontal axis:
Difference = Ai − Bi
against
Average = (Ai + Bi) / 2
Altman explains that this plot makes the magnitude and distribution of the differences easier to inspect and helps determine whether disagreement changes with the size of the measurement (Altman, 1991).
Typically, the plot can display horizontal lines representing:
- the mean difference; and
- the lower and upper limits of agreement.
This is the familiar Bland Altman plot.
Look for More Than the Mean Bias
A good agreement plot should not be interpreted simply by checking whether the mean-difference line is close to zero.
Inspect whether:
- the differences are centered around a systematic positive or negative bias;
- the spread of the differences is approximately stable across the measurement range;
- disagreement becomes larger as measurements increase;
- unusual observations appear; and
- the pattern suggests that a constant difference is an inadequate description of disagreement.
Altman specifically considers situations in which the scatter of differences increases with the magnitude of the measurement. In such circumstances, analysis on a transformed scale may be more appropriate; for proportional differences, logarithmic analysis can lead to limits interpretable in proportional rather than raw-unit terms (Altman, 1991).
The graph is not decorative. It is a diagnostic component of the measurement agreement analysis.
Why Testing Whether the Correlation Is Statistically Significant Answers the Wrong Question
Suppose a researcher tests:
H0: ρ = 0
and obtains a small P value.
What has been learned?
The evidence concerns whether the population linear correlation is compatible with zero under the assumptions of that inferential procedure. Rosner treats hypothesis testing for correlation as inference about correlation coefficients—that is, about association between continuous variables (Rosner, 2016).
But the method-comparison question is not:
“Is there any nonzero linear relationship between the measurements?”
It is:
“Are measurements from these two methods sufficiently close to one another for the intended use?”
Those are different hypotheses and different targets.
A statistically significant correlation can therefore coexist with unacceptable bias or wide individual discrepancies. Indeed, when both methods respond to genuine differences among subjects, evidence of positive association may be entirely unsurprising.
Testing the correlation against zero does not quantify:
- average measurement bias;
- the spread of within-subject differences;
- the likely discrepancy between methods for an individual;
- whether disagreement changes with measurement magnitude; or
- whether the observed disagreement is acceptable for the intended application.
A small P value for correlation therefore cannot rescue an agreement analysis that never examined agreement.
“No significant mean difference” is not proof of agreement either
Replacing correlation testing with a paired test of whether the mean difference equals zero does not solve the problem.
Altman warns explicitly against concluding that methods agree because their means are not significantly different. Large scatter in individual differences can make an important average bias statistically non-significant; paradoxically, worse agreement can make rejection of the zero-difference hypothesis less likely (Altman, 1991).
Agreement is principally an estimation problem, not a search for a convenient nonsignificant hypothesis test. Altman frames method comparison in terms of estimating the degree of agreement (Altman, 1991).
Repeatability Is Related to Agreement but Is Not the Same Thing
Agreement
Method comparison asks how closely different methods measure the same subjects.
Repeatability
Repeatability asks how closely the same method reproduces its measurements when measurements are repeated under relevant conditions.
If replicate measurements are available for each method, Altman recommends examining the differences between repeated measurements made with the same method. The standard deviation of these within-method differences can be used to compare repeatability (Altman, 1991).
This distinction matters because a method with poor repeatability cannot be expected to agree closely with another method.
Agreement between Method A and Method B therefore reflects, at least in part, the measurement variation inherent in the methods themselves.
Without replicate measurements, researchers cannot use an ordinary unreplicated method-comparison study to compare the repeatability of the methods (Altman, 1991).
Agreement Does Not Tell You Which Method Is Correct
Suppose two methods disagree substantially.
An agreement analysis can show that the discrepancy is large, but it does not automatically reveal which method is closer to the truth.
Altman emphasizes that method-comparison data generally cannot identify which method is more accurate when the true value is unknown. Poor agreement also does not imply that both methods are poor: one method could be inaccurate or poorly repeatable while the other performs well (Altman, 1991).
This is another reason to define the scientific objective before choosing the analysis.
Agreement
Agreement is about correspondence between methods.
Accuracy against a credible reference
Accuracy against a credible reference is a different question.
When Diagnostic Accuracy Is the Real Question
Not every study involving two tests or measurements is a method-agreement study.
Suppose an investigational medical test is intended to identify whether a person has a target condition, and disease status is established using an appropriate reference process.
The question is then not simply whether the numerical result of the new test agrees with the numerical result of another measurement method.
It is a classification or diagnostic-accuracy question.
Pepe's framework evaluates medical tests through disease-specific performance measures. For a continuous test, changing the threshold changes the true-positive fraction and false-positive fraction, and the ROC curve summarizes the attainable combinations of these quantities across thresholds (Pepe, 2003).
Diagnostic accuracy can therefore involve measures such as:
- sensitivity or true-positive fraction;
- specificity or false-positive fraction;
- predictive values where appropriate to the sampling design and population;
- diagnostic likelihood ratios; and
- ROC curves or AUC for suitable continuous or ordinal tests.
These measures answer questions about classification relative to disease status, not numerical agreement between two continuous measurement methods.
Agreement and diagnostic accuracy can lead to different study designs
If two instruments are intended to measure the same continuous quantity interchangeably, focus on paired numerical differences and agreement.
If a test is intended to distinguish disease from non-disease, focus on its classification performance relative to an appropriate disease or reference status.
Pepe also emphasizes that diagnostic-accuracy estimates depend on study design and the validity of the reference process. An imperfect reference standard, selective verification, or an unrepresentative spectrum of cases and controls can distort estimated test performance (Pepe, 2003).
A high correlation between an index test and another continuous measurement does not answer those diagnostic-accuracy questions either.
A Practical Correlation vs Agreement Workflow
- Define the research question. Is the target linear association, numerical agreement, repeatability, or diagnostic classification?
- Confirm the paired structure. The two measurements should be linked within the same subjects when the objective is direct method comparison.
- Plot the raw measurements. Include the line of equality when examining agreement.
- Calculate a difference for every subject. Choose and report a consistent direction, such as Method A minus Method B.
- Plot difference against average. Look for bias, changing variability, unusual observations, and dependence of disagreement on measurement magnitude.
- Estimate average bias. Report the mean difference in the original measurement units when that scale is appropriate.
- Quantify variation in individual differences. The standard deviation of the differences is central to evaluating individual disagreement.
- Estimate limits of agreement where their assumptions are suitable. Interpret them in the units relevant to the measurement.
- Judge substantive acceptability. Statistical calculations do not define whether the limits are clinically or scientifically tolerable.
- Assess repeatability separately when replicate measurements are available. Do not infer within-method repeatability from between-method correlation.
- Use diagnostic-accuracy methods instead when the scientific target is disease classification.
Common Mistakes
“The correlation is 0.95, so the methods agree”
No. A large correlation describes strong linear association. It does not quantify the numerical discrepancies between paired measurements (Altman, 1991).
“The correlation is statistically significant, so agreement is established”
No. The significance test addresses evidence about a correlation parameter, not whether within-subject differences are acceptably small (Rosner, 2016).
“The paired t test is nonsignificant, so the methods agree”
No. A test of zero mean difference addresses average bias, and failure to reject that null does not establish small individual differences. Wide disagreement can exist around an average difference near zero (Altman, 1991).
“The Bland Altman plot is just another correlation plot”
No. The difference-versus-average plot directly examines discrepancies between paired measurements and whether their magnitude changes across the measurement range (Altman, 1991).
“Narrow-looking limits mean the methods are interchangeable”
Not automatically. The acceptable magnitude of disagreement must be judged in the scientific or clinical context. Statistics alone cannot define acceptable agreement (Altman, 1991).
“Poor agreement proves both methods are inaccurate”
No. Without knowledge of the true value, ordinary method comparison generally cannot determine which method is closer to truth (Altman, 1991).
“A diagnostic test should be assessed by whether it correlates with the reference measurement”
Not when the target is diagnostic classification. Diagnostic accuracy concerns performance relative to disease or reference status and requires measures suited to that target, such as true- and false-positive fractions or ROC-based evaluation for continuous tests (Pepe, 2003).
Reporting Agreement Between Two Methods
A defensible method-comparison report should make the statistical target explicit.
For a standard continuous-method comparison, report:
- what quantity both methods are intended to measure;
- that measurements are paired within subjects;
- the direction used to calculate differences;
- the mean difference as an estimate of average bias;
- the variability of the paired differences;
- limits of agreement where appropriate;
- a difference-versus-average plot;
- any evidence that disagreement changes with measurement magnitude;
- whether replicate measurements were available to assess repeatability; and
- whether the observed magnitude of disagreement is acceptable for the intended use.
A correlation coefficient can be reported if linear association is itself of interest, but it should not be presented as evidence that the methods agree.
Bottom Line
The central lesson in correlation vs agreement is that statistical methods must match the research question.
Correlation asks whether two quantitative variables move together in a linear way. Agreement asks whether two measurements of the same quantity are sufficiently close. Because these are different targets, a high—and even highly statistically significant—correlation cannot establish measurement agreement (Altman, 1991; Rosner, 2016).
For continuous method comparison, preserve the paired structure and analyze the differences between methods. Estimate average bias, quantify the variation of those differences, examine a Bland Altman difference-versus-average plot, and use limits of agreement where supported by the data and assumptions. Then decide whether that degree of disagreement is acceptable in the substantive context (Altman, 1991).
When the real objective is to determine how well a test classifies disease relative to a reference status, the problem changes again. That is a diagnostic-accuracy question, requiring diagnostic-performance measures rather than either correlation or ordinary measurement-agreement statistics (Pepe, 2003).
Association
Decision: Correlation.
Interchangeability of measurements
Decision: Agreement analysis.
Disease classification
Decision: Diagnostic-accuracy analysis.
FAQs
What is the difference between correlation and agreement?
Correlation describes the strength of linear association between two variables. Agreement concerns how close paired measurements from two methods are to one another. Two methods can therefore correlate strongly while having substantial systematic or individual differences (Altman, 1991; Rosner, 2016).
Does high correlation mean two methods are interchangeable?
No. High correlation can occur because subjects vary greatly from one another even when the two methods disagree substantially within individual subjects. Interchangeability requires direct assessment of the paired differences and their practical magnitude (Altman, 1991).
What is the Bland Altman method?
The Bland Altman method evaluates agreement by analyzing paired differences between methods. The mean difference estimates average bias, while the standard deviation of the differences can be used to construct limits describing the range within which most individual method differences are expected to fall under the relevant assumptions (Altman, 1991).
What does a Bland Altman plot show?
It plots the difference between the two measurements against their average. This makes it possible to inspect average bias, the spread of disagreement, unusual observations, and whether the amount of disagreement changes with measurement magnitude (Altman, 1991).
What are limits of agreement?
Limits of agreement are boundaries based on the mean and standard deviation of paired differences. Altman describes approximate 95% limits as the mean difference ± 2 standard deviations of the differences for reasonably symmetric differences. Their substantive acceptability must be judged from the measurement context (Altman, 1991).
Is a nonsignificant paired t test evidence of agreement?
No. A paired test concerns whether the average difference is statistically distinguishable from zero. It does not establish that individual discrepancies are acceptably small. Methods can have little average bias but poor individual agreement (Altman, 1991).
Why is statistical significance of the correlation irrelevant to agreement?
A correlation significance test asks whether the population correlation differs from its null value, commonly zero. Agreement asks whether paired measurements are sufficiently close. A significant association therefore answers a different research question (Rosner, 2016).
How is repeatability different from agreement?
Repeatability concerns variation when the same method measures the same subject repeatedly. Agreement concerns differences between different methods. Replicate measurements are needed to evaluate and compare within-method repeatability directly (Altman, 1991).
Is method agreement the same as diagnostic accuracy?
No. Agreement concerns numerical correspondence between measurements. Diagnostic accuracy concerns how well a test distinguishes disease or another target condition relative to a reference status. Continuous diagnostic tests may be evaluated through threshold-specific true- and false-positive fractions and ROC curves rather than ordinary limits-of-agreement analysis (Pepe, 2003).
Should I report correlation alongside Bland Altman analysis?
You may report correlation if linear association is independently relevant to the scientific question, but it should not be interpreted as evidence of agreement. For interchangeability, the paired differences, bias, variation, graphical assessment, and limits of agreement are the relevant focus (Altman, 1991).
References
Altman, D. G. (1991). Practical statistics for medical research. Chapman & Hall.
Pepe, M. S. (2003). The statistical evaluation of medical tests for classification and prediction. Oxford University Press.
Rosner, B. (2016). Fundamentals of biostatistics (8th ed.). Cengage Learning.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.