Should You Categorize a Continuous Variable? Why Arbitrary Cutoffs Can Weaken Your Analysis
Categorizing continuous variables can discard information, reduce precision, create artificial boundaries, and weaken regression and prediction models. This Resource explains when to keep a predictor continuous, when to consider flexible modeling, and when a scientifically established threshold may justify categorization.
Age, blood pressure, BMI, laboratory measurements, and biomarker concentrations are measured on continuous scales for a reason: the observed values contain information about where each individual lies along that scale. Converting those measurements into categories such as “low/high,” “normal/abnormal,” or quartiles can make a table look simpler, but it changes the statistical question and imposes a new model on the data.
The default strategy should usually be:
Keep continuous → consider transformation or flexible modeling → categorize only when scientifically justified.
Harrell gives a detailed account of the problems caused by dichotomization: reduced precision and power, flat relationships within categories, artificial discontinuities at category boundaries, loss of full predictor information, and invalid ordinary inference when cutpoints are chosen using the outcome (Harrell, 2015). Steyerberg similarly treats dichotomization as problematic in prediction modeling and explicitly considers flexible continuous functions, including restricted cubic splines and fractional polynomials, as alternatives (Steyerberg, 2019).
The important distinction is therefore not simply continuous vs categorical variable. The real question is whether converting an inherently continuous measurement into categories represents the scientific relationship better—or merely makes the analysis appear easier.
The Core Decision: Preserve the Measurement Unless the Categories Have a Scientific Purpose
A continuous predictor contains information about differences throughout its observed range. Once it is reduced to categories, observations that were distinguishable can become statistically identical.
Suppose a predictor is divided at a threshold c. The model now treats everyone on one side of c as belonging to one category and everyone on the other side as belonging to another. The original distances between their measurements no longer enter through that categorized representation.
That creates two strong assumptions:
- Within a category, the modeled predictor effect is flat.
- At the category boundary, the modeled effect changes abruptly.
Harrell emphasizes both consequences. Categorization assumes constant response within intervals and a discontinuity when a boundary is crossed. He argues that such discontinuities are rarely plausible for ordinary biological predictors (Harrell, 2015).
This is why a cutoff is not merely a convenient coding decision. It defines the shape of the predictor–outcome relationship.
Why Categorizing Continuous Variables Loses Information
Dichotomization replaces measured values with group membership
If a continuous predictor is converted into “low” and “high,” the regression model no longer uses the exact predictor values. It uses membership in one of two groups.
Harrell identifies reduced precision and reduced power as direct consequences of dichotomization. Using several categories instead of two can represent the relationship somewhat more closely, but requires additional indicator variables and degrees of freedom while still producing flat regions and discontinuous jumps (Harrell, 2015).
Steyerberg likewise highlights potentially dramatic information loss from dichotomizing ordered predictors and specifically raises the problems of dichotomization for estimating predictor effects, controlling confounding, and making individualized predictions (Steyerberg, 2019).
Collecting a precise measurement and then replacing it with a coarse category can discard information that was already available.
More categories do not fully solve the problem
Creating three, four, or five groups may look more sophisticated than a high/low split, but the fundamental structure remains.
Each interval still assumes that individuals within it share the same categorized predictor contribution, while neighboring individuals on opposite sides of a boundary are treated as belonging to different groups.
Harrell also notes that multiple intervals require more parameters than fitting an appropriately smooth relationship and that wide outer intervals can leave substantial heterogeneity within categories (Harrell, 2015).
So the choice is not:
two categories vs many categories.
A more useful comparison is:
categories vs an appropriately specified continuous function.
Arbitrary Cutoffs Create Artificial Boundaries
Consider what a cutoff means mathematically.
If the boundary is c, observations immediately below and immediately above c are assigned to different groups. Yet observations much farther apart on the same side of the threshold are treated identically by the categorized predictor.
That is a strong model of the underlying relationship.
Harrell explicitly contrasts this with continuous modeling. When a predictor such as blood pressure is modeled continuously, clinically meaningful contrasts can be estimated between exact values. Categorization instead makes interpretation depend partly on the distribution of predictor values inside the categories (Harrell, 2015).
This matters because apparently simple group comparisons can conceal substantial within-group variation.
A cutoff should therefore not be defended merely because it produces an easy-to-read odds ratio or hazard ratio.
The “Optimal Cutoff” Problem
A particularly risky strategy is to try many possible cutpoints and retain the one that produces the most favorable association, smallest P value, greatest separation, or otherwise most attractive result.
The selected cutoff is then data-driven.
Harrell warns that when cutpoints are selected using information about the response, ordinary P values and confidence intervals are invalid because they fail to account for the multiplicity and uncertainty introduced by searching for the cutpoint. He specifically notes that ordinary P values can become too small and confidence intervals can fail to achieve their claimed coverage (Harrell, 2015).
This is the central optimal cutoff problem:
The final analysis can look as though one threshold had been specified in advance even though the data were searched to find it.
The search itself is part of the model-building procedure.
For prediction models, this also creates an overfitting problem. Steyerberg emphasizes that prediction-model specification should ideally be prespecified and that decisions driven by observed predictor–outcome relationships can make apparent model performance optimistic (Steyerberg, 2019).
Warning signs that your cutoff is data-driven
Be cautious when:
- many possible thresholds were examined before one was selected;
- the chosen value is the cutoff giving the smallest P value;
- the threshold was selected because it maximized separation between outcome groups;
- several versions of the predictor were tried and only the most favorable categorization is reported;
- the cutoff was chosen after examining outcome frequencies or regression coefficients;
- the final report presents the threshold as though it had been prespecified when it was actually discovered from the analysis;
- the same dataset is used both to discover the cutoff and to report its apparent predictive performance without accounting for that modeling step.
Key diagnostic question: Would this exact cutoff have been chosen if the outcome values had been hidden?
If not, the outcome has helped define the predictor representation.
A Clinically Established Threshold Is Different From a Statistical Convenience Cutpoint
Not every threshold is arbitrary.
Some scientific or clinical questions genuinely concern membership above or below a threshold. A threshold may define a treatment decision, diagnostic classification, eligibility rule, established clinical state, or another substantively meaningful distinction.
That situation should be distinguished from creating a cutoff because the statistical analysis appears easier afterward.
Altman repeatedly emphasizes aligning statistical analysis with the substantive research question and distinguishes clinically meaningful quantities from values selected merely for statistical convenience. His treatment of continuous relationships also shows that nonlinearity can be modeled directly rather than forcing a curved relationship into groups (Altman, 1991).
Scientifically established threshold
The boundary exists because the scientific or clinical question requires it.
Statistical convenience threshold
The boundary exists primarily because the analyst wants categories, a simpler table, a binary odds ratio, or a favorable statistical result.
Even when an established threshold is important clinically, retaining the original continuous measurement can still be valuable for modeling. A decision threshold and a predictor's statistical representation do not necessarily have to be identical.
When Categorization May Be Scientifically Defensible
Categorization deserves consideration when the categories themselves are part of the target question rather than an analytical shortcut.
Examples of defensible reasoning include situations in which:
- a threshold was established independently of the current outcome data;
- the scientific question explicitly concerns states defined by that threshold;
- the categories correspond to a genuine decision rule that will be applied in practice;
- a substantively meaningful discontinuity at the boundary is actually being hypothesized;
- the purpose is to communicate or implement a prespecified classification while the consequences of discarding continuous information are understood.
The justification should be stated explicitly.
“Categories are easier to interpret” is not, by itself, evidence that categorization represents the predictor–outcome relationship appropriately.
Continuous Does Not Mean Linear
One reason researchers categorize predictors is concern that ordinary regression assumes a straight-line relationship.
That concern is legitimate; the proposed solution often is not.
Keeping a variable continuous does not require fitting it as one untransformed linear term.
Altman describes polynomial regression as one approach to modeling a curved relationship between a continuous predictor and outcome rather than assuming linearity (Altman, 1991). Harrell develops flexible predictor modeling extensively, including regression splines, while Steyerberg discusses nonlinear continuous functions including restricted cubic splines and fractional polynomials (Harrell, 2015; Steyerberg, 2019).
The modeling decision should therefore be separated into two questions:
1. Should the measurement remain continuous?
Usually yes unless there is a scientific reason to categorize it.
2. What functional form should represent its relationship with the outcome?
That requires examination rather than an automatic assumption of linearity.
A Practical Framework for a Regression Continuous Predictor
-
Keep the predictor continuous
Begin with the original measurement scale rather than immediately creating groups.
This preserves the information available in the observed values and avoids imposing unmotivated boundaries.
-
Define the scientific contrast you need
A continuous predictor can still yield interpretable effects.
For a predictor represented adequately by a linear term, the regression coefficient can be expressed for a scientifically meaningful change rather than necessarily a one-unit change.
When the relationship is nonlinear, interpretation can instead focus on estimated outcomes, risks, or contrasts between meaningful predictor values.
Harrell specifically notes that continuous modeling allows comparisons between exact predictor values, whereas an odds ratio based on categories depends on the distribution of observations within those categories (Harrell, 2015).
-
Assess whether a simple linear functional form is adequate
Do not equate “continuous” with “linear.”
Examine whether the chosen model adequately represents the predictor–outcome relationship. The appropriate scale depends on the regression model: for example, a simple continuous term in logistic regression represents linearity on the model's linear-predictor scale, not necessarily a linear change in probability.
-
Consider transformation or flexible modeling when needed
If the relationship is not adequately represented by a simple linear term, consider a scientifically and statistically appropriate nonlinear function.
Supported options in the approved sources include polynomial terms, fractional polynomials, and spline functions. Harrell places particular emphasis on regression splines, and Steyerberg discusses restricted cubic splines and fractional polynomials for flexible continuous modeling (Harrell, 2015; Steyerberg, 2019).
Flexibility is not free: more complex functions require more information and should be compatible with the available sample size. The objective is not maximum flexibility but an adequate representation of the relationship.
-
Categorize only if the scientific justification survives scrutiny
If categorization is still proposed, document:
- where the boundary came from;
- whether it was chosen before examining outcomes;
- what scientific state or decision it represents;
- why a discontinuity at that value is meaningful;
- what information is being discarded;
- and whether the continuous version should also be retained for modeling or sensitivity analysis.
What Category Boundaries Assume
Categorization silently imposes a particular functional form.
For a dichotomized predictor, that form is essentially a step:
one modeled level below the cutoff → abrupt change at the cutoff → another modeled level above it.
With several categories, the result becomes a staircase.
Harrell explicitly identifies both components: flat effects within intervals and discontinuities at their boundaries (Harrell, 2015).
That means categorization does not eliminate assumptions about functional form.
It replaces a visible continuous-model assumption with a strong step-function assumption.
The appropriate comparison is therefore not “model assumption versus no model assumption.” Both approaches make assumptions; the question is which assumptions are scientifically credible and statistically useful.
Implications for Regression Models
Unnecessary categorization can affect regression analysis in several ways.
First, it can reduce precision and statistical power because information in the original measurements has been discarded (Harrell, 2015).
Second, it can misrepresent the functional relationship by imposing flat regions and artificial jumps.
Third, using several categories consumes additional degrees of freedom while retaining the basic limitations of interval-based modeling.
Fourth, categorizing a continuous adjustment variable can leave heterogeneity within categories. Harrell specifically notes the possibility of residual confounding when continuous variables are represented through broad intervals (Harrell, 2015).
Finally, interpretation may not become as clean as it appears. A category-based effect compares mixtures of predictor values inside each group, whereas a continuous model can estimate contrasts between explicitly specified values.
Implications for Prediction Models
The disadvantages become especially important when the objective is individualized prediction.
A patient or study participant arrives with an observed value—not merely a label saying that the value exceeds a cutoff.
Harrell makes this point explicitly: knowing the exact predictor value contains more information for predicting an individual's outcome than knowing only that the measurement falls above or below a threshold (Harrell, 2015).
Steyerberg's prediction-model framework likewise treats dichotomization of continuous predictors as problematic for individualized prediction and emphasizes flexible continuous modeling where appropriate (Steyerberg, 2019).
A prediction model built from arbitrary categories can therefore create implausible risk behavior: two nearly identical values on opposite sides of a boundary may receive different predictor contributions, while substantially different values within the same category receive the same contribution from that categorized variable.
If the cutoff itself was selected from the development data, the search procedure adds another source of overfitting and optimism.
Interpretation Without Forcing Categories
Researchers sometimes categorize because they want a simple statement such as:
“High values were associated with twice the odds.”
But continuous modeling does not prevent interpretable reporting.
Depending on the model and functional form, useful reporting approaches include:
- effects for a meaningful increment in the predictor;
- contrasts between two prespecified, scientifically meaningful values;
- predicted outcomes or risks at selected meaningful values;
- a graph of the fitted continuous relationship with uncertainty;
- contrasts derived from a nonlinear fitted function rather than individual spline coefficients.
Harrell explicitly illustrates the advantage of estimating an odds ratio between exact predictor settings rather than relying on an arbitrary threshold-based comparison (Harrell, 2015).
The objective should be interpretable estimation without unnecessary information destruction, not categorization merely to obtain one coefficient.
Questions to Ask Before Converting a Continuous Variable Into Groups
- What is the scientific reason for the boundary?
- Was the cutoff defined independently of the current outcome data?
- Would I use exactly this threshold in another dataset?
- Does the scientific mechanism plausibly imply an abrupt change at this value?
- Am I assuming that values within each category have the same modeled effect?
- How much predictor information will be discarded?
- Am I categorizing primarily because the software output or table will be easier to present?
- Could the predictor remain continuous and be interpreted through meaningful contrasts instead?
- Have I considered whether the relationship is nonlinear?
- Would transformation, polynomial terms, fractional polynomials, or spline functions represent the relationship more appropriately?
- If I searched several cutpoints, has that data-driven search been incorporated into inference or prediction-model validation?
- If prediction is the goal, does categorization unnecessarily reduce information available for individualized prediction?
If the main justification is “this cutoff gave the best result,” the analysis has moved from prespecified categorization into data-driven model searching.
The StatsAlly Decision Rule
| Decision | Preferred approach |
|---|---|
| Continuous measurement with no scientifically established boundary | Keep continuous |
| Continuous predictor adequately represented by a simple functional form | Model continuously |
| Continuous predictor with evidence or scientific expectation of nonlinearity | Consider transformation or flexible continuous modeling |
| Clinically or scientifically meaningful prespecified threshold | Categorization may be defensible for that specific purpose |
| Cutoff chosen because it maximizes statistical significance or apparent performance | Treat as data-driven model selection, not a prespecified threshold |
| Categories created only to make regression output easier to explain | Prefer continuous modeling and meaningful contrasts |
The hierarchy is:
Keep continuous → consider transformation/flexible modeling → categorize only when scientifically justified.
Bottom Line
Categorizing continuous variables is not a harmless formatting decision.
Dichotomization can reduce precision and power, discard information, impose flat relationships within groups, create artificial discontinuities at cutpoints, and weaken individualized prediction. Searching the same dataset for an “optimal” cutoff creates an additional data-dependent modeling problem and can invalidate ordinary inferential calculations if the search is ignored (Harrell, 2015).
The alternative is not to assume that every continuous predictor has a simple linear effect. Continuous predictors can be represented with transformations or flexible nonlinear functions when the scientific relationship requires them. Altman describes polynomial modeling of nonlinear continuous relationships, while Harrell and Steyerberg provide broader frameworks for flexible continuous predictor modeling, including spline-based approaches (Altman, 1991; Harrell, 2015; Steyerberg, 2019).
Preserve the continuous information first. Model its functional form deliberately. Introduce categories only when the categories themselves have a defensible scientific or clinical meaning.
Frequently Asked Questions
Should I categorize continuous variables before regression?
Usually not merely for convenience. Dichotomization reduces precision and power and imposes flat relationships within categories plus abrupt changes at boundaries. Continuous predictors can instead be modeled directly, with nonlinear functions when needed (Harrell, 2015).
Why is dichotomizing continuous variables a problem?
Dichotomization discards distinctions among values within each resulting group. Harrell identifies reduced precision and power, artificial discontinuities, loss of full predictor information, and potential residual confounding among its consequences (Harrell, 2015).
Is it acceptable to use a clinically established cutoff?
Potentially. A threshold that exists independently of the current outcome analysis and corresponds to the actual scientific or clinical question is fundamentally different from a cutoff invented for statistical convenience. Its use should still be justified explicitly, and a clinically useful decision threshold does not necessarily require discarding the continuous measurement from every statistical model.
What is wrong with finding the “optimal” cutoff from my data?
Searching multiple thresholds and reporting the most favorable one makes the cutoff part of the data-driven modeling process. Harrell states that when cutpoints are chosen using the response, ordinary P values and confidence intervals are invalid unless the cutpoint-search process is properly accounted for (Harrell, 2015).
Does keeping a predictor continuous mean assuming a linear relationship?
No. A continuous predictor can be modeled nonlinearly. Altman describes polynomial regression for curved relationships, while Harrell and Steyerberg discuss flexible approaches including spline functions; Steyerberg also discusses fractional polynomials (Altman, 1991; Harrell, 2015; Steyerberg, 2019).
How can I interpret a continuous predictor without creating high and low groups?
Use a scientifically meaningful predictor contrast. Depending on the fitted model, report an effect for a meaningful increment, compare two prespecified values, present predicted outcomes at meaningful values, or display the fitted continuous relationship. Harrell specifically notes that continuous modeling permits estimation between exact predictor settings rather than mixtures defined by arbitrary categories (Harrell, 2015).
Why is categorization particularly problematic for prediction models?
Prediction for an individual can use the exact observed predictor value. Categorization throws some of that information away and can assign the same predictor contribution to substantially different values within a category while creating an abrupt change for nearly identical values on opposite sides of a cutoff. Both Harrell and Steyerberg identify dichotomization as problematic for individualized prediction (Harrell, 2015; Steyerberg, 2019).
References
Altman, D. G. (1991). Practical statistics for medical research. Chapman & Hall.
Harrell, F. E., Jr. (2015). Regression modeling strategies: With applications to linear models, logistic and ordinal regression, and survival analysis (2nd ed.). Springer.
Steyerberg, E. W. (2019). Clinical prediction models: A practical approach to development, validation, and updating (2nd ed.). Springer. https://doi.org/10.1007/978-3-030-16399-0
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.