Levene’s Test: How to Check Equality of Variances
Levene’s test checks whether two or more groups have equal variances, helping researchers assess the homogeneity of variance assumption before certain statistical analyses. This guide explains how the test works, how to interpret its p-value, when to use Welch’s t-test or Welch’s ANOVA, and how Levene’s test compares with Brown-Forsythe and Bartlett’s test.
Statistics calculator and practical guide
Levene's Test Calculator for Equality of Variances
Use Levene's test to assess whether two or more independent groups have similar population variances. Enter or upload your data, run the test, and use the guide below to understand the hypotheses, assumptions, test statistic, p-value, and what to do when variances are unequal.
Run Levene's Test
Developer note: Mount the existing DataClue Levene's test calculator here. Preserve its CSV/Excel upload, sample-data option, variable selection, and analysis controls.
On this page
- What Levene's test measures
- Null and alternative hypotheses
- How to use the calculator
- How to interpret the result
- Levene's test formula
- Mean, median, and trimmed-mean versions
- Assumptions and requirements
- Worked example
- Alternatives and next steps
- Reproduce the test in Python
- How to report Levene's test
- Frequently asked questions
What is Levene's test?
Levene's test is a statistical test for homogeneity of variance, meaning equality of population variances across two or more groups. It is often used when a later analysis, such as a conventional independent-samples t-test or one-way ANOVA, relies on an equal-variance assumption.
The method was introduced by Howard Levene in 1960. Instead of comparing sample variances directly, it transforms each observation into an absolute deviation from a group center and then compares those deviations across groups. This construction makes Levene-type procedures less sensitive to departures from normality than Bartlett's test.
Null and alternative hypotheses
For k independent groups, Levene's test evaluates:
Null hypothesis (H0): σ12 = σ22 = ... = σk2
Alternative hypothesis (H1): at least one group variance differs from another.
A statistically significant result suggests evidence against equal variances. A non-significant result means the data do not provide enough evidence to conclude that the variances differ. It does not prove that all population variances are exactly equal.
How to use the Levene's test calculator
- Prepare your data. Use one row per observation. Include a numeric outcome column and a grouping column that identifies the independent groups.
- Paste data or upload a CSV/Excel file. Use the first row for variable names.
- Select the outcome variable. This should be the numeric variable whose variance you want to compare, such as test score, blood pressure, yield, or response time.
- Select the grouping variable. This variable defines the categories being compared, such as treatment group, classroom, machine, or experimental condition.
- Run the analysis. Review the Levene statistic, degrees of freedom, and p-value returned by the calculator.
- Interpret the result in context. Compare the p-value with your prespecified significance level, commonly α = 0.05, while also inspecting the data for outliers and unusual group distributions.
For the sample structure already shown on this DataClue page, Score can serve as the numeric outcome and Group as the categorical grouping variable.
How to interpret Levene's test
| Result | Statistical decision | Practical interpretation |
|---|---|---|
| p < α | Reject H0 | There is evidence that at least one population variance differs. |
| p ≥ α | Fail to reject H0 | There is insufficient evidence to conclude that the population variances differ. |
Do not interpret a large p-value as proof of equal variances. With small samples, the test may have limited power to detect real differences. It is good practice to inspect descriptive statistics and plots alongside the formal test.
Levene's test formula
Suppose there are k groups, a total sample size of N, and group i contains Ni observations. First define an absolute-deviation score for each observation:
Zij = |Yij − Ci|
Here, Ci is a chosen center for group i. In the original Levene test it is the group mean. Robust variants use the group median or a trimmed mean.
The test statistic can then be written as:
W = [(N − k)/(k − 1)] × [Σ Ni(Z̄i. − Z̄..)²] / [ΣΣ(Zij − Z̄i.)²]
Under the null hypothesis, the statistic is evaluated against an F distribution with k − 1 and N − k degrees of freedom. Conceptually, this is an ANOVA-style comparison applied to the absolute deviations rather than to the original observations.
Mean, median, or trimmed mean: which version should you use?
The choice of center affects the robustness and power of the test. NIST and SciPy describe three common variants:
| Center | Typical use | Notes |
|---|---|---|
| Mean | Approximately symmetric, moderate-tailed data | This is the original form proposed by Levene. |
| Median | Skewed or generally non-normal data | Often called the Brown-Forsythe modification and commonly chosen for robustness. |
| Trimmed mean | Heavy-tailed data | Reduces the influence of extreme observations by trimming observations before calculating the center. |
If you do not have a strong distributional reason to prefer the mean-based form, the median-centered version is a defensible robust choice for many practical datasets. The choice should be made based on the data-generating situation rather than by trying several versions and reporting whichever gives the preferred p-value.
Assumptions and data requirements
Levene's test is more robust to non-normality than Bartlett's test, but it still has important requirements:
- Independent observations: measurements should be independent within and across groups for the standard test.
- Independent groups: the usual form is designed for separate groups, not paired or repeated-measures observations.
- Numeric outcome: the variable whose spread is being compared should be measured on a meaningful quantitative scale.
- Categorical grouping variable: each observation must belong to one of the groups being compared.
- Enough observations to estimate spread: extremely small groups provide weak information about variance and can produce unstable conclusions.
Normality is not a strict requirement for using a Levene-type test in the way it is for Bartlett's test. However, skewness, heavy tails, outliers, imbalance, and very small samples can still affect test behavior, which is why the choice of mean, median, or trimmed-mean center matters.
Worked example using the sample scores
Consider the two score groups displayed in the sample data on this calculator page:
| Group A | Group B |
|---|---|
| 78 | 58 |
| 85 | 64 |
| 90 | 70 |
| 76 | 62 |
| 82 | 68 |
Illustrative median-centered Levene/Brown-Forsythe result:
W = 0.1125, p ≈ 0.746
At α = 0.05, p is greater than 0.05, so we fail to reject the null hypothesis of equal variances. For these ten observations, there is not enough evidence to conclude that the score variances differ between Group A and Group B.
This example is intentionally small and should be treated as a demonstration of interpretation, not as evidence that a non-significant Levene test establishes variance equality in every sample.
Levene's test vs. common alternatives
| Method | Use it when | Main consideration |
|---|---|---|
| Levene's test | You need a general test of equal variances across independent groups. | Less sensitive to non-normality than Bartlett's test. |
| Brown-Forsythe modification | Your data are skewed or contain outliers and you want a robust Levene-type test. | Uses deviations from group medians. |
| Bartlett's test | Data are well approximated by normal distributions. | Can have good performance under normality but is sensitive to departures from normality. |
| Fligner-Killeen test | You want a distribution-free-style robust alternative for comparing spread. | Useful when normality is doubtful. |
| Welch's t-test / Welch's ANOVA | Your actual research question is about means and equal variances are doubtful. | These compare means while allowing unequal variances; they are not tests of variance equality themselves. |
Reproduce Levene's test in Python
Researchers can reproduce the worked example with SciPy:
from scipy.stats import levene
group_a = [78, 85, 90, 76, 82]
group_b = [58, 64, 70, 62, 68]
statistic, p_value = levene(group_a, group_b, center="median")
print(statistic, p_value)
In SciPy, center="mean" gives the original Levene form, center="median" gives the median-centered Brown-Forsythe form, and center="trimmed" uses a trimmed mean.
How to report Levene's test
A concise report should state the version of the test when relevant, the test statistic, degrees of freedom, p-value, and interpretation.
Example wording:
“A median-centered Levene test did not provide evidence of unequal variances between the two groups, W(1, 8) = 0.11, p = .746.”
If the result is significant, report that there is evidence of variance heterogeneity rather than saying the assumption has “failed” without context. Then choose a downstream method appropriate to the research question and design.
Frequently asked questions
What does a significant Levene's test mean?
A significant result means the observed data provide evidence against the null hypothesis that all group variances are equal. It indicates that at least one population variance appears to differ, but it does not identify which specific pair of groups is responsible.
What does p > 0.05 mean in Levene's test?
At a 0.05 significance level, p > 0.05 means you fail to reject the equal-variance null hypothesis. This is an “insufficient evidence of a difference” conclusion, not proof that the population variances are identical.
Does Levene's test require normal data?
It is commonly used because it is less sensitive to departures from normality than Bartlett's test. The median-centered Brown-Forsythe modification is often preferred when distributions are skewed, while a trimmed-mean version can be useful for heavy-tailed data.
Can I use Levene's test for more than two groups?
Yes. Levene's test is defined for two or more independent groups. The null hypothesis states that all group population variances are equal.
Is Levene's test the same as the Brown-Forsythe test?
They belong to the same family of procedures. The original Levene test centers observations around each group mean. A widely used Brown-Forsythe modification centers them around the group median, improving robustness in many non-normal settings.
What should I do if Levene's test is significant before ANOVA?
If the research question concerns group means and unequal variances are plausible, consider a method designed for heteroscedastic data, such as Welch's ANOVA, rather than automatically proceeding with a conventional equal-variance ANOVA. The final choice should also account for independence, sample-size imbalance, distribution shape, and the study design.
Can Levene's test be used with paired or repeated-measures data?
The standard test assumes independent observations and independent groups, so it is not a direct fit for paired or repeated-measures designs. Those designs require methods that account for within-subject dependence.
Methodology and sources
This guide follows standard descriptions of Levene-type tests from the NIST/SEMATECH e-Handbook of Statistical Methods and the SciPy statistical computing documentation. The distinction among mean-, median-, and trimmed-mean-centered versions follows the original work by Levene and the robust modifications studied by Brown and Forsythe.
- NIST/SEMATECH: Levene Test for Equality of Variances
- SciPy documentation: scipy.stats.levene
- Levene, H. (1960). “Robust Tests for Equality of Variances.” In Contributions to Probability and Statistics, pp. 278–292.
- Brown, M. B., & Forsythe, A. B. (1974). “Robust Tests for the Equality of Variances.” Journal of the American Statistical Association, 69, 364–367.
Try it in DataClue
Ready to run Levene's Test?
Test the equality of variances across groups in a sample.
Run Levene's Test