Visual Tools
Calculators
Tables
Mathematical Keyboard
Converters
Other Tools


Chebyshev's Inequality






Chebyshev's Inequality


In many probability problems, the average of a random variable is known, but its full distribution is not.

Even in this limited setting, information about how much the variable varies around its mean can strongly restrict how often large deviations occur. Chebyshev's inequality captures this idea by relating deviation probabilities directly to the variance.

The result is deliberately general. It applies without assuming any particular distribution shape and provides guaranteed bounds on how far values can stray from the mean, based solely on variability.

Key Terms

Variance— Var⁡(X)\operatorname{Var}(X), the key input to Chebyshev's bound
Standard Deviation— σ\sigma, used to express the deviation threshold
Expected Value— μ=E[X]\mu = E[X], the center of the bound
Random Variable— must have finite mean and variance

See All Probability Definitions →


What Chebyshev's Inequality Applies To


Chebyshev's inequality applies to random variables under very broad conditions.

Specifically:
• the random variable must have a finite mean
• the variance must exist and be finite
• no assumptions are made about the distribution's shape

The inequality applies equally to discrete and continuous random variables.
There is no requirement of symmetry, boundedness, or independence.

Compared to Markov's inequality, Chebyshev uses more information — variance — and therefore provides a tighter bound, while still remaining distribution-free.

Statement of Chebyshev's Inequality


Let XX be a random variable with finite mean μ\mu and finite variance σ2\sigma^2.

For any real number k>0k > 0, Chebyshev's inequality states:

P(∣X−μ∣≥k σ)≤1k2\mathbb{P}\big(|X - \mu| \ge k\,\sigma\big) \le \frac{1}{k^2}


Equivalently, for any ε>0\varepsilon > 0,

P(∣X−μ∣≥ε)≤σ2ε2\mathbb{P}\big(|X - \mu| \ge \varepsilon\big) \le \frac{\sigma^2}{\varepsilon^2}


The bound depends only on the variance and makes no assumptions about the form of the distribution.
xPDFμ=10μ-aμ+a24681012141618
Chebyshev bound for a normal variable, mean 10, variance 4, deviation 3

The bound is the variance divided by the square of the deviation, here 4 over 9, and the shaded tails show the true probability of deviating by at least 3, which is much smaller. Two units of standard deviation against a deviation of three produces the same ratio for every distribution. Change the variance or the deviation and watch the bound move on the Chebyshev inequality visualizer.

The next section reads the inequality in words.

Chebyshev's Inequality Notation

Notation

Chebyshev's Inequality Notation

Chebyshev's marks refine Markov's in two moves: the tail turns two-sided, and the threshold changes units — from multiples of the mean to multiples of the standard deviation. This section fixes the deviation event, the bound's two equivalent spellings, and the flipped form that reads as a coverage guarantee. All of them are catalogued among the probability symbols.
The one-sided tail P(X≥a)P(X \ge a) and the ≤-as-guarantee reading come from Markov's inequality notation; the deviation event ∣X−μ∣>ε|X - \mu| > \varepsilon and the quantified ε\varepsilon from law of large numbers notation; μ\mu and σ\sigma from expected value and variance notation.
∣X−μ∣≥kσ|X - \mu| \ge k\sigma
X lands at least k standard deviations from its mean — in either direction
Two refinements in one event: the absolute value makes the tail two-sided — both directions at once, where Markov's X≥aX \ge a watches only one — and the threshold kσk\sigma is measured in units of the standard deviation, the origin of everyday "two-sigma, three-sigma" talk.
CasesThe same event in absolute units swaps kσk\sigma for a plain tolerance ε\varepsilon — the spelling the law of large numbers runs on; the two are interchangeable through ε=kσ\varepsilon = k\sigma, and which one a text uses signals whether it thinks in σ-units or in raw distance.
Do not confuseMarkov's kk. There the letter multiplies the mean, here the standard deviation — same letter, different unit, and carrying a bound from one threshold convention to the other silently changes what is being claimed.
≤1k2\le \dfrac{1}{k^2}, ≤σ2ε2\le \dfrac{\sigma^2}{\varepsilon^2}
at most one over k squared — at most variance over tolerance squared
One inequality, two dresses: substitute ε=kσ\varepsilon = k\sigma and each becomes the other. The σ\sigma-units spelling gives the clean distribution-free constants; the ε\varepsilon spelling exposes the mechanism — variance up top, tolerance squared below — that Why Chebyshev Improves on Markov below explains.
CasesThe bound only says something when k>1k > 1 — at k≤1k \le 1 the right side reaches 11 and the statement goes vacuous, Markov-style; the square is the visible improvement over Markov's 1/k1/k: twice the threshold buys four times the guarantee, not two.
Do not confuseEach other, mid-formula. Writing σ2/k2\sigma^2/k^2 or 1/ε21/\varepsilon^2 mixes the two spellings — the numerator and the denominator must come from the same convention, and the mismatch produces bounds that are silently off by a factor of σ2\sigma^2.
P(∣X−μ∣<kσ)≥1−1k2\mathbb{P}\big(|X - \mu| < k\sigma\big) \ge 1 - \dfrac{1}{k^2}
at least the fraction one minus one-over-k-squared lies within k standard deviations
The same fact flipped through the complement: the tail bound becomes a coverage guarantee, and both marks turn around together — ≥\ge becomes << inside, ≤\le becomes ≥\ge outside. In numbers: at least 75%75\% within 2σ2\sigma, at least 89%89\% within 3σ3\sigma, for every distribution with a variance.
CasesThis is the honest, distribution-free cousin of the normal distribution's 68–95–99.7 rule — Chebyshev promises less (0, 75, 89) because it assumes nothing about shape.
Do not confuseThe normal rule itself. Quoting "95% within two sigmas" for data nobody checked for normality claims the bell curve's privilege without its assumptions — Chebyshev's 75%75\% is what the variance alone can promise, and the gap between the two numbers is exactly what normality is worth.

What the Inequality Is Saying


Chebyshev's inequality states that a random variable with small variance cannot stray far from its mean very often.

The variance measures how spread out the values of the variable are. If this spread is limited, then large deviations from the mean must be rare. The inequality makes this idea precise by turning variability into a bound on deviation probabilities.

The result does not predict where values will cluster.
It only guarantees that extreme departures from the mean are constrained by the size of the variance.

For this reason, Chebyshev's inequality should be read as a guarantee, not as a description of typical behavior.

Why Chebyshev Improves on Markov


Markov's inequality uses only the expected value of a random variable and applies to non-negative quantities. As a result, it places very weak restrictions on how values are distributed.

Chebyshev's inequality incorporates additional information: the variance. By measuring how values spread around the mean, variance allows much stronger control over deviations in both directions.

In effect, Chebyshev's inequality applies Markov's idea to squared deviations from the mean. This extra structure tightens the bound while still avoiding any assumptions about the exact form of the distribution.

Chebyshev therefore refines Markov's inequality: it remains distribution-free, but achieves stronger bounds by using more information.

Typical Use Cases


Chebyshev's inequality is used when information about variability is available, but the exact distribution is unknown or unnecessary.

Common uses include:
• bounding the probability of large deviations from the mean
• analyzing the stability of averages and estimators
• providing guarantees in problems involving uncertainty and noise
• serving as a tool in theoretical arguments where distributional details are absent

In many applications, Chebyshev's inequality provides a first meaningful bound that can later be refined if stronger assumptions are introduced.

Role in Convergence and Limit Results


Chebyshev's inequality plays a central role in establishing convergence results in probability.

Because it controls the probability of deviations from the mean using variance, it is naturally suited to analyzing sequences of random variables, especially averages. When variances shrink, Chebyshev's inequality shows that large deviations become increasingly unlikely.

This makes the inequality a key tool in proofs of the Law of Large Numbers, where it is used to show that sample averages concentrate around their expected value.

In this role, Chebyshev's inequality acts as a bridge between basic moment information and asymptotic behavior.

Limitations of Chebyshev's Inequality


Although Chebyshev's inequality is widely applicable, the bounds it provides are often conservative.

Because it depends only on variance, the inequality ignores finer features of the distribution such as shape, symmetry, or tail behavior. As a result, the bound may be much larger than the true probability, especially for well-behaved distributions.

Chebyshev's inequality is also insensitive to how deviations occur. Two random variables with the same variance but very different distributions receive the same bound, even if one is far more concentrated than the other.

For tighter control, additional assumptions or more specialized inequalities are usually required.
xPDFμ=10μ-aμ+a24681012141618
Chebyshev bound for a uniform variable

For a uniform variable the true tail probability can be zero while the bound stays positive: the inequality knows only the variance and cannot see that the distribution has no tails at all. That gap is the price of generality. Switch between families and see how loose the bound gets on the Chebyshev inequality visualizer.

Why the inequality matters despite this is explained next.

Why Chebyshev's Inequality Matters


Chebyshev's inequality is the first probability bound that explicitly connects variability to deviation control.

By using variance, it shows that knowing how spread out a random variable is already imposes strong restrictions on how often extreme values can occur. This insight is fundamental in probability theory and underlies many results about stability and convergence.

Because it makes no assumptions about distribution shape, Chebyshev's inequality remains broadly applicable while still being strong enough to support key theoretical arguments.

Summary


Chebyshev's inequality provides an upper bound on the probability that a random variable deviates far from its mean.

It requires only finite mean and variance and applies to both discrete and continuous random variables. The resulting bound is distribution-free and stronger than what can be obtained from expectation alone.

For this reason, Chebyshev's inequality is a central tool in probability theory: simple, general, and essential for understanding convergence and variability.

Chebyshev's Inequality at a Glance

The table below collects the anatomy of Chebyshev&apos;s inequality — what it applies to, the moments it consumes, the bound it produces, how to read it, why it strengthens Markov, the role it plays, its key application, and where its sensitivity ends — into a single reference card.
Aspect Statement Note / example
Applies to any random variable with finite mean and finite variance no symmetry, shape, or boundedness assumed; discrete or continuous
Information used E[X] and Var(X) one extra moment beyond Markov
Statement P(|X − μ| ≥ kσ) ≤ 1 / k², for any k > 0 equivalently, P(|X − μ| ≥ ε) ≤ σ² / ε²
What it bounds two-sided deviation from the mean covers both tails simultaneously, unlike Markov
Plain reading small variance ⟹ large deviations from μ must be rare a guarantee, not a description of typical behaviour
Why it improves on Markov applies Markov to (X − μ)² and converts variance into a tail bound tighter bound while still distribution-free
Typical role bound deviations, analyze averages and estimators first meaningful bound when only mean and variance are known
Key application classical proof of the Law of Large Numbers Var(X̄n) = σ² / n shrinks ⟹ concentration around μ
Limitation sees only variance, not shape, symmetry, or tail behaviour two distributions with the same σ² receive the same bound

Chebyshev's Inequality FAQ

How much of a distribution lies within two standard deviations?

+
At least 75%, for any distribution with a finite variance. Chebyshev's coverage form says P(|X − μ| < kσ) ≥ 1 − 1/k², giving 75% at two sigma and about 89% at three. The familiar 95% figure belongs to the normal distribution, not to Chebyshev — quoting it for unchecked data claims the bell curve's privilege without its assumptions.Read more →

Why does Chebyshev's inequality use an absolute value?

+
To make the bound two-sided. |X − μ| ≥ kσ captures deviations in both directions at once — too far above the mean and too far below — where Markov's X ≥ a watches only one side. That symmetry is possible because variance measures spread around the mean rather than size, so no direction is privileged.Read more →

Is the bound 1/k² or σ²/ε²?

+
Both, but never mixed. They are the same inequality in two conventions: put ε = kσ and each becomes the other. The σ-units spelling gives the clean constants; the ε spelling exposes the mechanism, variance over squared tolerance. Writing σ²/k² or 1/ε² crosses the two and produces a bound silently off by a factor of σ².Read more →

Does k mean the same thing in Markov's and Chebyshev's inequalities?

+
No — same letter, different unit. In Markov's normalized form k multiplies the mean, so the threshold is k·E[X]. In Chebyshev's it multiplies the standard deviation, so the threshold is kσ. Carrying a value of k from one to the other silently changes what is being claimed.Read more →