Visual Tools
Calculators
Tables
Mathematical Keyboard
Converters
Other Tools


Markov Inequality






Markov's Inequality


In many situations, very little is known about a random variable beyond its average size.

Even without knowing the distribution, it is often possible to rule out extreme behavior. Markov's inequality does exactly this: it places an upper bound on the probability that a non-negative random variable exceeds a given level, using only its expected value.

The result is deliberately simple and broadly applicable. It trades precision for generality, providing a guaranteed bound under minimal assumptions.

Key Terms

Expected Value— E[X]E[X], the only quantity needed in Markov's bound
Random Variable— must be non-negative

See All Probability Definitions →


What Markov's Inequality Applies To


Markov's inequality applies to random variables that satisfy very minimal conditions.

Specifically:
• the random variable must be non-negative
• its expected value must exist and be finite
• no assumptions are made about the shape of the distribution

There is no requirement of symmetry, boundedness, or continuity.
The inequality holds equally for discrete and continuous random variables, as long as non-negativity is satisfied.

These minimal assumptions explain both the strength and the weakness of the result: it applies very broadly, but the bound it provides is often coarse.

Statement of Markov's Inequality


Let XX be a non-negative random variable with finite expected value E[X]\mathbb{E}[X].

For any real number a>0a > 0, Markov's inequality states:

P(X≥a)≤E[X]a\mathbb{P}(X \ge a) \le \frac{\mathbb{E}[X]}{a}


This inequality provides an upper bound on the probability that XX exceeds a given threshold, expressed solely in terms of its expectation.

No additional assumptions on the distribution of XX are required.
xPDFE[X]=10a=15010203040
Markov bound for an exponential variable, mean 10, threshold 15

The bound is the mean over the threshold, 10 over 15, and the shaded tail is the true probability of exceeding 15, which for the exponential family is e to the minus 1.5, about 0.22. The inequality holds for any non-negative variable with that mean, whatever its shape. Slide the threshold and watch bound and tail move on the Markov inequality visualizer.

What the inequality says in words is spelled out next.

Markov's Inequality Notation

Notation

Markov's Inequality Notation

This page is the first in the inequalities family to write its marks down, so the shared grammar starts here: the tail event, the bound chain that reads as a guarantee, and the normalized threshold that hands the story on to Chebyshev. All of them are catalogued among the probability symbols.
XX and the event-building braces come from random variable notation; E[X]E[X] from expected value notation; the blackboard register P\mathbb{P}, E\mathbb{E} from distribution notation.
P(X≥a)P(X \ge a)
the probability that X reaches the threshold a — a tail probability
The event "at least aa" wrapped in P(⋅)P(\cdot): everything from the threshold rightward, which is why it is called the right tail. The mirror P(X≤a)P(X \le a) is the left tail — and accumulated left tails are exactly what CDF notation packages as F(a)F(a), making the right tail 1−F(a)1 - F(a) wherever no atom sits at the boundary.
Cases≥\ge versus >> matters only where probability concentrates at the threshold itself: for continuous variables the two spellings agree, for discrete ones they differ by the atom P(X=a)P(X = a) — the same boundary bookkeeping the CDF's chosen ≤ settles on the left side.
Do not confuseA value of XX. P(X≥a)P(X \ge a) is a number about the variable, not a region of outcomes or a truncated variable — bounding it says nothing about which particular large values occur, only about their collective weight.
P(X≥a)≤E[X]a\mathbb{P}(X \ge a) \le \dfrac{\mathbb{E}[X]}{a}
the tail probability never exceeds the mean over the threshold
A probability on the left, a distribution-free recipe on the right, and ≤\le standing between them as a guarantee, not an estimate — the reading What the Inequality Is Saying below insists on. The bound is a ceiling: the truth can sit far beneath it.
CasesThe claim is quantified over every threshold — "for any a>0a > 0" — one inequality per choice of aa, all true at once; when the right side exceeds 11 the statement stays true and says nothing, which is the notation's way of admitting the threshold was too small for the information used.
Do not confuseAn approximation. Reading ≤\le as ≈\approx converts a worst-case ceiling into a prediction — the single most common misuse of the mark, and the reason the page calls the inequality a constraint.
P(X≥k E[X])≤1k\mathbb{P}(X \ge k\,\mathbb{E}[X]) \le \dfrac{1}{k}
at least k times the mean happens with probability at most one over k
The same inequality with the threshold rewritten in units of the mean: substituting a=k E[X]a = k\,\mathbb{E}[X] makes the bound dimensionless — no more than a fifth of the mass can sit at five times the average, whatever the distribution.
CasesThis units-of-the-mean spelling is the bridge to Chebyshev's inequality, which measures its threshold in units of the standard deviation instead — kσk\sigma rather than k E[X]k\,\mathbb{E}[X] — trading wider assumptions for a squared improvement in the bound.
Do not confuseA count. kk is a scale factor on the mean, not a number of events or trials — the same letter that counts successes elsewhere here stretches a threshold, and only the position next to E[X]\mathbb{E}[X] signals which job it holds.

What the Inequality Is Saying


Markov's inequality states that a non-negative random variable cannot take large values too frequently if its average size is small.

If the expected value of a quantity is limited, then the probability of observing values far above that average must also be limited. The larger the threshold chosen, the smaller the guaranteed upper bound on the probability of exceeding it.

The inequality does not attempt to predict how likely large values actually are.
It only guarantees that they cannot occur more often than the bound allows.

For this reason, Markov's inequality should be read as a constraint, not an approximation.

Why the Bound Is So General


Markov's inequality is extremely general because it relies on almost no information.

It uses only two facts:
• the random variable cannot take negative values
• its expected value exists

Nothing else about the distribution matters. The inequality does not depend on symmetry, spread, shape, or tail behavior. As a result, it applies equally to very different random mechanisms.

This generality comes at a cost.
Because the inequality ignores most of the structure of the distribution, the bound it produces is often far from tight.

Markov's inequality is therefore best understood as a baseline bound: it sets a limit that cannot be violated, but it rarely captures the true probability accurately.

Typical Use Cases


Markov's inequality is most often used when only minimal information about a random variable is available.

Common situations include:
• obtaining a quick upper bound on a tail probability
• reasoning about extreme outcomes without knowing a distribution
• providing a first step in theoretical arguments or proofs
• serving as a baseline before applying stronger inequalities

In practice, Markov's inequality is rarely the final result.
It is used to establish a guaranteed bound that can later be improved by introducing additional assumptions or information.

Relationship to Other Inequalities


Markov's inequality is the most basic member of a larger family of probability bounds.

It relies only on non-negativity and expectation, which makes it broadly applicable but weak. Other inequalities strengthen this bound by incorporating additional information about the random variable.

A direct refinement is Chebyshev's inequality, which applies Markov's inequality to squared deviations and uses variance to obtain a tighter bound. Further inequalities introduce higher moments or independence assumptions to sharpen the result even more.

In this sense, Markov's inequality serves as a starting point.
Many stronger probability inequalities can be viewed as extensions or refinements built on its underlying idea.

Limitations of Markov's Inequality


Although Markov's inequality always holds under its assumptions, the bounds it provides are often very loose.

Because it uses only the expected value, the inequality ignores how values are distributed around that average. As a result, the bound may be far larger than the true probability, especially when the random variable has light tails or is tightly concentrated.

Markov's inequality is also uninformative when the threshold is close to the expected value, since the bound may approach or exceed 1. In such cases, it provides little practical insight.

For these reasons, Markov's inequality is best viewed as a guarantee of what cannot happen too often, rather than a precise estimate of what does happen.
xPDFE[X]=10a=8010203040
A threshold below the mean: the bound exceeds 1 and says nothing

When the threshold is smaller than the mean the ratio is larger than 1, and a probability bound above 1 carries no information. Even for thresholds above the mean the bound is usually far from the truth because it ignores everything except the expected value. Move the threshold across the mean and watch the bound become useless on the Markov inequality visualizer.

Its importance lies elsewhere, as the next section argues.

Why Markov's Inequality Matters


Markov's inequality is the simplest non-trivial result that connects expectation to probability.

It shows that meaningful probabilistic statements can be made even when almost no information about a random variable is available. This idea lies at the core of many arguments in probability theory: before refining estimates, one must first establish absolute limits.

Because of its minimal assumptions, Markov's inequality appears repeatedly as a foundational tool. More advanced inequalities refine it, but none bypass the basic logic it introduces.

Summary


Markov's inequality provides an upper bound on the probability that a non-negative random variable exceeds a given level.

It requires only the existence of an expected value and makes no assumptions about distribution shape. The resulting bound is universal but often loose.

For this reason, Markov's inequality is best understood as a baseline result: simple, reliable, and foundational, but rarely the final word in probabilistic analysis.

Markov's Inequality at a Glance

The table below collects the anatomy of Markov's inequality — what it applies to, the single piece of information it consumes, the bound it produces, how to read it, the role it usually plays, and the situations where it stops being informative — into a single reference card.
Aspect Statement Note / example
Applies to non-negative random variables with finite expectation no symmetry, boundedness, or continuity assumed
Information used E[X] only no variance, no higher moments, no distribution shape
Statement P(X ≥ a) ≤ E[X] / a, for any a > 0 the bound depends only on E[X] and the threshold a
What it bounds upper tail of a non-negative variable one-sided; says nothing about P(X ≤ a)
Plain reading large values cannot occur too frequently if the average is small a constraint, not an approximation
Typical role quick baseline bound, first step in proofs rarely the final result; refined by stronger inequalities
Tightness very broad; usually loose ignores how mass is distributed around the mean
When it breaks down threshold a near E[X] makes the bound ≥ 1 and useless also loose for light-tailed or tightly concentrated X
Natural refinement apply Markov to (X − μ)² and use Var(X) this gives Chebyshev's inequality

Markov's Inequality FAQ

What is the 1/k form of Markov's inequality?

+
Substitute a = k·E[X] and the bound becomes dimensionless: P(X ≥ k·E[X]) ≤ 1/k. In words, no more than a fifth of the mass can sit at five times the average, no more than a tenth at ten times, and so on — for any non-negative variable, whatever its distribution. This is the form worth remembering.Read more →

Does Markov's inequality estimate the probability?

+
No — it is a ceiling, not a prediction. The ≤ is a guarantee that the true probability sits at or below E[X]/a, and it may sit far below. Reading ≤ as ≈ turns a worst-case constraint into an estimate, which is the single most common misuse of the inequality.Read more →

When does Markov's inequality tell you nothing?

+
Whenever the threshold is at or below the mean. If a ≤ E[X], then E[X]/a ≥ 1, and the bound says the probability is at most something ≥ 1 — true, but no news, since every probability already satisfies it. The inequality only becomes informative once the threshold sits meaningfully above the average.Read more →

Does it matter whether Markov's inequality uses ≥ or >?

+
Only for discrete variables. The two spellings differ by the atom P(X = a) sitting exactly at the threshold, which is zero for a continuous variable and can be substantial for a discrete one. Markov is normally stated with ≥, so the boundary value is included; switching to > gives a smaller left side and the bound still holds.Read more →