math combinatorics probability

Stirling's Approximation: A Formula That Tames Astronomical Factorials

Shuffle a deck of cards once, honestly, and you've almost certainly created an arrangement that has never existed before in the history of the universe and never will again. The number of possible orderings of 52 cards is 52! — "52 factorial," meaning 52 × 51 × 50 × ... × 2 × 1 — and that number is about 8.0658 × 10⁶⁷. Earth is roughly 4.5 billion years old, measured in seconds that's about 1.4 × 10¹⁷. Even if every person who has ever lived had been shuffling a new deck every second since the Big Bang, humanity wouldn't have scratched the surface of the possibilities. Factorials grow so viciously fast that they stop feeling like numbers and start feeling like jokes.

And yet physicists, statisticians, and computer scientists use expressions full of enormous factorials every single day — counting molecular arrangements in a gas, computing probabilities in coin flips, estimating binomial coefficients in information theory — without ever directly multiplying out a 67-digit monster. The reason they can get away with it is a 1730 formula from a Scottish mathematician that turns an unmanageable product into a clean, continuous curve: Stirling's approximation.

The Concept

Stirling's approximation says that for large n,

n! ≈ √(2πn) · (n/e)ⁿ

where e is Euler's number (≈2.71828) and π is, well, π. At first glance this looks like a strange thing to be true — the left side is a jagged, purely discrete product of whole numbers, and the right side is built from two of the smoothest constants in all of mathematics. But it works, and it works remarkably well, almost immediately.

A more commonly used cousin of the formula, especially in physics, takes the logarithm of both sides and drops the smaller terms:

ln(n!) ≈ n ln(n) − n

This stripped-down version is what you'll actually see scrawled on physics chalkboards, because in most of the places factorials show up in science, what you really want is the logarithm of a factorial (entropy, for instance, is fundamentally a logarithm of a count), and this version is almost embarrassingly simple to work with.

Where does it come from? The cleanest way to see it is to replace the discrete product n! with a continuous integral. There's a classical identity (first explored by Euler) that says

n! = ∫₀^∞ xⁿ e^(−x) dx

This integral is dominated by values of x near x = n — the integrand xⁿe^(−x) rises, peaks sharply around x = n, and falls away again, and the peak gets taller and narrower (in relative terms) as n grows. If you zoom in near that peak, the curve looks almost exactly like a bell curve — a Gaussian. The height of that peak gives you the (n/e)ⁿ term, and the width of the bell curve around it gives you the √(2πn) term. In other words, Stirling's formula is really a statement that factorials, viewed the right way, are secretly Gaussian. This "approximate a sharply peaked function by a Gaussian near its peak" trick is called Laplace's method, and it shows up constantly across physics and probability whenever something is being maximized or summed over an enormous number of possibilities.

Why It Matters

The formula earns its keep because factorials are everywhere that counting is involved, and counting problems explode combinatorially the moment you look away.

In physics, entropy is fundamentally a count. Ludwig Boltzmann's great insight — immortalized on his tombstone in Vienna's Zentralfriedhof as S = k log W (it was actually Max Planck who later wrote it in that precise symbolic form, crediting Boltzmann for the underlying idea) — says that the entropy S of a physical system is proportional to the logarithm of W, the number of microscopic arrangements ("microstates") consistent with what you can observe macroscopically. For something as ordinary as a container of gas, W involves the factorial of the number of molecules, and the number of molecules in even a thimbleful of gas is on the order of Avogadro's number, roughly 6 × 10²³. No computer that will ever exist could compute that factorial exactly. But physicists don't need to: they take ln(N!) ≈ N ln N − N, plug it into the entropy formula, and out pops the Sackur-Tetrode equation — the actual working formula (independently derived by Otto Sackur and Hugo Tetrode around 1912) that predicts the entropy of an ideal gas from its volume, temperature, and particle count, and that has been checked against real measurements ever since. The same trick, maximizing a huge count of arrangements, is how physicists derive the Maxwell-Boltzmann, Fermi-Dirac, and Bose-Einstein distributions that describe how energy spreads itself across particles in everything from stars to semiconductors.

In probability and statistics, Stirling's approximation is the engine behind one of the most important results in the field: the fact that flipping a fair coin a large number of times produces outcomes that cluster into a bell curve. This was first shown by Abraham de Moivre in the 1730s (and generalized later by Laplace), in what's now called the de Moivre–Laplace theorem — an early, special case of the Central Limit Theorem. The proof works by taking the binomial probability formula, which is stuffed with factorials, and applying Stirling's approximation to each one; the jagged binomial distribction smooths out into the familiar Gaussian curve. It's a small irony of history that de Moivre, while proving this theorem, independently worked out a version of the same factorial approximation — he just didn't nail down the exact constant √(2π) the way Stirling did, which is part of why Stirling's name, not de Moivre's, ended up on the formula.

In information theory, Stirling's approximation is what lets you estimate binomial coefficients — the number of ways to choose k items from n — without computing giant factorials directly. The result connects combinatorics directly to Shannon's entropy: the number of ways to arrange a sequence with a given proportion of symbols grows, to leading order, as 2 raised to n times the binary entropy of that proportion. This single fact underlies data compression theory and the "typical set" arguments that tell you how much you can compress a signal before you start losing information.

In random walks, the same approximation tells you that the probability of a wandering particle returning exactly to its starting point after 2n steps shrinks roughly as 1/√(πn) — slowly enough that, remarkably, a one-dimensional random walk is guaranteed to return home eventually (and a two-dimensional one too), while in three dimensions or more it generally wanders off and never comes back. That asymmetry between dimensions — famously summarized by mathematician Shizuo Kakutani as "a drunk man will find his way home, but a drunk bird may get lost forever" — traces directly back to how Stirling's approximation behaves as the exponent grows.

The Details

Let's see just how good this "approximation" really is, because the word undersells it. Here is the exact value of n! next to Stirling's estimate, with the relative error:

  • n = 1: exact = 1, Stirling ≈ 0.922 — off by 7.8%
  • n = 10: exact = 3,628,800, Stirling ≈ 3,598,696 — off by 0.83%
  • n = 20: exact ≈ 2.4329 × 10¹⁸, Stirling ≈ 2.4228 × 10¹⁸ — off by 0.42%
  • n = 52 (that card deck!): exact ≈ 8.06582 × 10⁶⁷, Stirling ≈ 8.05290 × 10⁶⁷ — off by 0.16%
  • n = 100: exact ≈ 9.33262 × 10¹⁵⁷, Stirling ≈ 9.32485 × 10¹⁵⁷ — off by 0.083%

Notice the pattern: the error keeps roughly halving every time n doubles, and it was already under 1% by the time n reached 10. For anything in physics, where n is typically in the trillions upon trillions, the "approximation" is accurate to more decimal places than any instrument could ever measure. And if you need even more precision, there's a refined version of the formula with a correction factor:

n! ≈ √(2πn) · (n/e)ⁿ · (1 + 1/(12n) + ...)

Adding just that one extra term drops the error at n = 10 from 0.83% down to about 0.0003% — accurate to four decimal places from a formula you could write on an index card.

It's worth sitting with how strange this is. The factorial function is about as "discrete" as mathematics gets: you can't take half a factorial of anything in the ordinary sense; it's defined only by repeated multiplication of whole numbers. Yet buried inside it, once n gets large, is a perfectly smooth exponential curve wearing a square-root correction as a hat — and that curve was hiding there even for n as small as 10. This is a small instance of a much bigger pattern in mathematics: that discreteness and smoothness are far more intertwined than they first appear, and that continuous tools (integrals, calculus, Gaussian curves) are often the fastest route to understanding something that looks purely combinatorial.

The historical story behind the formula has its own quiet drama. James Stirling was born in 1692 near the Scottish town of Stirling, studied at Balliol College, Oxford, and then lost his scholarship when his Jacobite family's politics caught up with him during the 1715 rebellion. He spent several years in Venice teaching mathematics — locals reportedly nicknamed him "the Venetian" — before returning to Britain with help from Isaac Newton, who sponsored his election to the Royal Society in 1726. Four years later, in 1730, Stirling published his approximation in a book called Methodus Differentialis. But he wasn't working in a vacuum: Abraham de Moivre, working on the probability problem described above, had already found the general shape of the same approximation and knew it involved some unknown constant. It was Stirling who pinned that constant down as √(2π) — and de Moivre said so himself, in print, crediting Stirling by name. So the formula's name reflects a genuinely collaborative (if slightly contested) discovery, not a solo one. Stirling later left academic mathematics almost entirely, spending the rest of his career managing lead mines in Leadhills, Scotland, where he modernized mine ventilation and cut miners' shifts to six hours — a surprising second act for a man whose name now sits permanently inside the machinery of thermodynamics and probability theory.

Takeaways

  • Stirling's approximation, n! ≈ √(2πn)(n/e)ⁿ, turns a wildly discrete, explosively growing product into a smooth, continuous curve built from e and π.
  • It's already accurate to under 1% by n = 10, and the error keeps shrinking the larger n gets — which is exactly when science needs it most, since real factorials in physics involve numbers like Avogadro's constant (~6×10²³).
  • It's the hidden engine behind Boltzmann's entropy formula, the Sackur-Tetrode equation for ideal gases, the de Moivre–Laplace theorem (an ancestor of the Central Limit Theorem), and the entropy bounds on binomial coefficients used in information theory.
  • The name is really a joint credit: Abraham de Moivre found the general shape of the formula while studying probability, and James Stirling determined the precise constant (√(2π)) that makes it exact.
  • It's a vivid reminder that discrete, combinatorial objects — the things you count one at a time — often have smooth, continuous structure hiding just beneath the surface, visible only once the numbers get large enough.

Resources: MacTutor History of Mathematics's biography of James Stirling; Wikipedia's articles on "Stirling's approximation," "Boltzmann's entropy formula," and "De Moivre–Laplace theorem"; the "That's Maths" blog post "Factorial 52: A Stirling Problem," for a fun deep dive on the card-shuffling number.