|

What is the Law of Large Numbers? A visual Python/Pandas example.

For this example we’ll look at a dataframe with salaries of baseball players in a given year. And we’ll use the Pandas, Numpy and Matplotlib libraries to makes sense of the numbers in a Histogram plots.

Let’s clarify what is meant by the mean of the baseball salaries we are reviewing and provide a deeper understanding of what that value means and how it can be corroborated. As you know a simple calculation using Python’s Numpy library (or a function in Excell for also) can provide the mean salary from this data set which would result in rounded 2,497,668.69 Dollars. As you can see by this next graph however the range of salaries varies greatly some star players make a lot more money than others. Please note that the values here are described as le7 which means tens of millions. So the upper range of salaries is in the 20+ million Dollar range.


You can see that the lower salaries though are a lot more frequent thus the mean (aka average) value is much lower. While that makes intuitive sense I wanted to provide a deeper understanding of how that number is derived by demonstrating a principle called the Law of Large Numbers. The law of large numbers means that if I take the mean of larger and larger samples of the data set we approach the population mean number.

So here you’ll see graphs with an increasing number of samples being provided to the 1000 sample mean calculations. Here’s what the calculation code looks like:

# Histogram1 Graph we take 1000 mean calculations of 5 random samples from the dataset
sample_means_5 = []
for i in range(1000):
  sample_means_5.append(baseball_salaries['salary'].sample(5).mean())
  pass

# Histogram2 Graph we take 1000 mean calculations of 10 random samples from the dataset
sample_means_10 = []
for i in range(1000):
  sample_means_10.append(baseball_salaries['salary'].sample(10).mean())
  pass

# And so on for 25,50,100...

Notice that the red line is the overall mean of all the samples and the green lines are the standard deviation marks above and below the mean. As you can see the numbers and appearance of the graph approximate the population histogram, mean and standard deviation. Each calculates 1000 means but each time from an increasing amount of samples starting a 5 each time all the way to 100. (5,10,25,50,100)

In this way we can see that the population mean that we calculated is approached in the last graph a lot more than in the first graph. The same is true for the standard deviation. I hope this helps in your understanding.

In this way we can see that the population mean that we calculated is approached in the last graph a lot more than in the first graph. See the numbers for each graph’s mean. (From Histogram1 through Histogram5)

2488890.7573999995
2487792.1818000004
2463101.6333999997
2501681.20268
2502497.11802

The same is true for the standard deviation. (From Histogram1 through Histogram5)

1533419.8924479042
1559738.6948010693
714972.2438367842
473714.0224874232
331423.4540842278

I hope this helps in your understanding.


 

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *