Question 1

What is variance and what does it measure?

Accepted Answer

Variance is a measure of how spread out the values in a dataset are from the mean. It quantifies the average squared deviation from the mean, giving greater weight to values that are farther from the center. A small variance indicates that data points cluster tightly around the mean, while a large variance indicates they are widely scattered. Variance is calculated by finding the mean, computing the squared difference of each value from the mean, and then averaging those squared differences. Variance is always non-negative, with zero variance indicating all values are identical. It serves as the foundation for many statistical techniques including hypothesis testing, confidence intervals, ANOVA, and regression analysis.

Question 2

Why do we use sample variance (n-1) instead of population variance (n)?

Accepted Answer

Sample variance divides by (n-1) instead of n to correct for a statistical bias called underestimation. When you calculate the mean from a sample and then measure deviations from that sample mean, the deviations tend to be smaller than they would be from the true population mean. This happens because the sample mean minimizes the sum of squared deviations for that particular sample. Dividing by (n-1) instead of n corrects this bias, producing an unbiased estimate of the population variance. This correction factor (n-1) is called degrees of freedom. The difference matters most for small samples; with n = 5, dividing by 4 versus 5 changes the result by 20%. For large samples (n greater than 100), the difference becomes negligible.

Question 3

What is the relationship between variance and standard deviation?

Accepted Answer

Standard deviation is simply the square root of variance. While both measure spread, they differ in units. If your data is in meters, variance is in meters squared, which is hard to interpret. Standard deviation brings the measurement back to the original units (meters), making it directly comparable to the data values. The empirical rule (68-95-99.7 rule) states that for normally distributed data, about 68% of values fall within one standard deviation of the mean, 95% within two, and 99.7% within three. This makes standard deviation an intuitive measure of typical deviation from the mean. Variance is preferred in mathematical derivations because it has nicer algebraic properties: the variance of a sum of independent variables equals the sum of their variances.

Question 4

What is the standard error of the mean and how does it differ from standard deviation?

Accepted Answer

Standard deviation measures the variability of individual observations within a dataset. Standard error of the mean (SEM) measures the precision of the sample mean as an estimate of the population mean. SEM equals the standard deviation divided by the square root of the sample size: SEM = SD/sqrt(n). Because of the square root relationship, quadrupling your sample size halves the standard error. SEM decreases with larger samples because the sample mean becomes a more precise estimate. Confidence intervals for the mean use SEM: a 95% CI is approximately mean plus or minus 2*SEM. Report standard deviation when describing the spread of individual values, and report SEM when describing the precision of the mean estimate. Confusing these two measures is a common error in research publications.

Variance and Standard Deviation Calculator

Formula

Worked Examples

Example 1: Exam Score Analysis

Example 2: Manufacturing Quality Control

Frequently Asked Questions

What is variance and what does it measure?

Why do we use sample variance (n-1) instead of population variance (n)?

What is the relationship between variance and standard deviation?

What is the standard error of the mean and how does it differ from standard deviation?

References