Question 1

What does R-squared tell you about the regression?

Accepted Answer

R-squared (coefficient of determination) measures the proportion of variance in the dependent variable that is explained by the independent variable(s). An R-squared of 0.85 means 85 percent of the variability in y is explained by the linear relationship with x, while 15 percent remains unexplained. R-squared ranges from 0 to 1, with higher values indicating better fit. However, R-squared always increases when more variables are added, even if they are irrelevant, which is why adjusted R-squared penalizes for unnecessary variables. A high R-squared does not prove causation and does not guarantee the model is appropriate; always examine residual plots for patterns that suggest nonlinearity or outlier influence.

Question 2

What is the standard error of the regression?

Accepted Answer

The standard error of the regression (also called residual standard error or root mean square error) measures the typical size of prediction errors. It is calculated as the square root of the sum of squared residuals divided by the degrees of freedom (n minus 2 for simple linear regression). A standard error of 0.5 means predictions are typically within about 0.5 units of actual values. Smaller standard errors indicate more precise predictions. The standard error is in the same units as the dependent variable, making it directly interpretable. It is used to construct confidence intervals for predictions and to calculate t-statistics for testing whether the slope and intercept are significantly different from zero.

Question 3

What assumptions must be met for valid linear regression?

Accepted Answer

Linear regression requires several assumptions for its statistical tests and confidence intervals to be valid. First, linearity: the true relationship between x and y is linear. Second, independence: observations are independent of each other (no autocorrelation). Third, homoscedasticity: the variance of residuals is constant across all x values. Fourth, normality: residuals are approximately normally distributed (most important for small samples). Fifth, no perfect multicollinearity in multiple regression. Violations of these assumptions can lead to biased estimates, incorrect standard errors, and misleading p-values. Diagnostic plots including residual plots, Q-Q plots, and leverage plots help assess whether assumptions are reasonably satisfied.

Question 4

When should you not use linear regression?

Accepted Answer

Linear regression is inappropriate in several situations. When the relationship is clearly nonlinear (curved scatter plot), polynomial or nonlinear regression may be needed. When the dependent variable is categorical (yes/no), logistic regression is appropriate instead. When data contains severe outliers, robust regression methods should be considered. When observations are correlated over time, time series models with autocorrelation terms are needed. When heteroscedasticity is present, weighted least squares or generalized least squares may be preferable. When multiple predictors are highly correlated (multicollinearity), regularization methods like ridge or lasso regression can help. Always plot the data first and examine residuals after fitting.

Regression Equation Calculator

Formula

Worked Examples

Example 1: Sales Forecasting from Advertising Spend

Example 2: Temperature and Energy Consumption

Frequently Asked Questions

What does R-squared tell you about the regression?

What is the standard error of the regression?

What assumptions must be met for valid linear regression?

When should you not use linear regression?

References