Zaymiey

📐 Mathematics  ·  Class 11  ·  JEE

Statistics - Practice Questions with Answers

68 free MCQs on Statistics with worked answers and explanations. Mean, median, mode, standard deviation, and data interpretation

Take the timed Statistics chapterwise test →

Below are 68 practice questions on Statistics, sorted Easy → Hard. Tap “Show answer & explanation” under any question to check yourself. Want the full theory first? Read the Statistics notes.

Normal Distribution: the 68-95-99.7 Rule68% within ±1σ95% within ±2σmean=median=mode (centre)

In a perfectly normal (bell-shaped) distribution, mean, median, and mode all coincide at the centre; about 68% of data falls within 1 standard deviation of the mean, 95% within 2, and 99.7% within 3 - the empirical rule used to judge how typical or extreme a value is.

Easy - 20 questions

Q1.

The average of a data set is called:

  • A Mode
  • B Median
  • C Mean
  • D Range
Show answer & explanation

Answer: C. Mean

Why: Mean (arithmetic average) = sum of all values / number of values.

Q2.

The middle value of an ordered data set is:

  • A Mean
  • B Median
  • C Mode
  • D Range
Show answer & explanation

Answer: B. Median

Why: The median is the middle value when data is arranged in order.

Q3.

The value that appears most often in a data set is:

  • A Mean
  • B Median
  • C Mode
  • D Standard deviation
Show answer & explanation

Answer: C. Mode

Why: Mode is the value(s) that occur most frequently in a data set.

Q4.

Find the mean of 5, 10, 15, 20, 25.

  • A 10
  • B 12
  • C 15
  • D 20
Show answer & explanation

Answer: C. 15

Why: Mean = (5+10+15+20+25)/5 = 75/5 = 15.

Q5.

Find the median of 3, 5, 7, 9, 11.

  • A 5
  • B 7
  • C 9
  • D 11
Show answer & explanation

Answer: B. 7

Why: For 5 values (odd count), median is the 3rd value in order = 7.

Q6.

Range of data 5, 8, 2, 11, 4 is:

  • A 6
  • B 7
  • C 8
  • D 9
Show answer & explanation

Answer: D. 9

Why: Range = maximum - minimum = 11 - 2 = 9.

Q7.

The mode of 4, 7, 7, 8, 9, 9, 9 is:

  • A 7
  • B 8
  • C 9
  • D 4
Show answer & explanation

Answer: C. 9

Why: 9 appears 3 times (most frequently), so mode = 9.

Q8.

Find the median of 2, 4, 6, 8 (even count).

  • A 4
  • B 5
  • C 6
  • D 7
Show answer & explanation

Answer: B. 5

Why: Even number of values: median = average of 2nd and 3rd = (4+6)/2 = 5.

Q9.

Data presented as a bar graph shows:

  • A Continuous data measured across a range
  • B Frequencies of different categories
  • C Change occurring gradually over time
  • D Probability values assigned to outcomes
Show answer & explanation

Answer: B. Frequencies of different categories

Why: Bar graphs display frequencies of categorical data using rectangular bars.

Q10.

What does a histogram show?

  • A Categorical data plotted as separate, non-adjacent bars
  • B Frequency distribution of continuous data in intervals
  • C A circular sector representation equivalent to a pie chart
  • D Only the averages of each class interval
Show answer & explanation

Answer: B. Frequency distribution of continuous data in intervals

Why: Histograms display frequency distributions of continuous data. Bars are joined (no gaps).

Q11.

If mean = 20 and there are 5 values, the sum is:

  • A 4
  • B 25
  • C 100
  • D 20
Show answer & explanation

Answer: C. 100

Why: Sum = mean × number of values = 20 × 5 = 100.

Q12.

Standard deviation measures:

  • A The central tendency of the data, like the mean does
  • B Spread (variability) of data around the mean
  • C The middle value of the data when arranged in order
  • D The single largest value present in the data set
Show answer & explanation

Answer: B. Spread (variability) of data around the mean

Why: Standard deviation measures how spread out values are from the mean. Higher SD = more spread.

Q13.

Which of these is a measure of central tendency?

  • A Range
  • B Variance
  • C Median
  • D Standard deviation
Show answer & explanation

Answer: C. Median

Why: Measures of central tendency (location): mean, median, mode. Others listed measure spread (dispersion).

Q14.

The cumulative frequency graph is also called:

  • A Histogram
  • B Frequency polygon
  • C Ogive
  • D Bar chart
Show answer & explanation

Answer: C. Ogive

Why: An ogive (cumulative frequency curve) plots cumulative frequency vs. class boundaries.

Q15.

Find the mean of 6, 6, 6, 6, 6.

  • A 0
  • B 6
  • C 30
  • D 5
Show answer & explanation

Answer: B. 6

Why: When all values are the same, mean = that value = 6.

Q16.

In a class of 10 students, the average score is 75. Total marks =

  • A 750
  • B 700
  • C 800
  • D 7500
Show answer & explanation

Answer: A. 750

Why: Total = mean × count = 75 × 10 = 750.

Q17.

Variance is the _____ of standard deviation.

  • A Square root
  • B Square
  • C Double
  • D Half
Show answer & explanation

Answer: B. Square

Why: Variance = (standard deviation)². Conversely, SD = square root of variance.

Q18.

Which graphical representation uses sectors of a circle?

  • A Histogram
  • B Bar graph
  • C Pie chart
  • D Line graph
Show answer & explanation

Answer: C. Pie chart

Why: A pie chart divides a circle into sectors proportional to the frequencies of categories.

Q19.

If median > mean, the distribution is:

  • A Symmetric
  • B Positively skewed
  • C Negatively skewed
  • D Normal
Show answer & explanation

Answer: C. Negatively skewed

Why: Negatively skewed (left skewed): long tail to the left, mode > median > mean.

Q20.

Mean deviation from mean of 10, 20, 30 is:

  • A 6.67
  • B 10
  • C 5
  • D 8
Show answer & explanation

Answer: A. 6.67

Why: Mean = 20. Deviations: |10-20|=10, |20-20|=0, |30-20|=10. Mean deviation = (10+0+10)/3 = 20/3 = 6.67.

Medium - 20 questions

Q21.

Coefficient of variation is used to:

  • A Find the mean value of a single data set
  • B Compare variability between datasets with different means
  • C Measure the correlation between two distinct variables
  • D Find the quartile boundaries of a frequency distribution
Show answer & explanation

Answer: B. Compare variability between datasets with different means

Why: CV = (SD/mean) × 100%. Allows comparison of variability between datasets with different scales.

Q22.

The median of grouped data uses the formula:

  • A l + (n/2 - F)/f × h
  • B l + (n - F)/f × h
  • C l + (f₁ - f₀)/(2f₁ - f₀ - f₂) × h
  • D sigma(fx)/sigma(f)
Show answer & explanation

Answer: A. l + (n/2 - F)/f × h

Why: Median (grouped) = l + [(n/2 - F)/f] × h, where l = lower limit, F = cumulative freq before median class, f = freq of median class, h = class width.

Q23.

In grouped data, the mode formula is:

  • A l + [(f₁-f₀)/(2f₁-f₀-f₂)] × h
  • B l + (n/2-F)/f × h
  • C Mean - Mode = 3(Mean-Median)
  • D sigma(fx)/n
Show answer & explanation

Answer: A. l + [(f₁-f₀)/(2f₁-f₀-f₂)] × h

Why: Mode (grouped) = l + [(f₁-f₀)/(2f₁-f₀-f₂)] × h, where f₁ = modal class freq, f₀ = prev class freq, f₂ = next class freq.

Q24.

Pearson's correlation coefficient r = 1 means:

  • A No correlation
  • B Perfect negative correlation
  • C Perfect positive correlation
  • D Some positive correlation
Show answer & explanation

Answer: C. Perfect positive correlation

Why: r = +1 indicates perfect positive linear correlation. As one variable increases, the other increases proportionally.

Q25.

IQR (Interquartile Range) = Q<sub>3</sub> - Q<sub>1</sub> measures:

  • A Total range from minimum to maximum
  • B Spread of middle 50% of data
  • C Average spread across the entire dataset
  • D Variance, the squared average deviation
Show answer & explanation

Answer: B. Spread of middle 50% of data

Why: IQR measures the range of the middle 50% of data, making it resistant to outliers.

Q26.

Standard deviation is preferred over mean deviation because:

  • A It is easier to calculate since it largely avoids using absolute values
  • B It gives more weight to extreme values and has algebraic tractability
  • C It tends to be smaller in value than the mean deviation in practice
  • D It equals the variance directly, making the two values interchangeable
Show answer & explanation

Answer: B. It gives more weight to extreme values and has algebraic tractability

Why: SD squares deviations (more sensitive to outliers), is algebraically tractable, and has favorable mathematical properties for further analysis.

Q27.

If all values in a dataset are multiplied by 2, the standard deviation:

  • A Stays the same
  • B Doubles
  • C Is halved
  • D Increases by 2
Show answer & explanation

Answer: B. Doubles

Why: SD scales with data. Multiplying all values by k multiplies SD by |k|. So SD doubles.

Q28.

The variance of 5, 5, 5, 5, 5 is:

  • A 5
  • B 1
  • C 0
  • D 25
Show answer & explanation

Answer: C. 0

Why: All values equal the mean (5). All deviations are zero. Variance = 0.

Q29.

Ogive is used to find:

  • A Mean
  • B Mode
  • C Median and quartiles
  • D Standard deviation
Show answer & explanation

Answer: C. Median and quartiles

Why: An ogive (cumulative frequency curve) is used to graphically determine median, quartiles, and percentiles.

Q30.

For data 3, 6, 9, 12, 15, the mean deviation from mean is:

  • A 2
  • B 3
  • C 4
  • D 5
Show answer & explanation

Answer: C. 4

Why: Mean = 9. Deviations: 6, 3, 0, 3, 6. Mean deviation = (6+3+0+3+6)/5 = 18/5 = 3.6. Closest to 4.

Q31.

If adding a constant k to each observation, mean:

  • A Remains the same
  • B Increases by k
  • C Multiplies by k
  • D Becomes 0
Show answer & explanation

Answer: B. Increases by k

Why: Mean shifts by the constant. New mean = original mean + k.

Q32.

The empirical relation between mean, median, and mode is:

  • A Mode = Mean + 2 Median
  • B Mean - Mode = 3(Mean - Median)
  • C Mode = 2 Mean - 3 Median
  • D Median = Mean - Mode
Show answer & explanation

Answer: B. Mean - Mode = 3(Mean - Median)

Why: Empirical formula: Mean - Mode = 3(Mean - Median). Or Mode = 3 Median - 2 Mean.

Q33.

For a negatively skewed distribution:

  • A Mean = Median = Mode
  • B Mean > Median > Mode
  • C Mean < Median < Mode
  • D Mode < Median < Mean
Show answer & explanation

Answer: C. Mean < Median < Mode

Why: Negative (left) skew: long tail to the left. Mean pulled left by extreme low values. Mean < Median < Mode.

Q34.

Step-deviation method for mean uses:

  • A Direct formula applying raw class marks without taking any shortcut
  • B Simplification by subtracting assumed mean and dividing by class width (h)
  • C Decomposition of the frequency expression into separate partial fractions
  • D Taking the logarithm of each class mark before averaging the result
Show answer & explanation

Answer: B. Simplification by subtracting assumed mean and dividing by class width (h)

Why: Step-deviation: ui = (xi - a)/h where a is assumed mean and h is class width. Mean = a + (sum(fiui)/sum(fi)) × h.

Q35.

Spearman's rank correlation is used when:

  • A Data is normally distributed with a known population variance value
  • B Data is on ordinal scale or ranks, or non-linear relationship suspected
  • C Data has roughly equal variance across the observed groups
  • D Used mainly for large samples exceeding thirty observations
Show answer & explanation

Answer: B. Data is on ordinal scale or ranks, or non-linear relationship suspected

Why: Spearman rank correlation: non-parametric. Used for ordinal data or when linear relationship assumption fails.

Q36.

The box-plot whiskers extend to:

  • A Mean plus or minus two standard deviations
  • B Minimum and maximum values (or 1.5×IQR from Q<sub>1</sub>/Q<sub>3</sub>)
  • C Exactly the first and third quartiles, Q<sub>1</sub> and Q<sub>3</sub>
  • D Mean plus or minus one standard deviation
Show answer & explanation

Answer: B. Minimum and maximum values (or 1.5×IQR from Q<sub>1</sub>/Q<sub>3</sub>)

Why: Box-plot whiskers: conventionally extend to min/max values within 1.5 × IQR from Q<sub>1</sub> and Q<sub>3</sub>. Points beyond are outliers.

Q37.

For data with high variability, standard deviation is:

  • A Low
  • B Zero
  • C High
  • D Equal to mean
Show answer & explanation

Answer: C. High

Why: High variability = data spread widely from the mean = high standard deviation.

Q38.

The 50th percentile equals the:

  • A Mean
  • B Mode
  • C Median
  • D First quartile
Show answer & explanation

Answer: C. Median

Why: The 50th percentile = Q2 = median. It divides the data into equal halves.

Q39.

In a frequency distribution, class mark (midpoint) of interval 30-40 is:

  • A 30
  • B 35
  • C 40
  • D 70
Show answer & explanation

Answer: B. 35

Why: Class mark = (lower limit + upper limit)/2 = (30+40)/2 = 35.

Q40.

Mean of 10 numbers is 15. If one number 20 is replaced by 30, new mean is:

  • A 15
  • B 16
  • C 17
  • D 14
Show answer & explanation

Answer: B. 16

Why: Sum = 150. Remove 20, add 30: new sum = 160. New mean = 160/10 = 16.

Hard - 28 questions

Q41.

For grouped data, Pearson's first coefficient of skewness is:

  • A (Mean - Mode)/SD
  • B (Mean - Median)/SD
  • C 3(Mean - Median)/SD
  • D (Median - Mode)/SD
Show answer & explanation

Answer: C. 3(Mean - Median)/SD

Why: Pearson's first skewness = (Mean - Mode)/SD. For grouped data the mode is ill-defined, so the second formula 3(Mean - Median)/SD is used instead via the empirical relation Mode ≈ 3Median - 2Mean. Answer: 3(Mean - Median)/SD.

Q42.

The regression line y on x passes through:

  • A The origin point specifically, regardless of the data
  • B (x̄, ȳ) - the means
  • C All of the individual data points at once
  • D Just the first and last data point in the set
Show answer & explanation

Answer: B. (x̄, ȳ) - the means

Why: The regression line y on x is ȳ = b·x̄ + a, derived by minimising Σ(yᵢ - ŷᵢ)². Setting partial derivatives to zero gives the normal equations, whose solution always passes through the point of means (x̄, ȳ). Both regression lines share this point.

Q43.

If byx × bxy = r², where byx and bxy are regression coefficients, then r =

  • A byx × bxy
  • B sqrt(byx × bxy)
  • C (byx + bxy)/2
  • D byx/bxy
Show answer & explanation

Answer: B. sqrt(byx × bxy)

Why: By definition byx = r·(σy/σx) and bxy = r·(σx/σy). Their product: byx × bxy = r²·(σy/σx)·(σx/σy) = r². Therefore r = ±√(byx × bxy), with sign matching that of both coefficients.

Q44.

For a normal distribution N(μ, σ²), about 95% of data lies within:

  • A μ ± σ
  • B μ ± 2σ
  • C μ ± 3σ
  • D μ ± 1.5σ
Show answer & explanation

Answer: B. μ ± 2σ

Why: Empirical 68-95-99.7 rule: P(μ−σ < X < μ+σ) ≈ 68%, P(μ−2σ < X < μ+2σ) ≈ 95%, P(μ−3σ < X < μ+3σ) ≈ 99.7%. Standard normal Z-table gives P(−1.96 < Z < 1.96) = 95%, confirming μ ± 2σ.

Q45.

Partial correlation measures:

  • A Correlation removing effect of a third variable
  • B The simple two-variable correlation with no adjustment applied
  • C The standardized coefficient from a multiple regression model
  • D The cross-correlation between two time-shifted signals
Show answer & explanation

Answer: A. Correlation removing effect of a third variable

Why: Partial correlation: measures linear relationship between two variables while controlling for one or more other variables.

Q46.

The variance of combined data of two groups (n₁, x̄₁, σ₁²) and (n₂, x̄₂, σ₂²) is:

  • A (n₁σ₁²+n₂σ₂²)/(n₁+n₂), which overlooks the shift in means between the two groups
  • B A more complex formula involving d₁² and d₂² (distances from combined mean)
  • C (σ₁²+σ₂²)/2, a simple unweighted average that ignores group sizes
  • D σ₁ × σ₂, the simple product of the two individual standard deviations
Show answer & explanation

Answer: B. A more complex formula involving d₁² and d₂² (distances from combined mean)

Why: Combined mean x̄ = (n₁x̄₁+n₂x̄₂)/(n₁+n₂). Let d₁ = x̄₁−x̄ and d₂ = x̄₂−x̄. Combined variance = [n₁(σ₁²+d₁²) + n₂(σ₂²+d₂²)]/(n₁+n₂). The d² terms capture the shift between group means and the combined mean.

Q47.

Moment generating function M(t) of a random variable X is defined as:

  • A E[X]
  • B E[X²]
  • C E[e<sup>tX</sup>]
  • D E[tX]
Show answer & explanation

Answer: C. E[e<sup>tX</sup>]

Why: MGF: M(t) = E[e<sup>tX</sup>]. Expanding e<sup>tX</sup> = 1 + tX + t²X²/2! + … gives M(t) = 1 + tE[X] + t²E[X²]/2! + …. The n-th derivative M⁽ⁿ⁾(0) = E[Xⁿ], the n-th raw moment. Answer: E[e<sup>tX</sup>].

Q48.

Stratified random sampling ensures:

  • A Every individual in the population has an exactly equal chance
  • B Proportional representation from each subgroup (stratum)
  • C Only random selection, with strata playing no role at all
  • D Equal sample size drawn from every group, regardless of its size
Show answer & explanation

Answer: B. Proportional representation from each subgroup (stratum)

Why: Stratified sampling: population divided into strata, samples taken from each stratum proportionally or as needed. Ensures representation.

Q49.

The harmonic mean of 2, 4, 8 is approximately:

  • A 3.43
  • B 4.00
  • C 2.18
  • D 5.33
Show answer & explanation

Answer: A. 3.43

Why: Harmonic mean HM = n / Σ(1/xᵢ). Here: 1/2 + 1/4 + 1/8 = 4/8 + 2/8 + 1/8 = 7/8. HM = 3 ÷ (7/8) = 3 × 8/7 = 24/7 ≈ 3.43. Useful for rates/speeds where equal time (not equal distance) is spent.

Q50.

Laspeyres price index uses base year quantities. It tends to:

  • A Underestimate inflation
  • B Overestimate inflation
  • C Give exact inflation
  • D Equal Paasche index always
Show answer & explanation

Answer: B. Overestimate inflation

Why: Laspeyres (base-year weights) tends to overestimate inflation because it ignores substitution away from expensive goods. Paasche (current-year weights) tends to underestimate.

Q51.

In regression analysis, R² (coefficient of determination) measures:

  • A The correlation coefficient itself, without squaring it
  • B Proportion of variance in Y explained by X
  • C The slope of the fitted regression line
  • D The raw sum of squared residuals from the regression
Show answer & explanation

Answer: B. Proportion of variance in Y explained by X

Why: SS_total = Σ(yᵢ−ȳ)², SS_residual = Σ(yᵢ−ŷᵢ)². R² = 1 − SS_residual/SS_total = SS_regression/SS_total. R² ∈ [0,1]; a value of 0.85 means 85% of variance in Y is explained by the regression model.

Q52.

Tchebychev (Chebyshev) inequality states that for any distribution, P(|X - μ| ≥ kσ) ≤:

  • A 1/k
  • B 1/k²
  • C
  • D 1/(2k)
Show answer & explanation

Answer: B. 1/k²

Why: Chebyshev's inequality (no normality needed): P(|X−μ| ≥ kσ) ≤ 1/k² for any k > 1. Proof uses Markov's inequality on (X−μ)². Equivalently, at least 1−1/k² of data lies within k standard deviations of the mean.

Q53.

The standard error of mean (SEM) for sample size n from population with SD σ is:

  • A σ/n
  • B σ × n
  • C σ/√n
  • D σ²/n
Show answer & explanation

Answer: C. σ/√n

Why: For n iid observations each with variance σ², Var(X̄) = Var(ΣXᵢ/n) = nσ²/n² = σ²/n. So SEM = SD(X̄) = σ/√n. Doubling n halves the SEM, reflecting improved precision of the sample mean estimate.

Q54.

The mean of the first five natural numbers is:

  • A 3
  • B 5
  • C 2.5
  • D 15
Show answer & explanation

Answer: A. 3

Why: Mean = (1 + 2 + 3 + 4 + 5)/5 = 15/5 = 3.

Q55.

The median of the data 2, 4, 6, 8, 10 is:

  • A 6
  • B 5
  • C 8
  • D 4
Show answer & explanation

Answer: A. 6

Why: For five ordered values, the median is the third (middle) one, which is 6.

Q56.

The mode of the data 2, 3, 3, 4, 5 is:

  • A 3
  • B 2
  • C 4
  • D 5
Show answer & explanation

Answer: A. 3

Why: The mode is the most frequently occurring value; 3 appears twice, more than any other.

Q57.

The range of the data 5, 8, 12, 20 is:

  • A 15
  • B 12
  • C 20
  • D 8
Show answer & explanation

Answer: A. 15

Why: Range = largest − smallest = 20 − 5 = 15.

Q58.

If the mean of five observations is 10, their total sum is:

  • A 50
  • B 10
  • C 15
  • D 2
Show answer & explanation

Answer: A. 50

Why: Sum = mean × number of observations = 10 × 5 = 50.

Q59.

The standard deviation of a data set is the ___ of its variance:

  • A the square root
  • B the square
  • C the reciprocal
  • D double the value
Show answer & explanation

Answer: A. the square root

Why: Standard deviation = √(variance).

Q60.

The variance of a set of data is always:

  • A non-negative
  • B negative
  • C exactly zero
  • D below the mean
Show answer & explanation

Answer: A. non-negative

Why: Variance is a mean of squared deviations, so it can never be negative.

Q61.

The variance of the first 10 natural numbers is:

  • A 8.25
  • B 9.5
  • C 10
  • D 7.5
Show answer & explanation

Answer: A. 8.25

Why: Variance of first n natural numbers = (n² − 1)/12 = 99/12 = 8.25.

Q62.

If every observation in a data set is increased by 5, the variance:

  • A increases by 5
  • B increases by 25
  • C remains unchanged
  • D increases by 10
Show answer & explanation

Answer: C. remains unchanged

Why: Adding a constant shifts the data without changing spread, so the variance is unchanged.

Q63.

The mean of 100 observations is 50. Later one observation 100 was found to be misread as 10. The corrected mean is:

  • A 50.9
  • B 49.1
  • C 59
  • D 50.09
Show answer & explanation

Answer: A. 50.9

Why: Corrected total increases by 90, so the mean rises by 90/100 = 0.9 to 50.9.

Q64.

The standard deviation of the data 2, 4, 6, 8, 10 is:

  • A 2
  • B 2√2
  • C 3
  • D 4
Show answer & explanation

Answer: B. 2√2

Why: Mean = 6; squared deviations 16, 4, 0, 4, 16 sum to 40; variance 8; SD = √8 = 2√2.

Q65.

If the standard deviation of x₁, x₂, ..., xₙ is σ, the standard deviation of 3x₁, 3x₂, ..., 3xₙ is:

  • A σ
  • B
  • C
  • D σ/3
Show answer & explanation

Answer: B. 3σ

Why: Scaling each observation by 3 scales the standard deviation by |3|, giving 3σ.

Q66.

For a data set with standard deviation 5 and mean 25, the coefficient of variation is:

  • A 5%
  • B 20%
  • C 25%
  • D 0.2%
Show answer & explanation

Answer: B. 20%

Why: CV = (SD/mean)·100 = (5/25)·100 = 20%.

Q67.

For a moderately skewed distribution with mean 10 and median 8, the mode is:

  • A 2
  • B 4
  • C 6
  • D 12
Show answer & explanation

Answer: B. 4

Why: Using mode = 3·median − 2·mean = 24 − 20 = 4.

Q68.

Two groups have sizes 20 and 30 with means 30 and 20 respectively. Their combined mean is:

  • A 25
  • B 24
  • C 26
  • D 22
Show answer & explanation

Answer: B. 24

Why: Combined mean = (20·30 + 30·20)/50 = 1200/50 = 24.