In addition to measures of central tendency, descriptive statistics provide tools for describing the data distribution. Measures of variability include variance, standard deviation (SD), and range, each of which provides a different description of the data (Bhandari, 2021). First, the range is the difference between the maximum and minimum values in a set — the larger this measure, the greater the distance between the marginal values. Second, variance is a measure of the scatter of the data relative to the mean: the higher the variance, the more they deviate from the mean and from each other. Finally, SD defines the measure by which each value in the distribution deviates from the mean on average — SD is higher for data with greater variance (NIH, 2020).
Measures of variability are essential for descriptive analysis, but their use also involves some risks. Variance, SD, and range provide sufficient information to describe the spread in the data distribution and are commonly used to assess it: the higher the measures, the greater the spread. On the other hand, these measures are sensitive to outliers, which means that extreme values negatively affect calculation quality (Zhu, 2023).
In the distribution (10, 20, 30, 40, 50, 1000), the correct solution for analysis would be to remove the outlier (x = 1000).The measures would be s2 = 250.0, sd = 15.8, range = 40. If the outlier was not removed, the measures of variability change dramatically: s2 = 157,016.7, sd = 396.3, range = 990. The range is not a sufficiently informative measure since only the maximum and minimum points are used, and the nature of the data in the middle is not explored. For a distribution (0, 1, 1, 1, 1, 2, 2, 2, 2, 500), the range would be 500, and this result would say nothing about what patterns exist between the edge points.
The SD is the most helpful characteristic because it provides more accurate information about the data’s spread, uses the same units of measure, and has a smoother effect on outliers due to the square root. SD also relates better to measures of central tendency (e.g., the mean) because it has the same units of measure and shows the data’s spread around the mean. In general, if there are many modes in the distribution, the measures of spread (except range) will be lower.
References
Bhandari, S. (2023). Variability | Calculating range, IQR, variance, standard deviation. Scribbr.
NIH. (2020). Common terms and equations. NIH.
Zhu, A. (2023). How do outliers affect statistical inference? Medium.