Year11-MATH-2-1-6 Univariate Data Analysis

In Year 11 Mathematics, Univariate Data Analysis involves exploring and describing characteristics of one variable at a time, using tools like frequency tables, histograms, box plots, measures of center (mean, median, mode), and spread (range, interquartile range, standard deviation) to understand its distribution, shape, location, and identify patterns or outliers, forming the foundation before looking at relationships between variables.

Key Concepts & Techniques:

  • What it is: Analyzing single data sets (e.g., heights of students, number of siblings, test scores) to understand patterns within that variable.
  • Data Types: Working with both categorical (e.g., favourite colour, gender) and numerical (discrete/continuous, e.g., age, height) data.
  • Displaying Data:
    • Frequency Tables: Organizing counts and percentages of data values or grouped intervals.
    • Histograms & Column Charts: Visualizing frequency distributions for numerical and categorical data, respectively.
    • Stem-and-Leaf Plots: Showing individual data points while preserving order, useful for seeing distribution and outliers.
    • Box Plots (Box-and-Whisker Plots): Displaying the five-number summary (min, Q1, median, Q3, max) and highlighting outliers.
  • Describing Data:
    • Measures of Centre: Mean (average), Median (middle value), Mode (most frequent).
    • Measures of Spread: Range, Interquartile Range (IQR), Variance, Standard Deviation (how spread out the data is).
  • The Statistical Investigation Process: Using univariate analysis as a fundamental step to describe data before moving to more complex bivariate (two-variable) analysis, helping to form research questions and understand data quality. 

Why it’s important for Year 11:

It builds foundational skills in data interpretation, allowing students to:

  • Understand the basic properties of a data set.
  • Identify unusual data points (outliers).
  • Prepare data for more advanced analyses, like correlation or regression (bivariate data). 

Frequency tables in 11th-grade mathematics are  univariate data analysis to organize and summarize data collected on a single variable. They display the frequency (count) of each value or range of values within a dataset, providing a clear picture of how often different outcomes occur. 

Key concepts and components include:

  • Frequency: The number of times a particular value or category appears in a dataset.
  • Categories or Classes: Data is grouped into specific categories or intervals (called class intervals, especially for continuous data).
  • Tally Marks (Optional): A preliminary method for counting occurrences before converting them into numerical frequencies.
  • Relative Frequency: The proportion or percentage of times a value occurs, calculated by dividing the frequency of a category by the total number of data points.
  • Cumulative Frequency: The running total of frequencies, showing the number of data points up to a particular value or class interval.
  • Data Visualization: Frequency tables are foundational for creating visual representations like histograms, bar charts, and frequency polygons, which help interpret data distributions. 


Histograms

  • Data Type: Histograms are used to display the frequency distribution of continuous numerical data (quantitative data), such as height, weight, age, or test scores.
  • Appearance: The columns (bars) in a histogram are drawn adjacent to each other, with no gaps between them, to emphasize that the data flows along a continuous range. The horizontal axis (x-axis) is a continuous number line divided into intervals, known as “bins” or “classes”.
  • Purpose: They help visualize the shape, center, and spread of the data distribution, allowing for the identification of patterns like skewness, symmetry, or the presence of outliers or multiple peaks (bimodal distribution). 

Column Charts (Bar Graphs)

  • Data Type: Column charts (or bar graphs) are used to represent categorical or discrete data, where each bar represents a distinct, separate category (qualitative data), such as favorite colors, types of fruit, or different store locations.
  • Appearance: The bars in a column chart have gaps between them, which reinforces the idea that the categories are independent and discrete, with no numerical relationship between them. The order of the bars can often be rearranged (e.g., alphabetically or by size) without affecting the meaning of the data.
  • Purpose: They are primarily used for comparing values across different categories to see which category is the largest or smallest. 

Measures of Centre

Measures of centre, in 11th grade univariate data analysis, are  central position or typical value of a dataset. These measures help in summarizing the data and identifying where most values lie. 

The primary measures of centre taught at this level are the:

  • Mean: The arithmetic average, calculated by summing all the values in a dataset and dividing by the number of values. It is sensitive to extreme values (outliers).
  • Median: The middle value of a dataset when it is ordered from least to greatest. If there is an even number of values, the median is the average of the two middle numbers. It is less affected by outliers than the mean.
  • Mode: The value that appears most frequently in a dataset. A dataset may have one mode (unimodal), more than one mode (multimodal), or no mode at all if all values appear with the same frequency. 

Understanding these measures is crucial for interpreting the distribution and characteristics of statistical data.

Measures of spread (or dispersion)  .

In 11th-grade mathematics, the key measures of spread typically covered include : 

  • Range: The simplest measure, calculated as the difference between the maximum and minimum values in a dataset.
  • Interquartile Range (IQR): The difference between the upper quartile (Q3) and the lower quartile (Q1), representing the spread of the middle 50% of the data.
  • Variance: A measure of the average squared deviation of each data point from the mean, indicating how far each number is from the mean.
  • Standard Deviation (SD): The square root of the variance, providing a measure of spread in the same units as the original data, making it easier to interpret. 

Standard deviation is a statistical measure used in univariate data analysis to quantify the amount of variation or dispersion within a set of data points . In 11th-grade mathematics, it is typically taught as a way to understand how spread out the numbers in a dataset are from their average (mean) . 

Key Concepts

  • Measure of Spread: Standard deviation is a key indicator of the data’s distribution. A low standard deviation indicates that most values are clustered closely around the mean, while a high standard deviation suggests the values are widely spread out .
  • Relationship to the Mean: It provides context for the mean. Knowing the average height of students in a class is more meaningful if you also know the standard deviation; it tells you if most students are close to that average height or if there are many very tall and very short students .
  • Square Root of Variance: Mathematically, the standard deviation is defined as the square root of the variance. Variance measures the average squared difference from the mean, and taking the square root brings the measure back to the original units of the data . 

Calculation (Simplified Steps)

The calculation of the standard deviation involves several steps :

  1. Find the Mean: Calculate the average of the dataset.
  2. Calculate Deviations: Subtract the mean from each data point.
  3. Square the Deviations: Square each result from step 2 to eliminate negative numbers.
  4. Find the Variance: Calculate the mean of these squared deviations.
  5. Take the Square Root: The square root of the variance is the standard deviation

In 11th grade, students often learn both how to perform this calculation manually and how to use a calculator’s built-in statistical functions to find it efficiently.

Standard deviation is an indicator of data variability. It is calculated by squaring the difference between each data point and the mean, summing the sum, dividing by the number of data points (variance), and then taking the square root. Using the example of class heights, it quantifies how far each student’s height varies from the mean (variance). A larger variance indicates more tall and short students, while a smaller variance indicates that everyone is closer to the mean.

How to Calculate Standard Deviation (Example: Class A, Class B)
[Example Data]
Class A (average 170cm): 160cm, 170cm, 180cm
Class B (average 170cm): 165cm, 170cm, 175cm
Note: Both classes have a mean of 170cm, but the variance is different.

Calculate the difference (deviation) from the mean.
Class A: (160-170=-10), (170-170=0), (180-170=10)
Class B: (165-170=-5), (170-170=0), (175-170=5)

Squaring the deviations.
Class A: (10)2=100,02=0,102=100(-10)^2 =100, 0^2=0, 10^2 = 100
Class B: (5)2=25,02=0,52=25(-5)^2=25, 0^2=0, 5^2=25

Calculate the sum of the squared deviations and divide by the number of data points (3) to calculate the variance.
Class A: (100+0+100) ÷ 3 = 200 ÷ 3 = 66.67 (variance).
Class B: (25+0+25) ÷ 3 = 50 ÷ 3 = 16.67 (Variance)

Calculate the square root of the variance (standard deviation)
Class A: 66.67\sqrt{66.67}≈ 8.16 (standard deviation)
Class B: 16.67\sqrt{16.67}≈ 4.08 (standard deviation)

What we can learn from the results
Class A has a larger standard deviation than Class B (approximately 8.16 vs. approximately 4.08), indicating greater variation in height (more tall and short children).

Standard Deviation and Deviation Score
Standard deviation is the basis for calculating deviation scores. A deviation score of 60 is the mean + one standard deviation, and a deviation score of 40 is the mean – one standard deviation.
In this way, standard deviation allows us to specifically grasp the spread (dispersion) of data, which cannot be determined by the mean alone.

To calculate a deviation value from standard deviation (SD), use the formula “(individual score – average score) ÷ standard deviation x 10 + 50.” This formula calculates the standard deviation value by calculating how many standard deviations your score is away from the average score (50), multiplying that by 10, and adding 50. The smaller the dispersion (standard deviation) in the distribution of test scores, the more a small difference from the average score will be reflected in the deviation value.

Deviation Value Calculation Formula
Deviation Value = (your score – average score) ÷ standard deviation × 10 + 50

Calculation Procedure and Example
Check your score, average score, and standard deviation.
Calculate the difference (deviation) between your score and the average score.
Divide this difference by the standard deviation to find the standardized value (normalized variate) (how many standard deviations your score is from the mean).
Multiply this value by 10 and add 50.

Example:
Your score: 80
Average score: 60
Standard deviation: 20
Deviation: 80 – 60 = 20
Standardization: 20 ÷ 20 = 1
Standard deviation: 1 × 10 + 50 = 60
In this example, the standard deviation is 60. Since your score is 20 points higher (one standard deviation higher) than the average score (60), the standard deviation is 50 + 10 = 60.

********************************************************************************************************************************************************

11年生 数学 2-1-6 科目2:一般 ユニット2 – 6章:単変量データ分析

11年生の数学では、単変量データ分析では、度数分布表、ヒストグラム、箱ひげ図、中心値(平均値、中央値、最頻値)、広がり(範囲、四分位範囲、標準偏差)などのツールを用いて、一度に1つの変数の特性を探索・記述し、その分布、形状、位置を理解し、パターンや外れ値を特定します。これは、変数間の関係性を考察する前の基礎となります。

主要な概念と手法:

概要:単一のデータセット(例:生徒の身長、兄弟姉妹の数、テストの点数)を分析し、その変数内のパターンを理解すること。

データの種類:カテゴリデータ(例:好きな色、性別)と数値データ(離散・連続、例:年齢、身長)の両方を扱います。

データの表示:

度数分布表:データ値またはグループ化された区間の個数と割合を整理します。

ヒストグラムと縦棒グラフ:数値データとカテゴリデータの度数分布をそれぞれ視覚化します。

幹葉図:順序を維持しながら個々のデータポイントを表示します。分布や外れ値を確認するのに役立ちます。

箱ひげ図:5つの数値の要約(最小値、第1四分位範囲、中央値、第3四分位範囲、最大値)を表示し、外れ値を強調表示します。

データの説明:

中心の尺度:平均値(平均)、中央値(中央値)、最頻値(最頻値)

広がりの尺度:範囲、四分位範囲(IQR)、分散、標準偏差(データの広がり具合)

統計調査プロセス:より複雑な二変量解析に進む前に、単変量解析をデータ記述の基本ステップとして用い、研究課題の設定とデータ品質の理解を支援します。

11年生にとって重要な理由:

データ解釈の基礎スキルを構築し、生徒は以下のことが可能になります。

データセットの基本的な特性を理解する。

異常なデータポイント(外れ値)を特定する。

相関分析や回帰分析(二変量データ)などのより高度な分析のためにデータを準備する。

>>度数分布表 :

単変量データ分析において、単一の変数について収集されたデータを整理・要約するために使用されるツールです。データセット内の各値または値の範囲の頻度(カウント)を表示し、異なる結果がどの程度頻繁に発生するかを明確に示します。

主要な概念と構成要素は以下のとおりです。

頻度:データセット内で特定の値またはカテゴリが出現する回数。

カテゴリまたはクラス:データは特定のカテゴリまたは間隔(特に連続データの場合はクラス間隔と呼ばれます)にグループ化されます。

集計マーク(オプション):数値頻度に変換する前に、出現回数を数えるための予備的な方法。

相対頻度:値が出現する割合またはパーセンテージ。カテゴリの頻度をデータポイントの総数で割ることによって算出されます。

累積度数:度数の累計。特定の値または階級区間までのデータポイントの数を示します。

データの視覚化:度数表は、ヒストグラム、棒グラフ、度数ポリゴンなどの視覚的表現を作成するための基礎であり、データ分布の解釈に役立ちます。

>> ヒストグラム

データの種類:ヒストグラムは、身長、体重、年齢、テストの点数など、連続した数値データ(定量データ)の頻度分布を表すために使用されます。

外観:ヒストグラムの縦棒(バー)は、データが連続した範囲に沿って流れていることを強調するために、互いに隙間なく隣接して描画されます。横軸(X軸)は、「ビン」または「クラス」と呼ばれる区間に分割された連続した数値線です。

目的:ヒストグラムは、データ分布の形状、中心、広がりを視覚化するのに役立ち、歪度、対称性、外れ値や複数のピーク(二峰性分布)の存在などのパターンを識別するのに役立ちます。

>> 縦棒グラフ(棒グラフ)

データの種類:縦棒グラフ(棒グラフ)は、カテゴリデータまたは離散データを表すために使用されます。各バーは、好きな色、果物の種類、店舗の所在地など、明確に区別されたカテゴリ(定性データ)を表します。

外観:縦棒グラフの棒グラフには間隔があり、各カテゴリが独立して離散的であり、数値的な関係がないという印象を与えます。棒グラフの順序は、データの意味に影響を与えることなく、多くの場合、並べ替えることができます(例:アルファベット順、サイズ順など)。

目的:縦棒グラフは主に、異なるカテゴリ間で値を比較し、どのカテゴリが最大または最小であるかを確認するために使用されます。

>> 中心尺度

11年生の単変量データ分析における中心尺度とは、データセットの中心位置または典型的な値を表す単一の値です。これらの尺度は、データを要約し、ほとんどの値がどこに位置しているかを特定するのに役立ちます。

このレベルで教えられる主要な中心尺度は次のとおりです。

平均値:データセット内のすべての値を合計し、値の数で割ることによって算出される算術平均。極端な値(外れ値)の影響を受けやすいです。

中央値:データセットを最小から最大の順に並べたときの中央の値。値が偶数個の場合、中央値は中央の2つの数値の平均です。平均値よりも外れ値の影響を受けにくいです。

最頻値:データセットで最も頻繁に出現する値。データセットには、1つの最頻値(単峰性)、複数の最頻値(多峰性)、またはすべての値が同じ頻度で出現する場合は最頻値が存在しない場合があります。

これらの尺度を理解することは、統計データの分布と特性を解釈する上で非常に重要です。

広がり(または散布度)の尺度

単変量データセット内のデータ点が中心付近でどのように分布しているか、またはどのようにばらついているかを表します。11年生の数学で一般的に扱われる広がりの主要な尺度には、以下のものがあります 。

範囲:データセット内の最大値と最小値の差として計算される最も単純な尺度 。

四分位範囲IQR):上位四分位(Q3)と下位四分位(Q1)の差で、データの中央50%の広がりを表します 。

分散:各データ点の平均からの偏差の二乗平均値で、各数値が平均からどれだけ離れているかを示します 。

標準偏差:分散の平方根で、元のデータと同じ単位で広がりの尺度を提供し、解釈を容易にします 。

標準偏差は、単変量データ分析において、データポイント集合内の変動または分散の大きさを定量化するために使用される統計的尺度です。高校2年生(11年生)の数学では、通常、データセット内の数値が平均値からどの程度広がっているかを理解する方法として教えられます。

主要概念

広がりの尺度:標準偏差は、データの分布を示す重要な指標です。標準偏差が低い場合、ほとんどの値が平均値の周辺に密集していることを示し、標準偏差が高い場合、値が広く広がっていることを示します。

平均との関係:平均の背景を提供します。クラスの生徒の平均身長を知ることは、標準偏差も知っていればより意味を持ちます。標準偏差を知ることで、ほとんどの生徒が平均身長に近いのか、それとも非常に背の高い生徒と非常に背の低い生徒が多いのかがわかります。

分散の平方根:数学的には、標準偏差は分散の平方根として定義されます。分散は平均値からの差の二乗平均値を測るもので、平方根を取ることでデータの元の単位に戻ります。

計算(簡略化された手順)

標準偏差の計算には、いくつかの手順があります。

1. 平均値を求める:データセットの平均を計算します。

2. 偏差を計算する:各データポイントから平均値を引きます。

3. 偏差を二乗する:手順2の結果をそれぞれ二乗して、負の数を消去します。

4. 分散を求める:これらの二乗偏差の平均を計算します。

5. 平方根を取る:分散の平方根が標準偏差です。

11年生では、生徒はこの計算を手作業で行う方法と、電卓に組み込まれている統計関数を使って効率的に計算する方法の両方を学ぶことがよくあります。

計算例

標準偏差は、データのばらつきを示す指標で、平均値からの各データの差を二乗して合計し、データの数で割り(分散)、その平方根を取ることで計算されます。クラスの身長で例えるなら、平均身長から個々の生徒の身長がどれくらい離れているか(ばらつき具合)を数値化するもので、分散が大きいほど背が高い子・低い子が多く、小さいほど皆が平均身長に近いことを意味します。 

標準偏差の計算手順(例:Aクラス、Bクラス)

  1. 【データ例】
    • Aクラス (平均170cm): 160cm, 170cm, 180cm
    • Bクラス (平均170cm): 165cm, 170cm, 175cm
    • : どちらも平均は170cmですが、ばらつきが違います。
  2. 平均からの差(偏差)を求める
    • Aクラス: (160-170=-10), (170-170=0), (180-170=10)
    • Bクラス: (165-170=-5), (170-170=0), (175-170=5)
  3. 偏差を二乗する
  4. Aクラス: (10)2(-10)^2=100, 020^2=0, 10210^2=100
  5. Bクラス: (5)2(-5)^2=25, 020^2=0, 525^2=25
  1. 二乗した偏差の合計を求め、データの数(3)で割る(分散を求める)
    • Aクラス: (100+0+100) ÷ 3 = 200 ÷ 3 = 66.67 (分散)
    • Bクラス: (25+0+25) ÷ 3 = 50 ÷ 3 = 16.67 (分散)
  2. 分散の平方根を求める(標準偏差)
  3. Aクラス: 66.67\sqrt{66.67}≈ 8.16 (標準偏差) 
  4. Bクラス:  16.67\sqrt{16.67}≈ 4.08 (標準偏差)

結果からわかること

  • Aクラスの方がBクラスよりも標準偏差が大きい(約8.16 vs 4.08ため、身長のばらつきが大きい(背の高い子・低い子が多い)と言えます。 

標準偏差と偏差値

  • 標準偏差は偏差値計算の基礎です。偏差値60は平均+標準偏差1つ分、偏差値40は平均-標準偏差1つ分という関係になります。 

このように、標準偏差を使うと平均値だけでは分からない、データの広がり具合(散らばり具合)を具体的に把握できるのです。 

偏差値の計算式

  • 偏差値 = (自分の得点 平均点) ÷ 標準偏差 × 10 + 50 

計算手順と例

  1. 自分の得点、平均点、標準偏差を確認する
  2. 得点と平均点の差(偏差)を計算する
  3. その差を標準偏差で割り、標準化した値(基準化変量)を求める(平均から標準偏差の何倍離れているか)。
  4. その値に10を掛け、50を足す。 

:

  • 自分の得点: 80点
  • 平均点: 60点
  • 標準偏差: 20点
  1. 偏差: 80 – 60 = 20点
  2. 標準化: 20 ÷ 20 = 1
  3. 偏差値: 1 × 10 + 50 = 60 

この例では、偏差値は60となります。平均点(60点)より20点高い(標準偏差1個分高い)ため、偏差値は50+10=60になるという計算です。 


Posted

in

by

Tags: