0% found this document useful (0 votes)

25 views26 pages

Basic Statistical Descriptions of Data

The document discusses basic statistical descriptions of data, focusing on central tendency, dispersion, and graphical representations. It covers key concepts such as mean, median, mode, variance, standard deviation, and various visual tools like boxplots and scatter plots. Additionally, it explains distance measures for different data types and the importance of understanding data distributions and correlations.

Uploaded by

visuvaan

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

25 views26 pages

Basic Statistical Descriptions of Data

Uploaded by

visuvaan

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

You are on page 1/ 26

Basic Statistical Descriptions of Data

 Motivation
 To better understand the data: central tendency,
variation and spread
 Data dispersion characteristics
 median, max, min, quantiles, outliers, variance, etc.
 Numerical dimensions correspond to sorted intervals
 Data dispersion: analyzed with multiple granularities
of precision
 Boxplot or quantile analysis on sorted intervals
 Dispersion analysis on computed measures
 Folding measures into numerical dimensions
 Boxplot or quantile analysis on the transformed cube
1
Measuring the Central Tendency
 Mean (algebraic measure) (sample vs. population): 1 n
x   xi   x
Note: n is sample size and N is population size. n i 1 N
n
Weighted arithmetic mean:
w x

i i
 Trimmed mean: chopping extreme values x i 1
n
 Median: w
i 1
i

 Middle value if odd number of values, or average of

the middle two values otherwise
 Estimated by interpolation (for grouped data):
n / 2  ( freq )l
median  L1  ( ) width
 Mode freq median
 Value that occurs most frequently in the data
 Unimodal, bimodal, trimodal
 Empirical formula: mean  mode  3  (mean  median)
2
Symmetric vs. Skewed Data
 Median, mean and mode of symmetric
symmetric, positively and
negatively skewed data

positively skewed negatively skewed

January 26, 2025 Data Mining: Concepts and Techniques 3

Measuring the Dispersion of Data
 Quartiles, outliers and boxplots
 Quartiles: Q1 (25th percentile), Q3 (75th percentile)
 Inter-quartile range: IQR = Q3 – Q1
 Five number summary: min, Q1, median, Q3, max
 Boxplot: ends of the box are the quartiles; median is marked; add
whiskers, and plot outliers individually
 Outlier: usually, a value higher/lower than 1.5 x IQR
 Variance and standard deviation (sample: s, population: σ)
 Variance: (algebraic, scalable computation)
1 n 1 n 2 1 n
 [ xi  ( xi ) 2 ]
n n
1 1
s  ( xi  x )         xi   2
2 2 2 2 2
( x )
n  1 i 1 n  1 i 1 n i 1 N i 1
i
N i 1

 Standard deviation s (or σ) is the square root of variance s2 (or σ2)

4
Boxplot Analysis

 Five-number summary of a distribution

 Minimum, Q1, Median, Q3, Maximum
 Boxplot
 Data is represented with a box
 The ends of the box are at the first and third
quartiles, i.e., the height of the box is IQR
 The median is marked by a line within the
box
 Whiskers: two lines outside the box extended
to Minimum and Maximum
 Outliers: points beyond a specified outlier
threshold, plotted individually

5
Properties of Normal Distribution Curve

 The normal (distribution) curve

 From μ–σ to μ+σ: contains about 68% of the

measurements (μ: mean, σ: standard deviation)

 From μ–2σ to μ+2σ: contains about 95% of it
 From μ–3σ to μ+3σ: contains about 99.7% of it

6
Graphic Displays of Basic Statistical Descriptions

 Boxplot: graphic display of five-number summary

 Histogram: x-axis are values, y-axis repres. frequencies
 Quantile plot: each value xi is paired with fi indicating
that approximately 100 fi % of data are  xi
 Quantile-quantile (q-q) plot: graphs the quantiles of
one univariant distribution against the corresponding
quantiles of another
 Scatter plot: each pair of values is a pair of coordinates
and plotted as points in the plane

7
Quantile Plot
 Displays all of the data (allowing the user to assess both
the overall behavior and unusual occurrences)
 Plots quantile information
 For a data xi data sorted in increasing order, fi
indicates that approximately 100 fi% of the data are
below or equal to the value xi

Data Mining: Concepts and Techniques 8

Quantile-Quantile (Q-Q) Plot
 Graphs the quantiles of one univariate distribution against the
corresponding quantiles of another
 View: Is there is a shift in going from one distribution to another?
 Example shows unit price of items sold at Branch 1 vs. Branch 2 for
each quantile. Unit prices of items sold at Branch 1 tend to be lower
than those at Branch 2.

9
Scatter plot
 Provides a first look at bivariate data to see clusters of
points, outliers, etc
 Each pair of values is treated as a pair of coordinates and
plotted as points in the plane

10
Positively and Negatively Correlated Data

 The left half fragment is positively

correlated
 The right half is negative correlated

11
Uncorrelated Data

12
Similarity and Dissimilarity
 Similarity
 Numerical measure of how alike two data objects are

 Value is higher when objects are more alike

 Often falls in the range [0,1]

 Dissimilarity (e.g., distance)

 Numerical measure of how different two data objects

are
 Lower when objects are more alike

 Minimum dissimilarity is often 0

 Upper limit varies

 Proximity refers to a similarity or dissimilarity

13
Data Matrix and Dissimilarity Matrix
 Data matrix
 n data points with p  x11 ... x1f ... x1p 
 
dimensions  ... ... ... ... ... 
x ... xip 
 Two modes
... xif
 i1 
 ... ... ... ... ... 
x ... xnf ... xnp 
 n1 
 Dissimilarity matrix
 0 
 n data points, but
 d(2,1) 0 
registers only the  
 d(3,1) d ( 3,2) 0 
distance  
 A triangular matrix  : : : 
d ( n,1) d ( n,2) ... ... 0
 Single mode

14
Proximity Measure for Nominal Attributes

 Can take 2 or more states, e.g., red, yellow, blue,

green (generalization of a binary attribute)
 Method 1: Simple matching
 m: # of matches, p: total # of variables
d (i, j)  p 
p
m

 Method 2: Use a large number of binary attributes

 creating a new binary attribute for each of the
M nominal states

15
Proximity Measure for Binary Attributes
Object j
 A contingency table for binary data
Object i

 Distance measure for symmetric

binary variables:
 Distance measure for asymmetric
binary variables:
 Jaccard coefficient (similarity
measure for asymmetric binary
variables):
 Note: Jaccard coefficient is the same as “coherence”:

16
Dissimilarity between Binary Variables
 Example
Name Gender Fever Cough Test-1 Test-2 Test-3 Test-4
Jack M Y N P N N N
Mary F Y N P N P N
Jim M Y P N N N N

 Gender is a symmetric attribute

 The remaining attributes are asymmetric binary
 Let the values Y and P be 1, and the value N 0
01
d ( jack , m ary)   0.33
2 01
11
d ( jack , jim )   0.67
111
1 2
d ( jim , m ary)   0.75
11 2
17
Standardizing Numeric Data
x
z   
 Z-score:
 X: raw score to be standardized, μ: mean of the population, σ:
standard deviation
 the distance between the raw score and the population mean in
units of the standard deviation
 negative when the raw score is below the mean, “+” when above
 An alternative way: Calculate the mean absolute deviation
sf  1
n (| x1 f  m f |  | x2 f  m f | ... | xnf  m f |)
where m  1 (x  x  ...  x )
n 1f 2 f xif  m f
.
f nf

zif  sf
 standardized measure (z-score):
 Using mean absolute deviation is more robust than using standard
deviation

18
Example:
Data Matrix and Dissimilarity Matrix
Data Matrix
point attribute1 attribute2
x1 1 2
x2 3 5
x3 2 0
x4 4 5

Dissimilarity Matrix
(with Euclidean Distance)

x1 x2 x3 x4
x1 0
x2 3.61 0
x3 5.1 5.1 0
x4 4.24 1 5.39 0

19
Distance on Numeric Data: Minkowski Distance
 Minkowski distance: A popular distance measure

where i = (xi1, xi2, …, xip) and j = (xj1, xj2, …, xjp) are two
p-dimensional data objects, and h is the order (the
distance so defined is also called L-h norm)
 Properties
 d(i, j) > 0 if i ≠ j, and d(i, i) = 0 (Positive definiteness)
 d(i, j) = d(j, i) (Symmetry)
 d(i, j)  d(i, k) + d(k, j) (Triangle Inequality)
 A distance that satisfies these properties is a metric
20
Special Cases of Minkowski Distance
 h = 1: Manhattan (city block, L1 norm) distance
 E.g., the Hamming distance: the number of bits that are

different between two binary vectors

d (i, j) | x  x |  | x  x | ... | x  x |
i1 j1 i2 j 2 ip jp

 h = 2: (L2 norm) Euclidean distance

d (i, j)  (| x  x |2  | x  x |2 ... | x  x |2 )
i1 j1 i2 j 2 ip jp

 h  . “supremum” (Lmax norm, L norm) distance.

 This is the maximum difference between any component

(attribute) of the vectors

21
Example: Minkowski Distance
Dissimilarity Matrices
point attribute 1 attribute 2 Manhattan (L1)
x1 1 2
L x1 x2 x3 x4
x2 3 5 x1 0
x3 2 0 x2 5 0
x4 4 5 x3 3 6 0
x4 6 1 7 0
Euclidean (L2)
L2 x1 x2 x3 x4
x1 0
x2 3.61 0
x3 2.24 5.1 0
x4 4.24 1 5.39 0

Supremum
L x1 x2 x3 x4
x1 0
x2 3 0
x3 2 5 0
x4 3 1 5 0
22
Ordinal Variables

 An ordinal variable can be discrete or continuous

 Order is important, e.g., rank
 Can be treated like interval-scaled
 replace xif by their rank rif {1,..., M f }
 map the range of each variable onto [0, 1] by replacing
i-th object in the f-th variable by
rif 1
zif 
M f 1
 compute the dissimilarity using methods for interval-
scaled variables

23
Attributes of Mixed Type

 A database may contain all attribute types

 Nominal, symmetric binary, asymmetric binary, numeric,
ordinal
 One may use a weighted formula to combine their effects
 pf  1 ij( f ) dij( f )
d (i, j) 
 pf  1 ij( f )
 f is binary or nominal:
dij(f) = 0 if xif = xjf , or dij(f) = 1 otherwise
 f is numeric: use the normalized distance
 f is ordinal
 Compute ranks rif and rif  1
zif 
 Treat zif as interval-scaled M f 1
24
Cosine Similarity
 A document can be represented by thousands of attributes, each
recording the frequency of a particular word (such as keywords) or
phrase in the document.

 Other vector objects: gene features in micro-arrays, …

 Applications: information retrieval, biologic taxonomy, gene feature
mapping, ...
 Cosine measure: If d1 and d2 are two vectors (e.g., term-frequency
vectors), then
cos(d1, d2) = (d1  d2) /||d1|| ||d2|| ,
where  indicates vector dot product, ||d||: the length of vector d

25
Example: Cosine Similarity
 cos(d1, d2) = (d1  d2) /||d1|| ||d2|| ,
where  indicates vector dot product, ||d|: the length of vector d

 Ex: Find the similarity between documents 1 and 2.

d1 = (5, 0, 3, 0, 2, 0, 0, 2, 0, 0)
d2 = (3, 0, 2, 0, 1, 1, 0, 1, 0, 1)

d1d2 = 5*3+0*0+3*2+0*0+2*1+0*1+0*1+2*1+0*0+0*1 = 25
||d1||= (5*5+0*0+3*3+0*0+2*2+0*0+0*0+2*2+0*0+0*0)0.5=(42)0.5
= 6.481
||d2||= (3*3+0*0+2*2+0*0+1*1+1*1+0*0+1*1+0*0+1*1)0.5=(17)0.5
= 4.12
cos(d1, d2 ) = 0.94

Class BSC Book Statistics All Chpter Wise Notes
66% (50)
Class BSC Book Statistics All Chpter Wise Notes
128 pages
Maths Grade 10 Term 3 Topics
No ratings yet
Maths Grade 10 Term 3 Topics
6 pages
Statistics: a QuickStudy Laminated Reference Guide
From Everand
Statistics: a QuickStudy Laminated Reference Guide
BarCharts Publishing, Inc.
No ratings yet
02 Data
No ratings yet
02 Data
35 pages
Mod 4 Types of Data in Cluster Analysis
No ratings yet
Mod 4 Types of Data in Cluster Analysis
31 pages
IT326 - Ch2
No ratings yet
IT326 - Ch2
44 pages
2 2 Data
No ratings yet
2 2 Data
27 pages
DM-Knowing Your Data
No ratings yet
DM-Knowing Your Data
56 pages
02data Edited v2
No ratings yet
02data Edited v2
43 pages
9-2 Data Analysis and Pre-Processing Part 2 PDF
No ratings yet
9-2 Data Analysis and Pre-Processing Part 2 PDF
27 pages
2 1 Data
No ratings yet
2 1 Data
22 pages
CPSC 4830 2025summer Lecture 2
No ratings yet
CPSC 4830 2025summer Lecture 2
42 pages
Getting To Know Your Data
No ratings yet
Getting To Know Your Data
78 pages
Lectur 4 Basic Statistical Descriptions of Data
No ratings yet
Lectur 4 Basic Statistical Descriptions of Data
44 pages
Chapter 2: Getting To Know Your Data
No ratings yet
Chapter 2: Getting To Know Your Data
30 pages
Lec 2
No ratings yet
Lec 2
26 pages
02 Data
No ratings yet
02 Data
41 pages
Data Mining:: Concepts and Techniques
100% (1)
Data Mining:: Concepts and Techniques
63 pages
Lecture 2 - Exploratory Data Analysis
No ratings yet
Lecture 2 - Exploratory Data Analysis
35 pages
Lecture 2
No ratings yet
Lecture 2
62 pages
Module 1
No ratings yet
Module 1
64 pages
02 Kinds of Data
No ratings yet
02 Kinds of Data
41 pages
Concepts and Techniques: - Chapter 2
No ratings yet
Concepts and Techniques: - Chapter 2
36 pages
02 Data
No ratings yet
02 Data
36 pages
Chapter 2
No ratings yet
Chapter 2
65 pages
02 Data
No ratings yet
02 Data
66 pages
CH 2
No ratings yet
CH 2
68 pages
Concepts and Techniques: - Chapter 2
No ratings yet
Concepts and Techniques: - Chapter 2
29 pages
Data Analysts-1
No ratings yet
Data Analysts-1
65 pages
Unit 3 Data Preprocessing - Data
No ratings yet
Unit 3 Data Preprocessing - Data
90 pages
VIPDMTheory Chapter 2
No ratings yet
VIPDMTheory Chapter 2
56 pages
Lect 3
No ratings yet
Lect 3
51 pages
Concepts and Techniques: - Chapter 2
No ratings yet
Concepts and Techniques: - Chapter 2
65 pages
Data Mining (DM) : Lecture 3: Know Your Data
No ratings yet
Data Mining (DM) : Lecture 3: Know Your Data
53 pages
Concepts and Techniques: - Chapter 2
No ratings yet
Concepts and Techniques: - Chapter 2
65 pages
02 Data
No ratings yet
02 Data
62 pages
Data and Metrics
No ratings yet
Data and Metrics
35 pages
1 L2 Intro DAM
No ratings yet
1 L2 Intro DAM
27 pages
02know Your Data-Lecture2-3
No ratings yet
02know Your Data-Lecture2-3
53 pages
Data Preprocessing Data Basics
No ratings yet
Data Preprocessing Data Basics
86 pages
02 KnowYourData
No ratings yet
02 KnowYourData
44 pages
02 Data
No ratings yet
02 Data
24 pages
Data Warehousing and Data Mining
No ratings yet
Data Warehousing and Data Mining
46 pages
CS 591.03 Introduction To Data Mining Instructor: Abdullah Mueen
No ratings yet
CS 591.03 Introduction To Data Mining Instructor: Abdullah Mueen
52 pages
02 Data
No ratings yet
02 Data
65 pages
02data Part2
No ratings yet
02data Part2
34 pages
Lec.02 Getting To Know Your Data
No ratings yet
Lec.02 Getting To Know Your Data
62 pages
Chapter - 2 Data Mining
No ratings yet
Chapter - 2 Data Mining
21 pages
02know Your Data Lecture2 3
No ratings yet
02know Your Data Lecture2 3
53 pages
DWDM Unit-2
No ratings yet
DWDM Unit-2
19 pages
DM Unit-1-1
No ratings yet
DM Unit-1-1
56 pages
Module No 2 - Part 2 - Compressed - Compressed
No ratings yet
Module No 2 - Part 2 - Compressed - Compressed
46 pages
Data Type, Data Chart, Descriptive Statistics
No ratings yet
Data Type, Data Chart, Descriptive Statistics
65 pages
02 Data
No ratings yet
02 Data
64 pages
CH 2
No ratings yet
CH 2
35 pages
02 Data
No ratings yet
02 Data
65 pages
Slides of Lecture 2 of CS3319 SJTU
No ratings yet
Slides of Lecture 2 of CS3319 SJTU
35 pages
Data Mining 1
No ratings yet
Data Mining 1
29 pages
Transportation Data Mining: Chapter 2. Getting To Know Your Data
No ratings yet
Transportation Data Mining: Chapter 2. Getting To Know Your Data
77 pages
2 Knowing Data & Visualization
No ratings yet
2 Knowing Data & Visualization
51 pages
Learn Statistics Fast: A Simplified Detailed Version for Students
From Everand
Learn Statistics Fast: A Simplified Detailed Version for Students
Hesbon R.M
No ratings yet
Co-Clustering: Models, Algorithms and Applications
From Everand
Co-Clustering: Models, Algorithms and Applications
Gérard Govaert
No ratings yet
Grade 11 Data Handling
No ratings yet
Grade 11 Data Handling
31 pages
Choudhury Et Al 2017 Ready To Use Therapeutic Food Made From Locally Available Food Ingredients Is Well Accepted by
No ratings yet
Choudhury Et Al 2017 Ready To Use Therapeutic Food Made From Locally Available Food Ingredients Is Well Accepted by
11 pages
Unit II Notes
No ratings yet
Unit II Notes
36 pages
Maths Model For Grade 12 (NS) 2017
No ratings yet
Maths Model For Grade 12 (NS) 2017
7 pages
Gse Mathematics-Glossary-K-12
No ratings yet
Gse Mathematics-Glossary-K-12
10 pages
Seminar 3 Measures of Dispersion With Answers
No ratings yet
Seminar 3 Measures of Dispersion With Answers
7 pages
05.1 Data Organization PRESENTATION
No ratings yet
05.1 Data Organization PRESENTATION
19 pages
Measure of Position For Ungrouped Data
No ratings yet
Measure of Position For Ungrouped Data
10 pages
CHAPTER5 Assessment Answer
No ratings yet
CHAPTER5 Assessment Answer
20 pages
Assignment (1) SOlution
No ratings yet
Assignment (1) SOlution
15 pages
Measure of Dispersion (Range Quartile & Mean Deviation)
No ratings yet
Measure of Dispersion (Range Quartile & Mean Deviation)
55 pages
Module 2 - Exploratory Data Analysis (EDA) : Central Tendency and Variability
No ratings yet
Module 2 - Exploratory Data Analysis (EDA) : Central Tendency and Variability
56 pages
Gis Manual
No ratings yet
Gis Manual
22 pages
Yulu Business Case Study
No ratings yet
Yulu Business Case Study
42 pages
Unit 1 Data Acquisition
No ratings yet
Unit 1 Data Acquisition
62 pages
AP Stats Chapter 2 Homework Assignment
No ratings yet
AP Stats Chapter 2 Homework Assignment
4 pages
Elements of Statistics BCA Sem-I.
No ratings yet
Elements of Statistics BCA Sem-I.
46 pages
3) S1 Representation and Summary of Data - Dispersion
No ratings yet
3) S1 Representation and Summary of Data - Dispersion
27 pages
106 Data Science
No ratings yet
106 Data Science
11 pages
Module 2
No ratings yet
Module 2
75 pages
Lecture01 Describing Data Ver2
No ratings yet
Lecture01 Describing Data Ver2
67 pages
Mathematics Standard Year 11 Topic Guide Statistical Analysis
No ratings yet
Mathematics Standard Year 11 Topic Guide Statistical Analysis
12 pages
Ymzv Further Mathematics Bound Reference
No ratings yet
Ymzv Further Mathematics Bound Reference
30 pages
Measures of Spread
No ratings yet
Measures of Spread
5 pages
16.miralles 2025 Padel WIMU
No ratings yet
16.miralles 2025 Padel WIMU
11 pages
FIT1043 - Lecture 3 - 2024
No ratings yet
FIT1043 - Lecture 3 - 2024
69 pages
General Education 2024 Vol 7 Questionnaire
No ratings yet
General Education 2024 Vol 7 Questionnaire
45 pages