0% found this document useful (0 votes)

54 views13 pages

Machine Learing Algorithms

There are three main types of machine learning algorithms: supervised learning, unsupervised learning, and reinforcement learning. Supervised learning uses labeled data to predict outcomes, unsupervised learning finds hidden patterns in unlabeled data, and reinforcement learning learns from interactions with an environment. Some common algorithms discussed include linear regression, logistic regression, decision trees, support vector machines, naive bayes, and k-means clustering.

Uploaded by

Nexgen Technology

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as DOCX, PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

54 views13 pages

Machine Learing Algorithms

Uploaded by

Nexgen Technology

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as DOCX, PDF, TXT or read online on Scribd

You are on page 1/ 13

Broadly, there are 3 types of Machine Learning

Algorithms
1. Supervised Learning

How it works: This algorithm consist of a target / outcome variable (or dependent variable)
which is to be predicted from a given set of predictors (independent variables). Using these
set of variables, we generate a function that map inputs to desired outputs. The training
process continues until the model achieves a desired level of accuracy on the training data.
Examples of Supervised Learning: Regression, Decision Tree, Random Forest, KNN,
Logistic Regression etc.

2. Unsupervised Learning

How it works: In this algorithm, we do not have any target or outcome variable to predict /
estimate. It is used for clustering population in different groups, which is widely used for
segmenting customers in different groups for specific intervention. Examples of
Unsupervised Learning: Apriori algorithm, K-means.

3. Reinforcement Learning:

How it works: Using this algorithm, the machine is trained to make specific decisions. It
works this way: the machine is exposed to an environment where it trains itself continually
using trial and error. This machine learns from past experience and tries to capture the best
possible knowledge to make accurate business decisions. Example of Reinforcement
Learning: Markov Decision Process

List of Common Machine Learning Algorithms

Here is the list of commonly used machine learning algorithms. These algorithms can be
applied to almost any data problem:

1. Linear Regression
2. Logistic Regression
3. Decision Tree
4. SVM
5. Naive Bayes
6. kNN
7. K-Means
8. Random Forest
9. Dimensionality Reduction Algorithms
10. Gradient Boosting algorithms
1. GBM
2. XGBoost
3. LightGBM
4. CatBoost

1. Linear Regression
It is used to estimate real values (cost of houses, number of calls, total sales etc.) based on
continuous variable(s). Here, we establish relationship between independent and dependent
variables by fitting a best line. This best fit line is known as regression line and represented
by a linear equation Y= a *X + b.

The best way to understand linear regression is to relive this experience of childhood. Let us
say, you ask a child in fifth grade to arrange people in his class by increasing order of weight,
without asking them their weights! What do you think the child will do? He / she would
likely look (visually analyze) at the height and build of people and arrange them using a
combination of these visible parameters. This is linear regression in real life! The child has
actually figured out that height and build would be correlated to the weight by a relationship,
which looks like the equation above.

In this equation:

 Y – Dependent Variable
 a – Slope
 X – Independent variable
 b – Intercept

These coefficients a and b are derived based on minimizing the sum of squared difference of
distance between data points and regression line.

Look at the below example. Here we have identified the best fit line having linear equation
y=0.2811x+13.9. Now using this equation, we can find the weight, knowing the height of a
person.
Linear Regression is mainly of two types: Simple Linear Regression and Multiple Linear
Regression. Simple Linear Regression is characterized by one independent variable. And,
Multiple Linear Regression(as the name suggests) is characterized by multiple (more than 1)
independent variables. While finding the best fit line, you can fit a polynomial or curvilinear
regression. And these are known as polynomial or curvilinear regression.

Here’s a coding window to try out your hand and build your own linear regression model in
Python:

2. Logistic Regression
Don’t get confused by its name! It is a classification not a regression algorithm. It is used to
estimate discrete values ( Binary values like 0/1, yes/no, true/false ) based on given set of
independent variable(s). In simple words, it predicts the probability of occurrence of an event
by fitting data to a logit function. Hence, it is also known as logit regression. Since, it
predicts the probability, its output values lies between 0 and 1 (as expected).

Again, let us try and understand this through a simple example.

Let’s say your friend gives you a puzzle to solve. There are only 2 outcome scenarios – either
you solve it or you don’t. Now imagine, that you are being given wide range of puzzles /
quizzes in an attempt to understand which subjects you are good at. The outcome to this
study would be something like this – if you are given a trignometry based tenth grade
problem, you are 70% likely to solve it. On the other hand, if it is grade fifth history question,
the probability of getting an answer is only 30%. This is what Logistic Regression provides
you.

Coming to the math, the log odds of the outcome is modeled as a linear combination of the
predictor variables.

odds= p/ (1-p) = probability of event occurrence / probability of not event

occurrence
ln(odds) = ln(p/(1-p))
logit(p) = ln(p/(1-p)) = b0+b1X1+b2X2+b3X3....+bkXk

Above, p is the probability of presence of the characteristic of interest. It chooses parameters

that maximize the likelihood of observing the sample values rather than that minimize the
sum of squared errors (like in ordinary regression).

Now, you may ask, why take a log? For the sake of simplicity, let’s just say that this is one of
the best mathematical way to replicate a step function. I can go in more details, but that will
beat the purpose of this article.
Build your own logistic regression
model in Python here and check the accuracy:

Furthermore..

There are many different steps that could be tried in order to improve the model:

 including interaction terms

 removing features
 regularization techniques
 using a non-linear model

3. Decision Tree
This is one of my favorite algorithm and I use it quite frequently. It is a type of supervised
learning algorithm that is mostly used for classification problems. Surprisingly, it works for
both categorical and continuous dependent variables. In this algorithm, we split the
population into two or more homogeneous sets. This is done based on most significant
attributes/ independent variables to make as distinct groups as possible. For more details, you
can read: Decision Tree Simplified.
source: statsexchange

In the image above, you can see that population is classified into four different groups based
on multiple attributes to identify ‘if they will play or not’. To split the population into
different heterogeneous groups, it uses various techniques like Gini, Information Gain, Chi-
square, entropy.

The best way to understand how decision tree works, is to play Jezzball – a classic game from
Microsoft (image below). Essentially, you have a room with moving walls and you need to
create walls such that maximum area gets cleared off with out the balls.

So, every time you split the room with a wall, you are trying to create 2 different populations
with in the same room. Decision trees work in very similar fashion by dividing a population
in as different groups as possible.

More: Simplified Version of Decision Tree Algorithms

Let’s get our hands dirty and code our own decision tree in Python!

SVM (Support Vector Machine)

It is a classification method. In this algorithm, we plot each data item as a point in n-
dimensional space (where n is number of features you have) with the value of each feature
being the value of a particular coordinate.

For example, if we only had two features like Height and Hair length of an individual, we’d
first plot these two variables in two dimensional space where each point has two co-ordinates
(these co-ordinates are known as Support Vectors)

Now, we will find some line that splits the data between the two differently classified groups
of data. This will be the line such that the distances from the closest point in each of the two
groups will be farthest away.

In the example shown above, the line which splits the data into two differently classified
groups is the black line, since the two closest points are the farthest apart from the line. This
line is our classifier. Then, depending on where the testing data lands on either side of the
line, that’s what class we can classify the new data as.

More: Simplified Version of Support Vector Machine

Think of this algorithm as playing JezzBall in n-dimensional space. The tweaks in the
game are:
 You can draw lines/planes at any angles (rather than just horizontal or vertical as in
the classic game)
 The objective of the game is to segregate balls of different colors in different rooms.
 And the balls are not moving.

Try your hand and design an SVM model in Python through this coding window:

Naive Bayes
It is a classification technique based on Bayes’ theorem with an assumption of independence
between predictors. In simple terms, a Naive Bayes classifier assumes that the presence of a
particular feature in a class is unrelated to the presence of any other feature. For example, a
fruit may be considered to be an apple if it is red, round, and about 3 inches in diameter. Even
if these features depend on each other or upon the existence of the other features, a naive
Bayes classifier would consider all of these properties to independently contribute to the
probability that this fruit is an apple.

Naive Bayesian model is easy to build and particularly useful for very large data sets. Along
with simplicity, Naive Bayes is known to outperform even highly sophisticated classification
methods.

Bayes theorem provides a way of calculating posterior probability P(c|x) from P(c), P(x) and
P(x|c). Look at the equation below:

Here,

 P(c|x) is the posterior probability of class (target) given predictor (attribute).

 P(c) is the prior probability of class.
 P(x|c) is the likelihood which is the probability of predictor given class.
 P(x) is the prior probability of predictor.

Example: Let’s understand it using an example. Below I have a training data set of weather
and corresponding target variable ‘Play’. Now, we need to classify whether players will play
or not based on weather condition. Let’s follow the below steps to perform it.

Step 1: Convert the data set to frequency table

Step 2: Create Likelihood table by finding the probabilities like Overcast probability = 0.29
and probability of playing is 0.64.

Step 3: Now, use Naive Bayesian equation to calculate the posterior probability for each
class. The class with the highest posterior probability is the outcome of prediction.

Problem: Players will pay if weather is sunny, is this statement is correct?

We can solve it using above discussed method, so P(Yes | Sunny) = P( Sunny | Yes) * P(Yes)
/ P (Sunny)

Here we have P (Sunny |Yes) = 3/9 = 0.33, P(Sunny) = 5/14 = 0.36, P( Yes)= 9/14 = 0.64

Now, P (Yes | Sunny) = 0.33 * 0.64 / 0.36 = 0.60, which has higher probability.

Naive Bayes uses a similar method to predict the probability of different class based on
various attributes. This algorithm is mostly used in text classification and with problems
having multiple classes.

Code a Naive Bayes classification model in Python:

kNN (k- Nearest Neighbors)

It can be used for both classification and regression problems. However, it is more widely
used in classification problems in the industry. K nearest neighbors is a simple algorithm that
stores all available cases and classifies new cases by a majority vote of its k neighbors. The
case being assigned to the class is most common amongst its K nearest neighbors measured
by a distance function.

These distance functions can be Euclidean, Manhattan, Minkowski and Hamming distance.
First three functions are used for continuous function and fourth one (Hamming) for
categorical variables. If K = 1, then the case is simply assigned to the class of its nearest
neighbor. At times, choosing K turns out to be a challenge while performing kNN modeling.

More: Introduction to k-nearest neighbors : Simplified.

KNN can easily be mapped to our real lives. If you want to learn about a person, of whom
you have no information, you might like to find out about his close friends and the circles he
moves in and gain access to his/her information!

Things to consider before selecting kNN:

 KNN is computationally expensive

 Variables should be normalized else higher range variables can bias it
 Works on pre-processing stage more before going for kNN like an outlier, noise removal

Python Code

K-Means
It is a type of unsupervised algorithm which solves the clustering problem. Its procedure
follows a simple and easy way to classify a given data set through a certain number of
clusters (assume k clusters). Data points inside a cluster are homogeneous and heterogeneous
to peer groups.

Remember figuring out shapes from ink blots? k means is somewhat similar this activity.
You look at the shape and spread to decipher how many different clusters / population are
present!
How K-means forms cluster:

1. K-means picks k number of points for each cluster known as centroids.

2. Each data point forms a cluster with the closest centroids i.e. k clusters.
3. Finds the centroid of each cluster based on existing cluster members. Here we have new
centroids.
4. As we have new centroids, repeat step 2 and 3. Find the closest distance for each data point
from new centroids and get associated with new k-clusters. Repeat this process until
convergence occurs i.e. centroids does not change.

How to determine value of K:

In K-means, we have clusters and each cluster has its own centroid. Sum of square of
difference between centroid and the data points within a cluster constitutes within sum of
square value for that cluster. Also, when the sum of square values for all the clusters are
added, it becomes total within sum of square value for the cluster solution.

We know that as the number of cluster increases, this value keeps on decreasing but if you
plot the result you may see that the sum of squared distance decreases sharply up to some
value of k, and then much more slowly after that. Here, we can find the optimum number of
cluster.
Python Code

Random Forest
Random Forest is a trademark term for an ensemble of decision trees. In Random Forest,
we’ve collection of decision trees (so known as “Forest”). To classify a new object based on
attributes, each tree gives a classification and we say the tree “votes” for that class. The forest
chooses the classification having the most votes (over all the trees in the forest).

Each tree is planted & grown as follows:

1. If the number of cases in the training set is N, then sample of N cases is taken at
random but with replacement. This sample will be the training set for growing the
tree.
2. If there are M input variables, a number m<<M is specified such that at each node, m
variables are selected at random out of the M and the best split on these m is used to
split the node. The value of m is held constant during the forest growing.
3. Each tree is grown to the largest extent possible. There is no pruning.

For more details on this algorithm, comparing with decision tree and tuning model
parameters, I would suggest you to read these articles:

1. Introduction to Random forest – Simplified

2. Comparing a CART model to Random Forest (Part 1)
3. Comparing a Random Forest to a CART model (Part 2)
4. Tuning the parameters of your Random Forest model

Python Code:

Dimensionality Reduction Algorithms

In the last 4-5 years, there has been an exponential increase in data capturing at every
possible stages. Corporates/ Government Agencies/ Research organisations are not only
coming with new sources but also they are capturing data in great detail.
For example: E-commerce companies are capturing more details about customer like their
demographics, web crawling history, what they like or dislike, purchase history, feedback and
many others to give them personalized attention more than your nearest grocery shopkeeper.

As a data scientist, the data we are offered also consist of many features, this sounds good for
building good robust model but there is a challenge. How’d you identify highly significant
variable(s) out 1000 or 2000? In such cases, dimensionality reduction algorithm helps us
along with various other algorithms like Decision Tree, Random Forest, PCA, Factor
Analysis, Identify based on correlation matrix, missing value ratio and others.

To know more about this algorithms, you can read “Beginners Guide To Learn Dimension
Reduction Techniques“.

Gradient Boosting Algorithms

10.1. GBM

GBM is a boosting algorithm used when we deal with plenty of data to make a prediction
with high prediction power. Boosting is actually an ensemble of learning algorithms which
combines the prediction of several base estimators in order to improve robustness over a
single estimator. It combines multiple weak or average predictors to a build strong predictor.
These boosting algorithms always work well in data science competitions like Kaggle, AV
Hackathon, CrowdAnalytix.

More: Know about Boosting algorithms in detail

. XGBoost

Another classic gradient boosting algorithm that’s known to be the decisive choice between
winning and losing in some Kaggle competitions.

The XGBoost has an immensely high predictive power which makes it the best choice for
accuracy in events as it possesses both linear model and the tree learning algorithm, making
the algorithm almost 10x faster than existing gradient booster techniques.

The support includes various objective functions, including regression, classification and
ranking.

One of the most interesting things about the XGBoost is that it is also called a regularized
boosting technique. This helps to reduce overfit modelling and has a massive support for a
range of languages such as Scala, Java, R, Python, Julia and C++.

Supports distributed and widespread training on many machines that encompass GCE, AWS,
Azure and Yarn clusters. XGBoost can also be integrated with Spark, Flink and other cloud
dataflow systems with a built in cross validation at each iteration of the boosting process.

LightGBM
LightGBM is a gradient boosting framework that uses tree based learning algorithms. It is
designed to be distributed and efficient with the following advantages:

 Faster training speed and higher efficiency

 Lower memory usage
 Better accuracy
 Parallel and GPU learning supported
 Capable of handling large-scale data

The framework is a fast and high-performance gradient boosting one based on decision tree
algorithms, used for ranking, classification and many other machine learning tasks. It was
developed under the Distributed Machine Learning Toolkit Project of Microsoft.

Since the LightGBM is based on decision tree algorithms, it splits the tree leaf wise with the
best fit whereas other boosting algorithms split the tree depth wise or level wise rather than
leaf-wise. So when growing on the same leaf in Light GBM, the leaf-wise algorithm can
reduce more loss than the level-wise algorithm and hence results in much better accuracy
which can rarely be achieved by any of the existing boosting algorithms.

Also, it is surprisingly very fast, hence the word ‘Light’.

Max Born, Albert Einstein-The Born-Einstein Letters-Macmillan (1971)
100% (1)
Max Born, Albert Einstein-The Born-Einstein Letters-Macmillan (1971)
132 pages
Web Development Using Dotnet Internship Report
100% (1)
Web Development Using Dotnet Internship Report
37 pages
Machine Learning Interview Questions
From Everand
Machine Learning Interview Questions
Tech Interviews
4.5/5 (2)
Wiring C11
No ratings yet
Wiring C11
12 pages
2-Machine Learning Algorithms
No ratings yet
2-Machine Learning Algorithms
16 pages
Commonly Used Machine Learning Algorithms
No ratings yet
Commonly Used Machine Learning Algorithms
27 pages
Broadly, There Are 3 Types of Machine Learning Algorithms.
No ratings yet
Broadly, There Are 3 Types of Machine Learning Algorithms.
33 pages
Machinelearning Algorithm Basics2 NOTES
No ratings yet
Machinelearning Algorithm Basics2 NOTES
72 pages
Machine Learning
100% (3)
Machine Learning
46 pages
Essentials of Machine Learning Algorithms
No ratings yet
Essentials of Machine Learning Algorithms
15 pages
Commonly Used Machine Learning Algorithms
No ratings yet
Commonly Used Machine Learning Algorithms
38 pages
41 Machine Learning Algorithms I
No ratings yet
41 Machine Learning Algorithms I
8 pages
Commonly Used Machine Learning Algorithms (With Python and R Codes)
No ratings yet
Commonly Used Machine Learning Algorithms (With Python and R Codes)
19 pages
ML Algorithms
No ratings yet
ML Algorithms
12 pages
Machine Learning
No ratings yet
Machine Learning
53 pages
Lecture - 2 & 3
No ratings yet
Lecture - 2 & 3
62 pages
Interview Preparing - ML Draft
No ratings yet
Interview Preparing - ML Draft
12 pages
PID5108657
No ratings yet
PID5108657
8 pages
CS601 - Machine Learning - Unit 1 - Notes - 1672759748
No ratings yet
CS601 - Machine Learning - Unit 1 - Notes - 1672759748
13 pages
Unit 3 Machine Learning
No ratings yet
Unit 3 Machine Learning
12 pages
Chapter Four
No ratings yet
Chapter Four
75 pages
Unit 3
No ratings yet
Unit 3
61 pages
ML-Unit-2
No ratings yet
ML-Unit-2
6 pages
Supervised Learning Notes
No ratings yet
Supervised Learning Notes
13 pages
ML - Unit - 1
No ratings yet
ML - Unit - 1
47 pages
Machine Learning - Regression Notes
No ratings yet
Machine Learning - Regression Notes
9 pages
Learn Machine Learning in One Lesson Book
No ratings yet
Learn Machine Learning in One Lesson Book
8 pages
M2 - Supervised Machine Learning
No ratings yet
M2 - Supervised Machine Learning
79 pages
Machine Learning Algorithms For Breast Cancer Prediction
No ratings yet
Machine Learning Algorithms For Breast Cancer Prediction
8 pages
Machine Learning With Real Life Project: by - Rishabh Gaur
100% (2)
Machine Learning With Real Life Project: by - Rishabh Gaur
26 pages
v0_ML
No ratings yet
v0_ML
53 pages
Unit1 6thsemCS
No ratings yet
Unit1 6thsemCS
22 pages
ML Unit 2
No ratings yet
ML Unit 2
33 pages
ML & DL Notes
No ratings yet
ML & DL Notes
30 pages
Machine Learning
No ratings yet
Machine Learning
22 pages
Unit III
No ratings yet
Unit III
5 pages
(English (Auto-Generated) ) All Machine Learning Algorithms Explained in 17 Min (DownSub - Com)
No ratings yet
(English (Auto-Generated) ) All Machine Learning Algorithms Explained in 17 Min (DownSub - Com)
19 pages
Supervised ML
No ratings yet
Supervised ML
69 pages
Ca10bd6d De86 4bae 9427 c60d433d2076 Supervised Learning
No ratings yet
Ca10bd6d De86 4bae 9427 c60d433d2076 Supervised Learning
17 pages
AIML
No ratings yet
AIML
30 pages
Mechine Learning
No ratings yet
Mechine Learning
106 pages
Unit 1 - Machine Learning
No ratings yet
Unit 1 - Machine Learning
17 pages
Machine Learning Strategies
No ratings yet
Machine Learning Strategies
59 pages
Machine Learning Models
No ratings yet
Machine Learning Models
11 pages
Supervised Learning
No ratings yet
Supervised Learning
24 pages
Unit 1 Machine Learning - PDF Lands
No ratings yet
Unit 1 Machine Learning - PDF Lands
5 pages
UNIT1
No ratings yet
UNIT1
38 pages
Unit-5 MECH 3-2
No ratings yet
Unit-5 MECH 3-2
14 pages
ML Models
No ratings yet
ML Models
21 pages
Super Cheatsheet Machine Learning
100% (1)
Super Cheatsheet Machine Learning
15 pages
Machine Learning
No ratings yet
Machine Learning
33 pages
3.popular Machine Learning Algorithm
No ratings yet
3.popular Machine Learning Algorithm
11 pages
Supervised Learning
No ratings yet
Supervised Learning
46 pages
Slide 1
No ratings yet
Slide 1
29 pages
ML Unit-4
No ratings yet
ML Unit-4
20 pages
11 Most Common Machine Learning Algorithms Explained in A Nutshell by Soner Yıldırım Towards Data Science
No ratings yet
11 Most Common Machine Learning Algorithms Explained in A Nutshell by Soner Yıldırım Towards Data Science
16 pages
Machine Learning For Beginners PDF
No ratings yet
Machine Learning For Beginners PDF
29 pages
Technical Report
No ratings yet
Technical Report
5 pages
Machine Learning
No ratings yet
Machine Learning
100 pages
Machine Learning
100% (6)
Machine Learning
115 pages
Module 1 & 2
No ratings yet
Module 1 & 2
21 pages
Regression Models: by Mayuri Bhandari
No ratings yet
Regression Models: by Mayuri Bhandari
64 pages
Lecture 3
No ratings yet
Lecture 3
51 pages
Impact of Nutritional Factors in Blood Glucose
No ratings yet
Impact of Nutritional Factors in Blood Glucose
16 pages
A Study On Portfolio Analysis of Banking Sector1
No ratings yet
A Study On Portfolio Analysis of Banking Sector1
110 pages
C
No ratings yet
C
37 pages
A Transfer Learning Approach To Breast Cancer
No ratings yet
A Transfer Learning Approach To Breast Cancer
11 pages
A Study On Organisational Culture and Its Impact On Employees Behaviour
No ratings yet
A Study On Organisational Culture and Its Impact On Employees Behaviour
12 pages
Rice Quality Analysis Using Machine Learning
No ratings yet
Rice Quality Analysis Using Machine Learning
2 pages
Road Accident Prediction System Using Deep Learning
No ratings yet
Road Accident Prediction System Using Deep Learning
1 page
Pythonpython
No ratings yet
Pythonpython
6 pages
Stress Detection Using Machine Learning
100% (1)
Stress Detection Using Machine Learning
1 page
Web Server Log Analysis Sysytem
No ratings yet
Web Server Log Analysis Sysytem
3 pages
Secure Reversible Image Data Hiding For Secure Data Sharing
No ratings yet
Secure Reversible Image Data Hiding For Secure Data Sharing
2 pages
Internship Report Data Science
100% (1)
Internship Report Data Science
58 pages
Ieee 2022-23 Cse Titles
No ratings yet
Ieee 2022-23 Cse Titles
5 pages
Ieee 2022-23 Machin Learning Title
No ratings yet
Ieee 2022-23 Machin Learning Title
2 pages
Ieee 2022-2023 Eee Projeects
No ratings yet
Ieee 2022-2023 Eee Projeects
13 pages
Embedded System Ieee Projects Year 2021
No ratings yet
Embedded System Ieee Projects Year 2021
4 pages
Ieee Embedded 2022-23
No ratings yet
Ieee Embedded 2022-23
7 pages
Ieee 2022-23 Aggriculture
No ratings yet
Ieee 2022-23 Aggriculture
3 pages
Ieee 2022-23 Deep Learning Titles
No ratings yet
Ieee 2022-23 Deep Learning Titles
3 pages
LSTM
No ratings yet
LSTM
2 pages
FFS: Flood Forecasting System Based On Integrated Big and Crowd Source Data by Using Deep Learning Techniques
No ratings yet
FFS: Flood Forecasting System Based On Integrated Big and Crowd Source Data by Using Deep Learning Techniques
13 pages
Screenshot 2024-08-29 at 14.26.16
No ratings yet
Screenshot 2024-08-29 at 14.26.16
2 pages
Junior Cert History Notes
No ratings yet
Junior Cert History Notes
56 pages
Performance Evaluation of Mutual Funds Exploring The Economic
No ratings yet
Performance Evaluation of Mutual Funds Exploring The Economic
17 pages
Chapter11 PDF
No ratings yet
Chapter11 PDF
10 pages
OrionMX User Manual
No ratings yet
OrionMX User Manual
261 pages
Rejuvenating The Marketing Mix
No ratings yet
Rejuvenating The Marketing Mix
14 pages
Lesson 2 Working With Text
No ratings yet
Lesson 2 Working With Text
16 pages
Hallite Catalogue
No ratings yet
Hallite Catalogue
374 pages
Role of ERP Systems in Improving Human Resources Management Processes
No ratings yet
Role of ERP Systems in Improving Human Resources Management Processes
16 pages
National Building Code 2024
50% (2)
National Building Code 2024
27 pages
Very Hungry Caterpillar Homework
100% (1)
Very Hungry Caterpillar Homework
8 pages
7.A. Grammar and Vocabulary
No ratings yet
7.A. Grammar and Vocabulary
2 pages
Electric Vehicles and Their Impact To The Electric Grid in Isolated Systems
No ratings yet
Electric Vehicles and Their Impact To The Electric Grid in Isolated Systems
6 pages
Story Telling (The Little Mermaid) Tugas Bahasa Inggris
No ratings yet
Story Telling (The Little Mermaid) Tugas Bahasa Inggris
7 pages
A&P Unit I
No ratings yet
A&P Unit I
4 pages
Module-4-Carpentry-7-8-Perform-basic-preventive-maintenance - FOR STUDENT
0% (1)
Module-4-Carpentry-7-8-Perform-basic-preventive-maintenance - FOR STUDENT
15 pages
Fundamental Principles of Project Management
No ratings yet
Fundamental Principles of Project Management
5 pages
HMCVersion2150 10march2021
No ratings yet
HMCVersion2150 10march2021
1,546 pages
Sip Terms Definition
No ratings yet
Sip Terms Definition
7 pages
TP SW24M G2 POE+ - POE - Switch
No ratings yet
TP SW24M G2 POE+ - POE - Switch
2 pages
Gra Test U6-10
No ratings yet
Gra Test U6-10
2 pages
The Life of A Roman Soldier: by Calum Johnson
No ratings yet
The Life of A Roman Soldier: by Calum Johnson
24 pages
IVRI PHD Entrance Information Bulletin 2011-12
No ratings yet
IVRI PHD Entrance Information Bulletin 2011-12
76 pages
Tableau Certified Data Analyst: Beta Exam Guide
No ratings yet
Tableau Certified Data Analyst: Beta Exam Guide
15 pages
Crim. Pro Outline
No ratings yet
Crim. Pro Outline
39 pages
DXRB
No ratings yet
DXRB
8 pages
Coffee Corner - Geotechnical Analysis (AM) - Key Points On Earthquake Analysis
No ratings yet
Coffee Corner - Geotechnical Analysis (AM) - Key Points On Earthquake Analysis
3 pages
Grade 8, The Best Christmas Gift
No ratings yet
Grade 8, The Best Christmas Gift
18 pages

Machine Learing Algorithms

Uploaded by

Machine Learing Algorithms

Uploaded by

Broadly, there are 3 types of Machine Learning

List of Common Machine Learning Algorithms

Again, let us try and understand this through a simple example.

odds= p/ (1-p) = probability of event occurrence / probability of not event

Above, p is the probability of presence of the characteristic of interest. It chooses parameters

 including interaction terms

More: Simplified Version of Decision Tree Algorithms

SVM (Support Vector Machine)

More: Simplified Version of Support Vector Machine

 P(c|x) is the posterior probability of class (target) given predictor (attribute).

Step 1: Convert the data set to frequency table

Problem: Players will pay if weather is sunny, is this statement is correct?

Code a Naive Bayes classification model in Python:

kNN (k- Nearest Neighbors)

More: Introduction to k-nearest neighbors : Simplified.

Things to consider before selecting kNN:

 KNN is computationally expensive

1. K-means picks k number of points for each cluster known as centroids.

How to determine value of K:

Each tree is planted & grown as follows:

1. Introduction to Random forest – Simplified

Dimensionality Reduction Algorithms

Gradient Boosting Algorithms

More: Know about Boosting algorithms in detail

 Faster training speed and higher efficiency

Also, it is surprisingly very fast, hence the word ‘Light’.

You might also like