Comprehensive Machine Learning Algorithms Tutorial

Linear Regression and Working Principle

Linear Regression is a supervised learning algorithm used to predict a continuous dependent variable from one or more independent variables.

For Simple Linear Regression, the relationship between x and y is represented by:

where:

  • y = predicted value
  • x = input feature
  • b0 = intercept
  • b1 = slope

Calculation of Slope

Intercept

Therefore, the regression equation becomes:

Working

  1. Collect training data.
  2. Assume a linear relationship between input and output.
  3. Calculate the regression coefficients.
  4. Fit the best line through the data.
  5. Use the equation to predict unknown values.

Example

Consider:

(1,2), (2,4), (3,6), (4,8)

The relationship is:

Assumptions

Important assumptions include:

  • Linear relationship between variables.
  • Errors are independent.
  • Constant error variance.
  • Errors are appropriately modeled, commonly with an approximately normal distribution for inference.

Complexity

For fitting ordinary least squares with one feature, the computational cost is generally low and depends on the implementation; prediction for one observation is:

O(1)

Conclusion

Linear Regression is mainly used for predicting continuous numerical values such as price, salary, demand, or temperature.


Logistic Regression vs Linear Regression

Logistic Regression is a supervised learning algorithm mainly used for classification, especially binary classification.

Instead of directly predicting a continuous value, it predicts the probability that an observation belongs to a class.

Logistic and Sigmoid Function

The sigmoid function is:

where:

Therefore:

The output lies between:

Classification

A threshold such as 0.5 may be used:

Example

Suppose a model predicts:

P(pass) = 0.82

Using threshold 0.5:

0.82 > 0.5

Therefore, the prediction is:

Pass


Linear vs Logistic Regression Comparison

Linear RegressionLogistic Regression
Predicts continuous valuesMainly predicts class probabilities
Output can be any real valueOutput lies between 0 and 1
Uses linear equationUses sigmoid/logistic function
Example: house priceExample: pass/fail
Usually estimated using least squaresCommonly estimated using maximum likelihood

Applications

  • Spam detection
  • Disease classification
  • Loan approval
  • Customer churn prediction

Conclusion

Logistic Regression is appropriate when the target is categorical, particularly for binary classification.


Bayes Theorem and Naive Bayes Classifier

Bayesian Learning uses probability theory to predict the most likely hypothesis or class based on observed data.

Bayes Theorem

where:

  • P(H|D) = posterior probability
  • P(D|H) = likelihood
  • P(H) = prior probability
  • P(D) = evidence

Naive Bayes

Naive Bayes assumes that features are conditionally independent given the class.

For features:

we calculate:

The class with the largest posterior probability is selected.

Steps

  1. Calculate prior probability for each class.
  2. Calculate conditional probabilities of features.
  3. Multiply the probabilities.
  4. Compare the values for all classes.
  5. Select the class with maximum probability.

Example

Suppose we want to classify an email as:

  • Spam
  • Not Spam

Features:

  • Contains “offer”
  • Contains “free”

Calculate:

and

The class having the larger posterior probability is selected.

Advantages

  • Simple and fast.
  • Works well with high-dimensional data.
  • Requires relatively small training data.
  • Useful for text classification.

Limitation

The assumption that features are conditionally independent is often unrealistic.

Conclusion: Naive Bayes is a simple probabilistic classifier based on Bayes’ theorem and conditional independence.


Support Vector Machine Overview

Support Vector Machine (SVM) is a supervised learning algorithm mainly used for classification.

Its objective is to find a decision boundary, called a hyperplane, that separates classes while maximizing the margin.

Hyperplane

For a linear classifier:

where:

  • w = weight vector
  • x = feature vector
  • b = bias

Support Vectors

Support vectors are the training points closest to the decision boundary.

They determine the position and orientation of the optimal separating hyperplane.


Margin

The margin is the distance between the decision boundary and the nearest training points of the classes.

SVM attempts to maximize this margin.

Kernels

For non-linearly separable data, kernel functions can map data into a higher-dimensional feature space.

Important kernels are:

1. Linear Kernel

2. Polynomial Kernel

3. Gaussian/RBF Kernel

Advantages

  • Effective in high-dimensional spaces.
  • Can model nonlinear decision boundaries using kernels.
  • Effective when classes have a clear margin.

Issues

  • Choice of kernel is important.
  • Kernel parameters need appropriate tuning.
  • Training can become expensive for very large datasets.
  • Sensitive to feature scaling.

Conclusion: SVM classifies data by finding a suitable maximum-margin decision boundary, with kernels enabling nonlinear classification.


Decision Tree Learning and ID3 Algorithm

A Decision Tree is a supervised learning model that represents decisions in the form of a tree.

  • Internal node → attribute/test
  • Branch → outcome
  • Leaf → class/prediction

ID3 Algorithm

ID3 constructs a decision tree using Information Gain.

Steps

  1. Start with the complete training dataset.
  2. Calculate entropy of the dataset.
  3. Calculate information gain for each available attribute.
  4. Select the attribute with the highest information gain.
  5. Make it the decision node.
  6. Divide the dataset according to its values.
  7. Repeat the process recursively.
  8. Stop when the data becomes sufficiently pure or no useful attribute remains.

Entropy

For a dataset S:

For binary classification:

Information Gain

where A is an attribute and Sv is the subset corresponding to value v.

Example

Suppose two attributes are:

  • Outlook
  • Wind

If:

Gain(Outlook) = 0.25

and

Gain(Wind) = 0.10

then ID3 selects:

Outlook

as the root attribute.

Advantages

  • Easy to understand.
  • Easy to visualize.
  • Can handle categorical attributes.
  • Useful for classification.

Issues

  • Can overfit.
  • Can become very large.
  • ID3 may favor attributes with many distinct values.

Conclusion: ID3 uses entropy and information gain recursively to construct a decision tree.


Entropy and Information Gain Calculation

Suppose a dataset contains:

  • 6 positive examples
  • 4 negative examples

Total: 10

Step 1: Calculate Initial Entropy

Using:

Therefore:

Suppose an attribute divides the dataset into two subsets.

Subset 1

6 records:

  • 4 positive
  • 2 negative

Subset 2

4 records:

  • 2 positive
  • 2 negative

Weighted Entropy

Information Gain

The attribute with the highest Information Gain should be selected by ID3.


KNN Algorithm Guide

k-Nearest Neighbour (k-NN) is an instance-based learning algorithm used for classification and regression.

It predicts the output of a new data point using the k closest training examples.

Algorithm

  1. Select the value of k.
  2. Calculate the distance between the query point and all training points.
  3. Sort the points according to distance.
  4. Select the k nearest points.
  5. For classification, use majority voting.
  6. For regression, commonly use the average of the neighbors’ target values.

Euclidean Distance

For two points:

and

distance is:

Example

Suppose k = 3.

The three nearest neighbors have classes:

Class A
Class A
Class B

Majority class is:

Class A

Therefore the new point is classified as Class A.

Advantages

  • Simple to understand.
  • No explicit model-training phase.
  • Can handle nonlinear decision boundaries.

Limitations

  • Prediction can be computationally expensive for large datasets.
  • Sensitive to feature scaling.
  • Choice of k affects performance.
  • Sensitive to irrelevant features and noisy data.

Conclusion: k-NN makes predictions based on local similarity and is therefore called an instance-based/lazy learning method.


Locally Weighted Regression Overview

Locally Weighted Regression (LWR) is a non-parametric regression technique in which nearby training examples receive greater importance than distant examples when making a prediction.

Basic Idea

Traditional Linear Regression attempts to fit one global model to the entire dataset.

LWR instead builds a local model around the query point.

Steps

  1. Select a query point xq.
  2. Calculate the distance between xq and training points.
  3. Assign higher weights to nearby points.
  4. Assign lower weights to distant points.
  5. Fit a weighted regression model.
  6. Use the local model to predict the output.

A commonly used weight is Gaussian:

where τ controls the width of the neighborhood.

Comparison

Linear RegressionLWR
Global modelLocal model
Same model for all pointsModel changes with query point
ParametricNon-parametric
Less computationally expensive at prediction timeMore computationally expensive at prediction time
Assumes a global relationshipCan capture local relationships

Applications

  • Nonlinear regression
  • Local trend estimation
  • Data where relationships vary across regions

Conclusion: LWR is useful when a single global regression model cannot adequately represent local patterns.

Concept Learning and Hypothesis Space

Concept Learning is the process of learning a general concept or function from a set of labeled training examples.

For example, suppose we want to learn the concept:

“A person is eligible for a loan.”

The training data contains examples classified as Yes or No.

1. Training Examples

The learner receives examples containing:

  • Input attributes
  • Corresponding target/class

2. Hypothesis

A hypothesis is a possible rule that describes the target concept.

3. Hypothesis Space

The hypothesis space is the set of all hypotheses that the learning algorithm can consider.

4. Generalization

The learner attempts to find a hypothesis that correctly classifies the training examples and also performs well on unseen examples.

5. Goal

The goal is to find:

h* ∈ H

such that h* provides the best representation of the target concept based on the available training data.

Example Data

WeatherHumidityPlay
SunnyHighNo
SunnyNormalYes
RainyNormalYes

The learner tries to discover a rule that separates Yes and No examples.

Conclusion: Concept learning forms the foundation of supervised learning by learning a general rule from labeled examples.