Comprehensive Machine Learning Algorithms Tutorial
Linear Regression and Working Principle
Linear Regression is a supervised learning algorithm used to predict a continuous dependent variable from one or more independent variables.
For Simple Linear Regression, the relationship between x and y is represented by:
where:
- y = predicted value
- x = input feature
- b0 = intercept
- b1 = slope
Calculation of Slope
Intercept
Therefore, the regression equation becomes:
Working
- Collect training data.
- Assume a linear relationship between input and output.
- Calculate the regression coefficients.
- Fit the best line through the data.
- Use the equation to predict unknown values.
Example
Consider:
(1,2), (2,4), (3,6), (4,8)
The relationship is:
Assumptions
Important assumptions include:
- Linear relationship between variables.
- Errors are independent.
- Constant error variance.
- Errors are appropriately modeled, commonly with an approximately normal distribution for inference.
Complexity
For fitting ordinary least squares with one feature, the computational cost is generally low and depends on the implementation; prediction for one observation is:
O(1)
Conclusion
Linear Regression is mainly used for predicting continuous numerical values such as price, salary, demand, or temperature.
Logistic Regression vs Linear Regression
Logistic Regression is a supervised learning algorithm mainly used for classification, especially binary classification.
Instead of directly predicting a continuous value, it predicts the probability that an observation belongs to a class.
Logistic and Sigmoid Function
The sigmoid function is:
where:
Therefore:
The output lies between:
Classification
A threshold such as 0.5 may be used:
Example
Suppose a model predicts:
P(pass) = 0.82
Using threshold 0.5:
0.82 > 0.5
Therefore, the prediction is:
Pass
Linear vs Logistic Regression Comparison
| Linear Regression | Logistic Regression |
|---|---|
| Predicts continuous values | Mainly predicts class probabilities |
| Output can be any real value | Output lies between 0 and 1 |
| Uses linear equation | Uses sigmoid/logistic function |
| Example: house price | Example: pass/fail |
| Usually estimated using least squares | Commonly estimated using maximum likelihood |
Applications
- Spam detection
- Disease classification
- Loan approval
- Customer churn prediction
Conclusion
Logistic Regression is appropriate when the target is categorical, particularly for binary classification.
Bayes Theorem and Naive Bayes Classifier
Bayesian Learning uses probability theory to predict the most likely hypothesis or class based on observed data.
Bayes Theorem
where:
- P(H|D) = posterior probability
- P(D|H) = likelihood
- P(H) = prior probability
- P(D) = evidence
Naive Bayes
Naive Bayes assumes that features are conditionally independent given the class.
For features:
we calculate:
The class with the largest posterior probability is selected.
Steps
- Calculate prior probability for each class.
- Calculate conditional probabilities of features.
- Multiply the probabilities.
- Compare the values for all classes.
- Select the class with maximum probability.
Example
Suppose we want to classify an email as:
- Spam
- Not Spam
Features:
- Contains “offer”
- Contains “free”
Calculate:
and
The class having the larger posterior probability is selected.
Advantages
- Simple and fast.
- Works well with high-dimensional data.
- Requires relatively small training data.
- Useful for text classification.
Limitation
The assumption that features are conditionally independent is often unrealistic.
Conclusion: Naive Bayes is a simple probabilistic classifier based on Bayes’ theorem and conditional independence.
Support Vector Machine Overview
Support Vector Machine (SVM) is a supervised learning algorithm mainly used for classification.
Its objective is to find a decision boundary, called a hyperplane, that separates classes while maximizing the margin.
Hyperplane
For a linear classifier:
where:
- w = weight vector
- x = feature vector
- b = bias
Support Vectors
Support vectors are the training points closest to the decision boundary.
They determine the position and orientation of the optimal separating hyperplane.
Margin
The margin is the distance between the decision boundary and the nearest training points of the classes.
SVM attempts to maximize this margin.
Kernels
For non-linearly separable data, kernel functions can map data into a higher-dimensional feature space.
Important kernels are:
1. Linear Kernel
2. Polynomial Kernel
3. Gaussian/RBF Kernel
Advantages
- Effective in high-dimensional spaces.
- Can model nonlinear decision boundaries using kernels.
- Effective when classes have a clear margin.
Issues
- Choice of kernel is important.
- Kernel parameters need appropriate tuning.
- Training can become expensive for very large datasets.
- Sensitive to feature scaling.
Conclusion: SVM classifies data by finding a suitable maximum-margin decision boundary, with kernels enabling nonlinear classification.
Decision Tree Learning and ID3 Algorithm
A Decision Tree is a supervised learning model that represents decisions in the form of a tree.
- Internal node → attribute/test
- Branch → outcome
- Leaf → class/prediction
ID3 Algorithm
ID3 constructs a decision tree using Information Gain.
Steps
- Start with the complete training dataset.
- Calculate entropy of the dataset.
- Calculate information gain for each available attribute.
- Select the attribute with the highest information gain.
- Make it the decision node.
- Divide the dataset according to its values.
- Repeat the process recursively.
- Stop when the data becomes sufficiently pure or no useful attribute remains.
Entropy
For a dataset S:
For binary classification:
Information Gain
where A is an attribute and Sv is the subset corresponding to value v.
Example
Suppose two attributes are:
- Outlook
- Wind
If:
Gain(Outlook) = 0.25
and
Gain(Wind) = 0.10
then ID3 selects:
Outlook
as the root attribute.
Advantages
- Easy to understand.
- Easy to visualize.
- Can handle categorical attributes.
- Useful for classification.
Issues
- Can overfit.
- Can become very large.
- ID3 may favor attributes with many distinct values.
Conclusion: ID3 uses entropy and information gain recursively to construct a decision tree.
Entropy and Information Gain Calculation
Suppose a dataset contains:
- 6 positive examples
- 4 negative examples
Total: 10
Step 1: Calculate Initial Entropy
Using:
Therefore:
Suppose an attribute divides the dataset into two subsets.
Subset 1
6 records:
- 4 positive
- 2 negative
Subset 2
4 records:
- 2 positive
- 2 negative
Weighted Entropy
Information Gain
The attribute with the highest Information Gain should be selected by ID3.
KNN Algorithm Guide
k-Nearest Neighbour (k-NN) is an instance-based learning algorithm used for classification and regression.
It predicts the output of a new data point using the k closest training examples.
Algorithm
- Select the value of k.
- Calculate the distance between the query point and all training points.
- Sort the points according to distance.
- Select the k nearest points.
- For classification, use majority voting.
- For regression, commonly use the average of the neighbors’ target values.
Euclidean Distance
For two points:
and
distance is:
Example
Suppose k = 3.
The three nearest neighbors have classes:
Class A Class A Class B
Majority class is:
Class A
Therefore the new point is classified as Class A.
Advantages
- Simple to understand.
- No explicit model-training phase.
- Can handle nonlinear decision boundaries.
Limitations
- Prediction can be computationally expensive for large datasets.
- Sensitive to feature scaling.
- Choice of k affects performance.
- Sensitive to irrelevant features and noisy data.
Conclusion: k-NN makes predictions based on local similarity and is therefore called an instance-based/lazy learning method.
Locally Weighted Regression Overview
Locally Weighted Regression (LWR) is a non-parametric regression technique in which nearby training examples receive greater importance than distant examples when making a prediction.
Basic Idea
Traditional Linear Regression attempts to fit one global model to the entire dataset.
LWR instead builds a local model around the query point.
Steps
- Select a query point xq.
- Calculate the distance between xq and training points.
- Assign higher weights to nearby points.
- Assign lower weights to distant points.
- Fit a weighted regression model.
- Use the local model to predict the output.
A commonly used weight is Gaussian:
where τ controls the width of the neighborhood.
Comparison
| Linear Regression | LWR |
|---|---|
| Global model | Local model |
| Same model for all points | Model changes with query point |
| Parametric | Non-parametric |
| Less computationally expensive at prediction time | More computationally expensive at prediction time |
| Assumes a global relationship | Can capture local relationships |
Applications
- Nonlinear regression
- Local trend estimation
- Data where relationships vary across regions
Conclusion: LWR is useful when a single global regression model cannot adequately represent local patterns.
Concept Learning and Hypothesis Space
Concept Learning is the process of learning a general concept or function from a set of labeled training examples.
For example, suppose we want to learn the concept:
“A person is eligible for a loan.”
The training data contains examples classified as Yes or No.
1. Training Examples
The learner receives examples containing:
- Input attributes
- Corresponding target/class
2. Hypothesis
A hypothesis is a possible rule that describes the target concept.
3. Hypothesis Space
The hypothesis space is the set of all hypotheses that the learning algorithm can consider.
4. Generalization
The learner attempts to find a hypothesis that correctly classifies the training examples and also performs well on unseen examples.
5. Goal
The goal is to find:
h* ∈ H
such that h* provides the best representation of the target concept based on the available training data.
Example Data
| Weather | Humidity | Play |
|---|---|---|
| Sunny | High | No |
| Sunny | Normal | Yes |
| Rainy | Normal | Yes |
The learner tries to discover a rule that separates Yes and No examples.
Conclusion: Concept learning forms the foundation of supervised learning by learning a general rule from labeled examples.
