Friday, July 31, 2026

Understanding Limits, Continuity and Differentiability

Relationship, Hierarchy, and Conditions for Non-Differentiability
Note

The discussion below explains the relationship between limits, continuity and differentiability, together with important examples illustrating when differentiability fails. All mathematical expressions have been formatted using pure HTML so that the article remains fully compatible with Blogger without requiring MathJax, KaTeX or JavaScript.
1. Limits are not the same as derivatives

A derivative is defined using a limit, but the existence of a limit alone does not imply the existence of a derivative.

f(a) = limh→0 f(a+h) − f(a) h

Notice that this is a very specific limit (called the difference quotient).

A function may have an ordinary limit

limx→a f(x)

while the derivative does not exist.

Example
Definition
f(x) = |x|

At x = 0,

limx→0 |x| = 0

exists.

Graph of f(x) = |x|
x y (0,0)
Observation

The graph is continuous, but there is a sharp corner at the origin. This geometric feature prevents the derivative from existing at x=0 even though the limit exists.

But

limh→0 |h| − 0 h

does not exist because

One-Sided Limit Value
Left-hand limit −1
Right-hand limit 1
Important
  • Left-hand derivative = −1
  • Right-hand derivative = 1
  • Since the two one-sided derivatives are different, the derivative does not exist.

Hence

Property Result
Limit exists
Derivative exists
Conclusion

Your first statement is correct. The existence of an ordinary limit does not guarantee the existence of a derivative. A derivative requires the existence of a very specific limit—the difference quotient—which is a much stronger condition than the ordinary limit.


2. Order of Strictness

The proper hierarchy is

Differentiable->   Continuous->   Limit Exists

Every differentiable function is continuous.

Every continuous function has a limit equal to the function value.

The converse of none of these is true.

Key Observation

Differentiability is the strongest property. Continuity is weaker than differentiability. The existence of a limit is weaker than continuity. Therefore,

Limit Exists ->Continuous (Stronger) ->   Differentiable (Strongest)
Therefore

(i) Limit exists

(ii) Function may still not be continuous because

limx→a f(x)

may exist but

f(a) ≠ limx→a f(x)

or f(a) may not even exist.

Example
f(x) =
{ 1,  x = 0
0,  x ≠ 0

Here

limx→0 f(x) = 0

exists,

but

f(0) = 1

Hence not continuous.

Graph of the Piecewise Function
x y (0,1) Limit = 0
Observation

Every point except the origin lies on the horizontal line y = 0. As x approaches 0 from either side, the function approaches 0. However, the actual value of the function at the origin is f(0)=1, shown by the filled red point. Therefore, the limit exists, but the function is not continuous.

(iii) Even if continuous, the derivative may not exist.

Example
f(x) = |x|

Continuous everywhere.

Not differentiable at 0.

Geometric Interpretation
Left slope = −1 Right slope = 1
Conclusion

The function is perfectly continuous, yet its graph contains a sharp corner. The left-hand and right-hand slopes are different, so the derivative does not exist at the origin.


3. When is a Function Not Differentiable?
(i) Not Continuous

Correct.

Differentiability always implies continuity.

Differentiable  ⇒  Continuous

Therefore,

Logical Consequence

If a function is discontinuous, then it cannot be differentiable.

Summary

Discontinuous  ⇒  Not Differentiable


(ii) Sharp Corner (Corner Point)

Correct.

Example

f(x) = |x|

Left slope

= −1

Right slope

= 1

Hence the derivative does not exist.

Corner Geometry
Sharp Corner
Observation

Although the graph is continuous, the slope changes abruptly at the corner. Since the left-hand and right-hand derivatives are unequal, the derivative does not exist.


(iii) Vertical Tangent

Mostly correct.

Example

f(x) = x1/3

The derivative is

f(x) = 1 3x2/3

At

x = 0

the derivative becomes infinite.

Important

Most elementary calculus books simply state that the derivative does not exist at this point. Geometrically, however, the tangent line is vertical.

Graph of f(x) = x1/3
Vertical Tangent (0,0) x y
Geometric Interpretation

Unlike a corner, the graph remains smooth. However, the tangent becomes perfectly vertical at the origin. Since the slope tends to positive infinity, the ordinary finite derivative does not exist.


(iv) Wild Oscillation

Correct.

Example

f(x) = x sin ( 1 x )

At 0, define

f(0)=0

The function is continuous.

But

f(0)= limh→0 f(h)-f(0) h

Substituting f(h)=h sin(1/h) and f(0)=0,

f(0)= limh→0 sin ( 1 h )

does not exist because

Observation

As h approaches zero, the quantity 1/h becomes arbitrarily large. Consequently, the value of sin(1/h) oscillates endlessly between −1 and 1 without approaching any single value. Therefore, the required limit does not exist.

Graph of f(x) = x sin(1/x)
(0,0) Oscillation continues indefinitely
Conclusion

The function itself approaches the origin continuously. However, its slope oscillates infinitely rapidly near the origin. Since the derivative limit fails to converge, the derivative does not exist.


4. Examples

Let's examine each one.

(i) f(x) = |x|

Correct.

Continuous.

Not differentiable at 0.

Reason: corner.

Key Point

The function is perfectly continuous because there is no break in its graph. However, the graph changes direction abruptly at the origin. The left-hand derivative is −1 while the right-hand derivative is 1. Since these one-sided derivatives are unequal, the derivative does not exist.

Visual Interpretation of the Corner
Left slope = −1 Right slope = 1 Corner
(ii) x sin(Ï€/x)

(assuming f(0)=0).

Continuous.

Not differentiable at 0.

Reason: oscillation.

Observation

Although the function itself approaches the origin continuously, its slope oscillates increasingly rapidly near the origin. Therefore, the derivative fails to converge even though the function remains continuous.

Graph of f(x) = x sin(Ï€/x)
Oscillation near the origin

(iii) Weierstrass Function

This is one of the most famous examples.

Properties
  • Continuous everywhere.
  • Differentiable nowhere.
Illustration of a Weierstrass-Type Curve
Continuous but Nowhere Differentiable
Important Observation

The graph never contains any breaks or jumps, so it is continuous everywhere. However, no matter how much the graph is magnified, it never becomes locally straight. Consequently, no unique tangent exists at any point.


(iv) Blancmange Function

Also called the Takagi function.

Properties
  • Continuous everywhere.
  • Differentiable nowhere.
Illustration of the Blancmange (Takagi) Function
Self-Similar Fractal Structure
Observation

The Blancmange function exhibits self-similarity at different scales. Although it is continuous everywhere, it has no well-defined derivative at any point.


(v) Koch Snowflake

This one needs a small correction.

Important Correction

The Koch snowflake is not a function. Instead, it is a fractal curve. Therefore, we do not normally discuss its differentiability as a function of the form y = f(x).

More precisely,

  • The boundary of the Koch snowflake is a continuous curve.
  • It has no well-defined tangent at any point.
  • Consequently, it is regarded as a nowhere-differentiable curve.
Illustration of the Koch Snowflake Boundary
Continuous Fractal Curve
Why it is Different

Unlike the previous examples, the Koch snowflake does not represent a single-valued function. Instead, it is studied as a geometric curve. Its boundary is continuous, but nowhere smooth, making it an important example in fractal geometry.


One More Important Example

There is another classic example.

f(x) = x2/3

At x = 0, the graph has a cusp.

The derivative approaches

  • +∞ from one side,
  • −∞ from the other side.

Therefore, the derivative does not exist.

Key Observation

A cusp is fundamentally different from a corner. The graph remains continuous, but the tangent direction changes infinitely rapidly, producing infinite slopes with opposite signs.


Corner vs Cusp vs Vertical Tangent

Although all three situations result in the derivative not existing, they are geometrically very different. Understanding these differences is essential in elementary calculus.

Summary
  • Corner: Finite one-sided derivatives exist but are unequal.
  • Cusp: One-sided derivatives become infinite with opposite signs.
  • Vertical Tangent: Both one-sided derivatives approach the same infinite value.
Visual Comparison
Corner Cusp Vertical Tangent Left slope ≠ Right slope −∞ and +∞ slopes Same infinite slope
Important Distinction

Although all three graphs are continuous, the derivative fails for different geometric reasons. A corner has two different finite slopes. A cusp has two infinite slopes with opposite signs. A vertical tangent has infinite slopes of the same sign. Recognizing these differences is extremely important when studying differentiability.


Summary Hierarchy

The relationship between limits, continuity and differentiability can now be summarized as follows.

Differentiable Continuous Limit Exists

Differentiable → Continuous → Limit Exists

Key Implication

Every differentiable function is automatically continuous. Likewise, every continuous function automatically has a limit equal to the function value. However, the converse implications are not true.

Important Note

A function may possess a limit without being continuous. Similarly, a function may be continuous without being differentiable. Therefore, each property is stronger than the one below it in the hierarchy.


Converse Implications are False

The hierarchy

Differentiable  ⇒  Continuous  ⇒  Limit Exists

does not work in the reverse direction.

Differentiable Continuous Limit Exists False False
Common Misconception

Many students mistakenly assume that if a function is continuous, it must also be differentiable. The absolute value function, f(x) = |x|, is an immediate counterexample. Likewise, the existence of a limit alone does not guarantee continuity.


Complete Summary

A function may fail to be differentiable because of the following reasons.

  1. Discontinuity ✓
  2. Corner (sharp turn) ✓
  3. Cusp ✓
  4. Vertical tangent ✓
  5. Wild oscillation ✓
  6. Fractal behavior (for example, the Weierstrass or Blancmange function) ✓
Appendix A — Common Misconceptions
Misconception Reality
If the limit exists, the function is continuous. False. The function value must also equal the limit.
If a function is continuous, it is differentiable. False. The graph may contain a corner, cusp or vertical tangent.
Infinite derivative means derivative exists. False. Elementary calculus requires a finite derivative.
Every continuous curve represents a function. False. The Koch snowflake is a continuous curve but not a function y=f(x).
Appendix B — How to Test Differentiability
Does the limit exist? Is f(a)=Limit? Continuous? Compare Left and Right Derivatives If equal → Differentiable Otherwise examine Corner • Cusp • Vertical Tangent • Oscillation
Revision Cheat Sheet
  • Derivative is defined using a limit.
  • Every differentiable function is continuous.
  • Every continuous function has a limit equal to its function value.
  • The converse of both statements is false.
  • Discontinuity always destroys differentiability.
  • Corner → unequal finite slopes.
  • Cusp → opposite infinite slopes.
  • Vertical tangent → same infinite slope.
  • Oscillation → derivative limit fails to converge.
  • Weierstrass and Blancmange functions are continuous everywhere but differentiable nowhere.
  • Koch snowflake is a fractal curve, not a function.
Practice Questions
  1. Give an example of a function whose limit exists but which is not continuous.
  2. Give an example of a function that is continuous but not differentiable.
  3. Why does |x| fail to be differentiable at x = 0?
  4. Differentiate between a corner and a cusp.
  5. What is meant by a vertical tangent?
  6. Explain why x sin(1/x) is continuous but not differentiable at x = 0.
  7. State the hierarchy relating limits, continuity and differentiability.
  8. Why is the Koch snowflake not considered a function?
  9. Name two famous functions that are continuous everywhere but differentiable nowhere.
  10. Give one real-life application where differentiability is important.

Thursday, July 30, 2026

Evaluation Metrics for Classification and Regression

Evaluation metrics gauge model performance depending on whether your target is continuous (regression) or categorical (classification). Regression metrics measure the magnitude of prediction errors, while classification metrics measure the correctness of category assignments and the trade-offs between different error types.

Regression Metrics

These measure prediction errors for continuous values:

  • MAE (Mean Absolute Error): Average absolute error; robust to outliers.
  • MSE (Mean Squared Error): Average squared error; heavily penalizes outliers.
  • RMSE (Root Mean Squared Error): Square root of MSE; interprets in target units while penalizing large errors.
  • R² (R-Squared): Measures variance explained by the model, typically ranging from 0 to 1.
Classification Metrics

These evaluate predictions of distinct, categorical labels:

  • Accuracy: Overall correct predictions; best for balanced classes.
  • Precision: Focuses on correctness of positive predictions; crucial when false positives are costly.
  • Recall (Sensitivity): Focuses on finding all actual positives; critical when false negatives are dangerous.
  • F1-Score: Harmonic mean of precision and recall for imbalanced data.
  • AUC-ROC: Measures discrimination capability across thresholds.

Wednesday, July 29, 2026

Example : Finding weights and biases using Normal Equation

Here is how the Normal Equation finds the perfect line for your data without iterative guessing.

1. Define the Scenario

Imagine we want to predict a house's Price (in thousands of dollars) based on its Size (in hundreds of square feet). We have three data points:

  • House 1: Size = 1, Price = 15
  • House 2: Size = 2, Price = 20
  • House 3: Size = 3, Price = 25

We want to fit a straight line equation:

Price = w × Size + b
2. Construct the Matrices

To use the Normal Equation, we must arrange our data into a feature matrix X and a target vector y.

To calculate the bias (b), we add a dummy feature column filled with 1s to our matrix X.

X =
1 1
1 2
1 3
,
y =
15
20
25

In matrix X, the first column represents the bias multiplier (1). The second column represents the house sizes.

3. Apply the Normal Equation

The Normal Equation solves for the parameter vector θ using:

θ = (XTX)−1 XT y

First, we calculate the transpose XT:

XT =
1 1 1
1 2 3

Next, we multiply XT by X:

XTX =
1 1 1
1 2 3
×
1 1
1 2
1 3
=
3 6
6 14
Next, we find the inverse of (XTX):
(XTX)−1 =
1
(3 × 14) − (6 × 6)
×
14 −6
−6 3


=
1
6
×
14 −6
−6 3


=
7
3
−1
−1
1
2
Next, we multiply XT by y:
XTy =
1 1 1
1 2 3
×
15
20
25
=
60
130
Finally, we multiply (XTX)−1 by XTy to get θ:
θ =
7
3
−1
−1
1
2
×
60
130

=
(7/3 × 60) + (−1 × 130)
(−1 × 60) + (1/2 × 130)


=
140 − 130
−60 + 65


=
10
5
✅ Final Answer

The Normal Equation outputs θ =

10
5

meaning the optimal bias (b) is 10 and the weight (w) is 5, yielding the perfect prediction line:

Price = 5 × Size + 10

Why the Dummy Feature Column is Always Filled with 1s
Important Note

Note that the dummy feature column is always initialized to all 1s.

Why It Must Be All 1s
  • Isolates the Bias: Any number multiplied by 1 remains itself.
b × 1 = b
  • Maintains Consistency: It ensures the bias term is added equally to every single data point.
  • Allows Matrix Math: It turns the separate addition step (wX + b) into a clean dot product multiplication.
What Happens With Other Numbers?
  • If you used 0s: The bias term b would be multiplied by 0 and completely disappear from the equation.
  • If you used 2s: The math would still solve, but the final output in your θ vector would be
b
2

instead of the actual bias.

Key Takeaway

By keeping it as a column of 1s, the Normal Equation can treat the bias exactly like any other weight weight, saving you from writing separate algebraic code for it.

The "Penalty" and "C" values in scikit-learn's Logistic Regression

In scikit-learn, the penalty and C parameters modify the standard cost function (loss function) of the LogisticRegression class during the optimization process. They do not change the core predictive sigmoid formula itself, but they directly dictate how the model weights (w) are calculated during training.

The core math breaks down as follows:

1. The Standard Cost Function (No Regularization)

Without any penalty (penalty=None), a binary logistic regression model optimizes the standard log-loss (binary cross-entropy) cost function, denoted as L(w):

L(w) = − Σi=1n  [ yi log(Å·i) + (1 − yi) log(1 − Å·i) ]

Where:

  • yi is the true label (0 or 1).
  • Å·i is the predicted probability generated by the sigmoid function:
1
1 + e−wTxi
2. How penalty Alters the Formula

The penalty parameter adds a regularization term to the cost function to restrict the size of the coefficients and prevent overfitting. It determines the shape of this constraint:

penalty='l2' (Ridge, Default)

Adds the squared magnitude of coefficients (L2 norm).

Penalty Term = ½ ‖w‖22 = ½ Σj=1m wj2
penalty='l1' (Lasso)

Adds the absolute magnitude of coefficients (L1 norm), which drives unimportant weights to exactly zero to create a sparse model.

Penalty Term = ‖w‖1 = Σj=1m |wj|
3. How C Alters the Formula

In traditional statistics, regularization strength is controlled by a parameter named λ (lambda), where a larger λ means more aggressive penalty enforcement:

Traditional Cost = L(w) + λ × (Penalty Term)

Instead of λ, scikit-learn uses C, which is defined as the inverse of regularization strength:

C = 1/λ

Crucially, scikit-learn multiplies C by the loss term rather than the penalty term. The objective function that scikit-learn minimizes looks like this:

Important
Scikit-Learn Cost = C × L(w) + Penalty Term
Directly Comparing the Behavior of C
Value of C Regularization Strength Effect on the Objective Function Model Behavior
Small C (e.g., 0.01) High The penalty term dominates; the data loss L(w) is largely ignored. Coefficients shrink heavily toward zero; protects against overfitting.
Large C (e.g., 1000) Low The data loss L(w) dominates; the penalty term is ignored. Model prioritizes fitting the training data exactly; risks overfitting.
Key Observation

Although C is defined as the inverse of the regularization strength (C = 1 / λ), scikit-learn's optimization objective is implemented as:

Scikit-Learn Cost = C × L(w) + Penalty Term

As a result:

  • A smaller C places relatively less emphasis on fitting the training data and therefore results in stronger regularization.
  • A larger C places greater emphasis on minimizing the data loss, resulting in weaker regularization.

Decision Trees - II

Topic 2: Basic Terminology of Decision Trees

This topic forms the foundation of everything that follows in Decision Trees. Once you understand this terminology, advanced topics such as splitting criteria, tree pruning, feature importance, Random Forests, and Gradient Boosting become much easier to understand.

Why Do We Need This Terminology?

Every field has its own vocabulary.

For example:

  • Biology uses terms such as cell, tissue, and organ.
  • Computer Networks use terms like router, switch, and gateway.
  • Decision Trees have their own terminology that describes the structure of the tree and how predictions are made.

Throughout the remaining topics, we'll repeatedly use these terms. Therefore, it's essential to become familiar with them now.

Topics Covered

In this lesson, we'll study the following terms:

  • Root Node
  • Decision Node
  • Leaf (Terminal) Node
  • Parent Node
  • Child Node
  • Branch
  • Subtree
  • Depth of a Tree
  • Height of a Tree
  • Levels

Example Decision Tree

We'll use the same Decision Tree throughout this topic so that each term becomes easy to understand.

This tree predicts whether a customer will buy a product.

Age ≥ 30? / \ No Yes │ │ Don't Buy Income ≥ £100k? / \ Yes No │ │ Buy Married? / \ Yes No │ │ Buy Don't Buy
Observation

Every prediction starts at the top of the tree and follows one path until a final decision is reached.

1. Root Node

Definition

The Root Node is the topmost node of a Decision Tree. It is the starting point for every prediction made by the model.

In Our Example

Age ≥ 30?

This question is the Root Node because it is the first question every customer must answer.

Why Is the Root Node Important?

Every prediction begins at the Root Node.

Suppose a customer has the following details:

  • Age = 35 years
  • Income = £120,000
  • Married = Yes

The algorithm does not immediately examine the customer's income or marital status.

Instead, it always begins by asking:

Is Age ≥ 30?

Only after answering this question does the algorithm move to the next appropriate branch of the tree.

Characteristics of the Root Node

  • There is exactly one Root Node in every Decision Tree.
  • The Root Node has no parent.
  • Every prediction always starts from the Root Node.
  • It usually contains the most informative feature selected by the learning algorithm.

How Is the Root Node Chosen?

The Root Node is not selected randomly.

During training, the Decision Tree algorithm evaluates all available features and chooses the one that produces the best separation of the training data.

We'll Learn This Soon

In later topics, we'll discover how algorithms such as ID3, C4.5, and CART mathematically decide which feature deserves to become the Root Node using concepts such as Entropy, Information Gain, and Gini Impurity.

Part 1 Summary

  • A Decision Tree has its own structural vocabulary.
  • We'll use one example tree throughout this topic.
  • The Root Node is the topmost node and the starting point of every prediction.
  • Every Decision Tree has exactly one Root Node.
  • The Root Node usually contains the feature that best separates the training data.

2. Decision Node

Definition

A Decision Node is any node in a Decision Tree that asks a question and divides the data into two or more branches.

Every time the tree reaches a Decision Node, it evaluates a condition based on one of the input features and decides which branch should be followed next.

Decision Nodes in Our Example Tree

Consider the Decision Tree introduced earlier.

Age ≥ 30? / \ No Yes │ │ Don't Buy Income ≥ £100k? / \ Yes No │ │ Buy Married? / \ Yes No │ │ Buy Don't Buy

The following nodes ask questions and therefore are Decision Nodes:

  • Age ≥ 30?
  • Income ≥ £100k?
  • Married?

Function of a Decision Node

The primary purpose of a Decision Node is to partition (split) the dataset into smaller and more homogeneous groups.

For example, suppose the tree asks:

Income ≥ £100,000?

This single question divides all customers into two groups:

Group 1

Customers with an income greater than or equal to £100,000.

Group 2

Customers with an income below £100,000.

This repeated splitting of data is the core mechanism behind every Decision Tree.

3. Leaf Node (Terminal Node)

Definition

A Leaf Node is a node that does not split any further.

Instead of asking another question, it stores the final prediction made by the Decision Tree.

Leaf Nodes in Our Example Tree

Age ≥ 30? / \ No Yes │ │ ► Don't Buy Income ≥ £100k? / \ Yes No │ │ ► Buy Married? / \ Yes No │ │ ► Buy ► Don't Buy

The highlighted predictions are the Leaf Nodes.

  • Don't Buy
  • Buy
  • Buy
  • Don't Buy

Why Are They Called Terminal Nodes?

The word terminal means ending point.

Once the algorithm reaches a Leaf Node, the prediction process stops. No further questions are asked.

Example

Suppose a customer has:

  • Age = 25 years

The prediction path becomes:

Age ≥ 30? │ No │ Don't Buy

Since the tree has reached a Leaf Node, the prediction is complete.

Leaf Nodes in Classification Trees

In a Classification Tree, each Leaf Node contains a class label.

Typical examples include:

Yes
No
Spam
Not Spam
Approve
Reject

Leaf Nodes in Regression Trees

In a Regression Tree, the Leaf Node stores a numerical value instead of a category.

Example Prediction

Predicted House Price = £450,000

Therefore, the contents of a Leaf Node depend on whether the Decision Tree is solving a classification problem or a regression problem.

Decision Node vs. Leaf Node

Decision Node Leaf Node
Asks a question. Stores the final prediction.
Splits the dataset. Does not split further.
Has one or more child nodes. Has no children.
Represents a decision. Represents the final outcome.

Part 2 Summary

  • A Decision Node asks a question and splits the dataset into smaller groups.
  • The repeated splitting of data is the core mechanism of every Decision Tree.
  • A Leaf (Terminal) Node stores the final prediction and does not split further.
  • Classification Trees store class labels in Leaf Nodes, while Regression Trees store numerical values.
  • The prediction process always ends at a Leaf Node.

4. Parent Node

Definition

A Parent Node is any node that has one or more child nodes.

Whenever a node branches into other nodes, it automatically becomes the parent of those nodes.

Parent Relationship in Our Example

Consider the following portion of our Decision Tree:

Age ≥ 30? / \ No Yes │ │ Don't Buy Income ≥ £100k?

Here, Age ≥ 30? is the Parent Node because it has two children:

  • Don't Buy
  • Income ≥ £100k?

Another Example

Now consider another part of the same tree:

Income ≥ £100k? / \ Yes No │ │ Buy Married? / \ Yes No │ │ Buy Don't Buy

Here, Income ≥ £100k? is also a Parent Node because it has two children:

  • Buy
  • Married?

Important Observation

A node is called a Parent Node only because it has child nodes.

Whether a node is a parent has nothing to do with its position in the tree. It depends entirely on whether it has descendants.

Remember

Every internal decision node is usually a Parent Node because it continues the prediction process by creating one or more branches.

Can a Node Be Both a Parent and a Child?

Yes.

This is one of the most important concepts to understand.

A node can simultaneously be:

  • a child of another node, and
  • a parent of additional nodes.

Example

Age ≥ 30? │ ▼ Income ≥ £100k? / \ Buy Married?

Here, Income ≥ £100k? plays two different roles:

  • It is a child of Age ≥ 30?.
  • It is a parent of Buy and Married?.

5. Child Node

Definition

A Child Node is any node that descends from a Parent Node.

Every node except the Root Node has exactly one parent.

Child Relationship in Our Example

Income ≥ £100k? / \ Buy Married?

Both Buy and Married? are Child Nodes because they originate from the Parent Node Income ≥ £100k?.

Family Tree Analogy

One of the easiest ways to remember these terms is to compare a Decision Tree with a family tree.

Family Tree Decision Tree
Grandparent Root Node
Parent Decision Node
Child Leaf Node (or another Decision Node)
Memory Tip

Just like people can be both someone's child and someone else's parent, an internal node in a Decision Tree can simultaneously be a Child Node and a Parent Node.

Quick Practice

Consider the following simplified tree:

Weather? / \ Sunny Rainy │ │ Don't Play Play

Can you identify the different node types?

  • Weather? → Root Node and Parent Node
  • Don't Play → Child Node and Leaf Node
  • Play → Child Node and Leaf Node

Part 3 Summary

  • A Parent Node has one or more child nodes.
  • A Child Node descends from a Parent Node.
  • Every node except the Root Node has exactly one parent.
  • An internal Decision Node can be both a Parent Node and a Child Node at the same time.
  • The family tree analogy is an excellent way to remember these relationships.

6. Branch

Definition

A Branch is the connection between a Parent Node and one of its Child Nodes.

Every branch represents the outcome of a decision.

Once a question is answered, the algorithm follows the corresponding branch to continue making the prediction.

Branches in Our Example Tree

Consider the Root Node:

Age ≥ 30? / \ No / \ Yes / \ Don't Buy Income ≥ £100k?

Here, there are two branches:

  • Left Branch → Answer = No
  • Right Branch → Answer = Yes

How Should We Interpret a Branch?

A branch represents a decision path.

For example, suppose the algorithm follows this branch:

Age ≥ 30? → Yes

This means:

"Continue analysing only those customers whose age is at least 30 years."

Every branch narrows down the dataset by applying one additional rule.

A Complete Decision Path

Multiple branches together form a Decision Path.

Age ≥ 30? │ Yes │ Income ≥ £100k? │ Yes │ Buy

This complete sequence of branches means:

Customer is at least 30 years old and earns at least £100,000 → Predict Buy

7. Subtree

Definition

A Subtree is any portion of a Decision Tree that itself forms a complete tree.

In other words, if you select any node together with all of its descendants, the resulting structure is called a Subtree.

Example of a Subtree

Consider the following part of our Decision Tree:

Income ≥ £100k? / \ Buy Married? / \ Yes No │ │ Buy Don't Buy

This entire structure is a Subtree.

Notice that it has its own root (Income ≥ £100k?), branches, internal nodes and leaf nodes.

Why Are Subtrees Important?

Subtrees play a major role in many Decision Tree algorithms and advanced machine learning techniques.

Technique How Subtrees Are Used
Tree Pruning Removes unnecessary subtrees to reduce overfitting.
Random Forest Builds many independent trees, each containing numerous subtrees.
Gradient Boosting Sequentially grows trees and improves predictions by learning from previous trees.

Visualising a Subtree

Imagine removing everything above the highlighted node. The remaining structure is still a valid Decision Tree.

Original Tree Age ≥ 30? │ ▼ Income ≥ £100k? │ ▼ Married?
Everything beginning at Income ≥ £100k? forms one complete subtree.

Real-Life Analogy

Think of a company organisation chart.

Chief Executive Officer (CEO)

├── Sales Department

├── Finance Department

└── Engineering Department

If we focus only on the Engineering Department together with all its teams, we obtain a smaller organisation chart.

That smaller organisation chart is analogous to a Subtree in a Decision Tree.

Part 4 Summary

  • A Branch connects a Parent Node to one of its Child Nodes.
  • Every branch represents the outcome of a decision.
  • Several branches together form a Decision Path.
  • A Subtree is any node together with all of its descendants.
  • Subtrees are fundamental to tree pruning, Random Forests and Gradient Boosting algorithms.

8. Depth of a Tree

Definition

The depth of a node is the number of edges between the Root Node and that particular node.

In simple terms, depth tells us how far a node is from the Root Node.

Important Rule

We count edges (connections), not nodes.

Example

Consider the following simplified Decision Tree.

Age ≥ 30? / \ Don't Buy Income ≥ £100k? / \ Buy Married? | Don't Buy

Let's determine the depth of each node.

Counting the Edges

Start from the Root Node and count the number of connections required to reach each node.

Root Node

Age ≥ 30?

Number of edges from the root: 0


Income ≥ £100k?

Age ≥ 30? │ Income ≥ £100k?

Number of edges: 1


Married?

Age ≥ 30? │ Income ≥ £100k? │ Married?

Number of edges: 2

Depth of Individual Nodes

Node Number of Edges from Root Depth
Age ≥ 30? 0 0
Don't Buy 1 1
Income ≥ £100k? 1 1
Buy 2 2
Married? 2 2
Don't Buy 3 3

Depth of the Entire Tree

The depth of a Decision Tree is defined as the maximum depth among all of its nodes.

In other words, find the deepest Leaf Node and count how many edges separate it from the Root Node.

Example

Suppose the deepest Leaf Node has a depth of 3.

Tree Depth = 3

Why Is Tree Depth Important?

Tree depth directly affects how complex a Decision Tree becomes.

Tree Depth Interpretation
Small Depth Simpler model that is easier to understand but may underfit the data.
Large Depth More powerful model that may capture complex patterns but can overfit the training data.

Choosing an appropriate tree depth is therefore an important part of building an effective Decision Tree model.

Common Beginner Mistakes

  • Counting nodes instead of edges.
  • Assuming the Root Node has a depth of 1. It always has a depth of 0.
  • Confusing the depth of a single node with the depth of the entire tree.

Interview Tip

Question:

What is the depth of the Root Node?

Answer:

The Root Node always has a depth of 0 because there are no edges between the Root Node and itself.

Part 5 Summary

  • The depth of a node is the number of edges between the Root Node and that node.
  • The Root Node always has a depth of 0.
  • The depth of a tree is the maximum depth among all of its nodes.
  • Tree depth influences model complexity and the risk of overfitting.
  • Always count edges, not nodes.

9. Height of a Tree

Definition

The height of a node is the number of edges on the longest path from that node to any Leaf Node.

Unlike Depth, which measures the distance from the Root Node downward, Height measures the distance from a node downward to its deepest descendant.

Remember

Height is always measured downward towards the Leaf Nodes.

Example

Consider the following Decision Tree.

Age ≥ 30? / \ Don't Buy Income ≥ £100k? / \ Buy Married? | Don't Buy

Let's calculate the height of each node.

Calculating Height

Leaf Node

A Leaf Node has no descendants.

Therefore,

Height = 0

Income ≥ £100k?

The longest path from this node reaches Married? and then the final Don't Buy leaf.

Height = 2

Age ≥ 30?

This is the Root Node.

The longest path to a Leaf Node contains 3 edges.

Height = 3

Height of Individual Nodes

Node Height
Don't Buy (Leaf) 0
Buy (Leaf) 0
Married? 1
Income ≥ £100k? 2
Age ≥ 30? (Root) 3

Height of the Entire Tree

The height of a Decision Tree is simply the height of its Root Node.

Example

If the Root Node has a height of 3, then the entire Decision Tree also has a height of 3.

Depth vs Height

Beginners often confuse these two terms because both involve counting edges.

Depth Height
Measured from the Root Node. Measured towards the Leaf Nodes.
Counts edges from Root to the current node. Counts edges from the current node to the deepest leaf.
Root Node always has depth 0. Every Leaf Node always has height 0.

10. Levels

Definition

A Level groups together all nodes having the same Depth.

Therefore:

Nodes at the same depth belong to the same level.

Example

Level 0 Age ≥ 30? ────────────── Level 1 Don't Buy Income ≥ £100k? ────────────── Level 2 Buy Married? ────────────── Level 3 Don't Buy

Nodes at Each Level

Level Nodes
0 Age ≥ 30?
1 Don't Buy, Income ≥ £100k?
2 Buy, Married?
3 Don't Buy

Part 6 Summary

  • Height measures the distance from a node to its deepest Leaf Node.
  • The height of every Leaf Node is 0.
  • The height of the entire tree equals the height of the Root Node.
  • Depth is measured from the Root, whereas Height is measured towards the deepest Leaf.
  • A Level contains all nodes having the same depth.

Putting Everything Together

We have now learned all the important structural terms used in a Decision Tree.

Let's combine everything into a single labelled tree so that you can clearly see how every concept fits together.

Fully Labelled Decision Tree

Level 0 Age ≥ 30? (Root Node) Depth = 0 Height = 3 ├── No ─────────────► Don't Buy │ (Leaf Node) │ Level 1 │ Depth = 1 │ Height = 0 │ └── Yes │ ▼ Income ≥ £100k? (Decision Node) Level 1 Depth = 1 Height = 2 ├── Yes ───────────► Buy │ (Leaf Node) │ Level 2 │ Depth = 2 │ Height = 0 │ └── No │ ▼ Married? (Decision Node) Level 2 Depth = 2 Height = 1 ├── Yes ───────────► Buy │ (Leaf Node) │ Level 3 │ Height = 0 │ └── No ───────────► Don't Buy (Leaf Node) Level 3 Height = 0

Identifying Every Component

Component Example in Our Tree
Root Node Age ≥ 30?
Decision Nodes Income ≥ £100k?, Married?
Leaf Nodes Buy, Don't Buy
Parent Node Income ≥ £100k?
Child Nodes Buy, Married?
Branches Yes / No connections between nodes
Subtree Entire tree beginning at "Income ≥ £100k?"

What is a Decision Path?

A Decision Path is the complete sequence of branches followed from the Root Node to a Leaf Node.

Every prediction produced by a Decision Tree follows exactly one decision path.

Age ≥ 30?
      │
     Yes
      │
Income ≥ £100k?
      │
     Yes
      │
     Buy

This entire route is one complete Decision Path.

Prediction Walkthrough

Suppose a customer has the following information:

  • Age = 40 years
  • Income = £120,000
  • Married = Yes

Step-by-Step Prediction

  1. Start at the Root Node.
  2. Ask: Age ≥ 30?

    Answer: Yes
  3. Move along the Yes Branch to Income ≥ £100k?
  4. Ask: Income ≥ £100k?

    Answer: Yes
  5. Follow the Yes Branch to the Leaf Node.
  6. Final Prediction:
    Buy

Another Example

Consider another customer.

  • Age = 24 years
  • Income = £90,000
  • Married = No

Prediction:

Age ≥ 30? Answer → No ↓ Leaf Node ↓ Don't Buy

Notice that once the tree reaches a Leaf Node, no additional questions are asked.

Overall Prediction Flow

Start

↓

Root Node

↓

Decision Node

↓

Decision Node

↓

...

↓

Leaf Node

↓

Prediction

Every prediction generated by a Decision Tree follows exactly this workflow.

One-Minute Revision

Term Quick Memory Trick
Root Node Starting point
Decision Node Asks a question
Leaf Node Final prediction
Branch Connection between nodes
Subtree A smaller tree inside the main tree
Depth Distance from the Root Node
Height Distance to the deepest Leaf Node
Level All nodes with the same depth

Frequently Asked Interview Questions

Interviewers frequently ask these fundamental questions to check whether you understand the basic structure and terminology of a Decision Tree. These concepts are essential before moving on to more advanced topics such as splitting criteria, pruning and ensemble learning.

Question 1

What is the difference between a Decision Node and a Leaf Node?

Answer

A Decision Node asks a question about one of the input features and splits the dataset into two or more branches.

A Leaf Node, also called a Terminal Node, does not ask any further questions. Instead, it stores the final prediction produced by the model.

Decision Node Leaf Node
Asks a question Gives the final prediction
Splits the data Does not split further

Question 2

Can a node be both a Parent Node and a Child Node?

Answer

Yes.

Internal Decision Nodes usually play both roles simultaneously.

  • They are the Child Node of the node above them.
  • They are also the Parent Node of the nodes below them.
Age ≥ 30? │ Income ≥ £100k? │ Married?

In this example, Income ≥ £100k? is:

  • Child of Age ≥ 30?
  • Parent of Married?

Question 3

What is the depth of the Root Node?

Answer

The Root Node always has a depth of 0.

This is because there are zero edges between the Root Node and itself.

Quick Memory Tip

Root → Depth = 0

Question 4

What is a Subtree?

Answer

A Subtree is any node together with all of its descendants.

Every subtree is itself a valid Decision Tree containing its own:

  • Root Node
  • Branches
  • Decision Nodes
  • Leaf Nodes

Quick-Fire Interview Round

Question Expected Answer
Where does every prediction start? Root Node
Which node stores the final prediction? Leaf Node
What connects two nodes? Branch
What is the depth of the Root Node? 0
What is the height of every Leaf Node? 0
What forms a Decision Path? Sequence of branches from Root to Leaf

Understanding Limits, Continuity and Differentiability

Relationship, Hierarchy, and Conditions for Non-Differentiability Note The discussion below explains the relationship between lim...