> For the complete documentation index, see [llms.txt](https://maheshwarappa-a.gitbook.io/data-science-interview/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://maheshwarappa-a.gitbook.io/data-science-interview/machine-learning/decision-tree.md).

# Decision Tree

## 1.**Explain Decision Tree algorithm**&#x20;

A decision tree is a supervised machine learning algorithm mainly used for **Regression and Classification**. It breaks down a data set into smaller and smaller subsets while at the same time an associated decision tree is incrementally developed. The final result is a tree with decision nodes and leaf nodes. A decision tree can handle both categorical and numerical data.

![](https://54486267-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MJ5sRvnbi6ybKs7Ho3A%2F-MJMQYXppML-OY7oJV0L%2F-MJMYP2iMklIs7txNFuI%2Fimage.png?alt=media\&token=dfc6c517-7097-4f60-b945-2c3a32f2983f)

## 2.**What are Entropy and Information gain in Decision tree algorithm?**

**Entropy:** Entropy in Decision Tree stands for homogeneity. If the data is completely homogenous, the entropy is 0, else if the data is divided (50-50%) entropy is 1.

![](https://54486267-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MJ5sRvnbi6ybKs7Ho3A%2F-MJMQYXppML-OY7oJV0L%2F-MJM_qNgt_FJoHMftXBE%2Fimage.png?alt=media\&token=3e9a4b63-5753-4b81-9521-638b1712a3fa)

**Information Gain:** Information Gain is the decrease/increase in Entropy value when the node is split.

![](https://54486267-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-MJ5sRvnbi6ybKs7Ho3A%2F-MJMQYXppML-OY7oJV0L%2F-MJM_uCMJlkOSTzhM9qm%2Fimage.png?alt=media\&token=62b5b0f4-69c3-4c6c-8e37-a89cd84a9a78)

An attribute should have the highest information gain to be selected for splitting

## **3.What is pruning in Decision Tree?**

**Pruning** is a technique in machine learning and search algorithms that reduces the size of **decision trees** by removing sections of the **tree** that provide little power to classify instances. So, when we remove sub-nodes of a decision node, this process is called  pruning[ ](https://www.edureka.co/blog/implementation-of-decision-tree/)or opposite process of splitting.
