{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Random Forest Model (Clearly Explained !!)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"Hello kagglers, Welcome in this notebook !!\n\nIn this notebook we are going to talk about **Random Forest Model**, Then we will apply what we learned on Titanic Dataset.\n","metadata":{}},{"cell_type":"markdown","source":"#### Table Of Content:\n- [1 - Introduction](#introduction)\n- [2 - Decision Trees](#dt)\n- [3 - Random Forest Definition](#rf_definition)\n- [4 - Bagging](#bagging)\n- [5 - Coding](#code)\n","metadata":{}},{"cell_type":"markdown","source":"<img src = \"https://weblog.wur.eu/spotlight/wp-content/uploads/sites/44/2020/05/spring_forest_franken_germany_shutterstock_276041312.jpg\" width=\"700\">","metadata":{}},{"cell_type":"markdown","source":"### 1 ) Introduction <a class=\"anchor\" id=\"introduction\"></a>\n\nA big part of machine learning is classification — we want to know what class (a.k.a. group) an observation belongs to. The ability to precisely classify observations is extremely valuable for various business applications like predicting whether a particular user will buy a product or forecasting whether a given loan will default or not.\n\nData science provides a plethora of classification algorithms such as logistic regression, support vector machine, naive Bayes classifier, and decision trees. But near the top of the classifier hierarchy is the random forest classifier (there is also the random forest regressor but that is a topic for another day).\n\nIn this post, we will examine how basic decision trees work, how individual decisions trees are combined to make a random forest, and ultimately discover why random forests are so good at what they do.","metadata":{}},{"cell_type":"markdown","source":"### 2 ) Decision Trees: <a class=\"anchor\" id=\"dt\"></a>\n\nA decision tree is a non-parametric supervised learning algorithm, which is utilized for both classification and regression tasks. It has a hierarchical, tree structure, which consists of a root node, branches, internal nodes and leaf nodes.\n\n<center><img src = \"https://www.ibm.com/content/dam/connectedassets-adobe-cms/worldwide-content/cdp/cf/ul/g/df/de/Decision-Tree.component.complex-narrative-xl.ts=1640802170790.png/content/adobe-cms/us/en/topics/decision-trees/jcr:content/root/table_of_contents/intro/complex_narrative/items/content_group_1423241468/image\" width=\"600\">\n","metadata":{}},{"cell_type":"markdown","source":"As you can see from the diagram above, a decision tree starts with a root node, which does not have any incoming branches. The outgoing branches from the root node then feed into the internal nodes, also known as decision nodes. Based on the available features, both node types conduct evaluations to form homogenous subsets, which are denoted by leaf nodes, or terminal nodes. The leaf nodes represent all the possible outcomes within the dataset. As an example, let’s imagine that you were trying to assess whether or not you should go surf, you may use the following decision rules to make a choice:","metadata":{}},{"cell_type":"markdown","source":"<center><img src = \"https://www.ibm.com/content/dam/connectedassets-adobe-cms/worldwide-content/cdp/cf/ul/g/10/3c/Decision-Tree-Example.component.complex-narrative-xl.ts=1640801899950.png/content/adobe-cms/us/en/topics/decision-trees/jcr:content/root/table_of_contents/intro/complex_narrative/items/content_group_1441304462/image\" width=\"600\">\n<img src = \"\" width:400/></center>\n","metadata":{}},{"cell_type":"markdown","source":"This type of flowchart structure also creates an easy to digest representation of decision-making, allowing different groups across an organization to better understand why a decision was made.\nDecision tree learning employs a divide and conquer strategy by conducting a greedy search to identify the optimal split points within a tree. This process of splitting is then repeated in a top-down, recursive manner until all, or the majority of records have been classified under specific class labels. Whether or not all data points are classified as homogenous sets is largely dependent on the complexity of the decision tree. Smaller trees are more easily able to attain pure leaf nodes—i.e. data points in a single class. However, as a tree grows in size, it becomes increasingly difficult to maintain this purity, and it usually results in too little data falling within a given subtree. When this occurs, it is known as data fragmentation, and it can often lead to **overfitting**. As a result, decision trees have preference for small trees, which is consistent with the principle of parsimony in Occam’s Razor; that is, “entities should not be multiplied beyond necessity.” Said differently, decision trees should add complexity only if necessary, as the simplest explanation is often the best. **To reduce complexity and prevent overfitting**, pruning is usually employed; this is a process, which removes branches that split on features with low importance. The model’s fit can then be evaluated through the process of cross-validation. Another way that decision trees can maintain their accuracy is by forming an **ensemble via a random forest algorithm**; this classifier predicts more accurate results, particularly when the individual trees are uncorrelated with each other.","metadata":{}},{"cell_type":"markdown","source":"### 3 ) Random Forest Definition: <a class=\"anchor\" id=\"rf_definition\"></a>\n","metadata":{"execution":{"iopub.status.busy":"2022-08-06T19:19:31.387820Z","iopub.execute_input":"2022-08-06T19:19:31.388299Z","iopub.status.idle":"2022-08-06T19:19:31.419902Z","shell.execute_reply.started":"2022-08-06T19:19:31.388206Z","shell.execute_reply":"2022-08-06T19:19:31.418307Z"}}},{"cell_type":"markdown","source":"Random forests or random decision forests is an **ensemble learning method** for classification, regression and other tasks that operates by constructing a multitude of decision trees at training time. For classification tasks, the output of the random forest is the class selected by most trees. For regression tasks, the mean or average prediction of the individual trees is returned. Random decision forests correct for decision trees habit of overfitting to their training set.Random forests generally outperform decision trees, but their accuracy is lower than gradient boosted trees citation needed. However, data characteristics can affect their performance.\n\nRandom forest, like its name implies, consists of a large number of individual decision trees that operate as an ensemble. Each individual tree in the random forest spits out a class prediction and the class with the most votes becomes our model’s prediction (see figure below).\n\n<center><img src = \"https://miro.medium.com/max/526/1*VHDtVaDPNepRglIAv72BFg.jpeg\" width: 400></center>\n\nThe fundamental concept behind random forest is a simple but powerful one - the wisdom of crowds. In data science speak, the reason that the random forest model works so well is:\n\n##### A large number of relatively uncorrelated models (trees) operating as a committee will outperform any of the individual constituent\n##### models.\n\nThe low correlation between models is the key. Just like how investments with low correlations (like stocks and bonds) come together to form a portfolio that is greater than the sum of its parts, uncorrelated models can produce ensemble predictions that are more accurate than any of the individual predictions. The reason for this wonderful effect is that the trees protect each other from their individual errors (as long as they don’t constantly all err in the same direction). While some trees may be wrong, many other trees will be right, so as a group the trees are able to move in the correct direction. So the prerequisites for random forest to perform well are:\n\n1 - There needs to be some actual signal in our features so that models built using those features do better than random guessing.  \n\n2 - The predictions (and therefore the errors) made by the individual trees need to have low correlations with each other.","metadata":{}},{"cell_type":"markdown","source":"### 4 ) Bagging: <a class=\"anchor\" id=\"bagging\"></a>\n\n**Bagging** (Bootstrap Aggregation) — Decisions trees are very sensitive to the data they are trained on — small changes to the training set can result in significantly different tree structures. Random forest takes advantage of this by allowing each individual tree to randomly sample from the dataset with replacement, resulting in different trees. This process is known as bagging.\n\nNotice that with bagging we are not subsetting the training data into smaller chunks and training each tree on a different chunk. Rather, if we have a sample of size N, we are still feeding each tree a training set of size N (unless specified otherwise). But instead of the original training data, we take a random sample of size N with replacement. For example, if our training data was [1, 2, 3, 4, 5, 6] then we might give one of our trees the following list [1, 2, 2, 3, 6, 6]. Notice that both lists are of length six and that “2” and “6” are both repeated in the randomly selected training data we give to our tree (because we sample with replacement).\n\n<center><img src = \"https://miro.medium.com/max/620/1*EemYMyOADnT0lJWSXmTDdg.jpeg\" width : 400></center>\n","metadata":{}},{"cell_type":"markdown","source":"**Feature Randomness** — In a normal decision tree, when it is time to split a node, we consider every possible feature and pick the one that produces the most separation between the observations in the left node vs. those in the right node. In contrast, each tree in a random forest can pick only from a random subset of features. This forces even more variation amongst the trees in the model and ultimately results in lower correlation across trees and more diversification.\n\nLet’s go through a visual example — in the picture above, the traditional decision tree (in blue) can select from all four features when deciding how to split the node. It decides to go with Feature 1 (black and underlined) as it splits the data into groups that are as separated as possible.\n\nNow let’s take a look at our random forest. We will just examine two of the forest’s trees in this example. When we check out random forest Tree 1, we find that it it can only consider Features 2 and 3 (selected randomly) for its node splitting decision. We know from our traditional decision tree (in blue) that Feature 1 is the best feature for splitting, but Tree 1 cannot see Feature 1 so it is forced to go with Feature 2 (black and underlined). Tree 2, on the other hand, can only see Features 1 and 3 so it is able to pick Feature 1.\n\n##### So in our random forest, we end up with trees that are not only trained on different sets of data (thanks to bagging) but also use\n##### different features to make decisions.\n\nAnd that, my dear reader, creates uncorrelated trees that buffer and protect each other from their errors.","metadata":{}},{"cell_type":"markdown","source":"### Coding: <a class=\"anchor\" id=\"code\"></a>","metadata":{}},{"cell_type":"code","source":"#=======================================================================================\n# Importing the libaries:\n#=======================================================================================\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\nimport os\nwarnings.filterwarnings(\"ignore\")\n\n#=======================================================================================","metadata":{"execution":{"iopub.status.busy":"2022-08-06T20:17:36.480061Z","iopub.execute_input":"2022-08-06T20:17:36.480458Z","iopub.status.idle":"2022-08-06T20:17:36.486658Z","shell.execute_reply.started":"2022-08-06T20:17:36.480424Z","shell.execute_reply":"2022-08-06T20:17:36.485625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:#00ADB5;\n           display:fill;\n           border-radius:5px;\n           background-color:#393E46;\n           font-size:20px;\n           font-family:sans-serif;\n           letter-spacing:0.5px\">\n        <p style=\"padding: 10px;\n              color:white;\">\n            <b>1 ) Importing the data:</b>\n        </p>\n</div>\n","metadata":{}},{"cell_type":"code","source":"#=======================================================================================\n# Reading the data:\n#=======================================================================================\n\ndef read_data():\n    train_data = pd.read_csv(\"/kaggle/input/titanic/train.csv\")\n    print(\"Train data imported successfully!!\")\n    print(\"-\"*50)\n    test_data = pd.read_csv(\"/kaggle/input/titanic/test.csv\")\n    print(\"Test data imported successfully!!\")\n    return train_data , test_data\n\ntrain_data , test_data = read_data()\ncombine = [train_data , test_data]\n\n#=======================================================================================","metadata":{"execution":{"iopub.status.busy":"2022-08-06T20:17:39.881330Z","iopub.execute_input":"2022-08-06T20:17:39.881728Z","iopub.status.idle":"2022-08-06T20:17:39.901047Z","shell.execute_reply.started":"2022-08-06T20:17:39.881697Z","shell.execute_reply":"2022-08-06T20:17:39.899826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Note:** This notebook doesn't go deep into details .. we just want to see how does Random Forest Work.\nYou can see my full solution from [here](https://www.kaggle.com/code/odaymourad/detailed-and-typical-solution-ensemble-modeling)\n\n**IMPORTANT:** You can skip this cell !!","metadata":{}},{"cell_type":"code","source":"# =====================================================================================\n# Data Processing: \n# =====================================================================================\n\n# -------------------------------------------------------------------------------------\n# Drop Unuseful features:\ntrain_data.drop(columns = [\"PassengerId\"] , inplace = True)\nfor dataset in combine:\n    dataset.drop(columns = [\"Ticket\" , \"Cabin\"] , inplace = True)\nprint(\"Dropping features Done !!\")\n# -------------------------------------------------------------------------------------\n# Filling missed values:\ntrain_data.Embarked.fillna(train_data.Embarked.dropna().max(), inplace=True)\nfor dataset in combine:\n    dataset['Embarked'] = dataset['Embarked'].dropna().map({'S':0,'C':1,'Q':2}).astype(int)\nfor dataset in combine:\n    dataset['Sex'] = dataset['Sex'].map( {'female': 1, 'male': 0} ).astype(int)    \n    guess_ages = np.zeros((2,3))\n\nfor dataset in combine:\n    for i in range(0, 2):\n        for j in range(0, 3):\n            guess_df = dataset[(dataset['Sex'] == i) & \\\n                                  (dataset['Pclass'] == j+1)]['Age'].dropna()\n\n            age_guess = guess_df.median()\n\n            # Convert random age float to nearest .5 age\n            guess_ages[i,j] = int( age_guess/0.5 + 0.5 ) * 0.5\n            \n    for i in range(0, 2):\n        for j in range(0, 3):\n            dataset.loc[ (dataset.Age.isnull()) & (dataset.Sex == i) & (dataset.Pclass == j+1),\\\n                    'Age'] = guess_ages[i,j]\n\n    dataset['Age'] = dataset['Age'].astype(int)\ntest_data.Fare.fillna(test_data.Fare.dropna().median() , inplace= True)\nprint(\"Filling Missed Data Completed !!\")\n# -------------------------------------------------------------------------------------\n# Data Engineering:\ntrain_data['AgeBand'] = pd.cut(train_data['Age'], 5)\ntrain_data[['AgeBand', 'Survived']].groupby(['AgeBand'], as_index=False).mean().sort_values(by='AgeBand', ascending=True)\nfor dataset in combine:    \n    dataset.loc[ dataset['Age'] <= 16, 'Age'] = 0\n    dataset.loc[(dataset['Age'] > 16) & (dataset['Age'] <= 32), 'Age'] = 1\n    dataset.loc[(dataset['Age'] > 32) & (dataset['Age'] <= 48), 'Age'] = 2\n    dataset.loc[(dataset['Age'] > 48) & (dataset['Age'] <= 64), 'Age'] = 3\n    dataset.loc[ dataset['Age'] > 64, 'Age']\ntrain_data.drop(['AgeBand'], axis=1 , inplace = True)\ntrain_data['FareBand'] = pd.qcut(train_data['Fare'], 4)\ntrain_data[['FareBand', 'Survived']].groupby(['FareBand'], as_index=False).mean().sort_values(by='FareBand', ascending=False)\nfor dataset in combine:\n    dataset.loc[ dataset['Fare'] <= 7.91, 'Fare'] = 0\n    dataset.loc[(dataset['Fare'] > 7.91) & (dataset['Fare'] <= 14.454), 'Fare'] = 1\n    dataset.loc[(dataset['Fare'] > 14.454) & (dataset['Fare'] <= 31), 'Fare']   = 2\n    dataset.loc[ dataset['Fare'] > 31, 'Fare'] = 3\n    dataset['Fare'] = dataset['Fare'].astype(int)\ntrain_data.drop(['FareBand'], axis=1 , inplace = True)\nfor dataset in combine:\n    dataset['FamilySize'] = dataset['SibSp'] + dataset['Parch'] + 1\ntrain_data.drop(['Parch', 'SibSp'], axis=1 , inplace = True)\ntest_data.drop(['Parch', 'SibSp'], axis=1 , inplace = True)    \ntrain_data[['FamilySize', 'Survived']].groupby(['FamilySize'], as_index=False).mean().sort_values(by='Survived', ascending=False)\nfor dataset in combine:\n    dataset['Single'] = dataset['FamilySize'].map(lambda s: 1 if s == 1 else 0)\n    dataset['SmallF'] = dataset['FamilySize'].map(lambda s: 1 if  s == 2  else 0)\n    dataset['MedF'] = dataset['FamilySize'].map(lambda s: 1 if 3 <= s <= 4 else 0)\n    dataset['LargeF'] = dataset['FamilySize'].map(lambda s: 1 if s >= 5 else 0)\n    \ntrain_data.drop(columns = [\"FamilySize\"] , inplace = True)\ntest_data.drop(columns = [\"FamilySize\"] , inplace = True)\nfor dataset in combine:\n    dataset['Title'] = dataset.Name.str.extract(' ([A-Za-z]+)\\.', expand=False)\nfor dataset in combine:\n    dataset['Title'] = dataset['Title'].replace(['Lady', 'Countess','Capt', 'Col',\\\n    'Don', 'Dr', 'Major', 'Rev', 'Sir', 'Jonkheer', 'Dona'], 'Rare')\n    dataset['Title'] = dataset['Title'].replace('Mlle', 'Miss')\n    dataset['Title'] = dataset['Title'].replace('Ms', 'Miss')\n    dataset['Title'] = dataset['Title'].replace('Mme', 'Mrs')\ntitle_mapping = {\"Mr\": 1, \"Miss\": 2, \"Mrs\": 3, \"Master\": 4, \"Rare\": 5}\nfor dataset in combine:\n    dataset['Title'] = dataset['Title'].map(title_mapping)\n    dataset['Title'] = dataset['Title'].fillna(0)\ntrain_data.drop(['Name'], axis=1 , inplace = True)\ntest_data.drop(['Name'], axis=1 , inplace = True) \nprint(\"Data Engineering Completed!!\")","metadata":{"execution":{"iopub.status.busy":"2022-08-06T20:18:27.813233Z","iopub.execute_input":"2022-08-06T20:18:27.813641Z","iopub.status.idle":"2022-08-06T20:18:27.956143Z","shell.execute_reply.started":"2022-08-06T20:18:27.813610Z","shell.execute_reply":"2022-08-06T20:18:27.954905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ===================================================================================\n# Modeling\n# ===================================================================================","metadata":{"execution":{"iopub.status.busy":"2022-08-06T20:27:12.194416Z","iopub.execute_input":"2022-08-06T20:27:12.194833Z","iopub.status.idle":"2022-08-06T20:27:12.199423Z","shell.execute_reply.started":"2022-08-06T20:27:12.194799Z","shell.execute_reply":"2022-08-06T20:27:12.198653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\nfrom sklearn.model_selection import GridSearchCV, cross_val_score, StratifiedKFold, learning_curve","metadata":{"execution":{"iopub.status.busy":"2022-08-06T20:26:14.956723Z","iopub.execute_input":"2022-08-06T20:26:14.957739Z","iopub.status.idle":"2022-08-06T20:26:15.280730Z","shell.execute_reply.started":"2022-08-06T20:26:14.957699Z","shell.execute_reply":"2022-08-06T20:26:15.279647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ==================================================================================\n# Preparing Data For Training:\n# ==================================================================================\n\nY_train = train_data[\"Survived\"]\nX_train = train_data.drop(labels = [\"Survived\"],axis = 1)\nTest = test_data.drop(labels = [\"PassengerId\"],axis = 1)\nprint(f\"X_train shape is = {X_train.shape}\" )\nprint(f\"Y_train shape is = {Y_train.shape}\" )\nprint(f\"Test shape is = {Test.shape}\" )","metadata":{"execution":{"iopub.status.busy":"2022-08-06T20:25:16.436879Z","iopub.execute_input":"2022-08-06T20:25:16.437823Z","iopub.status.idle":"2022-08-06T20:25:16.447193Z","shell.execute_reply.started":"2022-08-06T20:25:16.437785Z","shell.execute_reply":"2022-08-06T20:25:16.445982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cross validate model with Kfold stratified cross val\nkfold = StratifiedKFold(n_splits=10)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T20:26:20.036100Z","iopub.execute_input":"2022-08-06T20:26:20.036522Z","iopub.status.idle":"2022-08-06T20:26:20.041644Z","shell.execute_reply.started":"2022-08-06T20:26:20.036488Z","shell.execute_reply":"2022-08-06T20:26:20.040532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Random Forest Classifire:\n**n_estimatorsint:** (default=100): The number of trees in the forest.  \n\n**criterion:** {“gini”, “entropy”, “log_loss”}, default=”gini”: The function to measure the quality of a split. Supported criteria are “gini” for the Gini impurity and “log_loss” and “entropy” both for the Shannon information gain, see Mathematical formulation. Note: This parameter is tree-specific.\n\n**max_depth** (default=None): The maximum depth of the tree. If None, then nodes are expanded until all leaves are pure or until all leaves contain less than min_samples_split samples.\n\n**min_samples_split** (int or float, default=2): The minimum number of samples required to split an internal node:\n- If int, then consider min_samples_split as the minimum number.\n- If float, then min_samples_split is a fraction and ceil(min_samples_split * n_samples) are the minimum number of samples for each split.\n\n**min_samples_leaf:** (int or float, default=1): The minimum number of samples required to be at a leaf node. A split point at any depth will only be considered if it leaves at least min_samples_leaf training samples in each of the left and right branches. This may have the effect of smoothing the model, especially in regression.\n- If int, then consider min_samples_leaf as the minimum number.\n- If float, then min_samples_leaf is a fraction and ceil(min_samples_leaf * n_samples) are the minimum number of samples for each node.\n\n**min_weight_fraction_leaf** (float, default=0.0): The minimum weighted fraction of the sum total of weights (of all the input samples) required to be at a leaf node. Samples have equal weight when sample_weight is not provided.\n\n**max_features:** ({“sqrt”, “log2”, None}, int or float, default=”sqrt”): The number of features to consider when looking for the best split:\n- If int, then consider max_features features at each split.\n- If float, then max_features is a fraction and max(1, int(max_features * n_features_in_)) features are considered at each split.\n- If “auto”, then max_features=sqrt(n_features).\n- If “sqrt”, then max_features=sqrt(n_features).\n- If “log2”, then max_features=log2(n_features).\n- If None, then max_features=n_features.\n\n**max_leaf_nodes:** (int, default=None): Grow trees with max_leaf_nodes in best-first fashion. Best nodes are defined as relative reduction in impurity. If None then unlimited number of leaf nodes.\n\n**bootstrap:** (bool, default=True): Whether bootstrap samples are used when building trees. If False, the whole dataset is used to build each tree.\n\n**oob_score** (bool, default=False): Whether to use out-of-bag samples to estimate the generalization score. Only available if bootstrap=True.\n","metadata":{}},{"cell_type":"code","source":"# RFC Parameters tunning \nRFC = RandomForestClassifier(random_state = 1)\n\n## Search grid for optimal parameters\nrf_param_grid = {\n              \"max_features\": [3,6 , 10],\n              \"min_samples_split\": [2, 3, 5],\n              \"min_samples_leaf\": [2, 3, 5],\n              \"n_estimators\" :[100],\n              \"criterion\": [\"gini\", \"entropy\", \"log_loss\"],\n              \"max_depth\": [None],\n              \"bootstrap\": [False]}\n\ngsRFC = GridSearchCV(RFC,param_grid = rf_param_grid, cv=kfold, scoring=\"accuracy\")\ngsRFC.fit(X_train,Y_train)\n\nRFC_best = gsRFC.best_estimator_\n\n# Best score\ngsRFC.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-08-06T20:53:34.253559Z","iopub.execute_input":"2022-08-06T20:53:34.253966Z","iopub.status.idle":"2022-08-06T20:55:18.537116Z","shell.execute_reply.started":"2022-08-06T20:53:34.253935Z","shell.execute_reply":"2022-08-06T20:55:18.536244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### References:","metadata":{}},{"cell_type":"markdown","source":"1 - https://en.wikipedia.org/wiki/Random_forest    \n2 - https://www.ibm.com/topics/decision-trees#:~:text=A%20decision%20tree%20is%20a,internal%20nodes%20and%20leaf%20nodes  \n3 - https://towardsdatascience.com/understanding-random-forest-58381e0602d2  \n4 - https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestClassifier.html","metadata":{}}]}