{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":35332,"databundleVersionId":3723648,"sourceType":"competition"}],"dockerImageVersionId":30761,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"color:#016FD0;margin:0;font-size:32px;font-family:Georgia;text-align:center;display:fill;border-radius:5px;overflow:hidden;font-weight:600;\">Gradient Boosting Explained - Ensemble Learning</div>\n\n<div style=\"text-align:center\">\n    <img width=\"800\" alt=\"image\" src=\"https://github.com/user-attachments/assets/2856f2f7-cbae-486e-b613-864ea7ec6b28\">\n</div>\n<div style=\"text-align:center\">\n    <a href=\"https://unsplash.com/photos/black-and-white-book-on-gray-marble-table-y_HlzTZxiSY\">Photo from Unsplash</a>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">My Goal</div>\n\nWhen I was preparing for interviews, I realized that although I had practical experience with many machine learning algorithms, I often lacked a deep and comprehensive understanding of the theoretical aspects, particularly when interviewers probed into the more intricate concepts. To address this gap, I decided to create a series of notebooks focusing on the theory behind key machine learning techniques, starting with this one on Gradient Boosting and Ensemble Learning. This notebook aims to consolidate insights and explanations from various sources into a cohesive and detailed resource.\n\nMy goal is to assist others in the community who might be facing similar challenges. This series is designed to provide a robust foundation in the theory behind essential machine learning algorithms, enabling you to not only grasp these concepts more thoroughly but also approach interviews with greater confidence.✌️","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Ensemble learning</div>\n\n## <span style=\"color: #016FD0;\">Introduction</span>\n\nIt’s vital to an understanding of XGBoost to first grasp the machine learning concepts and algorithms that XGBoost builds upon: ***supervised machine learning***, ***decision trees***, ***ensemble learning***, and ***gradient boosting***.\n\n“Unity is strength”. This old saying expresses pretty well the underlying idea that rules the very powerful “ensemble methods” in machine learning. Roughly, ensemble learning methods, that often trust the top rankings of many machine learning competitions (including Kaggle’s competitions), **are based on the hypothesis that combining multiple models together can often produce a much more powerful model.** Ensemble learning gives credence to the idea of the “wisdom of crowds,” which suggests that the decision-making of a larger group of people is typically better than that of an individual expert. Similarly, ensemble learning refers to a group (or ensemble) of base learners, or models, which work collectively to achieve a better final prediction. A single model, also known as a base or weak learner, may not perform well individually due to high variance or high bias. However, when weak learners are aggregated, they can form a strong learner, as their combination reduces bias or variance, yielding better model performance. \n\n\nThe purpose of this post is to introduce various notions of ensemble learning. We will give the reader some necessary keys to well understand and use related methods and be able to design adapted solutions when needed. We will discuss some well known notions such as boostrapping, bagging, random forest, boosting, stacking and many others that are the basis of ensemble learning. In order to make the link between all these methods as clear as possible, we will try to present them in a much broader and logical framework that, we hope, will be easier to understand and remember.\n\n**Outline**\n\nIn the first section of this post we will present the notions of weak and strong learners and we will introduce three main ensemble learning methods: bagging, boosting and stacking. Then, in the second section we will be focused on bagging and we will discuss notions such that bootstrapping, bagging and random forests. In the third section, we will present boosting and, in particular, its two most popular variants: adaptative boosting (adaboost) and gradient boosting. Finally in the fourth section we will give an overview of stacking.\n\n- **Bagging**:\n    - Bootstrapping\n    - **Random forests**\n    \n- **Boosting**:\n    - Adaptative boosting (**adaboost**)\n    - **Gradient boosting**\n\n- **Stacking**\n\n\n## <span style=\"color: #016FD0;\">What is ensemble learning?</span>\n\nEnsemble methods is a machine learning technique that combines several base models in order to produce one optimal predictive model.\n\n**Ensemble learning is a machine learning paradigm where multiple models (often called “weak learners”) are trained to solve the same problem and combined to get better results.** The main hypothesis is that when weak models are correctly combined we can obtain more accurate and/or robust models.\n\nEach weak learner is fitted on the training set and provides predictions obtained. The final prediction result is computed by combining the results from all the weak learners. \n\nEnsemble Learning tries to capture complementary information from its different contributing models—that is, **an ensemble framework is successful when the contributing models are statistically diverse.**\n\nIn other words, models that display performance variation when evaluated on the same dataset are better suited to form an ensemble.\n\nFor example—\n\nDifferent models which make incorrect predictions on different sets of samples from the dataset should be ensembled. If two statistically similar models are ensembled (models that make wrong predictions on the same set of samples), the resulting model will only be as good as the contributing models. An ensemble won’t make any difference to the prediction ability in such a case.\n\n## <span style=\"color: #016FD0;\">Bias-variance tradeoff</span>\n\n**Single weak learner**\n\nIn machine learning, no matter if we are facing a classification or a regression problem, the choice of the model is extremely important to have any chance to obtain good results. This choice can depend on many variables of the problem: quantity of data, dimensionality of the space, distribution hypothesis… Most of the errors from a model’s learning are from three main factors: variance, noise, and bias.\n\nA low bias and a low variance, although they most often vary in opposite directions, are the two most fundamental features expected for a model. Indeed, to be able to “solve” a problem, we want our model to have enough degrees of freedom to resolve the underlying complexity of the data we are working with, but we also want it to have not too much degrees of freedom to avoid high variance and be more robust. This is the well known bias-variance tradeoff.\n\n<img width=\"851\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/f9266003-7018-4cc3-9b0f-15f0deefd0be\">\n\nIn ensemble learning theory, we call **weak learners** (or **base models**) models that can be used as building blocks for designing more complex models by combining several of them. Technically, a weak learner is a classifier that has a weak correlation with the actual value. Most of the time, these basics models perform not so well by themselves either because they have a high bias (low degree of freedom models, for example) or because they have too much variance to be robust (high degree of freedom models, for example). Then, the idea of ensemble methods is to try reducing bias and/or variance of such weak learners by combining several of them together in order to create a **strong learner** (or **ensemble model**) that achieves better performances.\n\n**Ensemble methods are frequently illustrated using decision trees as this algorithm can be prone to overfitting (high variance and low bias) when it hasn’t been pruned and it can also lend itself to underfitting (low variance and high bias) when it’s very small, like a decision stump, which is a decision tree with one level.** Remember, when an algorithm overfits or underfits to its training dataset, it cannot generalize well to new datasets, so ensemble methods are used to counteract this behavior to allow for generalization of the model to new datasets. While **decision trees** can exhibit high variance or high bias, **it’s worth noting that it is not the only modeling technique that leverages ensemble learning to find the “sweet spot” within the bias-variance tradeoff.**  \n\n## <span style=\"color: #016FD0;\">Why use ensemble learning?</span>\n\nEnsemble models are used for various tasks in machine learning, including **classification**, **regression**, and **anomaly detection**. They are particularly effective in scenarios where single models may struggle, such as when dealing with **noisy or complex datasets**. Ensembles are used to improve the robustness and generalization of machine learning models. By combining the predictions of multiple models, ensembles can reduce overfitting and improve performance on unseen data. ***Ensemble models have advantages, including improved predictive performance, reduced overfitting, and increased robustness. Ensembles can also provide more reliable predictions by capturing different aspects of the data and reducing the impact of individual model biases.***\n\nDeep Learning is used for solving complex pattern recognition tasks.\n\nHowever—\n\nSuch models require a large amount of labeled data (think millions of annotated images) to perform optimally.\n\nTherefore, sometimes we need to rely on pre-trained models for solving supervised learning tasks, i.e., a model already trained on a large dataset is re-used for the task at hand with a fewer data samples.\n\nWhat’s more, without customized models trained specifically for the task we want to perform we can be certain that our model will eventually underperform.\n\nFor example, in a multi-class image classification task, the pre-trained model we are using might not provide the optimal performance for all the classes in the dataset. \n\nSimilarly, a different pre-trained model might work well on some other classes of the same data. Thus, we need a method that can aggregate the performance of all such models and provide a better solution for all distributions of data.\n\nThis is where the concept of “Ensemble Learning” comes into play. And let me tell you—it's a real game changer.\n\nHere are some of the scenarios where ensemble learning comes in handy. **The need for ensemble learning arises in several problematic situations that can be both data-centric and algorithm-centric, like a scarcity/excess of data, the complexity of the problem, constraint in computational resources, etc.**\n\n1. **Can't choose an “optimal” model**\n    <img width=\"624\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/2f8aacae-fd63-4d26-9442-afc83471a11d\">\n\n1. **High Problem Complexity:** Sometimes, a problem can have a complex decision boundary, and it might become impossible for a single classifier to generate the appropriate boundary. For example, if we have a linear classifier and we try to tackle a problem with a parabolic (polynomial) decision boundary. One linear classifier obviously cannot do the job well. However, an ensemble of multiple linear classifiers can generate any polynomial decision boundary. An example of such a case is shown in the diagram below.\n\n    <img width=\"811\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/497951db-6b5f-4179-88e1-7c608065f83c\">\n\n1. **Information Fusion:** The most prevalent reason for using an ensemble learning model is information fusion for enhancing classification performance. That is, models that have been trained on different distributions of data pertaining to the same set of classes are employed during prediction time to get a more robust decision. For example, we may have trained one cat/dog classifier on high-quality images taken by a professional photographer. In contrast, another classifier has been trained on data using low-quality photos captured on mobile phones. When predicting a new sample, integrating the decisions from both these classifiers will be more robust and bias-free.\n\n### <span style=\"color: #016FD0;\">An example</span>\n\nTo better understand this definition lets take a step back into ultimate goal of machine learning and model building. This is going to make more sense as I dive into specific examples and why Ensemble methods are used.\n\nI will largely utilize Decision Trees to outline the definition and practicality of Ensemble Methods (**however it is important to note that Ensemble Methods do not only pertain to Decision Trees**).\n\nA Decision Tree determines the predictive value based on series of questions and conditions. For instance, this simple Decision Tree determining on whether an individual should play outside or not. The tree takes several weather factors into account, and given each factor either makes a decision or asks another question. In this example, every time it is overcast, we will play outside. However, if it is raining, we must ask if it is windy or not? If windy, we will not play. But given no wind, tie those shoelaces tight because were going outside to play.\n\n<img width=\"866\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/610c2724-a696-4f97-a67e-259a5dea6b2f\">\n\nDecision Trees can also solve quantitative problems as well with the same format. In the Tree to the left, we want to know wether or not to invest in a commercial real estate property. Is it an office building? A Warehouse? An Apartment building? Good economic conditions? Poor Economic Conditions? How much will an investment return? These questions are answered and solved using this decision tree.\n\n<img width=\"866\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/424c2ab6-56c3-449f-bdf9-69a995fb350a\">\n\n\nWhen making Decision Trees, there are several factors we must take into consideration: On what features do we make our decisions on? What is the threshold for classifying each question into a yes or no answer? In the first Decision Tree, what if we wanted to ask ourselves if we had friends to play with or not. If we have friends, we will play every time. If not, we might continue to ask ourselves questions about the weather. By adding an additional question, we hope to greater define the Yes and No classes.\n\n**This is where Ensemble Methods come in handy! Rather than just relying on one Decision Tree and hoping we made the right decision at each split, Ensemble Methods allow us to take a sample of Decision Trees into account,** calculate which features to use or questions to ask at each split, and **make a final predictor based on the aggregated results of the sampled Decision Trees.**\n\n\n***In Summary***\n\nThe goal of any machine learning problem is to find a single model that will best predict our wanted outcome. **Rather than making one model and hoping this model is the best/most accurate predictor we can make, ensemble methods take a myriad of models into account, and average those models to produce one final model. It is important to note that Decision Trees are not the only form of ensemble methods, just the most popular** and relevant in DataScience today.\n\n\n### <span style=\"color: #016FD0;\">When to use ensemble learning</span>\n\nYou can employ ensemble learning techniques when you want to improve the performance of machine learning models. For example to increase the accuracy of classification models or to reduce the mean absolute error for regression models. Ensembling also results in a more stable model. \n\nWhen your model is overfitting on the training set, you can also employ ensembling learning methods to create a more complex model. The models in the ensemble would then improve performance on the dataset by combining their predictions. \n\n### <span style=\"color: #016FD0;\">When ensemble learning works best</span>\n\nEnsemble learning works best when the base models are not correlated. For instance, you can train different models such as linear models, decision trees, and neural nets on different datasets or features. The less correlated the base models, the better. \n\nThe idea behind using uncorrelated models is that each may be solving a weakness of the other. They also have different strengths which, when combined, will result in a well-performing estimator. For example, creating an ensemble of just tree-based models may not be as effective as combining tree-type algorithms with other types of algorithms. ","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"color: #016FD0;\">Types of Ensemble Methods</span>\n\n**Combine weak learners:**\n\nIn order to set up an ensemble learning method, we first need to select our base models to be aggregated. Most of the time (including in the well known **bagging** and **boosting methods**) **a single base learning algorithm is used so that we have homogeneous weak learners that are trained in different ways.** The ensemble model we obtain is then said to be “homogeneous”. However, there also exist some methods that use different type of base learning algorithms: **some heterogeneous weak learners are then combined into an “heterogeneous ensembles model”.**\n\n***One important point is that our choice of weak learners should be coherent with the way we aggregate these models.*** If we choose base models with low bias but high variance, it should be with an aggregating method that tends to reduce variance whereas if we choose base models with low variance but high bias, it should be with an aggregating method that tends to reduce bias.\n\nThis brings us to the question of how to combine these models. We can mention **three major kinds of meta-algorithms that aims at combining weak learners:**\n\n\n1. **BAGGing, or Bootstrap AGGregating.** Bagging, that often considers **homogeneous weak learners**, **learns them independently from each other in parallel** and combines them following some kind of deterministic **averaging process**. BAGGing gets its name because it combines **Bootstrapping** and **Aggregation** to form one ensemble model. Given a sample of data, multiple bootstrapped subsamples are pulled. A Decision Tree is formed on each of the bootstrapped subsamples. After each subsample Decision Tree has been formed, an algorithm is used to aggregate over the Decision Trees to form the most efficient predictor. The image below will help explain:\n\n    <img width=\"799\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/a904ae43-d3ea-42de-afeb-035381cd1c6a\">\n\n    1. **Random Forest** Models. Random Forest Models **can be thought of as BAGGing**, with a slight tweak. When deciding where to split and how to make decisions, BAGGed Decision Trees have the full disposal of features to choose from. Therefore, although the bootstrapped samples may be slightly different, the data is largely going to break off at the same features throughout each model. In contrary, Random Forest models decide where to split based on a random selection of features. Rather than splitting at similar features at each node throughout, Random Forest models implement a level of differentiation because each tree will split based on different features. This level of differentiation provides a greater ensemble to aggregate over, ergo producing a more accurate predictor. Refer to the image for a better understanding. **Similar to BAGGing, bootstrapped subsamples are pulled from a larger dataset. A decision tree is formed on each subsample. HOWEVER, the decision tree is split on different features** (in this diagram the features are represented by shapes). Random forests are an ensemble learning technique that builds off of decision trees. Random forests involve creating multiple decision trees using bootstrapped datasets of the original data. The model then selects the mode (the majority) of all of the predictions of each decision tree. What’s the point of this? **By relying on a “majority wins” model, it reduces the risk of error from an individual tree.** ***Random Forest uses random feature selection, and the base algorithm of it is a decision tree algorithm.***\n    \n        <img width=\"800\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/5bf318be-2adf-48cb-8795-4b4b607faa27\">\n\n1. **Boosting,** that often considers **homogeneous weak learners**, **learns them sequentially in a very adaptative way** (a base model depends on the previous ones) and combines them following a deterministic strategy\n\n1. **Stacking**, that often considers **heterogeneous weak learners**, **learns them in parallel** and combines them by **training a meta-model to output a prediction based on the different weak models predictions**\n\n\n***Very roughly, we can say that bagging will mainly focus at getting an ensemble model with less variance than its components whereas boosting and stacking will mainly try to produce strong models less biased than their components (even if variance can also be reduced).***\n\nIn the following sections, we will present in details bagging and boosting (that are a bit more widely used than stacking and will allow us to discuss some key notions of ensemble learning) before giving a brief overview of stacking.\n\n<img width=\"849\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/727133c5-d35b-49c2-a132-e536692dca7a\">\n","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"color: #016FD0;\">Voting and Averaging Based Ensemble Methods</span>\n\nVoting and averaging are two of the easiest examples of ensemble learning in machine learning. They are both easy to understand and implement. Voting is used for classification and averaging is used for regression. In both methods, the first step is to create multiple classification/regression models using some training dataset. Each base model can be created using different splits of the same training dataset and same algorithm, or using the same dataset with different algorithms, or any other method. \n\n1. **Majority Voting:** Every model makes a prediction (votes) for each test instance and the final output prediction is the one that receives more than half of the votes. If none of the predictions get more than half of the votes, we may say that the ensemble method could not make a stable prediction for this instance. Although this is one of the more popular ensemble techniques, you may try the most voted prediction (even if that is less than half of the votes) as the final prediction. In some articles, you may see this method being called “plurality voting”.\n\n1. **Weighted Voting:** Unlike majority voting, where each model has the same rights, we can increase the importance of one or more models. In weighted voting you count the prediction of the better models multiple times. Finding a reasonable set of weights is up to you.\n\n1. **Simple Averaging:** In simple averaging method, for every instance of test dataset, the average predictions are calculated. This method often reduces overfit and creates a smoother regression model.  For classification, averaging can be applied to the predicted probabilities for a more confident prediction.\n\n1. **Weighted Averaging:** Weighted averaging is a slightly modified version of simple averaging, where the prediction of each model is multiplied by the weight and then their average is calculated. The weights can be assigned based on each model's performance on a validation set or tuned using grid or randomized search techniques. This allows models with higher performance to have a greater influence on the final prediction.\n\n<img width=\"825\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/4d8ed185-b57f-455f-b793-c5fb19f031f7\">\n\n<img width=\"998\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/d234ff9f-ae3c-47a2-8b04-dd8b3e8a228e\">\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Bagging</div>\n\nIn **parallel methods** we fit the different considered learners **independently** from each others and, so, it is possible to **train them concurrently**. The most famous such approach is “bagging” (standing for “bootstrap aggregating”) that aims at producing an ensemble model that is more robust than the individual models composing it. It is one of the applications of the Bootstrap procedure to a **high-variance machine learning algorithm,** *typically* **decision trees.**\n\n## <span style=\"color: #016FD0;\">Bootstrapping</span>\n\nLet’s begin by defining bootstrapping. This statistical technique consists in generating samples of size B (called bootstrap samples) from an initial dataset of size N by randomly drawing with replacement B observations.\n\n<img width=\"795\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/bcf2373d-73b5-452b-8798-7e3fec9f3fc3\">\n\nUnder some assumptions, these samples have pretty good statistical properties: in first approximation, they can be seen as being drawn both directly from the true underlying (and often unknown) data distribution and independently from each others. So, they can be considered as **representative** and **independent samples** of the true data distribution (almost i.i.d. samples). The hypothesis that have to be verified to make this approximation valid are twofold. \n\n- First, the size N of the initial dataset should be large enough to capture most of the complexity of the underlying distribution so that sampling from the dataset is a good approximation of sampling from the real distribution (representativity). \n- Second, the size N of the dataset should be large enough compared to the size B of the bootstrap samples so that samples are not too much correlated (independence). \n\nNotice that in the following, we will sometimes make reference to these properties (**representativity** and **independence**) of bootstrap samples: the reader should always keep in mind that this is only an approximation.\n\nBootstrap samples are often used, for example, to evaluate variance or confidence intervals of a statistical estimators. By definition, a statistical estimator is a function of some observations and, so, a random variable with variance coming from these observations. In order to estimate the variance of such an estimator, we need to evaluate it on several independent samples drawn from the distribution of interest. In most of the cases, considering truly independent samples would require too much data compared to the amount really available. We can then use bootstrapping to generate several bootstrap samples that can be considered as being “almost-representative” and “almost-independent” (almost i.i.d. samples). These bootstrap samples will allow us to approximate the variance of the estimator, by evaluating its value for each of them.\n\n<img width=\"1339\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/6c2f0758-a6bd-4115-9a2c-9e385d95dfd4\">\n\n\n***Bootstrapping is often used to evaluate variance or confidence interval of some statistical estimators.***\n\n\n**If our sample is small and if the mean has error in it. We can improve the estimate of the mean by using the bootstrap procedure:**\n\n- Create many random sub-samples of our dataset with replacement so that same sample can be selected more than once.\n- Compute the mean of each sub-sample.\n- Calculate the average of all of our collected means and refer that as our estimated mean for the data.\n\nThe name Bootstrap Aggregating, also known as “Bagging”, summarizes the key elements of this strategy. In the bagging algorithm, the first step involves creating multiple models. These models are generated using the same algorithm with random sub-samples of the dataset which are drawn from the original dataset randomly with bootstrap sampling method. In bootstrap sampling, some original examples appear more than once and some original examples are not present in the sample. If you want to create a sub-dataset with m elements, you should select a random element from the original dataset m times. And if the goal is generating n dataset, you follow this step n times. In bagging, each sub-samples can be generated independently from each other. So generation and training can be done in parallel.\n\n## <span style=\"color: #016FD0;\">Bagging Explained</span>\n\nWhen training a model, no matter if we are dealing with a classification or a regression problem, we obtain a function that takes an input, returns an output and that is defined with respect to the training dataset. Due to the theoretical variance of the training dataset (we remind that a dataset is an observed sample coming from a true unknown underlying distribution), the fitted model is also subject to variability: **if another dataset had been observed, we would have obtained a different model.**\n\nThe idea of bagging is then simple: we want to fit several independent models and “average” their predictions in order to obtain a model with a lower variance. However, we can’t, in practice, fit fully independent models because it would require too much data. So, we rely on the good “approximate properties” of bootstrap samples (representativity and independence) to fit models that are almost independent.\n\nFirst, we create multiple bootstrap samples so that each new bootstrap sample will act as another (almost) independent dataset drawn from true distribution. **Then, we can fit a weak learner for each of these samples and finally aggregate them such that we kind of “average” their outputs and, so, obtain an ensemble model with less variance that its components.** Roughly speaking, as the bootstrap samples are approximatively independent and identically distributed (i.i.d.), so are the learned base models. Then, “averaging” weak learners outputs do not change the expected answer but reduce its variance (just like averaging i.i.d. random variables preserve expected value but reduce variance). When we are bagging with decision trees, we are less worried about individual trees that leads to overfitting of the training data. Due to this reason and for efficiency, the individual decision trees are grown deep and the trees are not pruned or clipped. These trees will have both high variance and low bias which is beneficial. \n\nSo, assuming that we have L bootstrap samples (approximations of L independent datasets) of size B denoted\n\n<img width=\"1412\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/7dc2671c-bd4c-4b30-9798-dafb1059bd32\">\n\nThere are several possible ways to aggregate the multiple models fitted in parallel. For a regression problem, the outputs of individual models can literally be averaged to obtain the output of the ensemble model. For classification problem the class outputted by each model can be seen as a vote and the class that receives the majority of the votes is returned by the ensemble model (this is called hard-voting). Still for a classification problem, we can also consider the probabilities of each classes returned by all the models, average these probabilities and keep the class with the highest average probability (this is called soft-voting). Averages or votes can either be simple or weighted if any relevant weights can be used.\n\nFinally, we can mention that one of the big advantages of bagging is that **it can be parallelised**. As the different models are fitted independently from each others, intensive parallelisation techniques can be used if required. Bagging is similar to **Divide and conquer**. It is a group of predictive models run on multiple subsets from the original dataset combined together to achieve better accuracy and model stability.\n\n**Main Steps involved in bagging are (Bagging Steps):**\n\n1. **Creating multiple datasets:** Sampling is done with a replacement on the original data set and new datasets are formed from the original dataset. If there are N observations and M features in training data set. A sample from training data set is taken randomly with replacement.\n\n1. A subset of M features is selected randomly and the feature which gives the best split is used to split the node iteratively or repeatedly. Bootstrapping is a sampling technique in which we create subsets of observations from the original dataset, with replacement. The size of the subsets is usually the same as the size of the original set. t). The size of subsets created for bagging may be less than the original set.\n\n1. **Building multiple classifiers:** On each of these smaller datasets, a classifier is built, usually, the same classifier is built on all the datasets. The tree can be grown to the largest.\n\n1. **Combining Classifiers:** The predictions of all the individual classifiers are now combined to give a better classifier, usually with very less variance compared to before. Above mentioned steps are repeated n times and prediction is given based on the aggregation or average of predictions from n number of trees.\n \n<img width=\"1232\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/85e24b6a-7d23-4db6-b006-aa4f63b13241\">\n\n***Bagging consists in fitting several base models on different bootstrap samples and build an ensemble model that “average” the results of these weak learners.***\n\n\n\n**Advantages:**\n\n- It reduces over-fitting of the model. Using random subsets of data, the risk of overfitting is reduced and flattened by averaging the results of the sub-models.\n- It can handle higher dimensionality data very well.\n- Maintains accuracy even for missing data.\n\n**Example: To explain the basic scenario of bagging. Below is an analysis of the relationship between ozone and temperature (taken from Wikipedia)** \n\nThe relationship between temperature and ozone in this data set is apparently non-linear, based on the scatter plot. Instead of building a single prediction model from the complete data set, 100 samples of the data were drawn. Each sample is different from the original data set, yet resembles it in distribution and variability. Predictions from these 100 were then samples made across the range of the data. The first 10 predicted smooth fits appear as grey lines in the figure below. The lines are clearly very wiggly and they overfit the data — a result of the bandwidth being too small.\n\n<img width=\"1095\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/ce8a1225-2090-48fc-8289-ce6bc6229221\">\n\nBy taking the average of 100 smoothers, each fitted to a subset of the original data set, we arrive at one bagged predictor (red line). Clearly, the mean is more stable and there is less overfit.\n\nCommon Bagging algorithms:\n\n- Bagging meta-estimator\n- Random forest\n\n**Implementation:**\n\nIn the following, we will explore some useful python functions from the sklearn.ensemble library. The function called BaggingClassifierhas a few parameters which can be looked up in the documentation, but the most important ones are `base_estimator`, `n_estimators`, and `max_samples`.\n\n- `base_estimator:` You have to provide the underlying algorithm that should be used by the random subsets in the bagging procedure in the first parameter. This could be for example Logistic Regression, Support Vector Classification, Decision trees, or many more.\n\n- `n_estimators`: The number of estimators defines the number of bags you would like to create here and the default value for that is 10.\n\n- `max_samples`: The maximum number of samples defines how many samples should be drawn from X to train each base estimator. The default value here is one point zero which means that the total number of existing entries should be used. You could also say that you want only 80% of the entries by setting it to 0.8.\n\nAfter setting the scenes, this model object works like many other models and can be trained using the `fit()` procedure including X and y data from the training set. The corresponding predictions on test data can be done using `predict()`.\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Random forests</div>\n\n**Learning trees are very popular base models for ensemble methods. Strong learners composed of multiple trees can be called “forests”.** Trees that compose a forest can be chosen to be either shallow (few depths) or deep (lot of depths, if not fully grown). Shallow trees have less variance but higher bias and then will be better choice for sequential methods that we will described thereafter. Deep trees, on the other side, have low bias but high variance and, so, are relevant choices for bagging method that is mainly focused at reducing variance.\n\nThe **random forest** approach is a bagging method where **deep trees**, fitted on bootstrap samples, are combined to produce an output with lower variance. However, random forests also use another trick to make the multiple fitted trees a bit less correlated with each others: when growing each tree, instead of only sampling over the observations in the dataset to generate a bootstrap sample, we also **sample over features** and keep only a random subset of them to build the tree.\n\nSampling over features has indeed the effect that all trees do not look at the exact same information to make their decisions and, so, it reduces the correlation between the different returned outputs. Another advantage of sampling over the features is that it makes the decision **making process more robust to missing data:** observations (from the training dataset or not) with missing data can still be regressed or classified based on the trees that take into account only features where data are not missing. Thus, random forest algorithm combines the concepts of bagging and random feature subspace selection to create more robust models.\n\n<img width=\"1383\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/9c35429d-bbc9-4ee1-b1ea-f8395d4292d4\">\n\n- Random Forest is a technique in ensemble learning that utilizes a decision tree group to make predictions.\n\n- The key concept behind Random Forest is introducing randomness in tree-building to create diverse trees.\n\n- To create each tree, a random subset of the training data is sampled (with replacement), and a decision tree is trained on this subset.\n\n- Additionally, rather than considering all features, a random subset of features is selected at each tree node to determine the best split.\n\n- The final prediction of the Random Forest is made by aggregating the predictions of all the individual trees (e.g., averaging for regression, majority voting for classification).\n\n- Random Forests are robust against overfitting and perform well on many datasets. Compared to individual decision trees, they are also less sensitive to hyperparameters.\n\n***Random forest method is a bagging method with trees as weak learners. Each tree is fitted on a bootstrap sample considering only a subset of variables randomly chosen.***\n\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Boosting</div>\n\n## <span style=\"color: #016FD0;\">What is boosting?</span>\nIn sequential methods the different combined weak models are no longer fitted independently from each others. The idea is to fit models **iteratively** such that the training of model at a given step depends on the models fitted at the previous steps. “Boosting” is the most famous of these approaches and it produces an ensemble model that is in general less biased than the weak learners that compose it. Boosting is an ensemble method for improving the model predictions of any given learning algorithm. The idea of boosting is to train weak learners sequentially, each trying to correct its predecessor. Boosting is used to create a collection of predictors. It refers to a group of algorithms that will make use of weighted averages to convert weak learners into stronger learners. Unlike bagging which runs each model independently and then aggregate the outputs at the end without preference to any model. Boosting is all about teamwork, it runs each model and dictates what features the next model will focus on.\n\nBoosting methods work in the same spirit as bagging methods: we build a family of models that are aggregated to obtain a strong learner that performs better. However, unlike bagging that mainly aims at reducing variance, boosting is a technique that consists in fitting sequentially multiple weak learners in a very adaptative way: each model in the sequence is fitted giving more importance to observations in the dataset that were badly handled by the previous models in the sequence. Intuitively, each new model **focus its efforts on the most difficult observations to fit up to now,** so that we obtain, at the end of the process, a strong learner with lower bias (even if we can notice that boosting can also have the effect of reducing variance). Boosting, like bagging, can be used for regression as well as for classification problems.\n\nBeing **mainly focused at reducing bias**, the base models that are often considered for boosting are models with low variance but high bias. For example, if we want to use trees as our base models, we will choose most of the time shallow decision trees with only a few depths. Another important reason that motivates the use of low variance but high bias models as weak learners for boosting is that these models are in general less computationally expensive to fit (few degrees of freedom when parametrised). Indeed, as computations to fit the different models **can’t be done in parallel (unlike bagging),** it could become too expensive to fit sequentially several complex models.\n\nOnce the weak learners have been chosen, we still need to define:\n\n- **how they will be sequentially fitted** (what information from previous models do we take into account when fitting current model?) and\n- **how they will be aggregated** (how do we aggregate the current model to the previous ones?). We will discuss these questions in the two following subsections, describing more especially two important boosting algorithms: ***adaboost*** and ***gradient boosting***.\n\nIn a nutshell, these two meta-algorithms differ on how they create and aggregate the weak learners during the sequential process. \n\n- *Adaptive boosting updates the weights attached to each of the training dataset observations* whereas \n- *gradient boosting updates the value of these observations.*\n\nThis main difference comes from the way both methods try to solve the optimisation problem of finding the best model that can be written as a weighted sum of weak learners.\n\n<img width=\"1266\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/c7f28450-6524-4e66-a445-4aedaec5fd37\">\n\n***Boosting consists in, iteratively, fitting a weak learner, aggregate it to the ensemble model and “update” the training dataset to better take into account the strengths and weakness of the current ensemble model when fitting the next base model.*** Boosting never changes the previous predictor and only corrects the next predictor by learning from mistakes. Since **Boosting is greedy**, it is recommended to **set a stopping criterion such as model performance (early stopping) or several stages** (e.g. depth of tree in tree-based learners) to prevent overfitting of training data. \n\n<img width=\"905\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/59b7057a-2340-4604-8bd2-3a53fd791fef\">\n\nOnce the first model is built, the falsely classified/predicted points are taken in addition to the second bootstrapped sample to train the second model. Then, the ensemble model (models 1 and 2) are used against the test dataset and the process continues. Thus, the boosting algorithm combines a number of weak learners to form a strong learner. The individual models would not perform well on the entire dataset, but they work well for some part of the dataset. Thus, each model actually boosts the performance of the ensemble.\n\n**Boosting Steps:**\n\n- In the first step, draw a random subset of training samples d1 without replacement from the training set D to train a weak learner C1.\n- Draw second random training subset d2 without replacement from the training set and add 50 percent of the samples which were previously wrongly classified / misclassified to train a weak learner C2.\n- In the third step, find the training samples d3 in the training set D on which C1 and C2 disagree to train a third weak learner C3.\n\n\n**Advantages:**\n\n- Boosting supports different loss functions.\n- It can **work well with interactions.**\n- The key idea here is clearly to create models that are also able to predict the more difficult data entries. This can then lead to a better fit of the model and reduces the bias.\n \n**Disadvantages:**\n\n- **It prone to over-fitting.** In comparison to Bagging, this technique uses weighted voting or weighted averaging based on the coefficients of the models that are considered together with their predictions. Therefore, this model can reduce underfitting, but might also tend to overfit sometimes.\n- It requires careful tuning of different hyper-parameters.\n\nBoosting is a method used in machine learning to reduce errors in predictive data analysis. Data scientists train machine learning software, called machine learning models, on labeled data to make guesses about unlabeled data. A single machine learning model might make prediction errors depending on the accuracy of the training dataset. For example, if a cat-identifying model has been trained only on images of white cats, it may occasionally misidentify a black cat. Boosting tries to overcome this issue by training multiple models sequentially to improve the accuracy of the overall system.\n\nBoosting is an ensemble learning method that combines a set of weak learners into a strong learner to minimize training errors. Boosting algorithms can improve the predictive power of your data mining initiatives. Machine learning models can be weak learners or strong learners:\n\n- **Weak learners:** Weak learners have low prediction accuracy, similar to random guessing. They are prone to overfitting—that is, they can't classify data that varies too much from their original dataset. For example, if you train the model to identify cats as animals with pointed ears, it might fail to recognize a cat whose ears are curled.\n\n- **Strong learners:** Strong learners have higher prediction accuracy. Boosting converts a system of weak learners into a single strong learning system. For example, to identify the cat image, it combines a weak learner that guesses for pointy ears and another learner that guesses for cat-shaped eyes. After analyzing the animal image for pointy ears, the system analyzes it once again for cat-shaped eyes. This improves the system's overall accuracy.\n\nIn boosting, **a random sample of data is selected, fitted with a model and then trained sequentially** — that is, **each model tries to compensate for the weaknesses of its predecessor.** With each iteration, the weak rules from each individual classifier are combined to form one, strong prediction rule. Unlike many ML models which focus on high quality prediction done by a single model, boosting algorithms seek to improve the prediction power by training a sequence of weak models, each compensating the weaknesses of its predecessors.\n\n<img width=\"869\" alt=\"image\" src=\"https://github.com/user-attachments/assets/fa01602e-543a-4d47-ac9c-da986653c7fd\">\n\nTo understand Boosting, it is crucial to recognize that boosting is a generic algorithm rather than a specific model. Boosting needs you to specify a weak model (e.g. regression, shallow decision trees, etc) and then improves it.\n\nWith that sorted out, it is time to explore different definitions of weakness and their corresponding algorithms.\n\n## <span style=\"color: #016FD0;\">Bagging versus boosting</span>\n\nBagging and boosting are two main types of ensemble learning methods. \n\n- As highlighted in this [study](https://www.d.umn.edu/~rmaclin/cs5751/notes/opitz-jair99.pdf), the **main difference** between these learning methods **is the way in which they are trained.** **In bagging, weak learners are trained in parallel, but in boosting, they learn sequentially.** This means that a series of models are constructed and with each new model iteration, the weights of the misclassified data in the previous model are increased. This redistribution of weights helps the algorithm identify the parameters that it needs to focus on to improve its performance. AdaBoost, which stands for “adaptative boosting algorithm,” is one of the most popular boosting algorithms as it was one of the first of its kind. Other types of boosting algorithms include XGBoost, GradientBoost, and BrownBoost. \n\n- Another difference between bagging and boosting is in **how they are used.** For example, bagging methods are typically used on weak learners that exhibit high variance and low bias, whereas boosting methods are leveraged when low variance and high bias is observed. **While bagging can be used to avoid overfitting, boosting methods can be more prone** to [this](https://www.researchgate.net/publication/221112418_Avoiding_Boosting_Overfitting_by_Removing_Confusing_Samples) (link resides outside ibm.com) **although it really depends on the dataset. However, parameter tuning can help avoid the issue.**\n\n<img width=\"786\" alt=\"image\" src=\"https://github.com/user-attachments/assets/55e83729-ba7f-4c54-9c6f-df157c7ef24a\">\n<img width=\"1053\" alt=\"image\" src=\"https://github.com/user-attachments/assets/07536861-7ac0-445e-9dd4-e32103e0a3ed\">\n\nIf you are already familiar with Random Forests, the Gradient Boosting algorithm implemented in XGBoost is also an ensemble of decision trees. Those trees are poor models individually, but when they are grouped they can be really performant.\n\nThe difference between XGBoost and Random Forest lies in the way those trees are built and combined. Random Forest builds fully grown decision trees in parallel on subsamples of the data. Each tree is higly specialized to predict on its subsample and do not generalize well (high variance). By combining the predictions made by each individual tree, the Random Forest algorithm decreases variance and gives good performance.\n\nXGBoost on the other hand, builds really short and simple decision trees iteratively. Each tree is called a “weak learner” for their high bias. XGBoost starts by creating a first simple tree which has poor performance by itself. It then builds another tree which is trained to predict what the first tree was not able to, and is itself a weak learner too. The algorithm goes on by sequentially building more weak learners, each one correcting the previous tree until a stopping condition is reached, such as the number of trees (estimators) to build.\n\n\n## <span style=\"color: #016FD0;\">How does boosting work?</span>\n\nTo understand how boosting works, let's describe how machine learning models make decisions. **Although there are many variations in implementation, data scientists often use boosting with decision-tree algorithms:**\n\n### Decision trees\nDecision trees are data structures in machine learning that work by dividing the dataset into smaller and smaller subsets based on their features. The idea is that decision trees split up the data repeatedly until there is only one class left. For example, the tree may ask a series of yes or no questions and divide the data into categories at every step.\n\n### Boosting ensemble method\nBoosting creates an ensemble model by combining several weak decision trees sequentially. It assigns weights to the output of individual trees. Then it gives incorrect classifications from the first decision tree a higher weight and input to the next tree. After numerous cycles, the boosting method combines these weak rules into a single powerful prediction rule.\n\n\n### How is training in boosting done?\nThe training method varies depending on the type of boosting process called the boosting algorithm. However, an algorithm takes the following general steps to train the boosting model:\n\n1. **Step 1:** The boosting algorithm assigns equal weight to each data sample. It feeds the data to the first machine model, called the base algorithm. The base algorithm makes predictions for each data sample.\n\n1. **Step 2:** The boosting algorithm assesses model predictions and increases the weight of samples with a more significant error. It also assigns a weight based on model performance. A model that outputs excellent predictions will have a high amount of influence over the final decision.\n\n1. **Step 3:** The algorithm passes the weighted data to the next decision tree.\n\n1. **Step 4:** The algorithm repeats steps 2 and 3 until instances of training errors are below a certain threshold.\n\n## <span style=\"color: #016FD0;\">Types of boosting</span>\n\nBoosting methods are focused on **iteratively combining weak learners** to build a strong learner that can predict more accurate outcomes. As a reminder, a weak learner classifies data slightly better than random guessing. This approach can provide robust results for prediction problems, and can even outperform neural networks and support vector machines for tasks like image retrieval.  \n\nBoosting algorithms can differ in **how they create** and **aggregate weak learners** during the sequential process. Three popular types of boosting methods include: \n\n1. **Adaptive boosting or AdaBoost:** Yoav Freund and Robert Schapire are credited with the creation of the AdaBoost algorithm. Adaptive Boosting (AdaBoost) was one of the earliest boosting models developed. It adapts and tries to self-correct in every iteration of the boosting process. This method operates iteratively, identifying misclassified data points and adjusting their weights to minimize the training error. The model continues optimize in a sequential fashion until it yields the strongest predictor. AdaBoost initially gives the same weight to each dataset. Then, it automatically adjusts the weights of the data points after every decision tree. It gives more weight to incorrectly classified items to correct them for the next round. It repeats the process until the residual error, or the difference between actual and predicted values, falls below an acceptable threshold. You can use AdaBoost with many predictors, and it is typically not as sensitive as other boosting algorithms. **This approach does not work well when there is a correlation among features or high data dimensionality.** Overall, **AdaBoost is a suitable type of boosting for classification problems.**\n\n1. **Gradient boosting:** Building on the work of Leo Breiman, Jerome H. Friedman developed gradient boosting, which works by sequentially adding predictors to an ensemble with each one correcting for the errors of its predecessor. However, instead of changing weights of data points like AdaBoost, the gradient boosting trains on the residual errors of the previous predictor. The name, gradient boosting, is used since it ***combines the gradient descent algorithm and boosting method.*** The difference between AdaBoost and GB is that GB does not give incorrectly classified items more weight. Instead, GB software optimizes the loss function by generating base learners sequentially so that the present base learner is always more effective than the previous one. This method attempts to generate accurate results initially instead of correcting errors throughout the process, like AdaBoost. For this reason, **GB software can lead to more accurate results. Gradient Boosting can help with both classification and regression-based problems.**\n\n1. **Extreme gradient boosting or XGBoost:** XGBoost is an implementation of gradient boosting that’s **designed for computational speed and scale.** XGBoost leverages multiple cores on the CPU, allowing for learning to occur in parallel during training. **It is a boosting algorithm that can handle extensive datasets, making it attractive for big data applications.** The key features of XGBoost are parallelization, distributed computing, cache optimization, and out-of-core processing.\n\n## <span style=\"color: #016FD0;\">Benefits of boosting</span>\n\nThere are a number of key advantages and challenges that the boosting method presents when used for classification or regression problems. \nThe key benefits of boosting include:  \n\n1. **Strong prediction power:** usually boosting > bagging (random forrest) > decision tree\n\n1. **Resilient to overfitting - in a degree**\n\n1. **Implicit Feature Selection**\n\n1. **Ease of Implementation:** Boosting can be used with several hyper-parameter tuning options to improve fitting. **No data preprocessing is required,** and boosting algorithms like have built-in routines to handle missing data. In Python, the scikit-learn library of ensemble methods (also known as sklearn.ensemble) makes it easy to implement the popular boosting methods, including AdaBoost, XGBoost, etc. Boosting has **easy-to-understand** and **easy-to-interpret** algorithms that learn from their mistakes. In addition, most languages have built-in libraries to implement boosting algorithms with many parameters that can fine-tune performance.\n\n1. **Reduction of bias:** Boosting algorithms combine multiple weak learners in a sequential method, iteratively improving upon observations. This approach can help to reduce high bias, commonly seen in shallow decision trees and logistic regression models. \n\n1. **Computational Efficiency:** Since boosting algorithms only select features that increase its predictive power during training, it can help to reduce dimensionality as well as increase computational efficiency.  \n\n## <span style=\"color: #016FD0;\">Challenges of boosting</span>\n\nThe key challenges of boosting include:  \n\n1. **Vulnerability to outlier data:** Boosting models are vulnerable to outliers or data values that are different from the rest of the dataset. Because each model attempts to correct the faults of its predecessor, outliers can skew results significantly. Since each weak classifier is dedicated to fix its predecessors’ shortcomings, the model may pay too much attention to outliers.\n\n1. **Overfitting:** There’s some dispute in the research around whether or not boosting can help reduce overfitting or exacerbate it. We include it under challenges because in the instances that it does occur, predictions cannot be generalized to new datasets. **Note on overfitting** One key difference between random forests and gradient boosting decision trees is the number of trees used in the model. Increasing the number of trees in random forests does not cause overfitting. After some point, the accuracy of the model does not increase by adding more trees but it is also not negatively effected by adding excessive trees. You still do not want to add unnecessary amount of trees due to computational reasons but there is no risk of overfitting associated with the number of trees in random forests. However, the number of trees in gradient boosting decision trees is very critical in terms of overfitting. Adding too many trees will cause overfitting so it is important to stop adding trees at some point.\n\n1. **Intense computation:** Sequential training in boosting is hard to scale up. Since each estimator is built on its predecessors, boosting models can be computationally expensive, although XGBoost seeks to address scalability issues seen in other types of boosting methods. Boosting algorithms can be slower to train when compared to bagging **as a large number of parameters can also influence the behavior of the model.** Since each estimator is built on its predecessors, the process is hard to parallelize.\n\n1. **Real-time implementation:** You might also find it challenging to use boosting for real-time implementation because the algorithm is more complex than other processes. Boosting methods have high adaptability, so you can use a wide variety of model parameters that immediately affect the model's performance.\n\n## <span style=\"color: #016FD0;\">Applications of boosting</span>\n\nBoosting algorithms are well suited for artificial intelligence projects across a broad range of industries, including:  \n\n- **Healthcare:** Boosting is used to lower errors in medical data predictions, such as predicting cardiovascular risk factors and cancer patient survival rates. For example, research (link resides outside ibm.com) shows that ensemble methods significantly improve the accuracy in identifying patients who could benefit from preventive treatment of cardiovascular disease, while avoiding unnecessary treatment of others. Likewise, another study (link resides outside ibm.com) found that applying boosting to multiple genomics platforms can improve the prediction of cancer survival time. \n\n- **IT:** Gradient boosted regression trees are used in **search engines for page rankings,** while the Viola-Jones boosting algorithm is used for image retrieval. As noted by Cornell (link resides outside ibm.com), boosted classifiers allow for the computations to be stopped sooner when it’s clear in which way a prediction is headed. This means that a search engine can stop the evaluation of lower ranked pages, while image scanners will only consider images that actually contains the desired object.   \n\n- **Finance:** Boosting is used with deep learning models to automate critical tasks, including fraud detection, pricing analysis, and more. For example, boosting methods in credit card fraud detection and financial products pricing analysis (link resides outside ibm.com) improve the accuracy of analyzing massive data sets to minimize financial losses.  ","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">AdaBoost</div>\n\n## <span style=\"color: #016FD0;\">Definition of Weakness of AdaBoost</span>\n\nAdaBoost is a specific Boosting algorithm **developed for classification problems** (also called discrete AdaBoost). The weakness is identified by the weak estimator’s error rate:\n\nIn each iteration, AdaBoost identifies miss-classified data points, increasing their weights (and decrease the weights of correct points, in a sense) so that the next classifier will pay extra attention to get them right.The following figure illustrates how weights impact the performance of a simple decision stump(tree with depth 1)\n\n<img width=\"775\" alt=\"image\" src=\"https://github.com/user-attachments/assets/e99c097f-a364-436a-b530-5d3e11e2bdac\">\n\nNow with weakness defined, the next step is to figure out how to combine the sequence of models to make the ensemble stronger overtime.\n\n## <span style=\"color: #016FD0;\">Pseudocode</span>\n\nThere are several different algorithms proposed by researchers. Here I’ll introduce the most popular method called SAMME, a specific method that deals with multi-classification problems. AdaBoost trains a sequence of models with augmented sample weights, generating ‘confidence’ coefficients Alpha for individual classifiers based on errors. Low errors leads to large Alpha, which means higher importance in the voting.\n\n<img width=\"853\" alt=\"image\" src=\"https://github.com/user-attachments/assets/3755cbd9-a4b7-4774-8078-5fd4ac31b40b\">\n\n\n## <span style=\"color: #016FD0;\">How It Works</span>\n\nIn adaptative boosting (often called “adaboost”), we try to define our ensemble model as a weighted sum of L weak learners\n\n<img width=\"802\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/b9b991c4-03f5-4d5b-8862-68d4fece90a2\">\n\nFinding the best ensemble model with this form is a **difficult optimisation problem**. Then, instead of trying to solve it in one single shot (finding all the coefficients and weak learners that give the best overall additive model), we make use of an **iterative optimisation process** that is much more tractable, even if it can lead to a sub-optimal solution. More especially, we add the weak learners one by one, looking at each iteration for the best possible pair (coefficient, weak learner) to add to the current ensemble model. In other words, we define recurrently the (s_l)’s such that\n\n<img width=\"461\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/79875db8-5234-4a98-bf30-4c47d205c32b\">\n\nwhere c_l and w_l are chosen such that s_l is the model that fit the best the training data and, so, that is the best possible improvement over s_(l-1). We can then denote\n\n<img width=\"918\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/00d965ac-37d8-49ed-95d7-f51b43db528c\">\n\nwhere E(.) is the fitting error of the given model and e(.,.) is the loss/error function. Thus, instead of optimising “globally” over all the L models in the sum, we approximate the optimum by optimising “locally” building and adding the weak learners to the strong model one by one.\n\nMore especially, when considering a binary classification, we can show that the adaboost algorithm can be re-written into a process that proceeds as follow. First, it **updates the observations weights** in the dataset and train a new weak learner with a special focus given to the observations misclassified by the current ensemble model. Second, it **adds the weak learner to the weighted sum** according to an update coefficient that expresse the performances of this weak model: the better a weak learner performs, the more it contributes to the strong learner.\n\nSo, assume that we are facing a binary classification problem, with N observations in our dataset and we want to use adaboost algorithm with a given family of weak models. At the very beginning of the algorithm (first model of the sequence), all the observations have the same weights 1/N. Then, we repeat L times (for the L learners in the sequence) the following steps:\n\n- fit the best possible weak model with the current observations weights\n\n- compute the value of the update coefficient that is some kind of scalar evaluation metric of the weak learner that indicates how much this weak learner should be taken into account into the ensemble model\n\n- update the strong learner by adding the new weak learner multiplied by its update coefficient\n\n- compute new observations weights that expresse which observations we would like to focus on at the next iteration (weights of observations wrongly predicted by the aggregated model increase and weights of the correctly predicted observations decrease)\n\nRepeating these steps, we have then build sequentially our L models and aggregate them into a simple linear combination weighted by coefficients expressing the performance of each learner. Notice that there exists variants of the initial adaboost algorithm such that LogitBoost (classification) or L2Boost (regression) that mainly differ by their choice of loss function.\n\n<img width=\"1360\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/411230ec-d6e1-445f-9084-04824adc426c\">\n\n- ***Adaboost updates weights of the observations at each iteration. Weights of well classified observations decrease relatively to weights of misclassified observations. Models that perform better have higher weights in the final ensemble model.***\n\n- ***Mostly, decision tree algorithm is preferred as a base algorithm for Adaboost and in sklearn library the default base algorithm for Adaboost is decision tree*** (AdaBoostRegressor and AdaBoostClassifier). By default, decision trees are used as base estimators. In order to obtain better results, the parameters of the decision tree can be tuned. You can also tune the number of base estimators. \n\n## <span style=\"color: #016FD0;\">Parameters - Fine Tuning</span>\n\n\n- `base_estimator`: object, optional (default=None). The base estimator from which the boosted ensemble is built. If None, then the base estimator is DecisionTreeClassifier(max_depth=1)\n\n- `n_estimators`: integer, optional (default=50). The maximum number of estimators at which boosting is terminated. In case of perfect fit, the learning procedure is stopped early.\n\n- `learning_rate`: float, optional (default=1.) **Learning rate shrinks the contribution of each classifier by learning_rate.**\n\n- `algorithm`: {‘SAMME’, ‘SAMME.R’}, optional (default=’SAMME.R’) If ‘SAMME.R’ then use the SAMME.R real boosting algorithm. base_estimator must support calculation of class probabilities. If ‘SAMME’ then use the SAMME discrete boosting algorithm.\n\n- `random_state`: int, RandomState instance or None, optional (default=None)\n\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Gradient Boosting</div>\n\n## <span style=\"color: #016FD0;\">Definition of Weakness of GB</span>\n\nGradient boosting approaches the problem a bit differently. Instead of adjusting weights of data points, **Gradient boosting focuses on the difference between the prediction and the ground truth.**\n\n\n## <span style=\"color: #016FD0;\">Pseudocode of GB</span>\n\nGradient boosting **requires a differential loss function** and works for both regression and classifications. I’ll use a simple Least Square as the loss function (for regression). The algorithm for classifications shares the same idea, but the math is slightly more complicated. Following is a visualization of how weak estimators H are built over time. Each time we fit a new estimator (regression tree with max_depth =3 in this case) to the gradient of loss(LS in this case).\n\n<img width=\"880\" alt=\"image\" src=\"https://github.com/user-attachments/assets/e3a15904-1aab-4a80-8573-60d0fbd5aec1\">\n\n## <span style=\"color: #016FD0;\">High Level Overview</span>\n\nGradient boosting is a special case of boosting algorithm where errors are minimized by a gradient descent algorithm and produce a model in the form of weak prediction models e.g. decision trees.\n\nThe major difference between boosting and gradient boosting is **how both the algorithms update model (weak learners) from wrong predictions.**\n\n- **Gradient boosting adjusts weights by the use of gradient (a direction in the loss function) using an algorithm called Gradient Descent, which iteratively optimizes the loss of the model by updating weights.** Loss normally means the difference between the predicted value and actual value. \n    - For **regression algorithms**, we use **MSE (Mean Squared Error)** loss as an evaluation metric \n    - while for **classification problems,** we use **logarithmic loss.**\n\n<img width=\"793\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/5b137827-e5df-4f4c-b53c-63b2c510a76d\">\n\n<img width=\"913\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/17794016-5c94-4a69-93d0-72550097654d\">\n\nGradient boosting uses Additive Modeling in which a new decision tree is added one at a time to a model that minimizes the loss using gradient descent. Existing trees in the model remain untouched and thus slow down the rate of overfitting. The output of the new tree is combined with the output of existing trees until the loss is minimized below a threshold or specified limit of trees is reached.\n\nAdditive Modeling in mathematics is a breakdown of a function into the addition of N subfunctions. In statistical terms, it can be **thought of as a regression model in which response y is the arithmetic sum of individual effects of predictor variables x.**\n\nGradient boosted decision trees are a type of boosting algorithm that uses gradient descent. Like other boosting methodologies, gradient boosting starts with a weak learner to make predictions. The first decision tree in gradient boosting is called the base learner. Next, new trees are created in an additive manner based on the base learner’s mistakes. The algorithm then calculates the residuals of each tree’s predictions to determine how far off the model’s predictions were from reality. Residuals are the difference between the model’s predicted and actual values. The residuals are then aggregated to score the model with a loss function.\n\nIn machine learning, loss functions are used to measure a model's performance. The gradient in gradient boosted decision trees refers to gradient descent. Gradient descent is used to minimize the loss (i.e. to improve the model’s performance) when we train new models. Gradient descent is a popular optimization algorithm used to minimize the loss function in machine learning problems. Some examples of loss functions include mean squared error or mean absolute error for regression problems, cross-entropy loss for classification problems or custom loss functions may be developed for a specific use case and dataset.\n\n## <span style=\"color: #016FD0;\">Gradient boosting - How it Works</span>\n\nIn gradient boosting, the ensemble model we try to build is also a weighted sum of weak learners\n\n<img width=\"792\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/e40fb448-8eb6-48fe-ab85-0bbe0862f864\">\n\nJust as we mentioned for adaboost, finding the optimal model under this form is too difficult and an iterative approach is required. The main difference with adaptative boosting is in the definition of the sequential optimisation process. Indeed, gradient boosting **casts the problem into a gradient descent one:** at each iteration we fit a weak learner to the opposite of the gradient of the current fitting error with respect to the current ensemble model. Let’s try to clarify this last point. First, theoretical gradient descent process over the ensemble model can be written\n\n<img width=\"832\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/7cac0010-1ca8-4aa6-94ab-32076596b55d\">\n\nis the opposite of the gradient of the fitting error with respect to the ensemble model at step l-1. This (pretty abstract) opposite of the gradient is a function that can, in practice, only be evaluated for observations in the training dataset (for which we know inputs and outputs): these evaluations are called **pseudo-residuals** attached to each observations. Moreover, even if we know for the observations the values of these pseudo-residuals, we don’t want to add to our ensemble model any kind of function: we only want to add a new instance of weak model. So, the natural thing to do is to **fit a weak learner to the pseudo-residuals computed for each observation.** Finally, the coefficient c_l is computed following a one dimensional optimisation process (line-search to obtain the best step size c_l).\n\nSo, assume that we want to use gradient boosting technique with a given family of weak models. At the very beginning of the algorithm (first model of the sequence), the pseudo-residuals are set equal to the observation values. Then, we repeat L times (for the L models of the sequence) the following steps:\n\n- fit the best possible weak learner to pseudo-residuals (approximate the opposite of the gradient with respect to the current strong learner)\n\n- compute the value of the optimal step size that defines by how much we update the ensemble model in the direction of the new weak learner\n\n- update the ensemble model by adding the new weak learner multiplied by the step size (make a step of gradient descent)\n\n- compute new pseudo-residuals that indicate, for each observation, in which direction we would like to update next the ensemble model predictions\n\n\nRepeating these steps, we have then build sequentially our L models and aggregate them following a gradient descent approach. Notice that, while adaptative boosting tries to solve at each iteration exactly the “local” optimisation problem (find the best weak learner and its coefficient to add to the strong model), gradient boosting uses instead a gradient descent approach and can more easily be adapted to large number of loss functions. **Thus, gradient boosting can be considered as a generalization of adaboost to arbitrary differentiable loss functions.**\n\n<img width=\"1348\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/3b0d65f7-5c42-4e35-876b-5f88fcd0b5bd\">\n\n\n***Gradient boosting updates values of the observations at each iteration. Weak learners are trained to fit the pseudo-residuals that indicate in which direction to correct the current ensemble model predictions to lower the error.***\n\n\n1. **Gradient boosting trees:** Gradient tree boosting also combines a set of weak learners to form a strong learner. There are three main items to note, as far as gradient boosting trees are concerned:\n    - a differential **loss function** has to be used, \n    - **decision trees** are used as weak learners, \n    - it’s an additive model, so trees are added one after the other. **Gradient descent** is used to minimize the loss when adding subsequent trees. \n\n    1. **eXtreme Gradient Boosting:** eXtreme Gradient Boosting, popularly known as XGBoost, is a top gradient boosting framework. It’s based on an ensemble of weak decision trees. It can do parallel computations on a single computer. **The algorithm uses regression trees for the base learner.** It also has cross-validation built-in. Developers love it for its accuracy, efficiency, and feasibility.\n    \n    1. **LightGBM:** LightGBM is a gradient boosting algorithm based on tree learning. Unlike other tree-based algorithms that use depth-wise growth, LightGBM uses leaf-wise tree growth. Leaf-wise growth algorithms tend to converge faster than dep-wise-based algorithms. \n    \n    1. **CatBoost** is a depth-wise gradient boosting library developed by Yandex. It grows a balanced tree using oblivion decision trees. As you can see in the image below, the same features are used when making left and right splits at each level.  Researchers need Catboost for the following reasons:\n\n    - The ability to handle categorical features natively,\n    - Models can be trained on several GPUs,\n    - It reduces parameter tuning time by providing great results with default parameters,\n    - Models can be exported to Core ML for on-device inference (iOS),\n    - It handles missing values internally,\n    - It can be used for both regression and classification problems.","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"color: #016FD0;\">Intuition of Gradient Boosting</span>\n\n### <span style=\"color: #016FD0;\">Roadmap</span>\n\nGradient boosting machines (GBMs) are currently very popular and so it's a good idea for machine learning practitioners to understand how GBMs work. The problem is that understanding all of the mathematical machinery is tricky and, unfortunately, these details are needed to tune the hyper-parameters. (Tuning the hyper-parameters is required to get a decent GBM model unlike, say, Random Forests.) Our goal in this article is to explain the intuition behind gradient boosting, provide visualizations for model construction, explain the mathematics as simply as possible, and answer thorny questions such as why GBM is performing “gradient descent in function space.” We've split the discussion into three morsels and a FAQ for easier digestion.\n\n1. **Gradient boosting: Distance to target:** First, we examine the most common form of GBM that optimizes the mean squared error (MSE), also called the L2 loss or cost. (The mean squared error is the average of the square of the difference between the true targets and the predicted values from a set of observations, such as a training or validation set.) As we'll see, A GBM is a composite model that combines the efforts of multiple weak models to create a strong model, and each additional weak model reduces the mean squared error (MSE) of the overall model. We give a fully-worked GBM example for a simple data set, complete with computations and model visualizations.\n\n1. **Gradient boosting: Heading in the right direction:** ***Optimizing a model according to MSE makes it chase outliers because squaring the difference between targets and predicted values emphasizes extreme values.*** When we can't remove outliers, it's better to optimize the mean absolute error (MAE), also called the L1 loss or cost. (The MAE is the average of the absolute value of the difference between the true targets and the predicted values.) This second article shows the computations and visualizations for a GBM that optimizes MAE for the same data set as used in the first article to optimize MSE.\n\n1. **Gradient boosting performs gradient descent:** The previous two articles give the intuition behind GBM and the simple formulas to show how weak models join forces to create a strong regression model. No attempt was made to show how we can abstract out a generalized GBM that works for any loss function. This last article demonstrates that gradient boosting is really doing a form of gradient descent and, therefore, is in fact optimizing MSE or MAE depending on the direction vectors we use to train the weak models. The discussion relies on a bit of derivative calculus but it's an important read if you'd like to learn deeply how GBM works. (To brush up on your vectors and derivatives, you can check out [The Matrix Calculus You Need For Deep Learning](https://explained.ai/matrix-calculus/index.html).) We finish off by clearing up a number of confusion points regarding gradient boosting. As Ben Gorman points out in [A Kaggle Master Explains Gradient Boosting](https://www.gormanalysis.com/blog/gradient-boosting-explained/), “This is the part that gets butchered by a lot of gradient boosting explanations.” His blog post does a good job of explaining it, but we give our own perspective here.\n\n1. **Frequently asked questions (FAQ):** There are lots of confusing things about gradient boosting machines and we have collected and answered a number of questions from students in this piece.\n\n### <span style=\"color: #016FD0;\">Gradient boosting: Distance to target</span>\n\n<img width=\"778\" alt=\"image\" src=\"https://github.com/user-attachments/assets/3626a49d-eedc-42dd-ac66-ddb030e4c90f\">\n<img width=\"761\" alt=\"image\" src=\"https://github.com/user-attachments/assets/c887ab6e-588c-43e0-a60d-80b6fcb12578\">\n<img width=\"791\" alt=\"image\" src=\"https://github.com/user-attachments/assets/160c7bdf-5273-424d-a5fd-eec8347e5b56\">\n<img width=\"790\" alt=\"image\" src=\"https://github.com/user-attachments/assets/ae760682-514a-404a-9843-3f5395a409f9\">\n<img width=\"778\" alt=\"image\" src=\"https://github.com/user-attachments/assets/8905fd1a-2dfc-43ab-9698-c466febc6ed7\">\n<img width=\"772\" alt=\"image\" src=\"https://github.com/user-attachments/assets/70b70aa6-ff6f-4d6e-aabd-405ee766d6e1\">\n<img width=\"792\" alt=\"image\" src=\"https://github.com/user-attachments/assets/78337d37-7f58-42d3-b6e6-14c93baa7a19\">\n<img width=\"791\" alt=\"image\" src=\"https://github.com/user-attachments/assets/67421cc2-ab51-429e-b7f6-f974c45a9b18\">\n<img width=\"758\" alt=\"image\" src=\"https://github.com/user-attachments/assets/ee471cff-0723-4676-8d28-2f90f18382d8\">\n<img width=\"815\" alt=\"image\" src=\"https://github.com/user-attachments/assets/12c8ba0f-ee63-4721-b764-be7a62412b33\">\n\n### <span style=\"color: #016FD0;\">Gradient boosting: Heading in the right direction</span>\n\n<img width=\"790\" alt=\"image\" src=\"https://github.com/user-attachments/assets/f93cc3d7-98c0-47f5-9dd3-787e83509781\">\n<img width=\"787\" alt=\"image\" src=\"https://github.com/user-attachments/assets/b6e36015-0ec9-493d-ba3f-9471bb91e444\">\n<img width=\"791\" alt=\"image\" src=\"https://github.com/user-attachments/assets/0c3b4ab8-9fe1-4699-9e0b-73f22aff5669\">\n<img width=\"799\" alt=\"image\" src=\"https://github.com/user-attachments/assets/0549250d-57c4-4c35-bfe7-b25d5ccd6656\">\n<img width=\"791\" alt=\"image\" src=\"https://github.com/user-attachments/assets/1a80b7ad-7250-4821-949d-416b0bc0b230\">\n<img width=\"795\" alt=\"image\" src=\"https://github.com/user-attachments/assets/3eb73ef5-6447-4ff6-9080-0a010f02c18b\">\n<img width=\"777\" alt=\"image\" src=\"https://github.com/user-attachments/assets/e51b1751-2095-4e10-ac48-ebf14cf7c806\">\n\n### <span style=\"color: #016FD0;\">Gradient boosting performs gradient descent</span>\n\n<img width=\"800\" alt=\"image\" src=\"https://github.com/user-attachments/assets/f1e4cb01-cd9e-4124-92df-6fc1a125e28f\">\n<img width=\"769\" alt=\"image\" src=\"https://github.com/user-attachments/assets/fe6206ab-430a-457e-a21c-7c357a955243\">\n<img width=\"760\" alt=\"image\" src=\"https://github.com/user-attachments/assets/ade06b70-fd9e-4088-aea7-b8224eb852bc\">\n<img width=\"770\" alt=\"image\" src=\"https://github.com/user-attachments/assets/3d735601-8412-4eb6-8dee-bd6eb32fd215\">\n<img width=\"773\" alt=\"image\" src=\"https://github.com/user-attachments/assets/96c8b5d0-01a1-4800-9197-a74b2ca583d7\">\n<img width=\"796\" alt=\"image\" src=\"https://github.com/user-attachments/assets/f7d7951c-d859-4726-9cea-ebf5db606cb3\">\n<img width=\"791\" alt=\"image\" src=\"https://github.com/user-attachments/assets/76821093-19dd-488e-ba43-c56016787c00\">\n<img width=\"788\" alt=\"image\" src=\"https://github.com/user-attachments/assets/2ba543df-8e90-4459-b302-e5b2bcff8913\">\n<img width=\"778\" alt=\"image\" src=\"https://github.com/user-attachments/assets/09d9ef6d-d998-4c7c-b4da-e89069f04618\">\n<img width=\"794\" alt=\"image\" src=\"https://github.com/user-attachments/assets/2a075d80-52ec-4649-9fc5-28f5927fc8db\">\n\n### <span style=\"color: #016FD0;\">Frequently asked questions (FAQ)</span>\n\n<img width=\"789\" alt=\"image\" src=\"https://github.com/user-attachments/assets/6ba6b787-6390-4f0a-b1bc-ff99f28e09b9\">\n<img width=\"794\" alt=\"image\" src=\"https://github.com/user-attachments/assets/14c10888-3235-4414-82d9-0fb0e30dc7f3\">\n<img width=\"797\" alt=\"image\" src=\"https://github.com/user-attachments/assets/a8d0dc36-4d74-436c-8b84-6bc22cbada38\">\n\n\n\n### <span style=\"color: #016FD0;\">Recap</span>\n\n<img width=\"800\" alt=\"image\" src=\"https://github.com/user-attachments/assets/2e5f9b6d-7199-4214-9c23-89e0a838bfc0\">\n<img width=\"788\" alt=\"image\" src=\"https://github.com/user-attachments/assets/3d9b9cdf-0fe9-4f1f-a968-5729cd83bb24\">\n<img width=\"772\" alt=\"image\" src=\"https://github.com/user-attachments/assets/aa194bd0-2ccc-4453-a2f9-705321dba5e5\">\n<img width=\"851\" alt=\"image\" src=\"https://github.com/user-attachments/assets/39dcb322-01e4-4195-8310-4ac78870e8af\">\n<img width=\"810\" alt=\"image\" src=\"https://github.com/user-attachments/assets/6811d280-3d5f-4a57-82ed-80f92412c294\">\n<img width=\"836\" alt=\"image\" src=\"https://github.com/user-attachments/assets/f201bce0-3302-4db8-9725-c54f25eaddec\">\n<img width=\"793\" alt=\"image\" src=\"https://github.com/user-attachments/assets/2c3cadc2-b3ba-4d3e-bb96-517a44b6085d\">\n<img width=\"849\" alt=\"image\" src=\"https://github.com/user-attachments/assets/ffac1ba6-5ccb-41d4-8f16-b1e4ddbb1e14\">\n\nAs mentioned in the previous section, ν is learning rate ranging between 0 and 1 which controls the degree of contribution of the additional tree prediction γ to the combined prediction F𝑚. A smaller learning rate reduces the effect of the additional tree prediction, but it basically also reduces the chance of the model overfitting to the training data.\n\nSome of you might feel that all those maths are unnecessarily complex as the previous section showed the basic idea in a much simpler way without all those complications. The reason behind it is that gradient boosting is designed to be able to deal with any loss functions as long as it is differentiable and the maths we reviewed is a generalized form of gradient boosting algorithm with that flexibility. That makes the formula a little complex, but it is the beauty of the algorithm as it has huge flexibility and convenience to work on a variety of types of problems. For example, if your problem requires absolute loss instead of squared loss, you can just replace the loss function and the whole algorithm works as it is as defined above. In fact, popular gradient boosting implementations such as XGBoost or LightGBM have a wide variety of loss functions, so you can choose whatever loss functions that suit your problem (see the various loss functions available in [XGBoost](https://xgboost.readthedocs.io/en/stable/parameter.html#learning-task-parameters) or [LightGBM](https://lightgbm.readthedocs.io/en/latest/Parameters.html#objective)).\n\n### <span style=\"color: #016FD0;\">Gradient boosting VS Stochastic Gradient Descent</span>\n\nGradient boosting solves a different problem than stochastic gradient descent.\n\nWhen optimizing a model using SGD, the architecture of the model is fixed. What you are therefore trying to optimize are the parameters, P of the model (in logistic regression, this would be the weights). Mathematically, this would look like this:\n\n<img width=\"601\" alt=\"image\" src=\"https://github.com/user-attachments/assets/5d5296d6-ee64-4a41-8a57-bc0aeafd223f\">\n\nWhich means I am trying to find the best parameters P for my function F, where ‘best’ means that they lead to the smallest loss possible (the vertical line in F(x∣P) just means that once I’ve found the parameters P, I calculate the output of F given x using them).\n\nGradient boosting doesn’t assume this fixed architecture. In fact, the whole point of gradient boosting is to find the function which best approximates the data. It would be expressed like this:\n\n<img width=\"643\" alt=\"image\" src=\"https://github.com/user-attachments/assets/a32ceeee-165a-489c-b9c4-26f59ed2e796\">\n\nThe only thing that has changed is that now, in addition to finding the best parameters P, I also want to find the best function F. This tiny change introduces a lot of complexity to the problem; whereas before, the number of parameters I was optimizing for was fixed (my logistic regression model is defined before I start training it), now, it can change as I go through the optimization process if my function F changes.\n\nObviously, searching all possible functions and their parameters to find the best one would take far too long, so gradient boosting finds the best function F by taking lots of simple functions, and adding them together.\n\n**Where SGD trains a single complex model, gradient boosting trains an ensemble of simple models.**\n\nIt does this the following way:\n\nTake a very simple model h, and fit it to some data (x, y):\n\n<img width=\"710\" alt=\"image\" src=\"https://github.com/user-attachments/assets/d44d59c5-8e66-41b2-8319-a50f600b5242\">\n\nWhen I’m training my second model, I obviously don’t want it to uncover the same pattern in the data as this first model h; ideally, it would improve on the errors from this first prediction. This is the clever part (and the ‘gradient’ part): this prediction will have some error, Loss(y, ŷ ). The next model I am going to fit will be on the gradient of the error with respect to the predictions, ∂Loss/∂ŷ .\n\nTo think about why this is clever, lets consider mean squared error:\n\n<img width=\"745\" alt=\"image\" src=\"https://github.com/user-attachments/assets/327a8137-bce8-4816-ac00-cee059377008\">\nCalculating this gradient,\n\n<img width=\"619\" alt=\"image\" src=\"https://github.com/user-attachments/assets/3ac6dbd8-b1cf-4d9a-bd45-8fef6a508b36\">\n\nIf for one data point, y=1 and ŷ =0.6, then the error in this prediction is MSE(1,0.6)=0.16 and the new target for the model will be the gradient, (y−ŷ )=0.4. Training a model on this target,\n\n<img width=\"726\" alt=\"image\" src=\"https://github.com/user-attachments/assets/0ea04ee0-4a88-42a6-8c0a-c39c02d7666d\">\n\nNow, for this same data point, where y=1 (and for the previous model, ŷ =0.6, the model is being trained to on a target of 0.4. Say that it returns ŷ_1=0.3. The last step in gradient boosting is to add these models together. For the two models I’ve trained (and for this specific data point), then\n\n<img width=\"679\" alt=\"image\" src=\"https://github.com/user-attachments/assets/533dd945-3ccd-4d05-a1ee-3486778cc457\">\n\n**By training my second model on the gradient of the error with respect to the loss predictions of the first model, I have taught it to correct the mistakes of the first model.** This is the core of gradient boosting, and what allows many simple models to compensate for each other’s weaknesses to better fit the data.\n\nI don’t have to stop at 2 models; I can keep doing this over and over again, each time fitting a new model to the gradient of the error of the updated sum of models.\n\nAn interesting note here is that at its core, gradient boosting is a method for optimizing the function F, but it doesn’t really care about h (since nothing about the optimization of h is defined). This means that any base model h can be used to construct F.\n","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"color: #016FD0;\">Parameters - Fine Tuning of GB</span>\n\n- `loss`: \n    - Regression: {‘ls’, ‘lad’, ‘huber’, ‘quantile’}, optional (default=’ls’)\n    - Classification: loss : {‘deviance’, ‘exponential’}, optional (default=’deviance’)\n\n- `learning_rate`: float, optional (default=0.1)\n\n- `n_estimators`: int (default=100). Gradient boosting is fairly robust to over-fitting so a large number usually results in better performance.\n\n- `subsample`: float, optional (default=1.0) The fraction of samples to be used for fitting the individual base learners. If smaller than 1.0 this results in Stochastic Gradient Boosting. subsample interacts with the parameter n_estimators. **Choosing subsample < 1.0 leads to a reduction of variance and an increase in bias.**\n\n- `criterion`: string, optional (default=”friedman_mse”) The function to measure the quality of a split.\n\n\n## <span style=\"color: #016FD0;\">Pros and Cons</span>\n\n### Pros\n\n1. Highly **efficient on both classification and regression tasks**\n\n1. **More accurate predictions** compared to random forests.\n\n1. Can handle **mixed type of features and no pre-processing is needed**\n\n### Cons\n\n1. Requires **careful tuning of hyperparameters**\n\n1. **May overfit if too many trees are used** (n_estimators)\n\n1. **Sensitive to outliers**. A common thing often forgotten is that Gradient Boosting is very sensitive to outliers since every classifier is forced to fix the errors in the predecessor learners.\n\n## <span style=\"color: #016FD0;\">Gradient Boosting Methods</span>\n\n### <span style=\"color: #016FD0;\">Gradient Boosting Machine (GBM)</span>\nGBM combines predictions from multiple decision trees, and all the weak learners are decision trees. The key idea with this algorithm is that every node of those trees takes a different subset of features to select the best split. As it’s a Boosting algorithm, each new tree learns from the errors made in the previous ones. A Gradient Boosting Machine or GBM combines the predictions from multiple decision trees to generate the final predictions. Keep in mind that all the weak learners in a gradient-boosting machine are decision trees.\n\n### <span style=\"color: #016FD0;\">Extreme Gradient Boosting Machine (XGBM)</span>\n\nExtreme Gradient Boosting or XGBoost is another popular boosting algorithm. In fact, **XGBoost** is simply an **improvised version of the GBM algorithm!** The working procedure of XGBoost is the same as GBM. The trees in XGBoost are built sequentially, trying to correct the errors of the previous trees. But there are certain features that make XGBoost slightly better than GBM:\n\n- One of the most important points is that XGBM implements parallel preprocessing (at the node level) which makes it faster than GBM\n\n- **XGBoost** also **includes** a variety of **regularization techniques that reduce overfitting and improve overall performance.** You can select the regularization technique by setting the hyperparameters of the XGBoost algorithm\n\n- Additionally, if you are using the XGBM algorithm, you don’t have to worry about imputing missing values in your dataset. The **XGBM** model **can handle the missing values on its own.** During the training process, the model learns whether missing values should be in the right or left node.\n\n### <span style=\"color: #016FD0;\">Light Gradient Boosting Machine (LightGBM)</span>\n\nLightGBM **can handle huge amounts of data. It’s one of the fastest algorithms for both training and prediction.** It generalizes well, meaning that it can be used to solve similar problems. It scales well to large numbers of cores and has open-source code, so you can use it in your projects for free. The LightGBM boosting algorithm is becoming more popular by the day **due to its speed and efficiency.** LightGBM is **able to handle huge amounts of data with ease.** But keep in mind that this algorithm does not perform well with a small number of data points.\n\nLet’s take a moment to understand why that’s the case.\n\nThe trees in LightGBM have a leaf-wise growth, rather than a level-wise growth. After the first split, the next split is done only on the leaf node that has a higher delta loss.\n\nConsider the example I’ve illustrated in the below image:\n\n<img width=\"857\" alt=\"image\" src=\"https://github.com/user-attachments/assets/e2eeaafd-9802-47bb-a0f9-bcfc37b8a745\">\n\nAfter the first split, the left node had a higher loss and is selected for the next split. Now, we have three leaf nodes, and the middle leaf node had the highest loss. The leaf-wise split of the LightGBM algorithm enables it to work with large datasets.\n\nIn order to speed up the training process, LightGBM uses a histogram-based method for selecting the best split. For any continuous variable, instead of using the individual values, these are divided into bins or buckets. This makes the training process faster and lowers memory usage.\n\n\n### <span style=\"color: #016FD0;\">Categorical Boosting (CatBoost)</span>\n\nThe CatBoost algorithm, a specific form of gradient boosting, specializes in working with **very diverse kinds of data.** It **shines with categorical data, but it also does well with numeric data and with datasets that contain both kinds of variables.** Most gradient-boosting algorithms can work reasonably well with categorical variables, but CatBoost outperforms them because of how it handles these variables. Not only is CatBoost good at learning with the kinds of variables that many machine learning models struggle with, but it also learns efficiently from unlabeled data—that is, from data without any obvious outcomes or targets. As the name suggests, CatBoost is a boosting algorithm that can handle categorical variables in the data. Most machine learning algorithms cannot work with strings or categories in the data. Thus, converting categorical variables into numerical values is an essential preprocessing step.\n\nCatBoost can internally handle categorical variables in the data. These variables are transformed to numerical ones using various statistics on combinations of features.\n\n**Another reason why CatBoost is being widely used is that it works well with the default set of hyperparameters.** Hence, as a user, we do not have to spend a lot of time tuning the hyperparameters.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Stacking</div>\n\nStacking mainly differ from bagging and boosting on two points. \n\n1. First stacking often considers **heterogeneous weak learners** (different learning algorithms are combined) whereas bagging and boosting consider mainly homogeneous weak learners. \n1. Second, stacking learns to combine the base models **using a meta-model whereas bagging and boosting combine weak learners following deterministic algorithms.**\n\nAs we already mentioned, the idea of stacking is to learn several different weak learners and **combine them by training a meta-model to output predictions based on the multiple predictions returned by these weak models.** So, we need to define two things in order to build our stacking model: the L learners we want to fit and the meta-model that combines them.\n\nFor example, for a classification problem, we can choose as weak learners a KNN classifier, a logistic regression and a SVM, and decide to learn a neural network as meta-model. Then, the neural network will take as inputs the outputs of our three weak learners and will learn to return final predictions based on it.\n\nSo, assume that we want to fit a stacking ensemble composed of L weak learners. Then we have to follow the steps thereafter:\n\n- split the training data in two folds\n- choose L weak learners and fit them to data of the first fold\n- for each of the L weak learners, make predictions for observations in the second fold\n- fit the meta-model on the second fold, using predictions made by the weak learners as inputs\n\nIn the previous steps, we split the dataset in two folds because predictions on data that have been used for the training of the weak learners are **not relevant for the training of the meta-model.** Thus, an obvious drawback of this split of our dataset in two parts is that we only have half of the data to train the base models and half of the data to train the meta-model. In order to overcome this limitation, we can however follow some kind of “k-fold cross-training” approach (similar to what is done in k-fold cross-validation) such that all the observations can be used to train the meta-model: for any observation, the prediction of the weak learners are done with instances of these weak learners trained on the k-1 folds that do not contain the considered observation. In other words, it consists in training on k-1 fold in order to make predictions on the remaining fold and that iteratively so that to obtain predictions for observations in any folds. Doing so, we can produce relevant predictions for each observation of our dataset and then train our meta-model on all these predictions.\n\n\n<img width=\"934\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/51e96195-d424-4646-b667-a572b5583db8\">\n\n<img width=\"801\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/7727f726-b6b1-4276-81a1-b176801d1a9f\">\n\n***Stacking consists in training a meta-model to produce outputs based on the outputs returned by some lower layer weak learners.***\n\n- Instead of using simple methods like averaging or voting, **stacking trains a meta-model to learn how to combine the base models' predictions best**.\n\n- The base models can be diverse to capture different aspects of the data, and the meta-model learns to weight its predictions based on its performance.\n\n\n**Multi-levels Stacking**\n\nA possible extension of stacking is multi-level stacking. It consists in doing **stacking with multiple layers**. As an example, let’s consider a 3-levels stacking. In the first level (layer), we fit the L weak learners that have been chosen. Then, in the second level, instead of fitting a single meta-model on the weak models predictions (as it was described in the previous subsection) we fit M such meta-models. Finally, in the third level we fit a last meta-model that takes as inputs the predictions returned by the M meta-models of the previous level.\n\nFrom a practical point of view, notice that for each meta-model of the different levels of a multi-levels stacking ensemble model, we have to choose a learning algorithm that can be almost whatever we want (even algorithms already used at lower levels). We can also mention that **adding levels can either be data expensive** (if k-folds like technique is not used and, then, more data are needed) **or time expensive** (if k-folds like technique is used and, then, lot of models need to be fitted).\n\n<img width=\"1165\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/db137a29-5b18-44da-8a71-6c89032546f1\">\n\n\n***Multi-level stacking considers several layers of stacking: some meta-models are trained on outputs returned by lower layer meta-models and so on. Here we have represented a 3-layers stacking model.*** However, such practices become computationally very expensive for a relatively small boost in performance.‍\n\n<img width=\"974\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/2d00b796-3845-475c-b4bc-a1c2c392e368\">\n\n<img width=\"1025\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/eb857973-c2d5-44a1-8735-964e42fde486\">\n\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Blending</div>\n\n- Blending is similar to stacking but more straightforward.\n\n- Instead of a meta-model, blending uses a simple method like averaging or a linear model to combine the predictions of the base models.\n\n- Blending is often used in competitions where simplicity and efficiency are important.\n\n## <span style=\"color: #016FD0;\">Other</span>\n\nEspecially in deep learning, it is a costly operation, even with transfer learning. So, this ensemble learning method proposed in this paper trains only one deep learning model and saves the model snapshots at different training epochs. \n\nThe ensemble of these models generates a final ensemble prediction framework on the test data.\n\nThey proposed some modifications to the usual deep learning model training regime to ensure the diversity in the model snapshots. The model weights saved at these different epochs need to be significantly different to make the ensemble successful.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Takeaways</div>\n\nThe main takeaways of this post are the following:\n\n- ensemble learning is a machine learning paradigm where multiple models (often called weak learners or base models) are trained to solve the same problem and combined to get better performances\n\n- the main hypothesis is that if we combine the weak learners the right way we can obtain more accurate and/or robust models\n\n- in bagging methods, several instance of the same base model are trained in parallel (independently from each others) on different bootstrap samples and then aggregated in some kind of “averaging” process\n\n- the kind of averaging operation done over the (almost) i.i.d fitted models in bagging methods mainly allows us to obtain an ensemble model with a lower variance than its components: that is why base models with low bias but high variance are well adapted for bagging\n\n- in boosting methods, several instance of the same base model are trained sequentially such that, at each iteration, the way to train the current weak learner depends on the previous weak learners and more especially on how they are performing on the data\n\n- this iterative strategy of learning used in boosting methods, that adapts to the weaknesses of the previous models to train the current one, mainly allows us to get an ensemble model with a lower bias than its components: that is why weak learners with low variance but high bias are well adapted for boosting\n\n- in stacking methods, different weak learners are fitted independently from each others and a meta-model is trained on top of that to predict outputs based on the outputs returned by the base models\n\n- Although ensemble methods can help you win machine learning competitions by devising sophisticated algorithms and producing results with high accuracy, **it is often not preferred in the industries where interpretability is more important.** Nonetheless, the effectiveness of these methods are undeniable, and their benefits in appropriate applications can be tremendous. In fields such as healthcare, even the smallest amount of improvement in the accuracy of machine learning algorithms can be something truly valuable.\n\n\n<img width=\"1054\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/ab187de1-c83e-40a2-8a30-f17e17d52246\">\n\n**Similarities:**\n\n1. **Ensemble methods:** In a general view, the similarities between both techniques start with the fact that both are ensemble methods with the aim to use multiple learners over a single model to achieve better results.\n\n1. **Multiple samples & aggregation:** To do that, both methods generate random samples and multiple training data sets. It is also similar that Bagging and Boosting both arrive at the end decision by aggregation of the underlying models: either by calculating average results or by taking a voting rank.\n\n1. **Purpose:** Finally, it is reasonable that both aim to produce higher stability and better prediction for the data.\n\n**Differences:**\n\n1. **Data partition | whole data vs. bias:** While bagging uses random bags out of the training data for all models independently, boosting puts higher importance on misclassified data of the upcoming models. Therefore, the data partition is different here.\n\n1. **Models | independent vs. sequences:** Bagging creates independent models that are aggregated together. However, boosting updates the existing model with the new ones in a sequence. Therefore, the models are affected by previous builds.\n\n1. **Goal | variance vs. bias:** Another difference is the fact that bagging aims to reduce the variance, but boosting tries to reduce the bias. Therefore, bagging can help to decrease overfitting, and boosting can reduce underfitting.\n\n1. **Function | weighted vs. non-weighted:** The final function to predict the outcome uses equally weighted average or equally weighted voting aggregations within the bagging technique. Boosting uses weighted majority vote or weighted average functions with more weight to those with better performance on training data.\n\n\nIn this post we have given a basic overview of ensemble learning and, more especially, of some of the main notions of this field: bootstrapping, bagging, random forest, boosting (adaboost, gradient boosting) and stacking. Among the notions that were left aside we can mention for example the Out-Of-Bag evaluation technique for bagging or also the very popular “XGBoost” (that stands for eXtrem Gradient Boosting) that is a library that implements Gradient Boosting methods along with a great number of additional tricks that make learning much more efficient (and tractable for big dataset).\n\nFinally, we would like to conclude by reminding that ensemble learning is about combining some base models in order to obtain an ensemble model with better performances/properties. Thus, even if bagging, boosting and stacking are the most commonly used ensemble methods, variants are possible and can be designed to better adapt to some specific problems. This mainly requires two things: fully understand the problem we are facing… and be creative!\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Gradient Boosting Decision Trees (GBDT) - Gradient Boosted Trees</div>\n\nGradient boosted trees consider the special case where the simple model h is a decision tree.\n\n<img width=\"922\" alt=\"image\" src=\"https://github.com/user-attachments/assets/75082311-0e68-4718-b94a-c0bc40eb0cee\">\n\nIn this case, there are going to be 2 kinds of parameters P: the weights at each leaf, w, and the number of leaves T in each tree (so that in the above example, T=3 and w=[2, 0.1, -1]).\n\nWhen building a decision tree, a challenge is to decide how to split a current leaf. For instance, in the above image, how could I add another layer to the (age > 15) leaf? A ‘greedy’ way to do this is to consider every possible split on the remaining features (so, gender and occupation), and calculate the new loss for each split; you could then pick the tree which most reduces your loss.\n\n<img width=\"966\" alt=\"image\" src=\"https://github.com/user-attachments/assets/60c263b6-fb88-4a6d-8230-ff3a4e4e11e7\">\n\nIn addition to finding the new tree structures, the weights at each node need to be calculated as well, such that the loss is minimized. Since the tree structure is now fixed, this can be done analytically now by setting the loss function = 0 \n\nA Gradient Boosting Decision Trees (GBDT) is a decision tree ensemble learning algorithm similar to random forest, for classification and regression. Ensemble learning algorithms combine multiple machine learning algorithms to obtain a better model.\n\n- Both random forest and GBDT build a model consisting of multiple decision trees. The difference is in how the trees are built and combined.\n\n<img width=\"511\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/94d962ce-a1c9-47ac-9dc7-374c52ba08dc\">\n\n- Random forest uses a technique called bagging to build full decision trees in parallel from random bootstrap samples of the data set. The final prediction is an average of all of the decision tree predictions.\n\n- The term “gradient boosting” comes from the idea of “boosting” or improving a single weak model by combining it with a number of other weak models in order to generate a collectively strong model. Gradient boosting is an extension of boosting where the process of additively generating weak models is formalized as a gradient descent algorithm over an objective function. Gradient boosting sets targeted outcomes for the next model in an effort to minimize errors. Targeted outcomes for each case are based on the gradient of the error (hence the name gradient boosting) with respect to the prediction.\n\n- GBDTs iteratively train an ensemble of shallow decision trees, with each iteration using the error residuals of the previous model to fit the next model. The final prediction is a weighted sum of all of the tree predictions. Random forest “bagging” minimizes the variance and overfitting, while GBDT “boosting” minimizes the bias and underfitting.\n\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">XGBoost</div>\n\n## <span style=\"color: #016FD0;\">Introduction to XGBoost</span>\n\nXGBoost is one of the fastest implementations of gradient boosted trees.\n\nIt does this by tackling one of the major inefficiencies of gradient boosted trees: considering the potential loss for all possible splits to create a new branch (especially if you consider the case where there are thousands of features, and therefore thousands of possible splits). XGBoost tackles this inefficiency by looking at the distribution of features across all data points in a leaf and using this information to reduce the search space of possible feature splits.\n\nAlthough XGBoost implements a few regularization tricks, this speed up is by far the most useful feature of the library, allowing many hyperparameter settings to be investigated quickly. \n\nIn prediction problems involving unstructured data (images, text, etc.) artificial neural networks tend to outperform all other algorithms or frameworks. However, when it comes to small-to-medium structured/tabular data, decision tree based algorithms are considered best-in-class right now. Please see the chart below for the evolution of tree-based algorithms over the years.\n\n<img width=\"840\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/fed4527b-3d4a-4649-8f0e-7572d8e7b5d5\">\n\n**XGBoost** stands for e**X**treme **G**radient **Boost**ing and it’s an open-source implementation of the **gradient boosted trees algorithm.** It has been one of the most popular machine learning techniques in Kaggle competitions, due to its **prediction power** and **ease of use**. It is a supervised learning algorithm that can be used for **regression** or **classification tasks**.\n\nRegardless of its futuristic name, it’s actually not that hard to understand, as long as we first go through a few concepts: decision trees and gradient boosting. \n\n\n## <span style=\"color: #016FD0;\">What is XGBoost?</span>\n\n- XGBoost, which stands for Extreme Gradient Boosting, is a scalable, distributed **gradient-boosted decision tree (GBDT)** machine learning library. It provides parallel tree boosting and is the leading machine learning library for **regression**, **classification**, **recommendations**, and **ranking problems**.\n\ne**X**treme **G**radient **Boost**ing (XGBoost) is a scalable and improved version of the gradient boosting algorithm (terminology alert) designed for efficacy, computational speed and model performance. It is an open-source library and a part of the Distributed Machine Learning Community. XGBoost is a perfect blend of software and hardware capabilities designed to enhance existing boosting techniques with accuracy in the shortest amount of time. XGBoost is a gradient boosting algorithm that is widely used in data science. It is an implementation of gradient boosting that is designed to be highly efficient, flexible and portable. **XGBoost was originally developed** by Tianqi Chen, **as an improvement on the GBM algorithm.** The algorithm was designed with the following goals in mind. The main purpose of XGBoost is to improve the performance and speed of gradient boosting models, which are used for both classification and regression tasks. XGBoost is widely appreciated for its:\n\n- *To be highly efficient*,  **Speed and Efficiency** - It’s optimized for speed and performance, both in terms of training time and computational efficiency.\n-  **Scalability** - It can handle large datasets and can be distributed across clusters.\n- *To be flexible*, **Flexibility** - It supports custom optimization objectives and evaluation criteria, allowing for fine-tuned control over the model.\n- *To be portable*\n- *It is accurate.*, **High Predictive Power** - It often produces models with high accuracy, making it a top choice in data science competitions.\n\nXGBoost has been shown to outperform other machine learning algorithms in a variety of tasks, including **classification**, **regression** and **ranking**. XGBoost works by **training a number of decision trees.** Each tree is trained on a subset of the data, and the predictions from each tree are combined to form the final prediction. **XGBoost is an improvement on the GBM algorithm.** The main difference is that XGBoost uses a more **regularized model**, which helps to prevent overfitting.\n\n\nXGBoost is **a decision-tree-based ensemble Machine Learning algorithm** that uses a **gradient boosting framework**. XGBoost is an implementation of **gradient-boosting decision trees.** Gradient boosting is a ML algorithm that creates a series of models and combines them to create an overall model that is more accurate than any individual model in the sequence. It supports both regression and classification predictive modeling problems. To add new models to an existing one, it uses a gradient descent algorithm called gradient boosting. Gradient boosting is implemented by the XGBoost library, also known as multiple additive regression trees, stochastic gradient boosting, or gradient boosting machines.\n\nXGBoost algorithm was developed as a research project at the University of Washington. Tianqi Chen and Carlos Guestrin presented their paper at SIGKDD Conference in 2016 and caught the Machine Learning world by fire. Since its introduction, this algorithm has not only been credited with winning numerous Kaggle competitions but also for being the driving force under the hood for several cutting-edge industry applications. As a result, there is a strong community of data scientists contributing to the XGBoost open source projects with ~350 contributors and ~3,600 commits on [GitHub](https://github.com/dmlc/xgboost/). The algorithm differentiates itself in the following ways:\n\n1. A wide range of applications: Can be used to solve **regression**, **classification**, **ranking**, and user-defined prediction problems.  It's an open-source library that can train and test models on large amounts of data. It has been used in many domains, from predicting ad click-through rates to classifying high-energy physics events. \n\n1. XGBoost is particularly popular because it's so **fast**, and that speed comes at **no cost to accuracy!** XGBoost is designed for speed, **ease of use**, and performance on large datasets. It **does not require optimization of the parameters or tuning**, which means that it can be used immediately after installation without any further configuration. **Execution speed** is crucial because it's essential to working with large datasets. When you use XGBoost, there are **no restrictions on the size of your dataset,** so you can work with datasets that are larger than what would be possible with other algorithms. **Model performance is also essential because it allows you to create models that can perform better than other models.** XGBoost has been compared to different algorithms such as random forest (RF), gradient boosting machines (GBM), and gradient boosting decision trees (GBDT). These comparisons show that XGBoost outperforms these other algorithms in execution speed and model performance.\n\n1. **Portability:** Runs smoothly on Windows, Linux, and OS X. It's also used in production by organizations across various verticals, including finance and retail.\n\n1. **Languages:** Supports all major programming languages including C++, Python, R, Java, Scala, and Julia.\n\n1. **Cloud Integration:** Supports AWS, Azure, and Yarn clusters and works well with Flink, Spark, and other ecosystems.\n\n1. XGBoost is **open source**, so it's free to use, and it has a large and growing community of data scientists actively contributing to its development. The library was built from the ground up to be efficient, flexible, and portable.\n\n\n## <span style=\"color: #016FD0;\">Historical Background</span>\n\nDecision trees, in their simplest form, **are easy-to-visualize and fairly interpretable algorithms** but building intuition for the next-generation of tree-based algorithms can be a bit tricky. See below for a simple analogy to better understand the evolution of tree-based algorithms. Imagine that you are a hiring manager interviewing several candidates with excellent qualifications. Each step of the evolution of tree-based algorithms can be viewed as a version of the interview process.\n\n1. **Decision Tree:** Every hiring manager has a set of criteria such as education level, number of years of experience, interview performance. A decision tree is analogous to a hiring manager interviewing candidates based on his or her own criteria.\n\n1. **Bagging:** Now imagine instead of a single interviewer, now there is an interview panel where each interviewer has a vote. Bagging or bootstrap aggregating involves combining inputs from all interviewers for the final decision through a democratic voting process.\n\n1. **Random Forest:** It is a bagging-based algorithm with a key difference wherein only a subset of features is selected at random. In other words, every interviewer will only test the interviewee on certain randomly selected qualifications (e.g. a technical interview for testing programming skills and a behavioral interview for evaluating non-technical skills).\n\n1. **Boosting:** This is an alternative approach where each interviewer alters the evaluation criteria based on feedback from the previous interviewer. This ‘boosts’ the efficiency of the interview process by deploying a more dynamic evaluation process.\n\n1. **Gradient Boosting:** A special case of boosting where errors are minimized by gradient descent algorithm e.g. the strategy consulting firms leverage by using case interviews to weed out less qualified candidates\n\n1. **XGBoost:** Think of XGBoost as gradient boosting on ‘steroids’ (well it is called ‘Extreme Gradient Boosting’ for a reason!). It is a perfect combination of **software and hardware optimization** techniques to yield superior results using **less computing resources in the shortest amount of time.**\n\n\n## <span style=\"color: #016FD0;\">How XGBoost works</span>\n\nXGBoost minimizes a regularized (L1 and L2) objective function that combines a convex loss function (based on the difference between the predicted and target outputs) and a penalty term for model complexity (in other words, the regression tree functions). The training proceeds iteratively, adding new trees that predict the residuals or errors of prior trees that are then combined with previous trees to make the final prediction. It's called gradient boosting because it uses a gradient descent algorithm to minimize the loss when adding new models.\n\nBoosting is an ensemble method, meaning it’s a way of combining predictions from several models into one. It does that by taking each predictor sequentially and modelling it based on its predecessor’s error (giving more weight to predictors that perform better):\n\n1. Fit a first model using the original data\n1. Fit a second model using the residuals of the first model\n1. Create a third model using the sum of models 1 and 2\n\n**Gradient boosting is a specific type of boosting,** called like that because it minimises the loss function using a gradient descent algorithm.\n\nNow that you understand decision trees and gradient boosting, understanding XGBoost becomes easy: it is a gradient boosting algorithm that uses decision trees as its “weak” predictors. ***Beyond that, its implementation was specifically engineered for optimal performance and speed.***\n\nHistorically, **XGBoost has performed quite well for structured, tabular data.** If you are dealing with non-structured data such as images, neural networks are usually a better option.\n\nXGBoost works by implementing gradient boosting in a highly optimized way. Here’s a step-by-step explanation of the process:\n\n1. **Initialization** - The process begins with an initial prediction, which is usually the mean of the target values for regression or the mode for classification.\n\n1. **Boosting Iterations:**\n\n    - **Compute Residuals** - For each instance in the dataset, compute the difference (residual) between the actual target value and the current prediction.\n    - **Fit a Weak Learner** - Fit a weak learner (typically a decision tree) to the residuals. The goal of this learner is to predict the residuals of the previous model.\n    - **Update Model** - Add the predictions of the weak learner to the overall model. This updates the current predictions to reduce the residuals.\n    - **Shrinkage** - Apply a learning rate to shrink the contribution of each weak learner, which helps prevent overfitting.\n\n1. **Regularization:** XGBoost includes several regularization techniques to improve generalization and reduce overfitting:\n    \n    - **L1 (Lasso) Regularization** - Encourages sparsity in the model, which can lead to simpler models.\n    - **L2 (Ridge) Regularization** - Helps distribute the weights more evenly and prevents them from becoming too large.\n    - **Tree Pruning** - Prunes trees to prevent overfitting by limiting the maximum depth or using a minimum loss reduction threshold for a split to be added.\n\n1. **Handling Missing Values** - XGBoost can handle missing values internally by learning the best imputation strategy based on the training data.\n\n1. **Parallel and Distributed Computing** - XGBoost leverages parallel processing and can be distributed across multiple machines to handle large datasets efficiently.\n\n1. **Optimizations** - Various optimizations like cache awareness, out-of-core computing, and optimized data structures are used to enhance performance.\n\n## <span style=\"color: #016FD0;\">Why does XGBoost perform so well?</span>\n\nThe two reasons to use XGBoost are also the two goals of the project:\n\n1. **Execution Speed**\n1. **Model Performance**\n\nThe implementation of the algorithm was engineered for efficiency of compute time and memory resources. A design goal was to make the best use of available resources to train the model. \n\n**XGBoost** and **Gradient Boosting Machines (GBMs)** are both **ensemble tree methods** that apply the principle of boosting weak learners (**CARTs** generally) using the gradient descent architecture. However, XGBoost improves upon the base GBM framework through systems optimization and algorithmic enhancements.\n\n<img width=\"829\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/61c70ba1-1662-4588-8165-3586c0dc93ab\">\n\n***What makes XGBoost a go-to algorithm for winning Machine Learning and Kaggle competitions?***\n\n<img width=\"778\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/8fbd19ca-163f-48ce-8c7c-b549ad57c140\">\n\nIsn’t it interesting to see a single tool to handle all our boosting problems! Here are the features with details and how they are incorporated in XGBoost to make it robust. XGBoost can be used with a simple SKlearn API (used in this tutorial) or a more flexible native API (used in the upcoming advanced tutorial). It is also available for other languages such as R, Java, Scala, C++, etc. and **can run on distributed environments such as Hadoop and Spark.** A word of warning before going further with the tutorial, as it is well known, there is no free lunch in Machine Learning, and XGBoost is no exception to this rule. **It can sometimes be harder to tune or have a higher tendency to overfitting than a simpler model such as Random Forest and perform poorly with non structured data.**\n\n### System Optimization:\n\n1. **Parallelization:** XGBoost approaches the process of sequential tree building using parallelized implementation. This is possible due to the interchangeable nature of loops used for building base learners; the outer loop that enumerates the leaf nodes of a tree, and the second inner loop that calculates the features. This nesting of loops limits parallelization because without completing the inner loop (more computationally demanding of the two), the outer loop cannot be started. Therefore, to improve run time, the order of loops is interchanged using initialization through a global scan of all instances and sorting using parallel threads. This switch improves algorithmic performance by offsetting any parallelization overheads in computation. Tree learning needs data in a sorted manner. To cut down the sorting costs, data is divided into compressed blocks (each column with corresponding feature value). XGBoost sorts each block parallelly using all available cores/threads of CPU. This optimization is valuable since a large number of nodes gets created frequently in a tree. In summary, XGBoost parallelizes the sequential process of generating trees. **With XGBoost, trees are built in parallel, instead of sequentially like GBDT. It follows a level-wise strategy, scanning across gradient values and using these partial sums to evaluate the quality of splits at every possible split in the training set.** ***XGBoost can also be implemented in its distributed mode using tools like Apache Spark, Dask or Kubernetes.***\n\n1. **Tree Pruning:**  Uses a technique called **\"max_depth\"** to limit the depth of trees, thus **preventing overfitting.** The stopping criterion for tree splitting within GBM framework is greedy in nature and depends on the negative loss criterion at the point of split. XGBoost uses ‘max_depth’ parameter as specified instead of criterion first, and starts pruning trees backward. This ‘depth-first’ approach improves computational performance significantly. Pruning is a machine learning technique to reduce the size of regression trees by replacing nodes that don’t contribute to improving classification on leaves. The idea of pruning a regression tree is to **prevent overfitting of the training data.** The most efficient method to do pruning is Cost Complexity or Weakest Link Pruning which internally uses **mean square error**, k-fold cross-validation and learning rate. XGBoost creates nodes (also called splits) up to max_depth specified and starts pruning from backward until the loss is below a threshold. Consider a split that has -3 loss and the subsequent node has +7 loss, XGBoost will not remove the split just by looking at one of the negative loss. It will compute the total loss (-3 + 7 = +4) and if it turns out to be positive it keeps both.\n\n1. **Hardware Optimization:** This algorithm has been designed to make efficient use of hardware resources. This is accomplished by cache awareness by allocating internal buffers in each thread to store gradient statistics. Further enhancements such as ‘out-of-core’ computing optimize available disk space while handling big data-frames that do not fit into memory. XGBoost also has a block structure for parallel learning. It makes it easy to scale up on multicore machines or clusters. It also uses cache awareness, which helps reduce memory usage when training models with large datasets. Finally, XGBoost offers **out-of-core computing** capabilities using **disk-based data structures instead of in-memory ones during the computation phase.**  By cache-aware optimization, we store gradient statistics (direction and value) for each split node in an internal buffer of each thread and perform accumulation in a mini-batch manner. This helps to reduce the time overhead of immediate read/write operations and also prevent cache miss. Cache awareness is achieved by choosing the optimal size of the block (generally 2¹⁶). **XGBoost uses a cache-aware prefetching algorithm which helps reduce the runtime for large datasets. The library can run more than ten times faster than other existing frameworks on a single machine.** Due to its impressive speed, XGBoost can process billions of examples using fewer resources, making it a scalable tree boosting system.\n\n1. **Scalability:** XGBoost can run distributedly thanks to distributed servers and clusters like Hadoop and Spark, so you can process enormous amounts of data. It’s also available in many programming languages like C++, JAVA, Python, and Julia. \n\n### Algorithmic Enhancements:\n\n1. **Regularization:** It penalizes more complex models through both LASSO (L1) and Ridge (L2) regularization to prevent overfitting. XGBoost offers regularization, which allows you to control overfitting by introducing L1/L2 penalties on the weights and biases of each tree. This feature is not available in many other implementations of gradient boosting. **Built in regularization:** XGBoost includes regularization as part of the learning objective, unlike regular gradient boosting. Data may also be regularized through hyperparameter tuning. Using XGBoost’s built in regularization also allows the library to give better results than the regular scikit-learn gradient boosting package.\n\n1. **Sparsity Awareness:** - **Handling Missing Values** - Automatically learns the best way to handle missing data. XGBoost naturally admits sparse features for inputs by automatically ‘learning’ best missing value depending on training loss and handles different types of sparsity patterns in the data more efficiently. Data Sparsity refers to the condition where a large percentage of data within a dataset is missing or is set to zero. **Sparsity Aware Split Finding** — It is quite common that the data we gather has sparsity (a lot of missing or empty values) or becomes sparse after performing data engineering (feature encoding). To be aware of the sparsity patterns in the data, a default direction is assigned to each tree. XGBoost handles missing data by assigning them to default direction and finding the best imputation value so that it minimizes the training loss. Optimization here is to visit only missing values which make the algorithm run 50x faster than the naïve method.\n\n1. **Weighted Quantile Sketch:**  **Handles weighted data to manage instances with different importance levels.** XGBoost employs the distributed weighted Quantile Sketch algorithm to effectively **find the optimal split points among weighted datasets.** Another feature of XGBoost is its ability to handle sparse data sets using the weighted quantile sketch algorithm. This algorithm allows us to deal with non-zero entries in the feature matrix while retaining the same computational complexity as other algorithms like stochastic gradient descent.\n\n1. **Non-linearity:** XGBoost can detect and learn from non-linear data patterns.\n\n1. **Cross-validation:** The algorithm comes with built-in cross-validation method at each iteration, taking away the need to explicitly program this search and to specify the exact number of boosting iterations required in a single run. Cross validation is a statistical method to evaluate machine learning models on unseen data. It comes in handy when the dataset is limited and prevents overfitting by not taking an independent sample (holdout) from training data for validation. By reducing the size of training data, we are compromising with the features and patterns hidden in the data which can further induce errors in our model. This is similar to cross_val_score functionality provided by the scikit-learn library. XGBoost uses built-in cross validation function cv(): `xgb.cv()`\n\n1. **Customized Objective Function** — An objective function intends to maximize or minimize something. In ML, we try to minimize the objective function which is a combination of the loss function and regularization term. Optimizing the loss function encourages predictive models whereas optimizing regularization leads to smaller variance and makes prediction stable. Different objective functions available in XGBoost are:\n    - reg: **linear** for regression\n    - reg: **logistic**, and **binary: logistic** for binary classification\n    - multi: **softmax**, and multi: **softprob** for multiclass classification\n    \n    <img width=\"760\" alt=\"image\" src=\"https://github.com/user-attachments/assets/156e414d-f048-4a6d-8de7-ee65f209137b\">\n\n1. **Customized Evaluation Metric** — This is a metric used to monitor the model’s accuracy on validation data.\n    - **rmse** — Root mean squared error **(Regression)**\n    - **mae** — Mean absolute error **(Regression)**\n    - error — Binary classification error (Classification)\n    - logloss — Negative log-likelihood (Classification)\n    - auc — Area under the curve (Classification)\n\nWe used Scikit-learn’s ‘Make_Classification’ data package to create a random sample of 1 million data points with 20 features (2 informative and 2 redundant). We tested several algorithms such as Logistic Regression, Random Forest, standard Gradient Boosting, and XGBoost.\n\n<img width=\"955\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/22a629f8-f032-4d88-9479-3bd7b0e09147\">\n\n\nAs demonstrated in the chart above, XGBoost model has the best combination of prediction performance and processing time compared to other algorithms. Other rigorous benchmarking studies have produced similar results. No wonder XGBoost is widely used in recent Data Science competitions.\n\n“When in doubt, use XGBoost” — Owen Zhang, Winner of Avito Context Ad Click Prediction competition on Kaggle\n\n## <span style=\"color: #016FD0;\">Hyperparameter Tuning</span>\n\nHyperparameter tuning is important because the performance of a machine learning model is heavily influenced by the choice of hyperparameters. Choosing the right set of hyperparameters can lead to better model performance, while choosing the wrong set can lead to poor performance. Additionally, when a model has too many hyperparameters, it can be difficult or even impossible to find the best set of hyperparameters manually.\n\n**Different types of hyperparameters in XGBoost:**\n\nIn XGBoost, there are two main types of hyperparameters: **tree-specific** and **learning task-specific**.\n\n**What are XGBoost Parameters?**\n\nThe overall parameters have been divided into 3 categories by XGBoost authors:\n\n1. **General Parameters:** Guide the overall functioning\n1. **Booster Parameters:** Guide the individual booster (tree/regression) at each step\n1. **Learning Task Parameters:** Guide the optimization performed\n\n<img width=\"823\" alt=\"image\" src=\"https://github.com/user-attachments/assets/1ac9e743-aaec-43b6-a411-3c3573f421bc\">\n\n### <span style=\"color: #016FD0;\">General Parameters</span>\n\nThese define the overall functionality of XGBoost.\n\n1. **`booster [default=gbtree]`:** Select the type of model to run at each iteration. It has 3 options:\n    - gbtree: tree-based models\n    - gblinear: linear models\n    - dart\n    \n     is the boosting algorithm, for which you have 3 options: `gbtree`, `gblinear` or `dart`. Select the type of model to run at each iteration. The default option is `gbtree` , which is the version I explained in this article. `dart` is a similar version that uses dropout techniques to avoid overfitting, and `gblinear` uses generalized linear regression instead of decision trees. **You can choose only for Regression (not classification) models using XGBoost:**\n    - **Decision Tree Base Learning, gbtree:** tree-based models\n    - **Linear Base Learning, gblinear:** linear models\n\n1. **`silent [default=0]`:** Silent mode is activated is set to 1, i.e., no running messages will be printed. It’s generally good to keep it 0 as the messages might help in understanding the model.\n\n1. **`nthread [default to the maximum number of threads available if not set]`:** This is used for parallel processing, and the number of cores in the system should be entered. If you wish to run on all cores, the value should not be entered, and the algorithm will detect it automatically.\n\n1. **`num_estimators`:** sets the number of boosting rounds, which equals setting the number of boosted trees to use. **The greater this number, the greater the risk of overfitting (but low numbers can also lead to low performance).** The amount of estimators defines the maximum number of iterations at which the boosting is terminated. It is called the “maximum” number, because the algorithm will stop on its own, in case good performance is achieved earlier. This is how many subtrees h will be trained. I put this first because introducing early stopping is the most important thing you can do to prevent overfitting. The motivation for this is that at some point, XGBoost will begin memorizing the training data, and its performance on the validation set will worsen. At this point, you want to stop training more trees. **Note that if you use early stopping,** XGBoost will return the final model (as opposed to the one with the lowest validation score), but this is okay since the best model will be this final model minus the additional, overfitting subtrees which were trained. ***However, adding too many trees can lead to overfitting. Generally speaking, as  n_estimators goes up, the learning rate should go down.*** You can isolate the best model using `trained_model.best_ntree_limit` in your predict method, as below:\n    `results = best_xgb_model.predict(x_test, ntree_limit=best_xgb_model.best_ntree_limit)`\n\n1. **`num_boost_round` and `early_stopping_round` :** The number of boosting stages. Although num_estimators and num_boost_round remain quite the same, you should keep in mind that the num_boost_round should be re-tuned each time you update a parameter. The first parameter we will look at is not part of the params dictionary, but will be passed as a standalone argument to the training method. This parameter is called num_boost_round and corresponds to the number of boosting rounds or trees to build. Its optimal value highly depends on the other parameters, and thus it should be re-tuned each time you update a parameter. You could do this by tuning it together with all parameters in a grid-search, but it requires a lot of computational effort. Fortunately XGBoost provides a nice way to find the best number of rounds whilst training. Since trees are built sequentially, instead of fixing the number of rounds at the beginning, we can test our model at each step and see if adding a new tree/round improves performance. To do so, we define a test dataset and a metric that is used to assess performance at each round. If performance haven’t improved for N rounds (N is defined by the variable early_stopping_round), we stop the training and keep the best number of boosting rounds. \n    - We still need to pass a num_boost_round which corresponds to the maximum number of boosting rounds that we allow. We set it to a large value hoping to find the optimal number of rounds before reaching it, if we haven't improved performance on our test dataset in early_stopping_round rounds\n    - `num_boost_round:` number of boosting rounds. Here we will use a large number again and count on `early_stopping_rounds` to find the optimal number of rounds before reaching the maximum.\n\nThere are 2 more parameters that are set automatically by XGBoost, and you need not worry about them. Let’s move on to Booster parameters.\n\n\n### <span style=\"color: #016FD0;\">Booster, Tree-specific Hyperparameters</span>\n\nThough there are 2 types of boosters, **I’ll consider only tree booster here because it always outperforms the linear booster, and thus the latter is rarely used.**\n\nTree-specific hyperparameters **control the construction and complexity of the decision trees:**\n\n1. **`max_depth [default=6]`:** maximum depth of a tree. Deeper trees can capture more complex patterns in the data, but may also **lead to overfitting.** Sets the maximum depth of the decision trees. The greater this number, the less conservative the model becomes. **If set to 0, then there is no limit for trees’ depth.** The maximum depth of a tree, same as GBM. Used to control over-fitting as higher depth will allow model to learn relations very specific to a particular sample. Should be tuned using CV. **Typical values: 3–10** The maximum tree depth each individual tree h can grow to. The default value of 3 is a good starting point, and I haven’t found a need to go beyond a max_depth of 5, even with fairly complex data. It is the maximum number of nodes allowed from the root to the farthest leaf of a tree. Deeper trees can model more complex relationships by adding more nodes, but as we go deeper, splits become less relevant and are sometimes only due to noise, causing the model to overfit.\n\n1. **`max_leaf_nodes`:** The maximum number of terminal nodes or leaves in a tree. Can be defined in place of `max_depth`. Since binary trees are created, a depth of ’n’ would produce a maximum of 2^n leaves. If this is defined, GBM will ignore max_depth.\n\n1. **`min_child_weight [default=1]`:** minimum sum of instance weight (hessian) needed in a child. This can be used to control the complexity of the decision tree by preventing the creation of too small leaves.  Defines the minimum sum of weights of all observations required in a child. This is similar to `min_child_leaf` in GBM but not exactly. This refers to min “sum of weights” of observations while GBM has min “number of observations”. Used to control over-fitting. Higher values prevent a model from learning relations which might be highly specific to the particular sample selected for a tree. Too high values can lead to under-fitting hence, it should be tuned using CV. It is the minimum weight (or number of samples if all samples have a weight of 1) required in order to create a new node in the tree. A smaller min_child_weight allows the algorithm to create children that correspond to fewer samples, thus allowing for more complex trees, but again, more likely to overfit.\n\n1. **`subsample [default=1]`:** percentage of rows used for each tree construction. It corresponds to the fraction of observations (the rows) to subsample at each step. By default it is set to 1 meaning that we use all rows. **Instead of using the whole training set every time, we can build a tree on slightly different data at each step, which makes it less likely to overfit to a single sample or feature.** Lowering this value can **prevent overfitting** by training on a smaller subset of the data. is the size of the sample ratio to be used when training the predictors. Default is 1, meaning there is no sampling and we use the whole data. If set to 0.7, for instance, then 70% of the observations would be randomly sampled to be used in each boosting iteration (a new sample is taken for each iteration). **It can help to prevent overfitting.** The fraction of the training data that is used to train each tree. Same as the subsample of GBM. Denotes the fraction of observations to be randomly samples for each tree. **Lower values make the algorithm more conservative and prevent overfitting** but too small values might lead to under-fitting. **Typical values: 0.5–1**\n\n1. **`colsample_bytree:`** percentage of columns used for each tree construction. Lowering this value can **prevent overfitting** by training on a subset of the features. It corresponds to the fraction of features (the columns) to use. By default it is set to 1 meaning that we will use all features.\n\n1. **`max_delta_step [default=0]`:** In maximum delta step, we allow each tree’s weight estimation to be. If the value is set to 0, it means there is no constraint. If it is set to a positive value, it can help to make the update step more conservative. **Usually, this parameter is not needed, but it might help in logistic regression when class is extremely imbalanced.** *This is generally not used but you can explore further if you wish.*\n\n### <span style=\"color: #016FD0;\">Learning task-specific Hyperparameters</span>\n\nLearning task-specific hyperparameters **control the overall behavior of the model and the learning process.** These parameters are used to define the optimization objective and the metric to be calculated at each step.\n\n1. **`eta (also known as learning rate) [default=0.3]`:** step size shrinkage used in updates to **prevent overfitting.** Lower values make the model more robust by taking smaller steps. Finally, the learning rate controls how much the new model is going to contribute to the previous one. Normally there is a trade-off between the number of iterations and the value of the learning rate. In other words: when taking smaller values of the learning rate, you should consider more estimators, so that your base model (the weak classifier) continues to improve. Makes the model more robust by shrinking the weights on each step. **Typical final values to be used: 0.01–0.2** Each weight (in all the trees) will be multiplied by this value, so that:\n\n    <img width=\"738\" alt=\"image\" src=\"https://github.com/user-attachments/assets/d2749da9-37d8-4423-84aa-1353c3d29887\">\n    \n    I found that decreasing the learning rate very often lead to an improvement in the performance of the model (although it did lead to slower training times). Because of the additive nature of gradient boosted trees, I found getting stuck in local minima to be a much smaller problem then with neural networks (or other learning algorithms which use stochastic gradient descent). In practice, having a lower eta makes our model more robust to overfitting thus, usually, the lower the learning rate, the best. But with a lower eta, we need more boosting rounds, which takes more time to train, sometimes for only marginal improvements. (also known as the **“step size”** or the **“shrinkage”**), is the most important gradient boosting hyperparameter. In the XGBoost library, it is known as “eta”, should be a number between 0 and 1 and the default is 0.36. The learning rate determines the rate at which the boosting algorithm learns from each iteration. A lower value of eta means slower learning, as it scales down the contribution of each tree in the ensemble, thus helping to prevent overfitting. Conversely, a higher value of eta speeds up learning, but it may lead to overfitting if not carefully tuned.\n\n1. **`gamma [default=0]`:** minimum loss reduction required to make a further partition on a leaf node of the tree. **Higher values increase the regularization.** The minimum loss reduction required to make a split. A node is split only when the resulting split gives a positive reduction in the loss function. Gamma specifies the minimum loss reduction required to make a split. Makes the algorithm conservative. The values can vary depending on the loss function and should be tuned.\n\n1. **`reg_alpha`** and **`reg_lambda`**: `reg_alpha` and `reg_lambda` are **L1** and **L2** regularisation terms, respectively. The greater these numbers, the more conservative (**less prone to overfitting but might miss relevant information**) the model becomes. Recommended values lie between 0–1000 for both.\n    - `lambda` [default=1] **L2 regularization term on weights** (analogous to **Ridge** regression) This used to handle the regularization part of XGBoost. Though many data scientists don’t use it often, it should be explored to reduce overfitting.  **Higher values increase the regularization.**\n\n    - `alpha` [default=0] **L1 regularization term on weight** (analogous to **Lasso** regression). **Can be used in case of very high dimensionality** so that the algorithm runs faster when implemented. **Higher values increase the regularization.**\n    \n    - `reg_alpha` and `reg_lambda` control the L1 and L2 regularization terms, which in this case limit how extreme the weights at the leaves can become. These two regularization terms have different effects on the weights; L2 regularization (controlled by the lambda term) encourages the weights to be small, whereas L1 regularization (controlled by the alpha term) encourages sparsity — so it encourages weights to go to 0. **This is helpful in models such as logistic regression, where you want some feature selection, but in decision trees we’ve already selected our features, so zeroing their weights isn’t super helpful. For this reason, I found setting a high lambda value and a low (or 0) alpha value to be the most effective when regularizing.**\n    \n    <img width=\"692\" alt=\"image\" src=\"https://github.com/user-attachments/assets/7f42d5e8-2485-4efe-917f-81183160934f\">\n\n1. **`scale_pos_weight [default=1]`:** A value greater than 0 should be used in case of high-class imbalance as it helps in faster convergence.** A value greater than 0 should be used in case of high class imbalance as it helps in faster convergence. A typical value to consider: sum(negative instances) / sum(positive instances).\n\n1. **`objective [default=reg:linear]`:** This defines the loss function to be minimized. Mostly used values are:\n    - **binary:** **logistic** –logistic regression for binary classification returns predicted probability (not class)\n    - **multi:** **softmax** –multiclass classification using the softmax objective, returns predicted class (not probabilities), you also need to set an additional num_class (number of classes) parameter defining the number of unique classes\n    - **multi:** **softprob** –same as softmax, but returns predicted probability of each data point belonging to each class.\n\n1. **`eval_metric [ default according to objective ]`:** The evaluation metrics are to be used for validation data. The default values are rmse for regression and error for classification. Typical values are:\n    - rmse – root mean square error\n    - mae – mean absolute error\n    - logloss – negative log-likelihood\n    - error – Binary classification error rate (0.5 thresholds)\n    - merror – Multiclass classification error rate\n    - mlogloss – Multiclass logloss\n    - auc: Area under the curve\n\n1. **`seed [default=0]`:** The random number seed. It can be used for generating reproducible results and also for parameter tuning.\n\nNext, instantiate an XGBoost model and, depending on your use case, select which objective function you’d like to use via the “object” hyperparameter. For example, if you have a multi-class classification task, you should set the objective to “multi:softmax”5. Alternatively, if you have a binary classification problem, you can use the logistic regression objective “binary:logistic”. Now you can use your training set to train the model and predict classifications for the data set aside as the test set. Assess the performance of the model by comparing the predicted values with the test set’s actual values. You may use metrics such as accuracy, precision, recall or f-1 score to evaluate your model. You may also want to visualize your true positives, true negatives, false positives and false negatives using a\nconfusion matrix.\n\n### <span style=\"color: #016FD0;\">Overview of different techniques for tuning hyperparameters</span>\n\nThis is helpful because there are many, many hyperparameters to tune. Nearly all of them are designed to limit overfitting (no matter how simple your base models are, if you stick thousands of them together they will overfit). The best hyperparameters can be found using grid search and cross-validation methods, which will iterate through a dictionary of possible hyperparameter combinations.\n\nThe list of hyperparameters was super intimidating to me when I started working with XGBoost, so I am going to discuss the 4 parameters I have found most important when training my models so far (I have tried to give a slightly more detailed explanation than the documentation for all the parameters in the appendix).\n\nMy motivation for trying to limit the number of hyperparameters is that doing any kind of grid / random search with all of the hyperparameters XGBoost allows you to tune can quickly explode the search space. I’ve found it helpful to start with the 4 below, and then dive into the others only if I still have trouble with overfitting.\n\n- **n_estimators (and early stopping)**\n- **max_depth**\n- **learning rate**\n- **reg_alpha** and **reg_lambda**\n\nNote that the other parameters are useful, and worth going through if the above terms don’t help with regularization. However, I have found that exploring all the hyperparameters can cause the search space to explode, so this is a good place to start.\n\n1. **Grid search:** is one of the most widely used techniques for hyperparameter tuning. It involves specifying a set of possible values for each hyperparameter, and then training and evaluating the model for each combination of hyperparameter values. Grid search is simple to implement and can be efficient when the number of hyperparameters and their possible values is small. However, it can become computationally expensive as the number of hyperparameters and possible values increases. \n    <img width=\"819\" alt=\"image\" src=\"https://github.com/user-attachments/assets/d0fadb8e-9b22-4ce2-b016-7c572e59a4a5\">\n\n1. **Random search:** is a variation of grid search that randomly samples from the set of possible hyperparameter values instead of trying all combinations. This can be more efficient than grid search because it does not need to evaluate all possible combinations. However, it can still be computationally expensive, especially when the number of hyperparameters and possible values is large.\n    <img width=\"875\" alt=\"image\" src=\"https://github.com/user-attachments/assets/5a55d9eb-ac76-46d3-a5ee-afb5cf2fe09a\">\n\n1. **Bayesian optimization:** is a more sophisticated technique that uses Bayesian methods to model the underlying function that maps hyperparameters to the model performance. It tries to find the optimal set of hyperparameters by making smart guesses based on the previous results. Bayesian optimization is more efficient than grid or random search because it attempts to balance exploration and exploitation of the search space. It can also deal with the cases of large number of hyperparameters and large search space. However, it can be more difficult to implement than grid search or random search and may require more computational resources.\n    <img width=\"914\" alt=\"image\" src=\"https://github.com/user-attachments/assets/8708c1ef-2a46-4139-9724-d0aed40af277\">\n\n**Using XGBoost’s CV:** In order to tune the other hyperparameters, we will use the cv function from XGBoost. It allows us to run cross-validation on our training dataset and returns a mean MAE score. We need to pass it:\n\n- `params:` our dictionary of parameters.\n- our `dtrain` matrix.\n- `num_boost_round:` number of boosting rounds. Here we will use a large number again and count on early_stopping_rounds to find the optimal number of rounds before reaching the maximum.\n- `seed:` random seed. It's important to set a seed here, to ensure we are using the same folds for each step so we can properly compare the scores with different parameters.\n- `nfold:` the number of folds to use for cross-validation\n- `metrics:` the metrics to use to evaluate our model, here we use MAE.\n    \nAs you can see, we don’t need to pass a test dataset here. It’s because the cross-validation function is splitting the train dataset into nfolds and iteratively keeps one of the folds for test purposes. \n    <img width=\"926\" alt=\"image\" src=\"https://github.com/user-attachments/assets/168bf928-5f7c-4d23-b0bf-ff3aff986523\">\n \nThere are several techniques that can be used to tune the hyperparameters of an XGBoost model including grid search, random search and Bayesian optimization. Grid search is simple to implement but can be computationally expensive when the number of hyperparameters and possible values is large. Random search is more efficient than grid search but still can be computationally expensive. Bayesian optimization is the most sophisticated technique, which balances exploration and exploitation, but can be more difficult to implement and require more computational resources.\n\n<img width=\"808\" alt=\"image\" src=\"https://github.com/user-attachments/assets/a660cb17-4db1-493a-9c9a-493dea479f0a\">\n<img width=\"825\" alt=\"image\" src=\"https://github.com/user-attachments/assets/c4cc4c21-eedb-4e7a-a9ab-6724bd8e31b6\">\n<img width=\"1088\" alt=\"image\" src=\"https://github.com/user-attachments/assets/aed97fbe-3cac-4881-abfb-cb8b368b8ee9\">\n\nAs expected it took us more rounds to get there, but we improved our MAE from 4.31 to 3.90. Is that good? Well it depends what you compare it to. Noting that we got this improvement almost for free, without adding data or engineering features, simply by spending a bit of time tuning our model, then it’s not bad. But it’s good to notice that it did not transform a poor model (we are still off by 4 comments on average whilst our average number of comments is 7…) into an excellent one. This is quite common with Machine Learning, whilst it is important to “roughly” tune your model to get good results from it, it will only get you that far. ***And there is a point after which additional time spent tuning it only provides marginal improvements. When it’s the case, it’s usually worth looking more closely at the data to find better ways of extracting information, and/or try other algorithms instead of fine tuning your current model.***\n\n## <span style=\"color: #016FD0;\">Tips to control your XGBoost model</span>\n\nGradient boosting is a popular machine learning technique used throughout many industries because of its performance on many classes of problems. In gradient boosting small models—called “weak learners” because individually they do not fit well—are fit sequentially to residuals of the previous models. These weak models perform with excellent predictive accuracy when added together into an ensemble. **This performance comes at the cost of high model complexity which makes them hard to analyze and can lead to overfitting.** As a data scientist in the Model Risk Office, it is my job to make sure that models are fit appropriately and are analyzed for weaknesses.\n\n**In gradient boosting, you can control the complexity in two ways:**\n\n1. **Make each learner in the ensemble weaker.**\n1. **Have fewer learners in the ensemble.**\n\nOne of the most popular boosting algorithms is the gradient boosting machine (GBM) package XGBoost. XGBoost is a lighting-fast open-source package with bindings in R, Python, and other languages. Due to its popularity, there is no shortage of articles out there on how to use XGBoost. Even so, most articles only give broad overviews of how the code works. \n\nIn this deep dive into XGBoost I want to discuss what I see as a common misconception about XGBoost - that the complexity of each learner and the number of learners are just two sides of the same coin. I will discuss how these two methods of controlling complexity are not the same: **depth adds interactions and grows complexity faster than adding trees.**\n\nXGBoost has a few different modes, but the one I will focus on here uses tree models as the individual learners in the ensemble. The XGBoost documentation contains a great introduction which I will summarize below.\n\n1. **Tree building:** Start with some target—either continuous or binary—and some base level predictions (usually either 0 or probability of 0.5). The basic idea is to build a tree that predicts the residuals between our base predictions and the target. At each stage, we recompute the residuals and then build a new tree. We add this tree to the mode, recompute the residuals again, and repeat this process until we have a model that fits well.\n\n1. **Split enumeration:** The trees are built using binary splits—these are just threshold cuts in single features (unlike H2O and some other packages, XGBoost handles only continuous data—even categorical features are treated as continuous). Here would be an English-language example of a split: all customers with an account balance <1000 go right and all customers with an account balance >=1000 go left.  How to pick the place to split? XGBoost looks at which feature and split-point maximizes the gain. The maximum gain is found where the sum of the loss from the child nodes most reduces the loss in the parent node. For small datasets, XGBoost tries all split points (thresholds) given by the data values for each feature and records their gain. It then picks the feature and threshold combination with the largest gain. For larger datasets (by default any dataset with more than 4194303 rows), XGBoost proposes fewer candidate splits. The locations of these candidate splits are decided by the quantiles of the data (weighted by the Hessian).\n\n1. **Pruning:** During the tree building process, XGBoost automatically stops if there is a node without enough cover (the sum of the Hessians of all the data falling into that node) or if it reaches the maximum depth. After the trees are built, XGBoost does an optional 'pruning' step that, starting from the bottom (where the leaves are) and working its way up to the root node, looks to see if the gain falls below gamma (a tuning parameter—see below). If the first node encountered has a gain value lower than gamma, then the node is pruned and the pruner moves up the tree to the next node. If, however, the node has gain higher than gamma, the node is left and the pruner does not check the parent nodes.\n\nNow that we have the basics, let's look at the ways a model builder can control overfitting in XGBoost.\n\n### <span style=\"color: #016FD0;\">XGBoost regularization</span>\n\nThe four methods to control overfitting\n\nThere are many ways of controlling overfitting, but they can mostly be summed up in four categories:\n\n1. **Regularization:** The regularization parameters act directly on the weights:\n    - `lambda - L2 regularization.` This term is a constant that is added to the second derivative (Hessian) of the loss function during gain and weight (prediction) calculations. This parameter can both shift which splits are taken and shrink the weights.\n    - `alpha - L1 regularization.` This term is subtracted from the gradient of the loss function during the gain and weight calculations. Like the L2 regularization it affects the choice of split points as well as the weight size.\n    - `eta (learning_rate)` - Multiply the tree values by a number (less than one) to make the model fit slower and prevent overfitting.\n    - `max_delta_step` - The maximum step size that a leaf node can take. In practice, this means that leaf values can be no larger than max_delta_step * eta.\n\n1. **Pruning:** Pruning removes splits directly from the trees during or after the build process (see more below):\n    - `gamma (min_split_loss)` - A fixed threshold of gain improvement to keep a split. Used during the pruning step of XGBoost.\n    - `min_child_weight` - Minimum sum of Hessians (second derivatives) needed to keep a child node during partitioning. Making this larger makes the algorithm more conservative (although this scales with data size). Since the second derivatives are different in classification and regression, this parameter acts differently for the two contexts. In regression this is just floor on the number of data instances (rows) that a node needs to see. For classification this gives the required sum of p*(1-p), where p is the probability, for data that are split into that node. Since probability ranges from 0 to 1, p*(1-p) ranges from 0 to 0.25 and so you will need at least four times the min_child_weight rows in that node to keep it\n    - `max_depth` - The maximum depth of a tree. While not technically pruning, this parameter acts as a hard stop on the tree build process. Shallower trees are weaker learners and are less prone to overfit (but can also not capture interactions - see below).\n\n1. **Sampling:** Sampling makes the boosted trees less correlated and prevents some feature masking effects. This makes them less correlated and more robust to noise:\n    - `subsample` - Subsample rows of the training data prior to fitting a new estimator.\n    - `colsample_*(bytree, bylevel, bynode)` - Fraction of features to subsample at different locations in the tree building process.\n\n1. **Early stopping:** Early stopping monitors a metric on a holdout dataset and stops building the ensemble when that metric no longer improves:\n    - The XGBoost documentation details early stopping in Python. Note: this parameter is different than all the rest in that it is set during the training not during the model initialization. Early stopping is usually preferable to choosing the number of estimators during grid search.\n\n1. **The shrinkage (learning) rate.**\n\n\nThus, those parameters can be used to control the complexity of the trees. It is important to tune them together in order to find a good trade-off between model bias and variance\n\n\n**Determining model complexity:**\n\nLarger, more complex, models can be prone to overfitting, slower to score, and harder to interpret. For those reasons, knowing just how complex your model is can be beneficial. Once you've built a model using one of the above parameters, how can you figure out how complex it is? How can you compare the complexity of two different models? \n\nTypically, modelers only look at the parameters set during training. However, the structure of XGBoost models makes it difficult to really understand the results of the parameters. One way to understand the total complexity is to count the total number of internal nodes (splits). We can count up the number of splits using the XGBoost text dump:\n\n<img width=\"1077\" alt=\"image\" src=\"https://github.com/user-attachments/assets/dbb5f1fb-6732-45ce-a347-36d006fad822\">\n\n### <span style=\"color: #016FD0;\">Model complexity: Depth vs. number of trees</span>\n\nThere are two basic ways in which tree ensemble models can be complex:\n\n1. **Deeper trees (each estimator is more complicated)**: In practice, deeper trees tend to be more complex than shallower trees, even when we use more estimators.\n\n1. **More trees**\n\nThese two types of complexity are not simply two sides of the same coin. They give different behaviors to the ensemble. But let's first look at just the number of nodes. \n\nThe theoretical maximum number of nodes is: `n_estimators x 2^max_depth`. In practice, model builders should be using early stopping and pruning so the real numbers would be lower. T\n\n1. **Deeper tree: Takeaway 1**\n    - Although every situation is different, **deeper trees tend to add complexity in the form of extra leaves faster than shallower trees.** This is true even though ensembles built with deeper trees tend to have fewer trees.\n    - But complexity is not usually what model builders care about. What about the holdout AUROC for these models? In the figure below we see the results for many different models built on the same dataset, but with different tuning parameters. We can see that as complexity increases (counted by the total number of splits in the ensemble) the model performance increases up to around 1500 total splits. After that the complexity increases while the performance stays largely flat.\n\n1. **Deeper tree: Takeaway 2**\n    - Extra complexity can help fit better models, but often gives diminishing returns to hold-out performance.\n    - **Depth:**\n        - Complexity is not the only reason to be wary of making your trees deeper. **Deeper trees add interactions in a way that adding more trees does not. Adding depth adds complexity in two ways:**\n        - Allows the **possibility for more complicated interactions**\n        - Additional splits (more granular space partition)\n    - **Adding trees only adds to the second complexity. Interactions between features require a depth of trees that is deep enough to handle the interaction.** Simply adding more trees will not increase the complexity of the interactions. That is, if you have a maximum depth of two, then at most two variables can interact together. I will demonstrate this using my favorite toy example: a model with four features, two of which interact strongly in an 'x’-shaped function (y ~ x1 + 5x2 - 10x2*(x3 > 0)).\n\n1. **Depth key takeaway:**\n    - **Depth is sometimes necessary to capture interactions between features.**\n    - A corollary to this is that if there are no interactions in the dataset, then there is no need for deep trees.\n\n## <span style=\"color: #016FD0;\">Implementation Details</span>\n\nIn practice, using XGBoost involves the following steps:\n\n1. **Data Preparation** - Preparing your dataset by handling missing values, encoding categorical variables, and splitting into training and testing sets.\n1. **Model Training** - Using the XGBoost library (e.g., xgboost in Python) to define the model parameters and train the model on the training dataset.\n1. **Hyperparameter Tuning** - Tuning hyperparameters like the learning rate, maximum depth of trees, number of boosting rounds, and regularization parameters to optimize model performance.\n1. **Model Evaluation** - Evaluating the model using appropriate metrics (e.g., accuracy, precision, recall for classification; RMSE for regression) on the testing set.\n1. **Prediction** - Using the trained model to make predictions on new data.\n\n-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------\nXGBoost is highly flexible in terms of the data formats it can accept. This flexibility is one of the reasons it is so popular in the machine learning community. Here are some key points about XGBoost's flexibility with different dataframes and data formats:\n\n**Supported Data Formats:**\n\n1. **DMatrix:**\n    - Native Format - XGBoost uses **its own optimized data structure called DMatrix.** This format is highly efficient and is **designed to handle large datasets with sparse or dense features.** **`DMatrix`** is an internal data structure used by XGBoost which is optimized for memory efficiency and training speed. We need to transform our numpy array of data using DMatrix so that it can be later utilized in dtrain and dtest parameters of its inbuilt functions. **Convert your data into the DMatrix format that XGBoost expects3. DMatrix is XGBoost's internal data structure optimized for memory efficiency and training speed.**\n    - Features - Handles missing values, **supports group information for ranking tasks, and allows for custom weights for instances.**\n    - Instead of numpy arrays or pandas dataFrame, XGBoost uses DMatrices. A DMatrix can contain both the features and the target. If you already have loaded you data into numpy arrays X and y, you can create a DMatrix with: `xgb.DMatrix(X, label=y)`\n\n1. **Pandas DataFrame:**\n    - Convenient for Python Users - XGBoost can directly accept pandas DataFrames, which are widely used in the Python ecosystem for data manipulation and analysis.\n    - Usage - You can convert a pandas DataFrame to a DMatrix or use it directly in training and prediction functions.\n\n1. **NumPy Arrays:**\n    - Compatibility - XGBoost works seamlessly with numpy arrays, another common data format in Python for numerical operations.\n    - Usage - Similar to DataFrames, numpy arrays can be directly used or converted to DMatrix.\n\n1. **SciPy Sparse Matrices:**\n    - Efficiency - For datasets with a large number of features **but relatively few non-zero entries, SciPy sparse matrices are efficient in terms of memory and computation.**\n    - Usage - XGBoost natively supports SciPy sparse matrices.\n\nOther Characteristics:\n\n- **Integration with Other Data Handling Tools:**\n    - DataFrame Libraries: XGBoost can integrate with various DataFrame libraries like **dask** for distributed computing or **cuDF** for GPU-accelerated data processing.\n    - File Formats: XGBoost can read from and write to various file formats including CSV, LibSVM, and binary buffers.\n\n- **GPU Support in XGBoost:**\n    - **GPU-Accelerated Algorithms:**\n        - Tree Construction - XGBoost provides GPU implementations for various tree construction algorithms, including exact, hist (approximate), and external memory algorithms.\n        - Predictor - The GPU predictor can be used for faster inference on large datasets.\n    - **Supported GPU Hardware:**\n        - NVIDIA GPUs - XGBoost's GPU acceleration is designed to work with NVIDIA GPUs using the CUDA platform. CUDA-capable GPUs are required for GPU-accelerated training and inference.\n    - **GPU Parameters** - To enable GPU acceleration, you need to set the tree_method parameter to gpu_hist, gpu_exact, or gpu_hist_external.\n    - **Other Parameters** - Parameters such as gpu_id can be used to specify which GPU to use if multiple GPUs are available.\n    - **Benefits of GPU Support:**\n        - **Speed** - Training time can be significantly reduced with GPU acceleration, especially for large datasets and complex models.\n        - **Scalability** - GPU support allows XGBoost to handle larger datasets that might be infeasible to process on a CPU due to time or memory constraints.\n        - **Efficiency** - GPUs are particularly well-suited for the parallel nature of tree boosting algorithms, leading to more efficient computations.\n\n- **k-fold Cross-validation** — In k-fold cross-validation, data is shuffled and divided into k equal sized subsamples. One of the k subsamples is used as a test/validation set and remaining (k -1) subsamples are put together to be used as training data. Then we fit a model using training data and evaluate it using the test set. This process is repeated k times so that every data point stays in validation set exactly once. The k results from each model should be averaged to get the final estimation. The advantage of this method is that we significantly reduce bias, variance, and also increase the robustness of the model. k-fold Cross validation using sklearn in XGBoost :\n    ```python\n        from sklearn.model_selection import KFold, cross_val_score\n        kfold = KFold(n_splits=15)\n        xgboost_score = cross_val_score(xg_cl, X, y, cv=kfold)\n    ```\n- Model tuning in XGBoost can be implemented by cross-validation strategies like **GridSearchCV** and **RandomizedSearchCV**.\n\n    1. **Grid Search** — We pass on a parameter’s dictionary to the function and compare the cross-validation score for each combination of parameters (many to many) in the dictionary and return the set having the best parameters.\n\n    2. **Random Search** — We draw a random value during each iteration from the range of specified values for each hyperparameter searched over, and evaluate a model with those hyperparameters. After completing all iterations, it picks the hyperparameter configuration with the best score.\n\n- **Plot Importance Module:** XGBoost library provides a built-in function to plot features ordered by their importance. The function is `plot_importance(model)` and it takes the trained model as its parameter. The function gives an informative bar chart representing the significance of each feature and names them according to their index in the dataset. The importance is calculated based on an importance_type variable which takes the parameters\n    - weights (default) — tells the times a feature appears in a tree\n    - gain — is the average training loss gained when using a feature\n    - cover — which tells the coverage of splits, i.e. the number of times a feature is used weighted by the total training point that falls in that branch.\n\n    <img width=\"840\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/1b0408f6-2e24-4eaa-9f60-4fdfaeccf516\">\n\n- **plot_tree** allows to visualize the trees that were built by XGBoost\n    <img width=\"824\" alt=\"image\" src=\"https://github.com/user-attachments/assets/82eddedd-65d3-4ff0-865f-bf42b7554fb2\">\n\n- You can also access the characteristics of your model\n    <img width=\"798\" alt=\"image\" src=\"https://github.com/user-attachments/assets/d1bad444-e268-42b5-9b83-4803f920ad88\">\n\n\n## <span style=\"color: #016FD0;\">The native XGBoost API</span>\n\nAlthough the scikit-learn API of XGBoost (shown in the previous tutorial) is easy to use and fits well in a scikit-learn pipeline, it is sometimes better to use the native API. Advantages include:\n\n1. Automatically find the best number of boosting rounds\n1. Built-in cross validation\n1. Custom objective functions\n\n## <span style=\"color: #016FD0;\">XGBoost For Ranking Problems</span>\n\nXGBoost includes ranking algorithms as part of its functionality. **Ranking algorithms** are particularly useful in **information retrieval** and **recommendation systems** where the goal is to **rank items according to** their relevance to a query or **user preferences**. In XGBoost, these ranking algorithms are implemented to handle such tasks efficiently.\n\n### <span style=\"color: #016FD0;\">Ranking Algorithms in XGBoost</span>\n\nXGBoost **supports several objective functions** tailored for ranking tasks:\n\n1. **Pairwise Ranking:**\n    - Rank:pairwise:\n    - Objective - Optimizes the pairwise loss, which means the model tries to correctly rank pairs of items.\n    - Use Case - Useful for applications like search engines, where the goal is to rank search results.\n\n1. **LambdaRank:**\n\n    - **Rank:ndcg (Normalized Discounted Cumulative Gain):**\n        - Objective - Optimizes the NDCG metric, which measures the quality of the ranked list by considering the position of relevant items.\n        - Use Case - Suitable for tasks where the relevance of items is more critical at the top of the list, such as search engine result rankings.\n    \n    - **Rank:map (Mean Average Precision):**\n        - Objective - Optimizes the MAP metric, which evaluates the precision of the ranked list of items.\n        - Use Case - Appropriate for recommendation systems where the precision of the top-ranked items is crucial.\n\nHow Ranking Works in XGBoost:\n\n1. **Data Preparation** - For ranking tasks, the dataset needs to be prepared with group information indicating which items belong to the same query or user session.\n\n1. **Model Training** - During training, XGBoost uses the specified ranking objective to build an ensemble of trees that optimize the chosen ranking metric.\n\n1. **Prediction** - The model predicts scores for each item, which can then be used to rank the items accordingly.\n\n\n## <span style=\"color: #016FD0;\">Conclusion for XGBoost</span>\n\n### Why XGBoost?\n\nXGBoost gained significant favor in the last few years as a result of helping individuals and teams win virtually every Kaggle structured data competition. In these competitions, companies and researchers post data after which statisticians and data miners compete to produce the best models for predicting and describing the data.\n\nXGBoost has been integrated with a wide variety of other tools and packages such as **scikit-learn** for Python enthusiasts and caret for R users. In addition, XGBoost is integrated with distributed processing frameworks like **Apache Spark** and **Dask.**\n\nIt’s noteworthy for data scientists that XGBoost and XGBoost machine learning models have the premier combination of prediction performance and processing time compared with other algorithms. This has been borne out by various benchmarking studies and further explains its appeal to data scientists.\n\n\n### So should we use just XGBoost all the time?\n\nWhen it comes to Machine Learning (or even life for that matter), there is no free lunch. As Data Scientists, we must test all possible algorithms for data at hand to identify the champion algorithm. Besides, picking the right algorithm is not enough. We must also choose the right configuration of the algorithm for a dataset by tuning the **hyper-parameters**. Furthermore, there are several other considerations for choosing the winning algorithm such as **computational complexity**, **explainability**, and **ease of implementation**. This is exactly the point where Machine Learning starts drifting away from science towards art, but honestly, that’s where the magic happens!\n\nI hope you now understand how XGBoost works and how to apply it to real data. Despite XGBoost’s inherent performance, **hyperparameter tuning** and **feature engineering** can make a huge difference in your results.\n\n### What does the future hold?\n\nMachine Learning is a very active research area and already there are several viable alternatives to XGBoost. Microsoft Research recently released **LightGBM** framework for gradient boosting that shows great potential. **CatBoost** developed by Yandex Technology has been delivering impressive bench-marking results. It is a matter of time when we have a better model framework that beats XGBoost in terms of prediction performance, flexibility, explanability, and pragmatism. However, until a time when a strong challenger comes along, XGBoost will continue to reign over the Machine Learning world!\n\nXGBoost has become a go-to algorithm for many data scientists and machine learning practitioners due to its efficiency, flexibility, and ability to handle large-scale data. Its integration with multiple programming languages makes it a versatile tool in the machine learning toolkit. XGBoost's support for ranking algorithms makes it a powerful tool for applications requiring the ranking of items, such as search engines and recommendation systems. By providing different objective functions for ranking, XGBoost allows for flexible and efficient optimization tailored to various ranking metrics.\n\nXGBoost's flexibility with different data formats makes it highly adaptable and easy to use within various data science workflows. Whether you're working with pandas DataFrames, numpy arrays, or SciPy sparse matrices, XGBoost provides seamless integration and efficient data handling capabilities, making it a versatile choice for many machine learning tasks.\n\nXGBoost's adaptability with GPU support makes it a powerful tool for handling large datasets and complex models efficiently. By leveraging NVIDIA GPUs and CUDA, XGBoost can significantly speed up both training and prediction processes, making it suitable for a wide range of machine learning applications. **The ease of enabling GPU support through simple parameter settings ensures that users can quickly take advantage of the performance benefits offered by GPUs.**\n\n## <span style=\"color: #016FD0;\">XGBoost and RAPIDS</span>\n\nCPU-powered machine learning tasks with XGBoost can literally take hours to run. That’s because creating highly accurate, state-of-the-art prediction results involves the creation of thousands of decision trees and the testing of large numbers of parameter combinations. Graphics processing units, or GPUs, with their massively parallel architecture consisting of thousands of small efficient cores, can launch thousands of parallel threads simultaneously to supercharge compute-intensive tasks.\n\n<img width=\"692\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/0119d420-cbcb-4bae-b5d7-9814b32974cb\">\n\n\nNVIDIA developed NVIDIA RAPIDS™—an open-source data analytics and machine learning acceleration platform—or executing end-to-end data science training pipelines completely in GPUs. It relies on NVIDIA CUDA® primitives for low-level compute optimization, but exposes that GPU parallelism and high memory bandwidth through user-friendly Python interfaces. \n\n<img width=\"890\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/d397e0f0-049f-4d19-a024-f5e03efa51ac\">\n\nFocusing on common data preparation tasks for analytics and data science, RAPIDS offers a familiar DataFrame API that integrates with scikit-learn and a variety of machine learning algorithms without paying typical serialization costs. This allows acceleration for end-to-end pipelines—from data prep to machine learning to deep learning. RAPIDS also includes support for multi-node, multi-GPU deployments, enabling vastly accelerated processing and training on much larger dataset sizes.\n\nXGBoost now includes seamless, drop-in GPU acceleration. This significantly speeds up model training and improves accuracy for better predictions.\n\nXGBoost now builds on the GoAI interface standards to provide zero-copy data import from cuDF, cuPY, Numba, PyTorch, and others. The Dask API makes it easy to scale to multiple nodes or multiple GPUs, and the  RAPIDS Memory Manager (RMM) integrates with XGBoost, so you can share a single, high-speed memory pool. The GPU-accelerated XGBoost algorithm makes use of fast parallel prefix sum operations to scan through all possible splits, as well as parallel radix sorting to repartition data. It builds a decision tree for a given boosting iteration, one level at a time, processing the entire dataset concurrently on the GPU. \n\n**GPU-Accelerated, End-to-End Data Pipelines with Spark + XGBoost**\n\nNVIDIA understands that machine learning at scale delivers powerful predictive capabilities for data scientists and developers and, ultimately, to end users. But this at-scale learning depends upon overcoming key challenges to both on-premises and cloud infrastructure, like speeding up pre-processing of massive data volumes and then accelerating compute-intensive model training.\n\nNVIDIA’s initial release of spark-xgboost enabled training and inferencing of XGBoost machine learning models across Apache Spark nodes. This has helped make it a leading mechanism for enterprise-class distributed machine learning.\n\nGPU-Accelerated Spark XGBoost speeds up pre-processing of massive volumes of data, allows larger data sizes in GPU memory, and improves XGBoost training and tuning time.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Light GBM</div>\n\nThe development of Boosting Machines started from AdaBoost to today’s much-hyped XGBOOST. XGBOOST has become a de-facto algorithm for winning competitions at Kaggle, simply because it is extremely powerful. But given lots and lots of data, even XGBOOST takes a long time to train.\n\nHere comes…. Light GBM into the picture.\n\n<img width=\"913\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/84183476-d981-4f3d-853e-0b5cad4d1f0c\">\n\n\nThe most natural question that will come to your mind is — Why another boosting machine algorithm? Is it faster than XGBOOST?\n\nWell, you guessed it right !!! In this post, we will compare between Light GBM & XGBoost and their performance with an example.\n\n\n\n## <span style=\"color: #016FD0;\">What is Light GBM?</span> \n\nLight GBM is a fast, distributed, high-performance gradient boosting framework based on decision tree algorithm, used for ranking, classification and many other machine learning tasks.\n\nSince it is based on decision tree algorithms, it splits the tree leaf wise with the best fit whereas other boosting algorithms split the tree depth wise or level wise rather than leaf-wise. So when growing on the same leaf in Light GBM, the leaf-wise algorithm can reduce more loss than the level-wise algorithm and hence results in much better accuracy which can rarely be achieved by any of the existing boosting algorithms.\n\nBefore is a diagrammatic representation by the makers of the Light GBM to explain the difference clearly.\n\n### <span style=\"color: #016FD0;\">Structural Differences in LightGBM and XGBoost</span> \n\nLightGBM uses a novel technique of **Gradient-based One-Side Sampling (GOSS)** to filter out the data instances for finding a split value while XGBoost uses pre-sorted algorithm & Histogram-based algorithm for computing the best split. Here instances are observations/samples.\n\nIn simple terms, Histogram-based algorithm splits all the data points for a feature into discrete bins and uses these bins to find the split value of the histogram. While it is efficient than the pre-sorted algorithm in training speed which enumerates all possible split points on the pre-sorted feature values, it is still behind GOSS in terms of speed.\n\nSo what makes this GOSS method efficient?\nGOSS (Gradient Based One Side Sampling) is a novel sampling method which downsamples the instances on the basis of gradients. As we know instances with small gradients are well trained (small training error) and those with large gradients are undertrained. A naive approach to downsample is to discard instances with small gradients by solely focussing on instances with large gradients but this would alter the data distribution. In a nutshell, GOSS retains instances with large gradients while performing random sampling on instances with small gradients.\n\n## <span style=\"color: #016FD0;\">Advantages of Light GBM</span>\n\n1. **Faster training speed and higher efficiency:** Light GBM use histogram based algorithm i.e it buckets continuous feature values into discrete bins which fasten the training procedure.\n\n1. **Lower memory usage:** Replaces continuous values to discrete bins which result in lower memory usage.\n\n1. **Better accuracy than any other boosting algorithm:** It produces much more complex trees by following leaf wise split approach rather than a level-wise approach which is the main factor in achieving higher accuracy. However, it **can sometimes lead to overfitting which can be avoided by setting the max_depth** parameter.\n\n1. **Compatibility with Large Datasets:** It is capable of performing equally good with large datasets with a significant reduction in training time as compared to XGBOOST.\n\n\n## <span style=\"color: #016FD0;\">Parameters of light GBM</span>\n\n1. `num_leaves`: the number of leaf nodes to use. Having a large number of leaves will improve accuracy, but will also lead to overfitting.\n\n1. `min_child_samples`: the minimum number of samples (data) to group into a leaf. The parameter can greatly **assist with overfitting**: larger sample sizes per leaf will reduce overfitting (but may lead to under-fitting).\n\n1. `max_depth`: controls the depth of the tree explicitly. Shallower trees reduce overfitting.\n\n1. **Tuning for imbalanced data:**\n\n    - `scale_pos_weight`: The simplest way to account for imbalanced or skewed data is to add weight to the positive class examples. the weight can be calculated based on the number of negative and positive examples: sample_pos_weight = number of negative samples / number of positive samples.\n\n1. **Tuning for overfitting:**In addition to the parameters mentioned above the following parameters can be used to control overfitting:\n    \n    - `max_bin`: the maximum numbers of bins that feature values are bucketed in. A **smaller** max_bin **reduces overfitting.**\n\n    - `min_child_weight`: the minimum sum hessian for a leaf. In conjuction with min_child_samples, **larger values reduce overfitting.**\n\n    - `bagging_fraction` and `bagging_freq`: enables bagging (subsampling) of the training data. Both values need to be set for bagging to be used. The frequency controls how often (iteration) bagging is used. **Smaller fractions and frequencies reduce overfitting.**\n\n    - `feature_fraction`: controls the subsampling of features used for training (as opposed to subsampling the actual training data in the case of bagging). **Smaller fractions reduce overfitting.**\n\n    - `lambda_l1` and `lambda_l2`: controls L1 and L2 regularization.\n\n1. **Tuning for accuracy:** Accuracy may be improved by tuning the following parameters:\n\n    - `max_bin`: a larger max_bin increases accuracy.\n\n    - `learning_rate`: using a smaller learning rate and increasing the number of iterations may improve accuracy.\n\n    - `num_leaves`: increasing the number of leaves increases accuracy with a high risk of overfitting.\n\n\n## <span style=\"color: #016FD0;\">Conclusion</span>\n\nSo now let’s compare LightGBM with XGBoost by applying both the algorithms to a census income dataset and then comparing their performance.\n\nThere has been only a slight increase in accuracy, AUC score and a slight decrease in rsme score by applying XGBoost over LightGBM but there is a significant difference in the execution time for the training procedure. **Light GBM is very fast when compared to XGBOOST** and is **a much better approach when dealing with large datasets.**\n\nThis turns out to be a huge advantage when you are working on large datasets in limited time competitions.\n\n**End Notes**\nIn this post, I’ve tried to compare the performance of Light GBM vs XGBoost. One of the disadvantages of using this LightGBM is its narrow user base — but that is changing fast. This algorithm apart from being more accurate and time-saving than XGBOOST has been limited in usage due to less documentation available. However, this algorithm has shown far better results and has outperformed existing boosting algorithms.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">CatBoost</div>\n\n## <span style=\"color: #016FD0;\">Introduction</span>\n\nCatBoost originated in a Russian company named Yandex. It is one of the latest boosting algorithms out there as it was made available in 2017. There were many boosting algorithms like XGBoost, LightGBM etc. but none comes close to the CatBoost for several reasons.\n\nCatBoost means Categorical Boosting because it is designed to work on categorical data flawlessly, If you have Categorical data in your dataset\n\nHere are some features of the CatBoost, which makes it stand apart from all the other Boosting Algorithm.\n\n- High Quality without parameter tuning\n- Categorical Features support\n- The fast and scalable GPU version\n- Improved accuracy by reducing overfitting\n- Fast Predictions\n- Works well with less data\n\nFor the reason mentioned above CatBoost is beloved in recent Kaggle competitions, now let’s answer another question When to use this Amazing Boosting Algo?\n\n\n## <span style=\"color: #016FD0;\">When to Use the CatBoost Algorithm?</span>\n\nThere are two types of Data out there Heterogeneous data and Homogeneous data.\n\n- **Heterogeneous data:** It is any data with high variability of data types and formats. They can be ambiguous and low quality due to missing values, high data redundancy. example: Dataset to predict Credit Score\n\n- **Homogeneous data:** It is dataset made up of the things which are similar to each other, which means that the entire dataset is of same data type and formate. example: Dataset of Images, Video, Sound, Text.\n\nCatBoost Works well on Heterogeneous data.\n\nSo, if you have a dataset of Images, Text or Sound it is probably a good idea to use Neural Networks as Neural Networks are state of the art for these types of Homogeneous data.\n\nBut, if you have a classification problem with Heterogeneous data then using CatBoost will be the safest thing to do because CatBoost is able to outperform the majority of the Boosting algorithms in the first run.\n\n\n## <span style=\"color: #016FD0;\">Parameters of CatBoost</span>\n\n\nTo solve this problem after importing the dataset I applied the pre-processing using Standard Scaler, and after that came the main part where we apply CatBoost Classifier. Let’s see the classifier code and understand every-single-line.\n\nSo, first of all we have to import the CatBoost Classifier as shown in the first line of the code.\n\nAfter importing the CatBoost Library we will create our model, now let’s go through those parameters.\n\n- **Iterations:** 1000 iterations means the CatBoost algorithm will run 1000 times to minimize the loss function.\n\n- **Loss Function:** In here as we are classifying multiple classes we have to specify ‘Multiclass’. In the case of Binary classification, it is okay if we don't mention the Loss Function the algorithm will understand and perform binary classification.\n\n- **bootstrap_type:** This parameter affects the Regularization & speed of the algorithm aspects of choosing a split for a tree when building the tree structure. Here we have chosen Bayesian, But it is okay if we didn’t specify this parameter.\n\n- **eval_metric:** In here as we have to do multiclass classification we have chosen ‘Multiclass’ as eval_metric and when working with Binary Classification we don’t have to specify this parameter.\n\n- **leaf estimation iterations:** This parameter defines rules for calculating leaf values after selecting the tree structure, we have taken 100 but it is also okay to not specify this parameter.\n\n- **random strength:** It specifies how random do we want our gradient boosting trees to be from each other. It is okay if we didn’t specify.\n\n- **depth:** How deep do we want our tree to be I have specified 7 because it gave me the highest accuracy but it is okay not to specify it and let the CatBoost algorithm use its default value.\n\n- **l2 leaf regularization:** To specify the L2-regularization value, we have taken 5 but it’s not mandatory.\n\n- **learning rate:** It is very important but generally default CatBoost learning rate of 0.03 also works well.\n\n- **Bagging temperature:** Defines the settings of the Bayesian bootstrap. It is used by default in classification and regression modes. Use the Bayesian bootstrap to assign random weights to objects. Not mandatory to specify.\n\n- **task type:** It is very much recommended to use CatBoost algorithm with GPU only because with CPU CatBoost algorithm becomes quite slow.\n\nAfter changing the parameters and finding the best parameter we will fit the model and predict the output. **It is very easy to work with and CatBoost only needs several parameters to tune**. Then you will be able to easily get very good performance compared to other boosting algorithms. **In the CatBoost you can run the model with just specifying the dataset type (Binary or Multiclass classification) and still you will be able to get a very good score without any overfitting.**","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Comparison of Gradient Boosting Variants</div>\n\nGradient Boosting is an ensemble model of a sequential series of shallow Decision Trees. The single trees are weak learners with little predictive skill, but together, they form a strong learner with high predictive skill. In this section, we will discuss different implementations of Gradient Boosting. The focus is to give a high-level overview of different implementations and discuss the differences. For a more in-depth understanding of each framework, further literature is given.\n\n## <span style=\"color: #016FD0;\">Benefits of Gradient Boosting</span>\n\nGradient boosting is one of the most popular machine learning algorithms for tabular datasets. It is powerful enough to:\n\n1. **find any nonlinear relationship between your model target and features**\n1. to have great usability that **can deal with missing values,** **outliers,** and **high cardinality categorical values** on your features without any special treatment. \n\nWhile you can build barebone gradient boosting trees using some popular libraries such as **XGBoost** or **LightGBM** without knowing any details of the algorithm, you still want to know how it works when you start tuning hyper-parameters, customizing the loss functions, etc., to get better quality on your model.\n\n**CatBoost**, **LightGBM**, and **XGBoost** are all variations of gradient boosting algorithms. Now you’ve understood the difference between bagging and boosting, we can move on to the differences in how the algorithms implement gradient boosting. The frameworks share common features, which include:\n\n1. **Gradient Boosting:** All three libraries employ gradient boosting techniques, which involve building an ensemble of weak learners (typically decision trees) in a sequential manner. Each new learner aims to correct the errors made by the previous one, resulting in a strong combined model with improved performance.\n\n1. **Tree-based Learning:** XGBoost, CatBoost, and LightGBM utilize tree-based learning algorithms. They construct decision trees as their base learners and optimize them to minimize a loss function.\n\n1. **Handling of Large Datasets:** These frameworks have been designed to work efficiently with large datasets, providing faster training and prediction times compared to other gradient boosting libraries.\n\n1. **Regularization:** All three libraries incorporate regularization techniques, such as L1 and L2 regularization, to prevent overfitting and improve generalization.\n\n1. **Tuning Parameters:** XGBoost, CatBoost, and LightGBM offer various hyperparameters that can be tuned to optimize model performance. Common parameters include learning rate, tree depth, number of trees, and regularization terms.\n\n1. **Cross-Validation:** Each of these libraries supports built-in cross-validation, allowing users to assess model performance and select the optimal number of boosting rounds.\n\n1. **Parallelization:** XGBoost, CatBoost, and LightGBM can utilize parallel processing during the training phase, which helps to speed up the computation and reduce training time.\nDespite their similarities, each framework has its unique strengths and features, making them suitable for different use cases and data types.\n\n## <span style=\"color: #016FD0;\">1. Decision Trees</span>\n\n### <span style=\"color: #016FD0;\">Advantages</span>\n\n1. **Simplicity:** Decision trees are intuitive and easy to understand. The decision-making process can be visualized and interpreted, which is a significant advantage when we need to explain the model to non-technical stakeholders.\n\n1. **Versatility:** They can handle both categorical and numerical data, making them a flexible choice for a wide range of problems.\n\n1. **No need for data preprocessing:** Decision trees require relatively little data preparation. They **are not affected by outliers and missing values to the extent that some other algorithms are.**\n\n1. **Feature selection:** Decision trees inherently perform feature selection, **using the most informative features first.** This can be particularly beneficial when dealing with datasets with a large number of features.\n\n### <span style=\"color: #016FD0;\">Disadvantages</span>\n\n1. **Overfitting:** Decision trees can create overly complex trees that don’t generalize well to unseen data. They can create a tree that perfectly classifies the training data but performs poorly on the test data, a problem known as overfitting. Techniques like pruning, setting the minimum number of samples required at a leaf node, or setting the maximum depth of the tree can help mitigate overfitting.\n\n1. **Bias towards features with more levels:** Decision trees can be biased towards variables with more levels. Features with more unique values or categories may be favored over others, potentially leading to suboptimal trees.\n\n1. **Instability:** Decision trees are sensitive to small changes in the data. A slight variation can result in a drastically different tree. This problem is mitigated in the ensemble methods that we will discuss later, like Random Forests and XGBoost.\n\n1. **Difficulty in capturing complex relationships:** While decision trees work well for decisions that can be structured hierarchically, they can struggle with tasks where features interact in complex ways, or where the decision boundary is more complex than can be captured by simple splits.\n\n## <span style=\"color: #016FD0;\">2. Random Forest</span>\n\n### <span style=\"color: #016FD0;\">Advantages</span>\n\n1. **Robust to Overfitting:** Due to the random nature of the Random Forest (random sampling of data points and features), the model is less prone to overfitting than a single decision tree. Each individual tree gets a different view of the data, so the overall model can capture a broader picture of the data without as much risk of memorizing the training set.\n\n1. **Handles Large Datasets and Feature Spaces:** Random Forest can easily handle datasets with a high dimensionality (many features) and a large number of data points, making it a good choice for complex datasets.\n\n1. **Parallelizable:** The training of the individual trees can be done in parallel, leading to faster training times.\n\n1. **Feature Importance:** Random Forest can provide insights into which features are most important in making predictions.\n\n\n### <span style=\"color: #016FD0;\">Disadvantages</span>\n\n1. **Complexity:** A Random Forest model creates a lot of trees (as defined by the user), which can make the model more complex and computationally expensive than a single decision tree.\n\n1. **Less Interpretability:** While a single decision tree is easily interpretable, this is not the case with a Random Forest. The decision-making process of a Random Forest is not as straightforward to visualize or explain due to the aggregation of many trees.\n\n1. **Longer Prediction Time:** Due to the need to make predictions with each tree in the forest, prediction time can be longer compared to other models.\n\n## <span style=\"color: #016FD0;\">3. XGBoost</span>\n\nXGBoost stands for eXtreme Gradient Boosting and is an algorithm focusing on **optimizing computation speed and model performance.** XGBoost was first developed by Tianqi Chen. It became popular and famous for many winning solutions in Machine Learning competitions. **XGBoost can be used as a separate library, but also as an integration in sklearn.** This makes it easy to combine the variety of methods available in sklearn. **The XGBoosts algorithm is capable of handling large and complex datasets.**\n\neXtreme Gradient Boosting is an **efficient** and **scalable implementation** of gradient boosted decision trees. It offers **great performance, is parallelizable, and handles missing values well.**\n\n1. **Handling missing values:** XGBoost automatically learns the optimal direction to handle missing values during training. This reduces the need for manual imputation.\n1. **Regularization:** XGBoost has built-in L1 (Lasso) and L2 (Ridge) regularization techniques, which help prevent overfitting.\n\n1. **Custom objective functions:** Users can define custom objective functions and evaluation criteria, making it more adaptable to different problems.\n\n1. **Early stopping:** XGBoost can stop training if the model's performance on a validation set does not improve after a certain number of iterations, reducing computation time.\n\n\n### <span style=\"color: #016FD0;\">Advantages</span>\n\n1. **Performance:** XGBoost is renowned for its performance and efficiency. It often provides superior results compared to other machine learning algorithms, especially on datasets with a mix of categorical and numeric features.\n\n1. **Speed:** XGBoost is parallelizable, meaning that it can use multiple cores on the CPU to train models faster. It also includes a number of techniques that make it faster and more memory-efficient, such as column block for storing and sorting data, and histogram-based splitting for handling continuous variables.\n\n1. **Versatility:** XGBoost can handle a variety of data types, including numeric, categorical, and ordinal data. **It can also handle missing values internally,** reducing the need for extensive data preprocessing.\n\n1. **Robustness:** XGBoost incorporates **a built-in regularization parameter that helps to avoid overfitting.** This parameter can control model complexity, making the model more generalizable to unseen data.\n\n### <span style=\"color: #016FD0;\">Disadvantages</span>\n\n1. **Prone to Overfitting if Not Properly Tuned:** Although XGBoost has regularization parameters to control overfitting, it can still be prone to overfitting if not properly tuned. Careful tuning of parameters such as the `learning rate`, the `depth of the tree`, and the `number of trees` is necessary.\n\n1. **Requires Careful Tuning:** XGBoost has a number of hyperparameters that need to be set, and getting the right combination can be challenging. `Grid search` or `randomized search` methods are often required to find the optimal settings.\n\n1. **Less Interpretability:** Although individual trees can be interpreted, the final model, which consists of an ensemble of many trees, can be difficult to interpret compared to simpler models like linear regression or single decision trees.\n\n### <span style=\"color: #016FD0;\">Real-world applications of XGBoost</span>\n\nXGBoost’s high performance and versatility have led to its use in a wide variety of applications. Here are a few examples:\n\n1. **Learning to rank:** One of the most popular use cases for the XGBoost algorithm is as a ranker. In information retrieval, the goal of learning to rank is to serve users content ordered by relevance. In XGBoost, the XGBRanker is based on the LambdaMART algorithm11.\n\n1. **Anomaly Detection:** XGBoost can be used to detect unusual patterns that deviate from the norm. This is useful in cybersecurity, fraud detection, and fault detection.\n\n1. **Fraud Detection:** Financial institutions and banks can use XGBoost to identify suspicious transactions and patterns. By training the model with historical data, XGBoost can accurately classify fraudulent activities, helping businesses save millions of dollars and protect their reputation.\n\n1. **Healthcare:** XGBoost has been employed in predicting disease progression, such as diabetes or cancer, by analyzing complex datasets with multiple features. It allows healthcare professionals to identify high-risk patients and intervene proactively, ultimately improving patient outcomes.\n\n1. **Predictive Analytics:** Businesses use XGBoost to forecast future trends and make strategic decisions. For example, predicting customer churn, sales forecasting, or inventory prediction.\n\n1. **Natural Language Processing:** XGBoost is used in various natural language processing tasks, such as sentiment analysis or topic modeling.\n\n1. **Recommender Systems:** XGBoost can be used in recommender systems to predict user preferences and recommend products or services. It plays a key role in the development of personalized recommendation systems where they predict the likelihood of a user liking a particular item.\n\n1. **Medical Diagnostics:** XGBoost can be used to diagnose diseases by identifying patterns in patient data.\n\n1. **Banking:** Random Forests are used in banking to predict if a customer is likely to default on a loan, helping banks manage their risk.\n\n1. **E-commerce:** Online retailers use Random Forests to predict whether a customer will like a product or not based on their past behavior.\n    - **Advertisement click through rate prediction:** Researchers used an XGBoost trained model to determine how frequently online ads had been clicked in 10 days of click through data. The goal of the research was to measure the effectiveness of online ads and pinpoint which ads work well\n\n1. **Stock Market:** Random Forests are used to predict the behavior of stock markets based on historical data, helping investors make informed decisions.\n\n1. **Store sales prediction:** XGBoost may be used for predictive modeling, as demonstrated in this paper where sales from 45 Walmart stores were predicted using an XGBoost model13.\n\n\n## <span style=\"color: #016FD0;\">4. LightGBM</span>\n\nLightGBM is short for Light Gradient-Boosting Machine, and **was developed by Microsoft**. It has similar advantages as XGBoost and is also **able to handle large and complex datasets.**\nThe main difference between LightGBM and XGBoost is the way the trees are built. In LightGBM the trees are not grown level-wise, but leave-wise. Light Gradient Boosting Machine is a gradient boosting framework that uses tree-based learning algorithms. It's efficient and faster, can handle large datasets, and is optimized for parallel and GPU processing.\n\n1. **Histogram-based algorithms:** LightGBM uses histogram-based algorithms to **speed up training, which reduces memory usage and computation time.**\n\n1. **Exclusive Feature Bundling (EFB):** LightGBM can automatically bundle exclusive (non-overlapping) features, further reducing memory usage and speeding up training.\n\n1. **Leaf-wise tree growth:** Unlike other gradient boosting frameworks that grow trees level-wise, **LightGBM grows trees leaf-wise, which can result in a more accurate model.**\n\n1. **Support for large datasets:** LightGBM can handle large datasets efficiently, making it suitable for big data problems.\n\n### <span style=\"color: #016FD0;\">Real-world applications of LightGBM</span>\n\n1. **Smart Cities:** LightGBM's ability to efficiently handle large datasets makes it an ideal choice for managing smart city infrastructure. By analyzing data from sensors, traffic patterns, and public services, **LightGBM can optimize energy consumption, improve traffic flow, and enhance overall quality of life for urban residents.**\n\n1. **E-commerce:** LightGBM can be employed to analyze massive amounts of customer data in real-time, enabling e-commerce businesses to make personalized product recommendations, optimize pricing, and predict inventory levels. This, in turn, leads to increased customer satisfaction and higher revenue.\n\n## <span style=\"color: #016FD0;\">5. CatBoost</span>\n\nCatBoost was developed by Yandex, a Russian technology company and **has a special focus on how categorical values are treated in Gradient Boosting.** It was developed in 2016 and was based on previous projects focussing on Gradient Boosting algorithms. It was first released open-source in 2017. CatBoost is a gradient boosting library with categorical feature support, providing fast and accurate results. It can **handle categorical variables** effectively, **requires less parameter tuning,** and has built-in overfitting prevention.\n\n### <span style=\"color: #016FD0;\">Advantages CatBoost</span>\n\n1. **Handling categorical variables:** CatBoost **uses an efficient encoding scheme for categorical variables, called \"Ordered Target Statistics,\"** reducing the need for preprocessing.\n\n1. **Robustness to parameter choices:** CatBoost is **less sensitive to hyperparameter tuning and can often deliver good performance with default settings.**\n\n1. **Feature importance:** CatBoost provides **built-in tools for feature importance calculation,** making it easier to understand the influence of each feature on the model.\n\n1. **Model interpretability:** CatBoost supports **model interpretation techniques like SHAP (SHapley Additive exPlanations)** values to help explain predictions.\n\n\n### <span style=\"color: #016FD0;\">Real-world applications of CatBoost</span>\n\n1. **Marketing Analytics:** CatBoost excels at handling categorical data, making it a perfect choice for analyzing customer behavior and creating targeted marketing campaigns. By leveraging CatBoost, companies can optimize their marketing strategies, improve customer engagement, and boost sales.\n\n1. **Insurance:** CatBoost can help insurance companies more accurately price policies and assess risk by analyzing a wide range of categorical variables, such as customer demographics, vehicle type, and claim history. This leads to better risk management and increased profitability for insurers.\n\n## <span style=\"color: #016FD0;\">Comparisons and When to Use What</span>\n\nComparative Discussion of the Three Algorithms\n\n1. **Decision Trees:** Decision trees are simple and easy to understand. They work well for data with categorical features and provide interpretability that the other two models lack. However, decision trees are prone to overfitting and might not provide the level of accuracy that Random Forest and XGBoost can achieve.\n\n1. **Random Forest:** Random Forest is an ensemble of decision trees that averages the results to improve the final output. It’s more robust to overfitting than a single decision tree and handles large datasets and feature spaces effectively. However, Random Forests may not be as easy to interpret as single decision trees and can be computationally intensive for very large datasets or complex trees.\n\n1. **XGBoost:** XGBoost is a powerful and efficient algorithm that leverages the concept of boosting to build a strong model. It often provides the highest accuracy among these three models, but it can be more complex to tune and interpret.\n\nWhen to Use Which Algorithm\nChoosing between these three models typically depends on your specific needs, the size and type of your data, and the balance between interpretability and accuracy that you desire. Here are some general guidelines:\n\n- Use **Decision Trees** when interpretability is highly important. Decision trees are also good when you have categorical data or when you need a quick and simple model.\n\n- Use **Random Forests** when you need a better balance between interpretability and accuracy. Random Forests are also good when you have large datasets with many features.\n\n- Use **XGBoost** when your primary concern is performance and you have the resources to tune the model properly. XGBoost is also effective when you have a mix of categorical and numerical features, and when you have a large volume of data.\n    - **Use when:** The dataset has **missing values**, **requires regularization to prevent overfitting**, or **benefits from custom objective functions.**\n\n- **CatBoost:**\n    - **Use when:** The dataset has **many categorical features**, **model interpretability** is important, or you **prefer minimal hyperparameter tuning**.\n\n- **LightGBM:**\n    - **Use when:** The **dataset is large, has many features, or requires faster training times**\n\nIn the end, it’s always a good practice to try different models and use cross-validation to select the model that performs best on your specific task. Remember, the ‘best’ model is the one that best serves your needs, not necessarily the one with the highest accuracy. However, as we’ve discussed, no single algorithm is the best choice for every problem. The choice depends on the specific characteristics of your data, the problem at hand, and the balance between interpretability and prediction accuracy. Each of these algorithms shines in different scenarios, and understanding their intricacies is crucial in determining when to use which.\n\n### <span style=\"color: #016FD0;\">Feature Handling</span>\n\n1. **Gradient Boosting in sklearn:**\n    1. **GradientBoostingRegressor / GradientBoostingClassifier:**\n        - Numerical Features: Gradient Boosting in sklearn expects numerical input.\n        - Categorical Features: Gradient Boosting in sklearn does not natively support categorical features. They need to be encoded using e.g. one-hot encoding or label-encoding.\n        - Missing Values: Gradient Boosting in sklearn does not handle missing values. Missing values need to be removed or imputed before training the model.\n\n    1. **HistGradientBoostingRegressor / HistGradientBoostingClassifier:** Histogram-based Gradient Boosting by sklearn was inspired by LightGBM.\n        - Numerical & Categorical Features: The Histogram-based Gradient Booster natively supports numerical and categorical features. For more details, please check the documentation.\n        - Missing Values: The Histogram-based Gradient Booster can handle missing values natively. More explanations can be found in the documentation.\n\n1. **XGBoost:**\n    - Numerical Features: XGBoost supports directly numerical data.\n    - Categorical Features: Both the native environment and the sklearn interface support categorical features using the parameter enable_categorical. Examples of both interfaces can be found in the documentation.\n    - Missing Values: XGBoost natively supports missing values. **Note, that the treatment of missing values differs with the booster type used.** For more explanations, please refer to the documentation.\n\n1. **LightGBM:**\n    - Numerical Features: LightGBM supports numerical data.\n    - Categorical Features: LightGBM has built-in support for categorical features. Categorical features can be specified using the parameter categorical_feature. For more details, please refer to the documentation.\n    - Missing Values: LightGBM natively supports the handling of missing values. It can be disabled using the parameter use_missing=false. More information is available in the documentation.\n\n1. **CatBoost:**\n    - Numerical Features: CatBoost supports numerical data.\n    - Categorical Features: CatBoost is specifically designed to handle categorical features without needing to preprocess them into numerical formats. How this transformation is done, is described in the documentation\n    - Missing Values: CatBoost has robust handling for missing values, treating them as a separate category or using specific strategies for imputation during training. How missing values are treated depends on the feature type. More details can be found in the documentation\n\n<img width=\"800\" alt=\"image\" src=\"https://github.com/user-attachments/assets/2363a80f-90fa-4425-954e-cb9dbd598945\">\n\n\n### <span style=\"color: #016FD0;\">Overfitting</span>\n\n1. **sklearn:** Traditional Gradient Boosting and Histogram-based Gradient Boosting both use **regularization**, **learning rate adjustment**, **subsampling**, and **early stopping** to combat overfitting. However, Histogram-based Gradient Boosting’s binning process adds an extra layer of complexity reduction and efficiency, making it particularly effective for large datasets.\n\n1. **XGBoost:** Combines regularization, shrinkage, early stopping, and **tree pruning.**\n\n1. **LightGBM:** Employs leaf-wise growth with **depth limits**, regularization, and early stopping.\n\n1. **CatBoost:** CatBoost uses **Ordered Boosting** and regularization to prevent overfitting.\n\n### <span style=\"color: #016FD0;\">GPU Support</span>\n\nThe implementations of XGBoost, LightGBM, and CatBoost support GPU usage, which enhances performance and speed on large datasets compared to the sklearn implementations.\n\n### <span style=\"color: #016FD0;\">Tree Growths</span>\n\nTraditionally Decision Trees are grown level-wise. That means first a level is developed completely, such that all leaves are grown before moving to the next level. An alternative approach is leaf-wise tree growth. **In this case, this criterion is relaxed and the tree is grown with the highest loss reduction considering all leaves, which may result in unsymmetric, irregular trees of larger depth.** The different methods are illustrated in the plot below. **The leaf-wise tree is built using the global best split, while the level-wise growth only uses a local minimum for the next split, also leaf-wise growth is computationally more efficient and less memory intensive.**\n\n1. **Gradient Boosting in sklearn:**\n    - GradientBoostingRegressor / GradientBoostingClassifier: In the Gradient Boosting algorithm of sklearn the trees are grown level-wise, that is they are grown level by level, expanding all nodes at a given depth before moving deeper. **This produces balanced trees and can be slower and more memory-intensive.*\n    - HistGradientBoostingRegressor / HistGradientBoostingClassifier: In HistGradientBossting histograms are used to approximate the data distributions, and then the trees are grown level by level. This algorithm is faster and more efficient with large datasets compared to traditional Gradient Boosting.\n\n1. **XGBoost:** The trees in XGBoost are also grown level-wise, **but advanced algorithms are used to find the optimal split and regularization, which improves computational efficiency. XGBoost offers the possibility to use histogram-based splitting.**\n\n1. **LightGBM:** In LightGBM the trees are grown **leaf-wise**. This produces **unbalanced and deeper trees**, **but is in general faster and more efficient.**\n\n1. **CatBoost:** CatBoost uses symmetric tree growth. Symmetric trees are grown by splitting all leaves at the same depth identically.\n\n<img width=\"696\" alt=\"image\" src=\"https://github.com/user-attachments/assets/bcf7564d-334a-4972-81b9-abf31fbf326e\">\n\n## <span style=\"color: #016FD0;\">Summary</span>\n\nGradient Boosting is a popular and performant algorithm for both classification and regression tasks. In this post we compared different implementations of this algorithm, comparing the method the individual trees are grown, but also their ability to handle categorical features and missing values. The main characteristics are that **LightGBM uses a leaf-wise tree growth strategy in contrast to the other algorithm and CatBoost was specifically designed to handle categorical features.** However other algorithms also offer native support for categorical features. **The implementations of XGBoost, LightGBM, and CatBoost additionally support the usage of a GPU, which makes them especially suitable for large datasets.**\n\nXGBoost, CatBoost, and LightGBM are popular gradient boosting frameworks that share several common features. They all employ gradient boosting with tree-based learning algorithms, work efficiently with large datasets, incorporate regularization techniques, offer tunable hyperparameters, and support built-in cross-validation. Additionally, these libraries leverage parallelization during the training phase to reduce computation time. Despite their similarities, each framework has unique strengths, making them suitable for different use cases and data types.\n\n1. **XGBoost**\n    - Efficient and scalable implementation of gradient boosted decision trees\n    - High performance and parallelizable\n    - Handles missing values well\n\n1. **CatBoost**\n    - Designed specifically for handling categorical features\n    - Fast and accurate results with less parameter tuning\n    - Built-in overfitting prevention\n\n1. **LightGBM**\n    - Uses tree-based learning algorithms for efficiency and speed\n    - Optimized for large datasets and parallel/GPU processing\n    - Faster performance than XGBoost and CatBoost, especially on large dataset\n\n## <span style=\"color: #016FD0;\">XGBoost vs. CatBoost</span>\n\nCatBoost is another gradient boosting framework. Developed by Yandex in 2017, it specializes in handling categorical features without any need for preprocessing and generally performs well out-of-the-box without the need to perform extensive hyperparameter tuning8. Like XGBoost, CatBoost has built in support for handling missing data. CatBoost is especially useful for datasets with many categorical features. According to Yandex, the framework is used for search, recommendation systems, personal assistants, self-driving cars, weather prediction and other tasks.\n\n## <span style=\"color: #016FD0;\">XGBoost vs. LightGBM</span>\n\nLightGBM (Light Gradient Boosting Machine) is the final gradient boosting algorithm we will review. LightGBM was developed by Microsoft and first released in 20169. Where most decision tree learning algorithms grow trees depth-wise, LightGBM uses a leaf-wise tree growth strategy10. Like XGBoost, LightGBM exhibits fast model training speed and accuracy and performs well with large datasets.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Resources</div>\n\n1. [What is boosting?](https://www.ibm.com/topics/boosting)\n1. [What is boosting in machine learning?\n](https://aws.amazon.com/what-is/boosting/)\n1. [Boosting Algorithms Explained\n](https://towardsdatascience.com/boosting-algorithms-explained-d38f56ef3f30)\n1. [Ensemble Methods in Machine Learning: What are They and Why Use Them?](https://towardsdatascience.com/ensemble-methods-in-machine-learning-what-are-they-and-why-use-them-68ec3f9fef5f)\n1. [Ensemble methods: bagging, boosting and stacking](https://towardsdatascience.com/ensemble-methods-bagging-boosting-and-stacking-c9214a10a205)\n1. [Ensemble Learning, Bagging, and Boosting Explained in 3 Minutes\n](https://towardsdatascience.com/ensemble-learning-bagging-and-boosting-explained-in-3-minutes-2e6d2240ae21)\n1. [Bagging and Boosting\n](https://www.i2tutorials.com/machine-learning-tutorial/machine-learning-bagging-boosting/)\n1. [Ensemble Methods: Elegant Techniques to Produce Improved Machine Learning Results\n](https://www.toptal.com/machine-learning/ensemble-methods-machine-learning#:~:text=Ensemble%20methods%20are%20techniques%20that,than%20a%20single%20model%20would.)\n1. [Boosting and Bagging explained with examples !!!\n](https://medium.com/swlh/boosting-and-bagging-explained-with-examples-5353a36eb78d)\n1. [Ensemble Learning: Bagging and Boosting\n](https://towardsdatascience.com/ensemble-learning-bagging-and-boosting-23f9336d3cb0)\n1. [Ensemble methods: Bagging & Boosting\n](https://medium.com/swlh/difference-between-bagging-and-boosting-f996253acd22)\n1. [Ensemble Learning: From Basics to Advanced Techniques!\n](https://www.simplilearn.com/ensemble-learning-article)\n1. [The Complete Guide to Ensemble Learning + Papers](https://www.v7labs.com/blog/ensemble-learning)\n1. [A Comprehensive Guide to Ensemble Learning: What Exactly Do You Need to Know](https://neptune.ai/blog/ensemble-learning-guide)\n1. [Optimizing XGBoost: A Guide to Hyperparameter Tuning\n](https://medium.com/@rithpansanga/optimizing-xgboost-a-guide-to-hyperparameter-tuning-77b6e48e289d)\n1. [How to explain gradient boosting\n](https://explained.ai/gradient-boosting/)\n1. [XGBoost Algorithm: Long May She Reign!\n](https://towardsdatascience.com/https-medium-com-vishalmorde-xgboost-algorithm-long-she-may-rein-edd9f99be63d)\n1. [What is XGBoost? An Introduction to XGBoost Algorithm in Machine Learning\n](https://www.simplilearn.com/what-is-xgboost-algorithm-in-machine-learning-article)\n1. [XGBoost: theory and practice\n](https://towardsdatascience.com/xgboost-theory-and-practice-fb8912930ad6)\n1. [XGBoost: A Deep Dive into Boosting + CODE](https://medium.com/sfu-cspmp/xgboost-a-deep-dive-into-boosting-f06c9c41349)\n1. [XGBoost Algorithm Explained in Less Than 5 Minutes + CODE](https://medium.com/@techynilesh/xgboost-algorithm-explained-in-less-than-5-minutes-b561dcc1ccee)\n1. [What is XGBoost?\n](https://www.nvidia.com/en-us/glossary/xgboost/)\n1. [How XGBoost Works\n](https://docs.aws.amazon.com/sagemaker/latest/dg/xgboost-HowItWorks.html)\n1. [Getting started with XGBoost](https://blog.cambridgespark.com/getting-started-with-xgboost-3ba1488bb7d4)\n1. [All You Need to Know about Gradient Boosting Algorithm − Part 1. Regression\n](https://towardsdatascience.com/all-you-need-to-know-about-gradient-boosting-algorithm-part-1-regression-2520a34a502)\n1. [All You Need to Know about Gradient Boosting Algorithm − Part 2. Classification\n](https://towardsdatascience.com/all-you-need-to-know-about-gradient-boosting-algorithm-part-2-classification-d3ed8f56541e)\n1. [What is XGBoost?\n](https://mljar.com/glossary/xgboost/)\n1. [A Gentle Introduction to XGBoost for Applied Machine Learning](https://machinelearningmastery.com/gentle-introduction-xgboost-applied-machine-learning/)\n1. [Gradient Boosted Decision Trees-Explained\n](https://towardsdatascience.com/gradient-boosted-decision-trees-explained-9259bd8205af)\n1. [XGBoost — How does this work\n](https://medium.com/@prathameshsonawane/xgboost-how-does-this-work-e1cae7c5b6cb)\n1. [Decision Tree, Random Forest, and XGBoost: An Exploration into the Heart of Machine Learning\n](https://medium.com/@brandon93.w/decision-tree-random-forest-and-xgboost-an-exploration-into-the-heart-of-machine-learning-90dc212f4948)\n1. [XGBoost: Everything You Need to Know\n](https://neptune.ai/blog/xgboost-everything-you-need-to-know)\n1. [Gradient boosting: Tips to control your XGBoost model\n](https://www.capitalone.com/tech/machine-learning/how-to-control-your-xgboost-model/)\n1. [Gradient Boosting and XGBoost\n](https://medium.com/@gabrieltseng/gradient-boosting-and-xgboost-c306c1bcfaf5)\n1. [Overfitting, regularization, and early stopping ](https://developers.google.com/machine-learning/decision-forests/overfitting-gbdt)\n1. [A practical Guide on XGBoost hyperparameters tuning](https://www.kaggle.com/code/prashant111/a-guide-on-xgboost-hyperparameters-tuning)\n1. [XGBOOST vs LightGBM: Which algorithm wins the race !!! + CODE](https://towardsdatascience.com/lightgbm-vs-xgboost-which-algorithm-win-the-race-1ff7dd4917d#:~:text=LightGBM%20uses%20a%20novel%20technique,Here%20instances%20are%20observations%2Fsamples.)\n1. [Understanding CatBoost Algorithm](https://medium.com/analytics-vidhya/catboost-101-fb2fdc3398f3)\n1. [Hyperparameter tuning in XGBoost](https://blog.cambridgespark.com/hyperparameter-tuning-in-xgboost-4ff9100a3b2f)\n1. [What is XGBoost?](https://www.ibm.com/topics/xgboost)\n1. [4 Boosting Algorithms You Should Know: GBM, XGBoost, LightGBM & CatBoost](https://www.analyticsvidhya.com/blog/2020/02/4-boosting-algorithms-machine-learning/)\n1. [Gradient Boosting Variants - Sklearn vs. XGBoost vs. LightGBM vs. CatBoost\n](https://datamapu.com/posts/classical_ml/gradient_boosting_variants/)\n1. [Comparing the Titans of Machine Learning: XGBoost, CatBoost and LightGBM\n](https://www.linkedin.com/pulse/comparing-titans-machine-learning-xgboost-catboost-lightgbm-iljin/)\n1. [XGBoost Parameters Tuning: A Complete Guide with Python Codes\n](https://www.analyticsvidhya.com/blog/2016/03/complete-guide-parameter-tuning-xgboost-with-codes-python/)\n1. [When to Choose CatBoost Over XGBoost or LightGBM [Practical Guide] - Check also for RANKING\n](https://neptune.ai/blog/when-to-choose-catboost-over-xgboost-or-lightgbm)\n1. [XGBoost vs LightGBM: How Are They Different\n](https://neptune.ai/blog/xgboost-vs-lightgbm)","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">WIP: Work In Progress</div>","metadata":{}}]}