{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":35332,"databundleVersionId":3723648,"sourceType":"competition"},{"sourceId":3727003,"sourceType":"datasetVersion","datasetId":2213609}],"dockerImageVersionId":30732,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"color:#016FD0;margin:0;font-size:32px;font-family:Georgia;text-align:center;display:fill;border-radius:5px;overflow:hidden;font-weight:600;\">Mastering Tabular Data Classification</div>\n\n<div style=\"text-align:center\">\n    <img width=\"1065\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/912213f5-95a8-4720-90c4-b9bda8430238\">\n</div>\n<div style=\"text-align:center\">\n    <a href=\"https://unsplash.com/photos/a-person-holding-a-credit-card-and-a-cell-phone-aGkR0b7hgI8\">Photo from Unsplash</a>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:#016FD0;margin:0;font-size:32px;font-family:Georgia;text-align:center;display:fill;border-radius:5px;overflow:hidden;font-weight:600;\">Work in Progress...</div>\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Introduction</div>\n\nAccording to TransUnion credit card and personal loan data, delinquencies are expected to increase to levels not seen since 2010 in the United States.\n\nTransUnion forecasted severe credit card delinquencies to rise to 2.6% at the end of 2023 from 2.1% at the close of 2022. Unsecured personal loan delinquency rates will increase to 4.3% from 4.1% in the same timeframe.\n\nTherefore, credit default prediction has always been central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, and minimize risk and exposure, which leads to a better customer experience and sound business economics. By predicting which customers are at the highest risk of defaulting on their credit card accounts, issuers can take proactive steps to minimize risk and exposure.\n\nWe implemented this solution for one of the largest payment card issuers in the world. We are sharing practical tips for the successful execution of credit default prediction. \n\n## <span style=\"color: #016FD0;\">Project Description</span>\n\nThe approach involves analyzing a dataset containing credit card transactions and payment records for a group of customers over a 12-month period from 2021–2022.\n\nSpecifically, the data includes customer information from their last 1–13 statements. The analysis focuses on various features that are associated with credit card defaults, with the goal of accurately predicting if a customer is likely to default in the next 6 months. By analyzing these features, the aim is to identify patterns and trends that can be used to assess the risk of a customer defaulting on their credit card payments.\n\n## <span style=\"color: #016FD0;\">My Goal</span>\n\nThis notebook is designed to provide a thorough understanding of various concepts and techniques essential for mastering tabular data classification, especially when dealing with imbalanced datasets. Inspired by my own experiences during technical interviews for data science and machine learning roles, this notebook aims to provide a deep dive into essential concepts such as model selection, feature engineering, handling imbalanced datasets, and evaluation metrics. \n\nWhether you are a learner preparing for interviews or a practitioner seeking to enhance your skills, this resource offers valuable insights and practical techniques to help you tackle real-world data challenges effectively.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Import Python Libraries</div>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom typing import List\n\nimport os\n\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport matplotlib.colors\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\nfrom plotly.offline import init_notebook_mode\n\nimport plotly.figure_factory as ff\nimport plotly.graph_objects as go\n\nimport warnings\nwarnings.filterwarnings('ignore')\n","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:51:04.735836Z","iopub.execute_input":"2024-06-09T10:51:04.736906Z","iopub.status.idle":"2024-06-09T10:51:07.436095Z","shell.execute_reply.started":"2024-06-09T10:51:04.736844Z","shell.execute_reply":"2024-06-09T10:51:07.434227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Dataset Loading and Basic Exploration</div>\n\n## <span style=\"color: #016FD0;\">Time series data</span>\n\nA time series is a collection of data points gathered over a period of time and ordered chronologically. The primary characteristic of a time series is that it’s indexed or listed in time order, which is a critical distinction from other types of data sets. If you were to plot the points of time series data on a graph, and one of your axes would always be time.\n\nSometimes called time-stamped data, time series data is a time order indexed sequence of data points. Typically, these data points track change over time and consist of successive measurements over a fixed time interval made from the same source. Time-series data are observations obtained over time through repeated measurements and collected together. Expressed visually on a graph, one of the axes is always time when you plot the points in time from series data.\n\n\nTime series data is everywhere, since time is a constituent of everything that is observable. As our world gets increasingly instrumented, sensors and systems are constantly emitting a relentless stream of time series data. Such data has numerous applications across various industries.\n\n**Linear vs. nonlinear time series data** \n\nA linear time series is one where, for each data point Xt, that data point can be viewed as a linear combination of past or future values or differences. Nonlinear time series are generated by nonlinear dynamic equations. They have features that cannot be modelled by linear processes: time-changing variance, asymmetric cycles, higher-moment structures, thresholds and breaks.\n\n**What is time series analysis and Time Series Forecasting?**\n\nTime series analysis is the collection of data at specific intervals over a period to identify trends, seasonality, and residuals to aid in forecasting a future event. Time series analysis involves inferring what has happened to a series of data points in the past and attempting to predict future values. Analyzing time series data allows for extracting meaningful statistics and other data characteristics. As the name suggests, time series data is a collection of observations created by repeating measurements over time. Once you have that information, you can plot it on a graph and learn more about precisely what you’re tracking.\n\nTime series analysis is a specific way of analyzing a sequence of data points collected over an interval of time. In time series analysis, analysts record data points at consistent intervals over a set period of time rather than just recording the data points intermittently or randomly. However, this type of analysis is not merely the act of collecting data over time. \n\nWhat sets time series data apart from other data is that the analysis can show how variables change over time. In other words, time is a crucial variable because it shows how the data adjusts over the course of the data points as well as the final results. It provides an additional source of information and a set order of dependencies between the data. \n\nTime series analysis typically requires a large number of data points to ensure consistency and reliability. An extensive data set ensures you have a representative sample size and that analysis can cut through noisy data. It also ensures that any trends or patterns discovered are not outliers and can account for seasonal variance. Additionally, time series data can be used for forecasting—predicting future data based on historical data.\n\n***Time series analysis is used for non-stationary data—things that are constantly fluctuating over time or are affected by time. Industries like finance, retail, and economics frequently use time series analysis because currency and sales are always changing.***\n\n\n**Components of Time Series Data**\n\nTime series data is generally comprised of different components that characterize the patterns and behavior of the data over time. By analyzing these components, we can better understand the dynamics of the time series and create more accurate models. Four main elements make up a time series dataset:\n\n- **Trends:** show the general direction of the data, and whether it is increasing, decreasing, or remaining stationary over an extended period of time. Trends indicate the long-term movement in the data and can reveal overall growth or decline. For example, e-commerce sales may show an upward trend over the last five years.\n\n- **Seasonality:** refers to predictable patterns that recur regularly, like yearly retail spikes during the holiday season. Seasonal components exhibit fluctuations fixed in timing, direction, and magnitude. For instance, electricity usage may surge every summer as people turn on their air conditioners.\n\n- **Cycles:** demonstrate fluctuations that do not have a fixed period, such as economic expansions and recessions. These longer-term patterns last longer than a year and do not have consistent amplitudes or durations. Business cycles that oscillate between growth and decline are an example.\n\n- **Noise:** Finally, noise encompasses the residual variability in the data that the other components cannot explain. Noise includes unpredictable, erratic deviations after accounting for trends, seasonality, and cycles.\n\nIn summary, the key components of time series data are:\n\n1. **Trends:** Long-term increases, decreases, or stationary movement\n1. **Seasonality:** Predictable patterns at fixed intervals\n1. **Cycles:** Fluctuations without a consistent period\n1. **Noise:** Residual unexplained variability\n\nUnderstanding how these elements interact allows for deeper insight into the dynamics of time series data.\n\n**Info Box:💡** When researching and collecting data, it’s essential to know what kind of data you’re getting so you can interpret and analyze it well. Most of the time, there are two types of data in a research study:\n\n1. **Categorical data / Qualitative data:** Categorical data can be put in groups or categories using names or labels. Each piece of a categorical dataset, also known as qualitative data, may be assigned to only one category based on its qualities, and each category is mutually exclusive.  There are two primary categories of categorical data: \n    \n    1. **Nominal data:** This is the data category that names or labels its categories. It has features resembling a noun and is occasionally referred to as naming data.\n    1. **Ordinary data:** Elements with rankings, orders, or rating scales are included in this category of categorical data. Nominal data can be ordered and counted but not measured.\n\n1. **Numerical data / Quantitative data:** Data expressed in numerical terms rather than in natural language descriptions are called numerical data. Numerical data can be of two types:\n\n    1. **Discrete Data:** Countable numerical data are discrete data. They are mapped one-to-one to natural numbers, in other words. Age, the number of students in a class, the number of candidates in an election, etc., are a few examples of discrete data in general.\n    1. **Continuous Data:** This is an uncountable data type for numbers. A series of intervals on a natural number line is used to depict them. Student CGPA, height, and other continuous data types are a few examples.\n\n\n\n## <span style=\"color: #016FD0;\">Amex Dataset</span>\n\n- **Temporal Component:** A time series dataset involves data points collected or recorded at specific time intervals. In this case, the data includes multiple statement dates per customer_ID, which indicates that the data points are collected over time.\n\n- **Sequential Order:** The data is inherently ordered by time. Each customer's profile information is aggregated at different statement dates, implying that the data is sequentially ordered and the temporal order matters for analysis and prediction.\n\n- **Temporal Dependencies:** Time series data often exhibit dependencies between data points over time. For credit default prediction, a customer’s financial behavior in previous months (delinquency, spending, payments, balance, and risk variables) can influence the likelihood of default in the future.\n\nSpecific Features Indicating Time Series Nature:\n\n- **Monthly Customer Profiles:** The dataset contains aggregated profile features for each customer at each statement date, showing that the features change over time.\n\n- **Observation Window:** The target variable is based on an 18-month performance window and whether the customer defaults within 120 days after their latest statement. This long-term observation implies time-based trends and patterns are crucial for accurate predictions.\n\nIn summary, because the dataset includes features collected at multiple time points (statement dates) per customer and involves analyzing patterns over these time points, it is a time series dataset.\n\n## <span style=\"color: #016FD0;\">Resources</span>\n\n1. [Time Series Analysis: Definition, Types, Techniques, and When It's Used\n](https://www.tableau.com/learn/articles/time-series-analysis#definition)\n1. [Time Series Data Analysis: Definitions & Best Techniques in 2024](https://www.influxdata.com/what-is-time-series-data/)\n1. [Time Series Analysis Explained\n](https://www.sigmacomputing.com/resources/learn/what-is-time-series-analysis)\n1. [Time Series Data Definition](https://www.scylladb.com/glossary/time-series-data/)","metadata":{}},{"cell_type":"code","source":"%%time\ntrain_df = pd.read_feather('../input/amexfeather/train_data.ftr')","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:51:07.438875Z","iopub.execute_input":"2024-06-09T10:51:07.439976Z","iopub.status.idle":"2024-06-09T10:51:24.111853Z","shell.execute_reply.started":"2024-06-09T10:51:07.439915Z","shell.execute_reply":"2024-06-09T10:51:24.110562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Types of Approaches for Credit Default Prediction</div>\n\nFor a credit default prediction problem with a time series component, there are several approaches you can use, each with its own advantages and considerations.\n\nThe data includes the monthly credit card statements of each customer and hence many approaches can be used that leverage the property of time series or using\n\n1. Decision Tree models/ Neural Networks using aggregated features of customer’s data.\n\n1. Recurrent Neural Networks/Transformers with raw data of each customer’s data.\n\n## <span style=\"color: #016FD0;\">Time Series Models</span>\n\nModels that explicitly consider the temporal order of the data.\n\n1. **ARIMA (AutoRegressive Integrated Moving Average):** Suitable for univariate time series, ARIMA models can be used for feature engineering by creating lagged features.\n\n1. **Exponential Smoothing:** This method is useful for capturing trends and seasonality in time series data.\n\n## <span style=\"color: #016FD0;\">Traditional Machine Learning Models</span>\n\nThese models can be applied to the aggregated or transformed features.\n\n1. **Logistic Regression:** Simple and interpretable, logistic regression can be used if the features are aggregated appropriately (e.g., using summary statistics like means, max, min over time).\n\n1. **Decision Trees and Ensemble Methods (Random Forest, Gradient Boosting):** These models handle non-linearity and interactions between features well. You can use engineered features that capture temporal patterns. These models work well with engineered features that summarize the data.\n\n1. **Neural Networks:** Feedforward neural networks can be used on aggregated features to capture complex patterns in the data.\n\n\n## <span style=\"color: #016FD0;\">Deep Learning Models</span>\n\nThese models can capture complex patterns and dependencies in the data.\n\n1. **RNN (Recurrent Neural Networks):** Suitable for sequential data, RNNs can model dependencies over time but may suffer from vanishing gradient issues. These are designed to handle sequential data and can capture dependencies over time. They are suitable for working with raw, time-ordered data of each customer's credit card statements.\n\n1. **LSTM (Long Short-Term Memory):** A type of RNN designed to capture long-term dependencies and mitigate vanishing gradient problems.\n\n1. **GRU (Gated Recurrent Units):** Similar to LSTM but with a simpler architecture, GRUs are also effective in capturing temporal dependencies.\n\n## <span style=\"color: #016FD0;\">Hybrid and Advanced Models</span>\n\nCombining different techniques to leverage their strengths.\n\n1. **Temporal Convolutional Networks (TCN):** These networks use convolutional layers to capture temporal patterns and can be effective for time series data.\n\n1. **Attention Mechanisms:** Used in models like Transformers, attention mechanisms can help focus on relevant time steps, improving performance for long sequences. These models, which include mechanisms like attention, are effective at handling long-range dependencies and can process sequential data more efficiently than traditional RNNs. They can be used to analyze the raw time series data of each customer's transactions and behaviors.\n\n1. **Autoencoders:** These can be used for feature extraction from time series data, capturing important patterns and reducing dimensionality.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Exploratory Data Analysis - EDA</div>\n\n## <span style=\"color: #016FD0;\">Theoretical Background and Knowledge Tips 💡</span>\n\n- **What is a Probability Density Function (PDF)?:** \n    \n    A probability density function describes a probability distribution for a random, continuous variable. Use a probability density function to find the chances that the value of a random variable will occur within a range of values that you specify. More specifically, a PDF is a function where **its integral for an interval provides the probability of a value occurring in that interval.** For example, what are the chances that the next IQ score you measure will fall between 120 and 140? In statistics, PDF stands for probability density function.\n    \n    Unlike distributions for discrete random variables where specific values can have non-zero probabilities, the likelihood for a single value is always zero for a continuous variable. Consequently, the probability density function provides the chances of a value falling within a specified range for continuous variables. For a PDF in statistics, probability density refers to the likelihood of a value occurring within an interval length of one unit.\n    \n    In short, probability density functions can find non-zero likelihoods for a continuous random variable X falling within the interval [a, b]. Or, in statistical notation: P (A < X < B).\n    \n    If you need to find likelihoods for a **discrete variable**, use a **Probability Mass Function (PMF)** instead.\n    \n    Graphing a probability density function gives you a probability density plot. These graphs are great for understanding how a PDF in statistics calculates probabilities. The chart below displays the PDF for IQ scores, which is a probability density function of a normal distribution. On a probability density plot, the total area under the curve equals one, representing the total probability of 1 for the full range of possible values in the distribution. In other words, the likelihood of a value falling anywhere in the complete distribution curve is 1. Finding the chances for a smaller range of interest involves finding the portion of the area under the curve corresponding to that range. In our example, we need to find the percentage of the area that falls between 120 and 140.\n    <img width=\"750\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/518ae112-0dc3-4c96-a785-73032167f81d\">\n    <img width=\"752\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/f10a03ff-bc40-4e6a-ac28-cf2b575db93b\">\n    \n    In the example above, my statistical software finds that P IQ (120 < x < 140) = 0.08738. The probability of a value falling within the range of 120 to 140 is 0.08738. These graphs also illustrate why probability density functions find a zero likelihood for an individual value. Consider that the probability for a PDF in statistics equals an area. For a non-zero area, you must have both a non-zero height and a non-zero width because Height X Width = Area. In this context, the height is the curve’s height on the graph, while the width relates to the range of values. When you have a single value, you have zero width, which produces zero area. Hence, a zero chance!\n    \n    Have a look at this guide: [How to Identify the Distribution of Your Data](https://statisticsbyjim.com/hypothesis-testing/identify-distribution-data/), if you are interested in that topic. \n\n    While the normal distribution is the most common, it can fit only symmetrical distributions. Fitting skewed distributions requires using other probability density functions. There are a variety of other probability density functions that correspond with distributions of different shapes and properties. Each PDF has between 1-3 parameters that define its shape. Before using a PDF to find a probability, you must identify the correct function and parameter values for the population you are studying. To accomplish that, you’ll typically gather a random sample from that population and use a combination of parameter estimation techniques and hypothesis tests to identify the distribution and parameters that most likely produced your sample.\n    \n    A probability density function (PDF) describes how likely it is to observe some outcome resulting from a data-generating process. A PDF can tell us which values are most likely to appear versus the less likely outcomes. This will change depending on the shape and characteristics of the PDF. The probability density function is a measurement of how often investment returns fall within a specified range. PDFs are typically depicted on a graph, with a normal bell curve indicating neutral market risk, and a skewed curve at either end indicating greater or lesser risk-reward. Skewness is a shift of the taller portion of the curve to the right or left. If the curve is shifted to the left with a long tail on the right (right skew), analysts consider it to suggest there is a greater upside reward. If it is shifted to the right with a long tail to the left (left skew), analysts suggest that there is greater downside risk. The image below demonstrates normally distributed data with a bell curve. The data mean is the line in the middle, and the vertical lines are standard deviations, or how far data falls from the mean. The first two vertical lines on either side of the mean show that 68.5% of the data fall within +/-1 standard deviation from the mean. So, if this were a curve of normally distributed stock returns, you would see that 68.5% of the time, returns fall between the -1 SD and +1 SD lines and that market risk is neutral (there is no skew).\n    <img width=\"684\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/bf3ad0ab-5de7-4d04-980c-57db3312e9d5\">\n    \n    - A **right-skewed distribution** is longer on the right side of its peak than on its left. Right skew is also referred to as **positive skew**. You can think of skewness in terms of tails. A tail is a long, tapering end of a distribution. \n        - The peak (mode) is shifted to the left.\n        - There is a long tail on the right.\n        - Interpretation:\n            - Most returns are low or moderate.\n            - There is a small probability of very high returns (greater upside reward).\n            - Investors may face lower risk most of the time but have potential for high rewards.\n\n    - A **left-skewed distribution** is longer on the left side of its peak than on its right. In other words, a left-skewed distribution has a long tail on its left side. Left skew is also referred to as **negative skew.**\n        - The peak (mode) is shifted to the right.\n        - There is a long tail on the left.\n        - Interpretation:\n            - Most returns are moderate to high.\n            - There is a small probability of very low returns (greater downside risk).\n            - Investors may face higher risk of significant losses despite generally positive returns.\n    - Why Skewness Indicates Greater or Lesser Risk/Reward\n   \n        - **Greater Upside Reward (Right Skew)**\n            - Example: Emerging markets or tech startups.\n            - Potential for very high returns if things go well.\n            - Investment is attractive due to the possibility of significant gains, but typical returns might be low or moderate.\n            - Risk: Moderate day-to-day risk, with the hope of high rewards from outlier events.\n        - **Greater Downside Risk (Left Skew)**\n            - Example: Established companies with occasional large setbacks.\n            - Consistent moderate returns, but risk of significant losses due to rare adverse events.\n            - Risk: Regular returns are decent, but there's a higher probability of facing large losses.\n\n- **How to find the PDF of a Continuous Random Variable?** \n\n    If you have data and no prior information about the distribution, use empirical methods like **kernel density estimation (KDE)** to estimate the PDF. This involves using a smoothing kernel (like Gaussian) to estimate the PDF from a sample of data points. Kernel density estimation is a technique for estimation of probability density function that is a must-have enabling the user to better analyse the studied probability distribution than when using a traditional histogram. In statistics, kernel density estimation (KDE) is the application of kernel smoothing for probability density estimation, i.e., a non-parametric method to estimate the probability density function of a random variable based on kernels as weights. Kernel density estimation is a nonparametric model used for estimating probability distributions. Before diving too deeply into kernel density estimation, it is helpful to understand the concept of nonparametric estimation. **KDE is a non-parametric method used to estimate the PDF based on a finite data sample without assuming any underlying distribution. Unlike parametric methods, KDE does not assume a specific form for the distribution (like normal, exponential, etc.). It constructs the PDF directly from the data.**\n\n- **What is a KDE plot?:**\n\n    **Kernel Density Estimation (KDE)** is a non-parametric technique for visualizing the **probability density function** of a continuous random variable. Kernel Density Estimation (KDE) is a non-parametric way to estimate the probability density function (PDF) of a random variable. This statistical technique is used for smoothing data, particularly when the data is univariate or multivariate. Unlike parametric estimation methods, which assume a specific distribution shape for the data (such as normal distribution), KDE imposes no such assumption, making it a more flexible tool for understanding the underlying structure of the data.\n\n## <span style=\"color: #016FD0;\">Univariate Analysis</span>\n\n- While ‘uni’ means one, variate indicates a variable. Therefore, univariate analysis is a form of analysis that only involves a single variable.  In a practical setting, a univariate analysis means the analysis of a single variable (or column) in a dataset (data table). Univariate analysis explores each variable in a data set, separately. It looks at the range of values, as well as the central tendency of the values. It describes the pattern of response to the variable. It describes each variable on its own. Univariate analysis is a basic kind of analysis technique for statistical data. Here the data contains just one variable and does not have to deal with the relationship of a cause and effect. \n\n- Among all the forms of analytical methods that data analysts practice, univariate analysis is considered one of the basic forms of analysis. It is typically the first step to understanding a dataset. The idea of univariate analysis is to first understand the variables individually. Then, you move into analyzing two or more variables simultaneously.\n\n- A dataset (also known as a feature set or simply a table) is a multidimensional heterogeneous data structure. It is formed by combining multiple one-dimensional data structures that are homogeneous.\n\n- The univariate data is not categorized by data type but rather by the purpose they serve or their nature. In this sense, univariate data (i.e., a single column) can be divided into ID, Numerical, and Categorical. This classification is essential because different types of univariate analysis are required for each type.\n\nUnivariate data classifications are as follows:\n\n- **`ID`**: This data has no statistical or aggregative properties, and they are used to identify a subject uniquely. For example, the column ‘customer_ID’.\n\n- **`Numerical (Quantitative)`**: This data has statistical properties. They can be of two types- Discrete and Continuous. \n    - **`Discrete`**: This dataset has discrete values (i.e., cannot have decimals). For example- ‘No of Family Members.\n    - **`Continuous`**: This dataset can have numbers with decimals. For example- ‘Income’.\n\n- **`Categorical (Qualitative)`**: Categorical data deals with descriptions or categories. They have aggregative properties and are of two types- Ordinal and Nominal. (Note- categorical univariate data can have numeric datatype)\n    - **`Ordinal`**: These categories have an order. For example- ‘Designation’ where the order can be Manager, Sr Manager, CEO and cannot be any other.\n    - **`Nominal`**: These categories do not have any order. For example- ‘Location’ has mutually exclusive categories.\n    \nTypically, a univariate is data that belongs to any of the types mentioned above. Now, to analyze such a dataset, different types of univariate analysis techniques are used depending on the type of variable in question.\n\n***What is the use of Univariate analysis?***\n\nIt is used to:\n\n- Describe/summarize\n- Find patterns\n- Data preparation such as missing value imputation, outlier treatment, feature reduction, feature transformation, normalization, scaling, etc.\n\n***Types of Univariate Analysis***\n\n<img width=\"912\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/6c30ca3c-7b49-4dba-82f2-b6022cce632b\">\n\nThe primary purpose of univariate analysis is to describe data. Using different techniques, these descriptions are found. These techniques can be categorized into the following groups:\n\n1. Graphical\n1. Tables\n1. Descriptive statistics\n1. Inferential statistics (i.e., use of frequency distributions)\n\nEach of these techniques provides information about the data in a unique way. Typically, a data analyst uses more than one technique to form their opinion about the data they are dealing with, as this helps them make important decisions related to data preparation. Let’s understand each technique. \n\n\n1. ***Graphical analysis:*** Various types of graphs can be used to understand data. The standard type of graphs include-\n\n    - **Histograms:** A histogram displays the frequency of each value or group of values (bins) in numerical data. This helps in understanding how the values are distributed.\n    - **Boxplot:** A boxplot provides several important information such as minimum, maximum, median, 1st, and 3rd quartiles. It is beneficial in identifying **outliers** in the data.\n    - **Density Curve:** The density curve helps in understanding the shape of the data’s distribution. It helps answer questions such as if the data is bimodal, normally distributed, skewed, etc. \n    - **Bar Chart:** Bar Charts, mainly frequency bar charts, is a univariate chart used to find the frequency of the different categories of categorical data.\n    - **Pie Chart:** Frequency Pie charts convey similar information to bar charts. The difference is that they have a circular formation with each slice indicating the share of each category in the data.\n\n1. ***Univariate tables:*** Tables help in univariate analysis and are typically used with categorical data or numerical data with limited cardinality. Different types of tables include:\n\n    - **Frequency Tables:** Each unique value and its respective frequency in the data is shown through a table. Thus, it summarizes the frequency the way a histogram, frequency bar, or pie chart does but in a tabular manner.\n    - **Grouped Tables:** Rather than finding the count of each unique value, the values are binned or grouped, and the frequency of each group is reflected in the table. It is typically used for numerical data with high cardinality.\n    - **Percentage (Proportion) Tables:** Rather than showing the frequency of the unique values (or groups), such a table shows their proportion in the data (in percentage).\n    - **Cumulative Proportion Tables:** It is similar to the proportion table, with the difference being that the proportion is shown cumulatively. It is typically used with binned data having a distinct order (or with categorical ordinal data).\n    \n1. ***Univariate Statistics - Descriptive:*** Univariate analysis can be performed in a statistical setting. Two types of statistics can be used here- Descriptive and Inferential. As the name suggests, descriptive statistics are used to describe data. The statistics used here are commonly referred to as summary statistics. For instance, if you have to describe a cube, you have to ‘measure’ it. By measuring its length, breadth, and height, you can describe it. Similarly, these descriptive or univariate statistics have specific measures that help us in describing the data. These measures are-\n\n    - **Measure of Central Tendency:** Statistics such as mean, median, and mode are considered here. They help in summarizing all the data through a single central value.\n    - **Measure of Variability:** Analysts also need to understand how the data varies from the central point. To understand this, specific univariate statistics can be calculated, such as range, interquartile range, variance, standard deviation, etc.\n    - **Measure of Shape:** The shape of the data distribution can explain a great deal about the data as the shape can help in identifying the type of distribution followed by the data. Each of these distributions has specific properties that can be used to your advantage. By analyzing the shapes, you will know if the data is symmetrical, non-symmetrical, left or right-skewed, is suffering from positive or negative kurtosis, among other things. These descriptive statistics can be used for calculating things like missing value proportions, upper and lower limits for outliers, level of variance through the coefficient of variance, etc. \n\n1. ***Inferential Statistics:*** Often, the data you are dealing with is a subset (sample) of the complete data (population). Thus, the common question here is – Can the findings of the sample be extrapolated to the population? i.e., Is the sample representative of the population, or has the population changed? Such questions are answered using specific hypothesis tests designed to deal with such univariate data-based problems. ***Hypothesis tests*** help us answer crucial questions about the data and their relation with the population from where they are drawn. Several hypotheses or univariate testing mechanisms come in handy here, such as-\n\n    - **Z Test:** Used for numerical (quantitative) data where the sample size is greater than 30 and the population’s standard deviation is known.\n    - **One-Sample t-Test:** Used for numerical (quantitative) data where the sample size is less than 30 or the population’s standard deviation is unknown.\n    - **Chi-Square Test:** Used with ordinal categorical data\n    - **Kolmogorov-Smirnov Test:** Used with nominal categorical data\n \n     All such univariate testing methods generate p-values that can be used to accept or reject different types of hypotheses.\n\n### <span style=\"color: #016FD0;\">Numerical Features Analysis - Target Analysis</span>\n\n1. **Target Analysis:** Distribution by Target: Analyze the distribution of delinquency variables by the target variable. Analyzing the distribution of delinquency variables by the target variable is crucial for several reasons, especially in the context of credit default prediction:\n\n    1. **Understanding Feature-Target Relationships - Discriminatory Power:** This analysis helps determine whether the delinquency variables can effectively discriminate between the two classes (default and non-default). If the distributions differ significantly between the classes, it indicates that the variable is informative for predicting the target.\n    2. **Feature Engineering - New Insights:** It provides insights that can be used for feature engineering. For instance, if certain delinquency variables show distinct patterns for defaults, you can create new features or transformations to capture these patterns better.\n    3. **Modeling Insights - Feature Importance:** Understanding how features relate to the target variable can guide you in selecting which features to include in your model and how to preprocess them.\n    4. **Detecting Non-Linearity - Non-Linear Relationships:** Sometimes, the relationship between a feature and the target variable is non-linear. Visualizing the distributions can help you detect such patterns and decide whether non-linear transformations or non-linear models might be more appropriate.\n    5. **Handling Imbalances and Outliers - Class Imbalance:** This analysis can help you understand how well the feature values are represented across different classes, which is important in imbalanced datasets. It also helps in identifying outliers specific to each class.\n    6. **Improving Model Performance - Predictive Power:** Features that show a clear distinction between the target classes are likely to be strong predictors. Incorporating such features can enhance model performance and predictive accuracy.\n\nAnalyzing the distribution of the variables by the target variable, often called \"Target Analysis,\" is an important step in exploratory data analysis for several reasons. Here’s why this analysis is crucial and what insights it can provide:\n\n1. **Discriminatory Power - Separation of Classes:** If the distributions of a feature for different target values (e.g., 0 and 1) are well separated, it indicates that the feature has good discriminatory power. For example, if the KDE plots of a feature for target=0 and target=1 do not overlap much, it is a strong signal that this feature is useful. By examining how the distribution of a numerical feature changes with the target variable, you can identify which features are potentially important for predicting the target. If the distributions are different for different target values, it suggests that the feature might have predictive power. and **Feature Selection** features that show a strong relationship with the target variable are good candidates for inclusion in the predictive model. Conversely, features that show little to no differentiation by target may be less useful.\n\n1. **Distribution Patterns:** Pattern Recognition: By analyzing the shape, central tendency (mean, median), and spread (variance) of the distributions, you can identify patterns that are characteristic of different target classes. For instance, you might find that defaults (target=1) tend to have higher balances on average compared to non-defaults (target=0).\n\n1. **Outliers and Extremes:** Box plots can help identify outliers in the data. If outliers are more prevalent in one target class than another, it can give clues about the conditions that lead to the target event.\n\n1. **Feature Transformation:** Understanding the distribution can guide the need for transformations (e.g., log transformations, scaling) to make the data more suitable for modeling. For example, highly skewed distributions may benefit from a logarithmic transformation.\n\n### <span style=\"color: #016FD0;\">Understanding Histograms, Boxplot, and KDE Plots</span>\n\n#### <span style=\"color: #016FD0;\">KDE Plots</span>\n\nIn summary, a KDE plot is a way to draw a smooth curve over your data points to see where most of them are and how they are spread out. It's like turning your data into a nice, smooth hill to understand it better. A KDE plot (Kernel Density Estimate plot) is like a smooth, continuous line that helps you see where most of the data points are in your dataset. Instead of counting data points in bins like a histogram, it draws a smooth curve that shows where the data is concentrated.\n\n- How it Works:\n\n    - Kernel: Think of a kernel as a small bump or hill that you place on top of each marble. Each hill represents one piece of data.\n    - Adding the Hills: When you add up all these small hills, you get a big hill that shows where most of the marbles are. The higher the hill, the more marbles (data points) are in that area.\n    - Bandwidth: This is like how wide or narrow each hill is. If the hills are wide, the big hill will be smoother. If the hills are narrow, the big hill will have more peaks and valleys.\n\n- Why Use It: A KDE plot helps you see the shape of your data. It shows you where the data points are bunched up and where there are fewer data points. A KDE plot is like a histogram but with a continuous, smoothed curve. One of its key advantages, especially in cases with imbalanced classes, is that it can adjust the visibility of each class based on its proportion or ratio to the overall data. In other words, even if one class is significantly smaller than the other — say, in a 1:500 ratio — a KDE plot can still make that smaller class visible and comparable to the larger one. Our classes in now more comparable to each other using KDE plot. The larger the overlap, the less predictive power a feature possesses in telling between classes. Honestly, I use the KDE plot above to explain to my stakeholders regarding why certain features posses higher prediction power compared to others and why each feature was scored differently. Which eventually gets their buy-in for the model.\n\n## <span style=\"color: #016FD0;\">Resources</span>\n\n1. [Probability Density Function: Definition & Uses](https://statisticsbyjim.com/probability/probability-density-function/)\n1. [The Basics of Probability Density Function (PDF), With an Example\n](https://www.investopedia.com/terms/p/pdf.asp)\n1. [MUST SEE VIZ](https://mathisonian.github.io/kde/)\n1. [What is Kernel Density Estimation?\n](https://deepai.org/machine-learning-glossary-and-terms/kernel-density-estimation)\n1. [Kernel density estimation\n](https://en.wikipedia.org/wiki/Kernel_density_estimation#:~:text=In%20statistics%2C%20kernel%20density%20estimation,based%20on%20kernels%20as%20weights.)\n1. [The Fundamentals of Kernel Density Estimation\n](https://www.aptech.com/blog/the-fundamentals-of-kernel-density-estimation/)","metadata":{}},{"cell_type":"markdown","source":"Given that the dataset is large and contains multiple rows per customer, reducing the dataset to one row per customer for the initial EDA can help simplify the analysis and make it more manageable. However, it's important to ensure that the aggregation method used retains the essential characteristics of each customer's data. Here are a few approaches we could take:\n\n- **Mean/Median:** Calculate the mean or median of each feature for each customer.\n- **Max/Min:** Use the maximum or minimum value of each feature for each customer.\n- **Last Statement:** Use the values from the last statement date for each customer.\n- **Custom Aggregation:** Create new features based on a combination of the above methods or other statistical measures (e.g., standard deviation, sum).\n\n\n***Considerations:***\n\n1. **Preserving Temporal Information:** If you aggregate by taking the last statement, be mindful that you might lose important temporal trends. For modeling purposes, it's often beneficial to retain time-series data.\n\n1. **Feature Engineering:** Aggregating data can be a feature engineering opportunity. For instance, you might create features that capture trends over time (e.g., the difference between the first and last statement values).\n\n*For the purposes of this **Exploratory Data Analysis (EDA)**, we will focus on the **last statement for each customer**. This approach simplifies the dataset by reducing it to a single row per customer, thus making it more manageable for initial analysis. By leveraging the most recent data point, we aim to capture the most current snapshot of each customer's financial behavior and status.*\n","metadata":{}},{"cell_type":"code","source":"# Convert statement_date to datetime if it's not already\ntrain_df[\"S_2\"] = pd.to_datetime(train_df[\"S_2\"])\n\n# Sort the data by customer_ID and statement_date to ensure proper aggregation\ntrain_df = train_df.sort_values(by=[\"customer_ID\", \"S_2\"])\n\n# Aggregation: Using the last statement for each customer\neda_df = train_df.groupby(\"customer_ID\").last().reset_index()\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_cols = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68', \"D_87\"]\n# Remove variables that are not in the dataset (if necessary)\nd_variables = [col for col in eda_df.columns if (col.startswith(('D'))) & (col not in cat_cols)]\n\n\n# Calculate the percentage of null values for each feature\nnull_percentage = eda_df[d_variables].isnull().mean() * 100\n\n# Calculate the variance of each feature\nvariance = eda_df[d_variables].var()\n\n# Combine the results into a new DataFrame\nsummary_df = pd.DataFrame({\n    'Feature': eda_df[d_variables].columns,\n    'Null Percentage': null_percentage,\n    'Variance': variance\n}).reset_index(drop=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"eda_df[\"D_87\"].value_counts(normalize=True, dropna=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"summary_df.sort_values(by=[\"Null Percentage\"],ascending=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# Filter the DataFrame once\neda_df_default = eda_df[eda_df['target'] == 1]\neda_df_paid = eda_df[eda_df['target'] == 0]\n\n# Determine the grid size\nnum_vars = len(d_variables)\ncols = 4\nrows = (num_vars // cols) + (num_vars % cols > 0)\n\n\n# Create a subplot grid\nfig = make_subplots(rows=rows, cols=cols, subplot_titles=d_variables)\n\n# Define colors\ndefault_color = '#933A16'\npaid_color = '#016FD0'\n\n# Loop over each variable and create a KDE plot\nfor i, var in enumerate(d_variables):\n    d_default = eda_df_default[var].dropna()\n    d_paid = eda_df_paid[var].dropna()\n\n    kde_fig = ff.create_distplot(\n        [d_default, d_paid],\n        group_labels=['Default', 'Paid'],\n        show_hist=False,\n        show_rug=False,\n        colors=[default_color, paid_color]\n    )\n    \n    for trace in kde_fig['data']:\n        trace.showlegend = False  # Hide legend for all traces initially\n        fig.add_trace(trace, row=(i // cols) + 1, col=(i % cols) + 1)\n\n# Add one legend manually\nfig.add_trace(go.Scatter(x=[None], y=[None], mode='markers',\n                         marker=dict(size=10, color=default_color),\n                         legendgroup='Default', showlegend=True, name='Default'))\nfig.add_trace(go.Scatter(x=[None], y=[None], mode='markers',\n                         marker=dict(size=10, color=paid_color),\n                         legendgroup='Paid', showlegend=True, name='Paid'))\n\n# Update layout for the grid\nfig.update_layout(\n    title='KDE Plot of Multiple Delinquency Variables',\n    height=300*rows,  # Adjust height for readability\n    legend=dict(title='Target', itemsizing='constant'),\n    plot_bgcolor='white',\n    showlegend=True  # Show legend once\n)\n\n# Add axis titles manually for readability and remove grid lines\nfor i, var in enumerate(d_variables):\n    row = (i // cols) + 1\n    col = (i % cols) + 1\n    fig.update_xaxes(title_text=var, row=row, col=col, showgrid=False)\n    fig.update_yaxes(showgrid=False)\n    if row == rows:\n        fig.update_xaxes(title_text=var, row=row, col=col)\n    if col == 1:\n        fig.update_yaxes(title_text='Density', row=row, col=col)\n\n# Show the plot\nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del d_default\ndel d_paid\ndel eda_df","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Dealing with Imbalanced Classes</div>\n\n## <span style=\"color: #016FD0;\">Introduction</span>\n\n- **Imbalanced Data:** Imbalanced data refers to datasets where the target class has an uneven distribution of observations, i.e., one class label has a very high number of observations, and the other has a deficient number of observations. Class imbalance is generally normal in classification problems. But in some cases, this imbalance is quite acute, where the majority class’s presence is much higher than that of the minority class. **Data often demonstrates skewed distributions with a long tail. However, most of the machine learning algorithms currently in use were designed around the assumption of a uniform distribution over each target category (classification).** Imbalanced datasets are those where there is a severe skew in the class distribution, such as 1:100 or 1:1000 examples in the minority class to the majority class.\n\nIn real-life applications, we face many challenges where we only have uneven data representations in which the minority class is usually the more important one and hence we require methods to improve its recognition rates. This issue poses a serious challenge to predictive modeling because learning algorithms will be biased towards the majority class. \n\nImportant day-to-day tasks in our lives such as preventing malicious attacks, detecting life-threatening diseases, or handling rare cases in monitoring systems face extreme class imbalance with ratios ranging from 1:1000 up to 1:5000 and one must design intelligent systems that can adjust and overcome such extreme bias. As any seasoned data scientist or statistician will be aware of, datasets are rarely distributed evenly across attributes of interest. Let’s imagine we are tasked with discovering fraudulent credit card transactions — naturally, the vast majority of these transactions will be legitimate, and only a very small proportion will be fraudulent. Similarly, if we are testing individuals for cancer, or for the presence of a virus (COVID-19 included), the positive rate will (hopefully) be only a small fraction of those tested. More examples include:\n\n- An e-commerce company predicting which users will buy items on their platform\n- A manufacturing company analyzing produced materials for defects\n- Spam email filtering trying to differentiation ‘ham’ from ‘spam’\n- Intrusion detection systems examining network traffic for malware signatures or atypical port activity\n- Companies predicting churn rates amongst their customers\n- Number of clients who closed a specific account in a bank or financial organization\n- Prediction of telecommunications equipment failures\n- Detection of oil spills from satellite images\n- Insurance risk modeling\n- Hardware fault detection\n\nOne has usually much fewer datapoints from the adverse class. This is unfortunate as we care a lot about avoiding misclassifying elements of this class.\n\nIn actual fact, it is pretty rare to have perfectly balanced data in classification tasks. Oftentimes the items we are interested in analyzing are inherently ‘rare’ events for the very reason that they are rare and hence difficult to predict. This presents a curious problem for aspiring data scientists since many data science programs do not properly address how to handle imbalanced datasets given their prevalence in industry.\n\nMachine learning algorithms by default assume that data is balanced. In classification, this corresponds to a comparative number of instances of each class. Classifiers learn better from a balanced distribution. It is up to the data scientist to correct for imbalances, which can be done in multiple ways.\n\n*Different Types of Imbalance:*\n\n1. Between-Class\n1. Within-Class\n1. Further Complication of Imbalance\n\nThere are a couple more difficulties increased by imbalanced datasets. Firstly, we have class overlapping. This is not always a problem, but can often arise in imbalanced learning problems and cause headaches. Class overlapping is illustrated in the below dataset.\n\n<img width=\"919\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/62f590d0-ae95-49cd-87f9-f3e4f2f972ca\">\n\nClass overlapping occurs in normal classification problems, so what is the additional issue here? Well, the class more represented in overlap regions tends to be better classified by methods based on global learning (on the full dataset). This is because the algorithm is able to get a more informed picture of the data distribution of the majority class.\n\nIn contrast, the class less represented in such regions tends to be better classified by local methods. If we take k-NN as an example, as the value of k increases, it becomes increasingly global and increasingly local. It can be shown that performance for low values of k has better performance on the minority dataset, and lower performance at high values of k. This shift in accuracy is not exhibited for the majority class because it is well-represented at all points.\n\nThis suggests that local methods may be better suited for studying the minority class. One method to correct for this is the CBO Method. The CBO Method uses cluster-based resampling to identify ‘rare’ cases and resample them individually, so as to avoid the creation of small disjuncts in the learned hypothesis. This is a method of oversampling — a topic that we will discuss in detail in the following section.\n\n<img width=\"877\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/175e421c-791a-4fac-9bfe-07e12272fbca\">\n\n### <span style=\"color: #016FD0;\">Why is Imbalanced Data a Problem?</span>\n\nIf we explain it simply, the main problem with imbalanced dataset prediction is how accurately we predict both majority and minority classes. Let’s start with an example of disease diagnosis. Now, we will predict disease from an existing dataset where, for every 100 records, only five patients are diagnosed. So, the majority class is 95% with no disease, and the minority class is only 5% with the disease. Now, assume our model predicts that all 100 out of 100 patients have no disease.\n\nSometimes, when the records of a particular class are much more than those of another class, our classifier may get biased towards the prediction. In this case, the confusion matrix for the classification problem shows how well our model classifies the target classes, and we arrive at the model’s model’s accuracy from the confusion matrix. It is calculated based on the model’s total number of correct predictions divided by the total number of predictions. In the above case, it is (0+95)/(0+95+0+5)=0.95 or 95%. This means that the model fails to identify the minority class, yet the accuracy score of the model will be 95%.\n\nThus, our traditional approach to classifying and calculating model accuracy is ineffective in the case of an imbalanced dataset.\n\n\nImbalanced dataset is a problem because it can lead to biased models and inaccurate predictions. Here’s why:\n\n\n1. **Skewed Class Distribution:** Imbalanced dataset occurs when one class (the minority class) is significantly underrepresented compared to another class (the majority class) in a classification problem. This can **skew the model’s learning process because it may prioritize the majority class, leading to poor performance on the minority class.**\n\n1. **Biased Model Training:** Machine learning models aim to minimize errors, often measured by metrics like accuracy. In imbalanced datasets, a model can achieve high accuracy by simply predicting the majority class for all instances, ignoring the minority class completely. As a result, **the model is biased towards the majority class and fails to capture patterns in the minority class accurately.**\n\n1. **Poor Generalization:** Imbalanced data can result in models that generalize poorly to new, unseen data, especially for the minority class. Since the model hasn’t learned enough about the minority class due to its scarcity in the training data, it may struggle to make accurate predictions for instances belonging to that class in real-world scenarios.\n\n1. **Costly Errors:** ***In many real-world applications, misclassifying instances from the minority class can be more costly or have higher consequences than misclassifying instances from the majority class.*** Imbalanced data exacerbates this issue because the model tends to make more errors on the minority class, potentially leading to significant negative impacts.\n\n1. **Evaluation Metrics Misleading:** Traditional evaluation metrics like **accuracy can be misleading in imbalanced datasets. For instance, a model achieving high accuracy may perform poorly on the minority class, which is often the class of interest.** Using metrics like \n    - precision, \n    - recall, \n    - F1-score, or \n    - area under the ROC curve (AUC-ROC) can **provide a more nuanced understanding of the model’s performance across different classes.**\n\n\n## <span style=\"color: #016FD0;\">Techniques to Handle Imbalanced Data Set Problem</span>\n\nIn rare cases like fraud detection or disease prediction, **it is vital to identify the minority classes correctly.** So, the model should not be biased to detect only the majority class but should give equal weight or importance to the minority class, too. Here, I discuss some techniques to handle imbalanced dataset problem. **There is no correct or wrong method; different techniques work well with other problems.**\n\n\n<img width=\"847\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/c2917909-22ff-4bb7-94d2-7087e5489d94\">\n\nThere are several approaches to solving class imbalance problem before starting classification, such as:\n\n- More samples from the minority class(es) should be acquired from the knowledge domain.\n\n- Changing the loss function to give the failing minority class a higher cost.\n\n- Oversampling the minority class.\n\n- Undersampling the majority class.\n\n- Any combination of previous approaches.\n\nOversampling methods duplicate or create new synthetic examples in the minority class, whereas undersampling methods delete or merge examples in the majority class. Oversampling is the most often used approach (but not necessarily the best).\n\n### <span style=\"color: #016FD0;\">1. Choose Proper Evaluation Metric</span>\n\nThe first technique to handle imbalanced data is choosing a proper evaluation metric. The accuracy of a classifier is the total number of correct predictions divided by the total number of predictions. This may be good enough for a well-balanced class but not ideal for an imbalanced class problem. However, accuracy is not appropriate when the data is imbalanced. Because the model can achieve higher accuracy by just predicting accurately the majority class while performing poorly on the minority class which in most cases is the class we care about the most. Other metrics, such as precision, measure how accurate the classifier’s prediction of a specific class, and recall measures the classifier’s ability to identify a class.\n\nClassification accuracy is a metric that summarizes the performance of a classification model as the number of correct predictions divided by the total number of predictions.\n\nAccuracy = Correct Predictions / Total Predictions\n\nAchieving 90 percent classification accuracy, or even 99 percent classification accuracy, may be trivial on an imbalanced classification problem. Consider the case of an imbalanced dataset with a 1:100 class imbalance. Blind guess will give us a 99% accuracy score (by betting on majority class).\n\nThe rule of thumb is: accuracy never helps in imbalanced dataset.\n\nThe most common metrics to use for imbalanced datasets are:\n\n- F1 score\n- Precision\n- Recall\n- AUC score (AUC ROC)\n- Average precision score (AP)\n- G-Mean\n\n**It is good practice to track multiple metrics when developing a machine learning model as each highlights different aspects of model performance.**\n\n\n**Have a look in the Evaluation Metrics - Explained section!!!**\n\n### <span style=\"color: #016FD0;\">2. Resampling (Oversampling and Undersampling)</span>\n\nThe second technique used to handle the imbalanced data is used to upsample or downsample the minority or majority class. When we are using an imbalanced dataset, we can oversample the minority class using replacement. This technique used to handle imbalanced data is called oversampling. Similarly, we can randomly delete rows from the majority class to match them with the minority class which is called undersampling. After sampling the data we can get a balanced dataset for both majority and minority classes. So, when both classes have a similar number of records present in the dataset, we can assume that the classifier will give equal importance to both classes.\n\nIt has been observed that our target class is imbalanced. So, we’ll upsample the data so that the minority class matches the majority class.\n\nYou can balance your data by resampling them. The followings are two different techniques for resampling:\n\n- Upsampling (increase your minority class)\n- Downsample (decrease your majority class)\n\nFor both of these, we will use the Sklearn Resample function. **`Sklearn.utils resample`** can be used for both undersamplings the majority class and oversample minority class instances.\n\n<img width=\"939\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/f1367a21-133c-450f-85ed-64eabde59e5d\">\n\n<img width=\"960\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/8e32d2e2-7e09-4d52-8b19-c629b51c3ad6\">\n\nUndersampling performs poorly compared to oversampling when it comes to identifying the majority class. But besides that, it identifies the minority class better than oversampling. Is undersampling a good idea? \n\n**Advantages:**\n\n- Data scientists can balance the dataset and reduce the risk of their analysis or machine learning algorithm skewing toward the majority. Because without resampling, scientists might come up with what is known as the accuracy paradox where they run a classification model with 90% accuracy. On closer inspection, though, they will find the results are heavily within the majority class. \n\n- Fewer storage requirements and better run times for analyses. Less data means you or your business needs less storage and time to gain valuable insights. \n\n**Disadvantages:**\n\n- Removing enough majority examples to make the majority class the same or similar size to the minority class results in a significant loss of data.\n\n- The sample of the majority class chosen could be biased, meaning, it might not accurately represent the real world, and the result of the analysis may be inaccurate. Therefore, it can cause the classifier to perform poorly on real unseen data.\n\n- **Undersampling** is recommended by many statistical researchers but **is only good if enough data points are available on the undersampled class.** Also, since the majority class will end up with the same number of points as the minority class, the statistical properties of the distributions will become ‘looser’ in a sense. However, we have not artificially distorted the data distribution with this method by adding in artificial data points.\n\n**Because of these disadvantages, some scientists might prefer oversampling. It doesn’t lead to any loss of information, and in some cases, may perform better than undersampling. But oversampling isn’t perfect either. Because oversampling often involves replicating minority events, it can lead to overfitting.** “The combination of SMOTE and under-sampling performs better than plain under-sampling.” To balance these issues, certain scenarios might require a combination of both over and undersampling to obtain the most lifelike dataset and accurate results. \n\n### <span style=\"color: #016FD0;\">3. SMOTE</span>\n\nThe third technique to handle imbalanced data is the **Synthetic Minority Oversampling Technique or SMOTE,** which is another technique to oversample the minority class. Simply adding duplicate records of minority class often don’t adon’ty new information to the model. In SMOTE new instances are synthesized from the existing data. If we explain it in simple words, SMOTE looks into minority class instances and use **k nearest neighbor** to select a random nearest neighbor, and a synthetic instance is created randomly in feature space.\n\n<img width=\"931\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/c1c3f36f-1ec6-4949-8d4c-aad5867ac35a\">\n\nIn this approach, we synthesize new examples from the minority class. \n\nThere are several methods available to oversample a dataset used in a typical classification problem. But the most common **data augmentation technique** is known as Synthetic Minority Oversampling Technique or SMOTE for short. As the name suggests, SMOTE creates “synthetic” examples rather than over-sampling with replacement. Specifically, SMOTE works the following way. It starts by randomly selecting a minority class example and finding its k nearest minority class neighbors at random. Then a synthetic example is created at a randomly selected point in the line that connects two examples in feature space.\n\n<img width=\"990\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/0c7df933-74e2-463a-b264-7ce4d79508a1\">\n<img width=\"871\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/89cfbb38-88a5-4113-bc5c-59da7d5dbf26\">\n\nHow do we generate these samples? The most common way is to generate points that are close in dataspace proximity to existing samples or are ‘between’ two samples, as illustrated below. How does SMOTE work? SMOTE generates new samples in between existing data points based on their local density and their borders with the other class. Not only does it perform oversampling, but can subsequently use cleaning techniques (undersampling, more on this shortly) to remove redundancy in the end. Below is an illustration for how SMOTE works when studying class data.\n\n<img width=\"1509\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/6cf748ae-a73f-4f14-8ec6-dadb994ed6b8\">\n\nThe algorithm for SMOTE is as follows. For each minority sample:\n\n– Find its k-nearest minority neighbours\n\n– Randomly select j of these neighbours\n\n– Randomly generate synthetic samples along the lines joining the minority sample and its j selected neighbours (j depends on the amount of oversampling desired)\n\n\n**The created synthetic examples from SMOTE for the minority class when added to the training set, balance the class distributions and cause the classifier to create larger and less specific decision regions helping the classifier generalize better and mitigate overfitting, rather than smaller and more specific regions which will cause the model to overfit to the majority class.**\n\n<img width=\"949\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/89915424-09cf-417d-81aa-a0dab7be6097\">\n\n\nThis approach is inspired by data augmentation techniques that proved successful in handwritten character recognition where operations like rotation and skew were natural ways to perturb the training data. \n\n**Advantages:**\n\n- It improves the overfitting caused by random oversampling as synthetic examples are generated rather than a copy of existing examples. As you may have suspected, there are some downsides to adding false data points. **Firstly, you risk overfitting, especially if one does this for points that are noise — you end up exacerbating this noise by adding reinforced measurements. In addition, adding these values randomly can also contribute additional noise to our model.** Using random oversampling (with replacement) of the minority class has the effect of making the decision region for the minority class very specific. In a decision tree, it would cause a new split and often lead to overfitting. SMOTE’s informed oversampling generalizes the decision region for the minority class. As a result, larger and less specific regions are learned, thus, paying attention to minority class samples without causing overfitting.\n\n- No loss of information.\n\n- It’s simple.\n\n**Disadvantages:**\n\n- While generating synthetic examples, SMOTE does not take into consideration neighboring examples that can be from other classes. This can increase the overlapping of classes and can introduce additional noise.\n\n- SMOTE is not very practical for high-dimensional data.\n\n- **Overgeneralization.** SMOTE’s procedure can be dangerous since it blindly generalizes the minority area without regard to the majority class. This strategy is particularly problematic in the case of highly skewed class distributions since, in such cases, the minority class is very sparse with respect to the majority class, thus resulting in a greater chance of class mixture.\n\n- **Another potential issue is that SMOTE might introduce the artificial minority class examples too deeply in the majority class space.** This drawback can be resolved by hybridization: combining SMOTE with undersampling algorithms. One of the most famous of these is Tomek Links. Tomek Links are pairs of instances of opposite classes who are their own nearest neighbors. In other words, they are pairs of opposing instances that are very close together. Tomek’s algorithm looks for such pairs and removes the majority instance of the pair. The idea is to clarify the border between the minority and majority classes, making the minority region(s) more distinct. Scikit-learn has no built-in modules for doing this, though there are some independent packages (e.g., TomekLink, imbalanced-learn). Thus, Tomek’s algorithm is an undersampling technique that acts as a data cleaning method for SMOTE to regulate against redundancy. As you may have suspected, there are many additional undersampling techniques that can be combined with SMOTE to perform the same function. An additional example is Edited Nearest Neighbors (ENN). ENN removes any example whose class label differs from the class of at least two of their neighbor. ENN removes more examples than the Tomek links does and also can remove examples from both classes. Other more nuanced versions of SMOTE include Borderline SMOTE, SVMSMOTE, and KMeansSMOTE, and more nuanced versions of the undersampling techniques applied in concert with SMOTE are Condensed Nearest Neighbor (CNN), Repeated Edited Nearest Neighbor, and Instance Hardness Threshold.\n\n### <span style=\"color: #016FD0;\">4. Threshold Moving</span>\n\n**In the case of our classifiers, many times classifiers actually predict the probability of class membership. We assign those prediction’s abilities to a certain class based on a threshold which is usually 0.5, i.e. if the probabilities < 0.5 it belongs to a certain class, and if not it belongs to the other class.**\n\nFor imbalanced class problems, this default threshold may not work properly. We need to change the threshold to the optimum value so that it can efficiently separate two classes. Also, we can use **ROC Curves** and **Precision-Recall Curves** to find the optimal threshold for the classifier. We can also use a grid search method or search within a set of values to identify the optimal value.\n\nIn this method first, we will find the probabilities for the class label, then we’ll find the optimum threshold to map the probabilities to its proper class label. The probability of prediction can be obtained from a classifier by using predict_proba() method from sklearn.\n\n<img width=\"910\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/21048a3f-b91e-402c-9acc-e2f4eb074396\">\n\nHere, we get the optimal threshold in 0.3 instead of our default 0.5.\n\n### <span style=\"color: #016FD0;\">5. BalancedBaggingClassifier</span>\n\nWhen we try to use a usual classifier to classify an imbalanced dataset, the model favors the majority class due to its larger volume presence. A BalancedBaggingClassifier is the same as a sklearn classifier but with additional balancing. It includes an additional step to balance the training set at the time of fit for a given sampler. This classifier takes two special parameters, “sampling_strategy” and “replacement”. The sampling_strategy decides the type of resampling required (e.g., ‘majority’ – resample only the majority class, ‘all’ – resample all classes, etc.), and replacement decides whether it is going to be a sample with replacement or not.\n\nAn illustrative example is given below:\n\n<img width=\"962\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/1a3df3f4-f8eb-4e75-8fd2-ef0c6ccfc936\">\n\n### <span style=\"color: #016FD0;\">6. Algorithm approach - Cost-Sensitive Learning</span>\n\nThis approach concentrates on modifying existing models to alleviate their bias towards the majority groups. This requires good insight into the modified learning algorithm and precise identification of reasons for its failure in learning the representations of skewed distributions. \n\nThe most popular techniques are cost-sensitive approaches (**weighted learners**). Here, the given model is modified to incorporate varying penalties for each considered group of examples. **In other words, we use Focal loss where we assign a higher weight to the minority class in our cost function which will penalize the model for misclassifying the minority class while at the same time reducing the weight of the majority class, causing the model to pay more attention to the underrepresented class.** Thus, boosting its importance during the learning process.\n\nAnother interesting algorithm-level solution is to apply one-class learning or one-class classification(OCC for short) that focuses on the target group, creating a data description. This way we eliminate bias towards any group, as we concentrate only on a single set of objects.\n\nOCC can be useful in imbalanced classification problems because it provides techniques for outlier and anomaly detection. It does this by fitting the model on the majority class data and predicting whether new data belong to the majority class or belong to the minority class meaning it’s an outlier/anomaly. \n\nOCC problems usually are practical classification tasks where majority class data is easily available but minority class is hard, expensive, and even impossible to gather, i.e. work of an engine, fraudulent transactions, intrusion detection for the computer system, and so on.\n\n**Weighted Loss**\n\nOne other way to avoid having class imbalance is to **weight the losses** differently. To choose the weights, you first need to calculate the class frequencies. Changing the weights in the loss function allows data scientists to balance the contribution of each class. One way of doing this is by multiplying each example from each class by a class-specific weight factor so that the overall contribution of each class is the same. Thanks to the Sklearn, there is a built-in parameter called class_weight in most of the ML algorithms which helps you to balance the contribution of each class. For example, in the `Sklearn’s RandomForestClassifier` you have options as `balanced`, `balanced_subsample` or you can give weights for each class manually as a dictionary.\n\nWe have discussed sampling techniques and are now ready to discuss cost-sensitive learning. In many ways, the two approaches are analogous — the main difference being that in cost-sensitive learning we perform under- and over-sampling by altering the relative weighting of individual samples.\n\n- **Upweighting.** Upweighting is analogous to over-sampling and works by increasing the weight of one of the classes keeping the weight of the other class at one.\n\n- **Down-weighting.** Down-weighting is analogous to under-sampling and works by decreasing the weight of one of the classes keeping the weight of the other class at one.\n\nAn example of how this can be performed using sklearn is via the sklearn.utils.class_weight function and applied to any sklearn classifier (and within keras).\n\n```python\n\nfrom sklearn.utils import class_weight\nclass_weights = class_weight.compute_class_weight('balanced', np.unique(y_train), y_train)\nmodel.fit(X_train, y_train, class_weight=class_weights)\n\n```\n\nIn this case, we have set the instances to be ‘balanced’, meaning that we will treat these instances to have balanced weighting based on their relative number of points — this is what I would recommend unless you have a good reason for setting the values yourself. If you have three classes and wanted to weight one of them 10x larger and another 20x larger (because there are 10x and 20x fewer of these points in the dataset than the majority class), then we can rewrite this as:\n\n\n```\nclass_weight = {0: 0.1,\n                1: 1.,\n                2: 2.}\n                \n```\n\n```python\n\n#Fitting RandomForestClassifier\nfrom sklearn.ensemble import RandomForestClassifier\nmodel = RandomForestClassifier(n_estimators=1000, class_weight={0:0.10, 1:0.90})\n\n\n```\nAmazing🥳, now you learned a very common technique that you can apply almost every kind of data with every ML algorithms.\n\n### <span style=\"color: #016FD0;\">7. Data Augmentation</span>\n\nFrom now onwards, the techniques that I’m going to mention are suitable mostly for Computer Vision and images. There are ways to generate data when you are working with images.\n\nData Augmentation is known as a very powerful technique used to artificially create variations in existing images to expand an existing image data set. This creates new and different images from the existing image data set that represents a comprehensive set of possible images. Here are several ways that you can do to increase your dataset.\n\n1. Flipping\n1. Zoom\n1. Rotating\n1. Shearing\n1. Cropping\n1. And lastly generating new images with **Generative Adversarial Networks**.\n\nNote: Sometimes using these techniques can harm your model more than helping it. Let’s assume the same Lung X-Ray dataset case where we have many different X-rays. If you flip the images you will have a heart in the right which can be seen as a cist. Again you killed people =( So be careful about using these augmentation techniques.\n\n### <span style=\"color: #016FD0;\">8. Transfer Learning</span>\n\nIf you don’t have enough data to generalize your model you can use pre-trained models and tune them for your own task. This is called Transfer Learning and it is very common in the Computer Vision field. Since the training, huge models take so much time and resources, and people are getting pre-trained models like DenseNet and use them. The approach is as follows:\n\n- Find a model which suits your requirements\n- Re-use the model. This may involve using all or parts of the model or the first layers of the neural network. Since the complexity of features increases with the number of layers, **the first layers always have more generic features that can be used successfully.**\n- **Fine-tuning:** *By going deeper, the neural network gets specialized on the task so you may want to tune the last couple of layers for your own task which is called Fine-tuning.*\n\n\n## <span style=\"color: #016FD0;\">Conclusion</span>\n\nIn conclusion, dealing with imbalanced datasets in classification problems poses significant challenges that traditional approaches often fail to address effectively. The skewed distribution of classes can lead to biased models, inaccurate predictions, and poor generalization of new data. Moreover, the misleading nature of traditional evaluation metrics like accuracy exacerbates these issues, making adopting alternative metrics such as precision, recall, F1-score, or AUC-ROC crucial.\n\nTo overcome these challenges, various techniques can be employed, including proper selection of evaluation metrics, resampling methods like oversampling and undersampling, utilizing algorithms designed for imbalance such as SMOTE, employing ensemble methods like BalancedBaggingClassifier, and adjusting threshold values for optimal classification. Each technique offers unique advantages and may be more suitable depending on the specific characteristics of the dataset and the problem at hand. We know also how to increase our model’s performance with:\n\n- Using Sampling Techniques\n- changing the Class Weights of Loss Function\n- Using Data Augmentation Techniques\n- and using Transfer Learning (The last two of them are more specific to Computer Vision.)\n\n## <span style=\"color: #016FD0;\">Resources</span>\n\n1. [5 Techniques to Handle Imbalanced Data For a Classification Problem\n](https://www.analyticsvidhya.com/blog/2021/06/5-techniques-to-handle-imbalanced-data-for-a-classification-problem/#:~:text=When%20we%20are%20using%20an,class%20which%20is%20called%20undersampling.)\n1. [How to Deal With Imbalanced Classification and Regression Data\n](https://neptune.ai/blog/how-to-deal-with-imbalanced-classification-and-regression-data)\n1. [How to Handle Imbalance Data and Small Training Sets in ML\n](https://towardsdatascience.com/how-to-handle-imbalance-data-and-small-training-sets-in-ml-989f8053531d)\n\n\n\n\n\n1. https://www.tensorflow.org/tutorials/structured_data/imbalanced_data\n1. https://www.kaggle.com/code/tboyle10/methods-for-dealing-with-imbalanced-data\n1. https://semaphoreci.com/blog/imbalanced-data-machine-learning-python\n1. https://www.shiksha.com/online-courses/articles/handle-imbalanced-data-in-a-classification-problem/\n1. https://www.kaggle.com/code/marcinrutecki/best-techniques-and-metrics-for-imbalanced-dataset\n1. https://towardsdatascience.com/the-5-most-useful-techniques-to-handle-imbalanced-datasets-6cdba096d55a\n1. https://datasciencedojo.com/blog/techniques-to-handle-imbalanced-data/\n1. https://machinelearningmastery.com/what-is-imbalanced-classification/\n1. https://encord.com/blog/an-introduction-to-balanced-and-imbalanced-datasets-in-machine-learning/\n1. https://python.plainenglish.io/handling-class-imbalance-in-machine-learning-cb1473e825ce\n1. https://www.blog.trainindata.com/machine-learning-with-imbalanced-data/\n1. https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-022-01821-w#:~:text=Many%20studies%20have%20demonstrated%20that,and%20more%20comprehensive%20strong%20model.\n1. https://thecontentfarm.net/ensemble-techniques-for-handling-class-imbalance/\n1. https://moldstud.com/articles/p-machine-learning-engineering-challenges-in-handling-imbalanced-datasets","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Evaluation Metrics - Explained</div>\n\nFor binary classification problems, the confusion matrix defines the base for performance measures. Most of the performance metrics are derived from the confusion matrix, i.e., accuracy, misclassification rate, precision, and recall. A confusion matrix is the most basic form of assessment of a binary classifier. Given the prediction outputs of our classifier and the true response variable, a confusion matrix tells us how many of our predictions are correct for each class, and how many are incorrect. The confusion matrix provides a simple visualization of the performance of a classifier based on these factors.\n\n<img width=\"942\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/9cfc5719-6ee3-41ea-9ede-8f7ca3e44d0e\">\n\nArguably the most important five metrics for binary classification are: \n\n- (1) precision, \n- (2) recall, \n- (3) F1 score, \n- (4) accuracy, and\n- (5) specificity.\n\n#### <span style=\"color: #016FD0;\">Precision</span>\n\nPrecision provides us with the answer to the question ***“Of all my positive predictions, what proportion of them are correct?”.*** If you have an algorithm that predicts all of the positive class correctly but also has a large portion of false positives, the precision will be small. It makes sense why this is called precision since it is a measure of how ‘precise’ our predictions are.\n\n<img width=\"1018\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/4da36806-7225-40bc-ad22-0c3fe7eade3d\">\n\n#### <span style=\"color: #016FD0;\">Recall</span>\n\nRecall provides us with the answer to a different question ***“Of all of the positive samples, what proportion did I predict correctly?”***. Instead of false positives, we are now interested in false negatives. These are items that our algorithm missed, and are often the most egregious errors (e.g. failing to diagnose something with cancer that actually has cancer, failing to discover malware when it is present, or failing to spot a defective item). The name ‘recall’ also makes sense for this circumstance as we are seeing how many of the samples the algorithm was able to pick up on.\n\n<img width=\"1061\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/aaa97456-7b57-4d72-a99a-0c9b53f60bce\">\n\n\nFor an imbalanced class dataset, the F1 score is a more appropriate metric. It is the harmonic mean of precision and recall and the expression is –\n\n<img width=\"794\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/d16022ef-62c7-4a84-946b-7f8ecbd34758\">\n\nSo, if the classifier predicts the minority class but the prediction is erroneous and the **false-positive increases, the precision metric will be low, and so will the F1 score.** Also, if the classifier identifies the minority class poorly, i.e., more of this class wrongfully predicted as the majority class, then **false negatives will increase, so recall and F1 score will be low.** *The F1 score only increases if the number and prediction quality improve.* To balance the recall and precision, i.e., improving recall, while keeping precision low, the F-score is proposed as a harmonic mean of the precision and recall. Since the F-score weights, precision, and recall equally and balances both concerns, it is less likely to be biased to the majority or minority class.\n\n**F1 score keeps the balance between precision and recall and improves the score only if the classifier identifies more of a certain class correctly.**\n\n\nTo accommodate the minority class, the Receiver Operating Characteristic (ROC) curve is proposed as a measure over a range of tradeoffs between the True Positive (TP) Rate and False Positive (FP) Rate. Another important performance measure is Area Under the Curve (AUC) is a commonly used performance metric for summarizing the ROC curve in a single score. Moreover, AUC is not biased towards the model’s performance on either the majority or minority class, which makes this measure more appropriate when dealing with imbalanced data.\n\n<img width=\"869\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/42b78388-8eca-4d05-8d9b-d42cffa575aa\">\n\nThis section is focused on how to handle these kinds of situations and I am not going into details of how to measure your model but bear in mind to change the measurement metrics from Accuracy to Precision, Recall, F1 Score. But even they are not actually that great. I would recommend going for:AUC Curves, Average Precision while dealing with imbalance data. Here is an amazing github repository which goes into very details of these metrics by talking about a particular imbalance dataset example: [credit card frauds](https://fraud-detection-handbook.github.io/fraud-detection-handbook/Chapter_4_PerformanceMetrics/Assessment_RealWorldData.html).\n\n\n## <span style=\"color: #016FD0;\">Resources</span>\n\n\n1. https://towardsdatascience.com/guide-to-classification-on-imbalanced-datasets-d6653aa5fa23","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Feature Engineering</div>\n\nFor traditional approaches like decision trees, ensemble methods, and neural networks, effective feature engineering is crucial. Here are some suggested feature engineering techniques. We start by creating aggregated features for all the customers:\nsince the data contains categorical, numerical features. We process them separately.\n\n***Creating meaningful features to capture temporal patterns and trends.***\n\n1. **Lag Features:** Creating lagged variables to represent previous values of a feature.\n\n1. **Rolling Statistics:** Using moving averages, sums, or other statistics over a rolling window.\n\n1. **Difference Features:** Calculating the difference between consecutive time steps to capture trends.\n\n\nFor traditional machine learning approaches like XGBoost and neural networks (NNs), it's generally important to aggregate features per customer rather than using the raw data directly. Here are the key reasons why:\n\n1. **Model Expectations and Input Format - Single Row per Instance:** Traditional machine learning models typically expect each instance (i.e., each customer) to be represented as a single row in the dataset. This means all relevant features for a customer need to be combined into one row. **Fixed Input Size:** Models like XGBoost and feedforward neural networks require a fixed input size. Raw time-series data results in multiple rows per customer, leading to variable input sizes, which these models cannot handle directly.\n\n2. **Handling Temporal Dependencies - Temporal Dynamics:** While models like LSTM or GRU are designed to handle temporal sequences directly, traditional models are not. Aggregating features allows you to capture temporal dynamics in a form that these models can process.\n\n3. **Reduction of Noise and Dimensionality - Noise Reduction:** Aggregating features (e.g., calculating mean, median, std) can help reduce noise in the data, making patterns more apparent. **Dimensionality Reduction:** Raw time-series data can be high-dimensional, leading to issues like overfitting. Aggregation reduces the number of features to a manageable size.\n\n4. **Feature Enrichment - Derived Insights:** Aggregated features often provide derived insights that are more informative than raw data. For example, the trend (difference features), variability (standard deviation), and overall level (mean, max) are more meaningful for prediction than raw monthly balances.\n\n\n## <span style=\"color: #016FD0;\">Aggregated Statistical Features on Raw Data (numerical)</span>\n\nSummarize the time series data into meaningful statistics that capture the overall trends and behaviors.\n\n\n1. **Mean, Median, Standard Deviation:** Compute these statistics for each feature across the 1–13 statements per customer to summarize central tendencies and variability. Central tendency refers to the measure that represents the center or typical value of a dataset. It provides a summary statistic that reflects the most common or average characteristics of the data. The central tendency is often used to describe the overall behavior of the data and to compare different datasets.\n\n1. **Min and Max Values:** Identify the minimum and maximum values for each feature to capture the range of behavior over the period.\n\n1. **Sum and Count:** Aggregate total amounts (e.g., total spending, total payments) and count occurrences (e.g., number of delinquent months).\n\n1. **Last values of the features**, i.e, the last status of the customer","metadata":{}},{"cell_type":"code","source":"all_cols = [c for c in list(train_df.columns) if c not in [\"customer_ID\",\"S_2\", \"target\"]]\n\ncat_features = [\"B_30\",\"B_38\",\"D_114\",\"D_116\",\"D_117\",\"D_120\",\"D_126\",\"D_63\",\"D_64\",\"D_66\",\"D_68\"]\nnum_features = [col for col in all_cols if col not in cat_features]","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:51:37.956847Z","iopub.execute_input":"2024-06-09T10:51:37.957280Z","iopub.status.idle":"2024-06-09T10:51:37.963752Z","shell.execute_reply.started":"2024-06-09T10:51:37.957247Z","shell.execute_reply":"2024-06-09T10:51:37.962521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def compute_aggregated_features(df: pd.DataFrame, num_features: List[str]) -> pd.DataFrame:\n    \"\"\"\n    Compute aggregated features for numerical columns in the dataset.\n    \n    Args:\n        df (pd.DataFrame): Input dataframe containing customer data with numerical features.\n        num_features (List[str]): List of numerical feature column names to compute aggregations.\n    \n    Returns:\n        pd.DataFrame: DataFrame with aggregated features for each customer.\n    \"\"\"\n    # Group by customer_ID and compute the aggregated features\n    aggregated_df = df.groupby(\"customer_ID\")[num_features].agg(\n        [\"mean\", \"median\", \"std\", \"min\", \"max\", \"sum\", \"count\", \"last\", \"first\"]\n    )\n    \n    # Flatten the MultiIndex columns\n    aggregated_df.columns = [\"_\".join(x) for x in aggregated_df.columns]\n    aggregated_df = aggregated_df.reset_index()\n    \n    return aggregated_df","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:51:38.527220Z","iopub.execute_input":"2024-06-09T10:51:38.528109Z","iopub.status.idle":"2024-06-09T10:51:38.535507Z","shell.execute_reply.started":"2024-06-09T10:51:38.528071Z","shell.execute_reply":"2024-06-09T10:51:38.534047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"code","source":"agg_train_df = compute_aggregated_features(train_df, num_features)\nagg_train_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:51:39.690532Z","iopub.execute_input":"2024-06-09T10:51:39.690956Z","iopub.status.idle":"2024-06-09T10:54:58.302739Z","shell.execute_reply.started":"2024-06-09T10:51:39.690914Z","shell.execute_reply":"2024-06-09T10:54:58.301650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"agg_train_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:54:58.304667Z","iopub.execute_input":"2024-06-09T10:54:58.305038Z","iopub.status.idle":"2024-06-09T10:54:58.312023Z","shell.execute_reply.started":"2024-06-09T10:54:58.304995Z","shell.execute_reply":"2024-06-09T10:54:58.310944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color: #016FD0;\">Difference Features - Aggregated Statistical Features on Diff Data (numerical)</span>\n\nDifference features involve calculating the difference between the values of a feature at consecutive time steps. This helps in capturing the changes or trends over time, which can be crucial for understanding the dynamics of the data and improving the performance of predictive models.\n\n**Importance of Difference Features:**\n\n1. **Capture Trends and Patterns:** By calculating the differences between consecutive time points, you can identify trends and patterns in the data that might be indicative of certain behaviors, such as an increasing balance, decreasing payments, or escalating delinquency.\n\n1. **Highlight Temporal Changes:** Difference features can reveal sudden changes or anomalies in the data, such as a significant drop in payments or a spike in spending, which might be important for predicting events like defaults.\n\n1. **Reduce Non-Stationarity:** Many machine learning models assume that the data is stationary (i.e., its statistical properties do not change over time). By using difference features, you can sometimes transform a non-stationary time series into a stationary one, making it more suitable for modeling.\n\n1. **Simplify the Data:** Instead of using the raw values, which might be affected by seasonal effects or other external factors, the differences can simplify the data and focus on the relative changes over time.\n\n\nTo pass the difference features (or any other features) to a model like XGBoost, you typically need to aggregate these features at the customer level. This is because XGBoost and similar models expect a single row per instance (customer, in this case).\n\nHere's how you can aggregate the difference features and other features to create a single row per customer:\n\nStep-by-Step Approach:\n\n1. **Calculate Difference Features**\n\n1. **Aggregate Features:** Aggregate the difference features (and possibly the raw features) using statistical measures (e.g., mean, std, sum, min, max).\n\n1. **Prepare the Final Dataset:** Combine these aggregated features into a single DataFrame, where each row represents a customer.\n","metadata":{}},{"cell_type":"code","source":"# Sort the DataFrame by customer_ID and statement_date\ntrain_df = train_df.sort_values(by=[\"customer_ID\", \"S_2\"])","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:54:58.313318Z","iopub.execute_input":"2024-06-09T10:54:58.313703Z","iopub.status.idle":"2024-06-09T10:55:02.014367Z","shell.execute_reply.started":"2024-06-09T10:54:58.313667Z","shell.execute_reply":"2024-06-09T10:55:02.013110Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def compute_diff_agg_features(df: pd.DataFrame, num_features: List[str]) -> pd.DataFrame:\n    \"\"\"\n    Compute difference and aggregate features for numerical columns in the dataset.\n    \n    Args:\n        df (pd.DataFrame): Input dataframe containing customer data with numerical features.\n        num_features (List[str]): List of numerical feature column names to compute differences and aggregations.\n\n    Returns:\n        pd.DataFrame: DataFrame with aggregated difference features for each customer.\n    \"\"\"\n    # Create column names for the difference features\n    diff_num_features = [f\"diff_{col}\" for col in num_features]\n    \n    # Compute the difference between consecutive time steps for numerical features\n    diff_df = df.groupby(\"customer_ID\")[num_features].diff().add_prefix(\"diff_\")\n    \n    # Combine diff_df with customer_ID\n    diff_df = pd.concat([df['customer_ID'], diff_df], axis=1)\n    \n    # Aggregate the difference features by customer_ID\n    num_agg_diff_df = diff_df.groupby(\"customer_ID\")[diff_num_features].agg(\n        [\"mean\", \"median\", \"std\", \"min\", \"max\", \"sum\", \"count\"]\n    )\n    \n    # Flatten the MultiIndex columns\n    num_agg_diff_df.columns = ['_'.join(x) for x in num_agg_diff_df.columns]\n    \n    # Reset the index to make customer_ID a column again\n    num_agg_diff_df = num_agg_diff_df.reset_index()\n    \n    return num_agg_diff_df","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:55:02.017735Z","iopub.execute_input":"2024-06-09T10:55:02.018248Z","iopub.status.idle":"2024-06-09T10:55:02.026574Z","shell.execute_reply.started":"2024-06-09T10:55:02.018203Z","shell.execute_reply":"2024-06-09T10:55:02.025121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result_df = compute_diff_agg_features(train_df, num_features)","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:55:02.028115Z","iopub.execute_input":"2024-06-09T10:55:02.028477Z","iopub.status.idle":"2024-06-09T10:58:08.883212Z","shell.execute_reply.started":"2024-06-09T10:55:02.028446Z","shell.execute_reply":"2024-06-09T10:58:08.881963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result_df","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:08.884746Z","iopub.execute_input":"2024-06-09T10:58:08.885225Z","iopub.status.idle":"2024-06-09T10:58:09.015692Z","shell.execute_reply.started":"2024-06-09T10:58:08.885182Z","shell.execute_reply":"2024-06-09T10:58:09.014337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color: #016FD0;\">Other Feature Engineering Ideas</span>\n\n\n### Time-Based Features\n\n1. **Lag Features:** Include the values of features from previous statements (e.g., previous month’s balance, previous three months’ payment).\n\n1. **Rolling/Moving Averages:** Compute moving averages (e.g., 3-month moving average of spending) to smooth out short-term fluctuations and highlight longer-term trends.\n\n1. **Cumulative Sums:** Calculate cumulative sums (e.g., cumulative spending) to track running totals over time.\n\n### Trend and Seasonality Features\n\n1. **Trend Features:** Fit linear regression models to the time series data for each feature to extract trend slopes (e.g., increasing or decreasing spending patterns).\n\n1. **Seasonal Indicators:** Create features to capture seasonal effects (e.g., month of the year, quarter).\n\n### Interaction Features:\n\n1. **Multiplicative Features:** Generate features that capture interactions between variables (e.g., balance * delinquency).\n1. **Ratios:** Create ratios between related features (e.g., payment-to-balance ratio, spending-to-income ratio).\n\n### Domain-Specific Features:\n\n1. **Credit Utilization:** Calculate the proportion of credit used relative to the credit limit if available.\n\n1. **Delinquency Patterns:** Count the number of times delinquency variables indicate late payments within the last 6 or 12 months.\n\n### Time Since Last Event:\n\n1. **Time Since Last Default:** Days since the last default event, if applicable.\n\n1. **Time Since Last Payment:** Days since the last payment to capture recency of activity.\n\n### Categorical Feature Encoding:\n\n1. **One-Hot Encoding:** Convert categorical features into binary columns (e.g., encoding ‘B_30’, ‘D_114’).\n\n1. **Frequency Encoding:** Replace each categorical value with the frequency of its occurrence in the dataset.\n\n\n### Target Encoding\n\n1. **Target Mean Encoding:** Encode categorical variables based on the mean of the target variable (e.g., the default rate for each category).","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"color: #016FD0;\">Lag Features (numerical)</span>\n\nEach ML algorithm expects data as input that must be formatted in a specific way, and so time series datasets generally require some cleaning and feature engineering processes before they can generate useful insights. Time series datasets may have values that are missing or may contain outliers, hence the essential need for the data preparation and cleaning phase. Good time series data preparation produces clean and well-curated data, which leads to more practical and accurate predictions.\n\nData preparation is the practice of transforming raw data so that data scientists can run it through ML algorithms to discover insights and, eventually, make predictions. Additionally, because time series data has a temporal property, only specific statistical methodologies are appropriate for data prepared in this way. In this article, I walk you through the most important steps to prepare your time series data for forecasting models.\n\nIn working with time series, data scientists must construct the output of their model by identifying the variable that they need to predict at a future date (e.g., future number of sales next Monday) and then leverage historical data and feature engineering to create input variables that are used to make predictions for that future date. Feature engineering efforts mainly have two goals:\n\n1. **Creating the correct input dataset to feed the ML algorithm:** In this case, the purpose of feature engineering in time series forecasting is to create input features from historical row data and shape the dataset as a supervised learning problem.\n\n1. **Increasing the performance of ML models:** The second most important goal of feature engineering is about generating valid relationships between input features and the output feature or target variable to be predicted. In this way, the performance of ML models can be improved.\n\n\nLag features refer to the values of previous time steps in the time series. They can help capture the autocorrelation present in the data, which is the relationship between the current value and its past values.\n\nAdding lag features can improve a model's performance by allowing it to learn patterns from the past to predict future values. You can create lag features by shifting the original time series data by a specific number of periods, often called the \"lag order.\"\n\nLag features are values at prior timesteps that are considered useful because they are created on the assumption that what happened in the past can influence or contain a sort of intrinsic information about the future. For example, it can be beneficial to generate features for sales that happened in previous days at 4:00 p.m. if you want to predict similar sales at 4:00 p.m. the next day.\n\n\n<img width=\"682\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/ad56927d-b252-4887-b5f0-c66d6452f798\">\n\nUsing lagged features is a common technique in time-series analysis and machine-learning applications. By shifting the values of a variable backward or forward in time by a certain number of time periods, lagged features can capture temporal dependencies and trends in the data, providing valuable insights and improving the accuracy of predictive models.\n\nIn particular, lagged features are useful in predicting future values of a variable, as they can help identify patterns and relationships between variables over time. They are often used in fields such as finance, economics, and weather forecasting, where accurate predictions of future values are critical.\n\nSeveral methods exist for generating lagged features, including shifting the values of a variable by a fixed number of time periods or by using a rolling window of previous values. The choice of method will depend on the specific application and the data characteristics. \n\n### How to create lagged features?\n\nTo create lagged features, you need to have a time series data set that has a regular and consistent frequency, such as hourly, daily, weekly, or monthly. You also need to decide how many lags you want to use, and how to align them with your target variable. For example, if you want to predict the sales of tomorrow, you can use the sales of today and the sales of yesterday as lagged features. However, if you want to predict the sales of next month, you may need to use the sales of the previous months as lagged features, and align them with the target month. You can create lagged features using various tools and methods, such as pandas in Python, dplyr in R, or Excel formulas. The general idea is to shift the data by the desired number of time steps, and join it with the original data by the time index.\n\n### How to select lagged features?\n\nNot all lagged features are equally useful or relevant for your time series analysis. Some lagged features may have a strong correlation with your target variable, while others may have a weak or negative correlation. Some lagged features may capture important temporal dependencies, while others may introduce noise or multicollinearity. Therefore, you need to select the lagged features that have the most predictive power and the least redundancy. You can use various techniques and criteria to select lagged features, such as correlation analysis, mutual information, feature importance, or cross-validation. You can also use domain knowledge and intuition to select the lagged features that make sense for your problem and data.\n\n### How to use lagged features in models?\n\nOnce you have created and selected the lagged features, you can use them in various types of models for time series analysis, such as linear regression, logistic regression, decision trees, random forests, neural networks, or deep learning. You can treat the lagged features as independent variables, and the target variable as the dependent variable. You can also use the lagged features to create interaction terms, polynomial terms, or nonlinear transformations. You can use the models to perform tasks such as forecasting, classification, anomaly detection, or clustering. You can evaluate the models using metrics such as mean absolute error, mean squared error, accuracy, precision, recall, or F1-score.\n\n***Steps to Create and Use Lagged Features:***\n\n1. **Create Lagged Features:** Generate new columns in your dataset that represent the values of the features at previous time steps (lags).\n\n1. **Aggregate Lagged Features:** If needed, aggregate the lagged features to create a single row per customer.\n\n1. **Train Traditional Models:** Use the dataset with lagged features to train traditional machine learning models like XGBoost or feedforward neural networks.\n\n### What are the benefits and challenges of lagged features?\n\nThe use of lagged features in time series analysis can provide several advantages, such as improving the accuracy and performance of models by capturing temporal dependencies and patterns in data, or enhancing the interpretability and explainability of models by revealing causal relationships and effects of past values on current values. Additionally, lagged features can reduce the need for external sources as they use information already available in data. However, using lagged features can also pose some challenges, such as increasing the complexity and dimensionality of data which requires more computational resources and processing time; or introducing potential issues such as overfitting, underfitting, or multicollinearity which may affect the generalization and stability of models. Furthermore, careful selection and validation of lagged features is necessary for successful implementation, depending on the characteristics and context of data and problem.\n\n\n\n\n### Most creative ways to use \"Lag\" features by Kagglers\n\nOn this competition we get information about clients of AMEX over time. Most high scoring notebooks on this competiion focused on aggregating the information per client and create a single row of extracted features: One for each client.\n\nOne of such agg function is last.\n\nQuick examination revealed that the last feature is extreamly powerful at predicting if the client defaults or not (well.. make sense..). So I took this two steps further:\n\n1. **First feature:** Just like the last feature: I added a first feature.\n\n1. **\"Lag\" fearures:** to capture the change over time about each client I calculated two features for every first, last pair:\n    -  **Last - First:** The change since we first see the client to the last time we see the client.\n    - **Last / First:** The fractional difference since we first see the client to the last time we see the client.\n\n\n***Benefits:***\n\n1. **Temporal Dependencies:** Lagged features help traditional models capture temporal dependencies and trends.\n\n1. **Model Flexibility:** This approach allows you to use powerful traditional models without needing sequence-based models.\n\n1. **Enhanced Predictive Power:** Including lagged features often improves the model's ability to predict future outcomes by leveraging past information.\n","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"color: #016FD0;\">Resources</span>\n\n\n1. [https://hackernoon.com/must-know-base-tips-for-feature-engineering-with-time-series-data](https://hackernoon.com/must-know-base-tips-for-feature-engineering-with-time-series-data)\n1. [Introduction to feature engineering for time series forecasting\n](https://medium.com/data-science-at-microsoft/introduction-to-feature-engineering-for-time-series-forecasting-620aa55fcab0)\n1. [Practical Guide for Feature Engineering of Time Series Data\n](https://dotdata.com/blog/practical-guide-for-feature-engineering-of-time-series-data/)\n1. [How can you use lagged features to capture temporal dependencies in time series data?\n](https://www.linkedin.com/advice/0/how-can-you-use-lagged-features-capture-temporal-ks4kc#:~:text=Lagged%20features%20involve%20shifting%20original,day%20in%20the%20previous%20year.)","metadata":{}},{"cell_type":"code","source":"def calculate_change_features(lag_df: pd.DataFrame, num_features: List[str]) -> pd.DataFrame:\n    \"\"\"\n    Calculate change (last - first) and fractional change (last / first) features.\n    \n    Args:\n        lag_df (pd.DataFrame): DataFrame containing 'first_' and 'last_' features for each numerical column.\n        num_features (List[str]): List of numerical feature column names to create change features for.\n    \n    Returns:\n        pd.DataFrame: DataFrame with added change and fractional change features.\n    \"\"\"\n    for feature in num_features:\n        lag_df[f'{feature}_change'] = lag_df[f'{feature}_last'] - lag_df[f'{feature}_first']\n        lag_df[f'{feature}_frac_change'] = lag_df[f'{feature}_last'] / (lag_df[f'{feature}_first'] + 1e-10)  # Avoid division by zero\n    \n    return lag_df","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:09.017484Z","iopub.execute_input":"2024-06-09T10:58:09.017817Z","iopub.status.idle":"2024-06-09T10:58:09.024501Z","shell.execute_reply.started":"2024-06-09T10:58:09.017790Z","shell.execute_reply":"2024-06-09T10:58:09.023162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"agg_train_df = calculate_change_features(agg_train_df, num_features)","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:09.025921Z","iopub.execute_input":"2024-06-09T10:58:09.026301Z","iopub.status.idle":"2024-06-09T10:58:12.587541Z","shell.execute_reply.started":"2024-06-09T10:58:09.026270Z","shell.execute_reply":"2024-06-09T10:58:12.586316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"agg_train_df","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:12.589042Z","iopub.execute_input":"2024-06-09T10:58:12.589429Z","iopub.status.idle":"2024-06-09T10:58:12.737558Z","shell.execute_reply.started":"2024-06-09T10:58:12.589396Z","shell.execute_reply":"2024-06-09T10:58:12.736255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"color: #016FD0;\">Categorical Feature Encoding</span>\n\n1. **One-Hot Encoding:** Convert categorical features into binary columns (e.g., encoding ‘B_30’, ‘D_114’).\n\n1. **Frequency Encoding:** Replace each categorical value with the frequency of its occurrence in the dataset.\n\n1. Compute mean, standard deviation (std), sum, and last value for each one-hot encoded categorical feature across all data.\n","metadata":{}},{"cell_type":"code","source":"def one_hot_and_aggregate(df: pd.DataFrame, cat_features: List[str]) -> pd.DataFrame:\n    \"\"\"\n    Perform One-Hot Encoding on categorical features and calculate mean, std, sum, and last value\n    for each one-hot encoded feature grouped by customer_ID.\n    \n    Args:\n        df (pd.DataFrame): Input dataframe containing categorical features.\n        cat_features (List[str]): List of categorical feature column names to be one-hot encoded and aggregated.\n    \n    Returns:\n        pd.DataFrame: DataFrame with aggregated one-hot encoded features.\n    \"\"\"\n    # Perform One-Hot Encoding\n    df_one_hot = pd.get_dummies(df, columns=cat_features, drop_first=False)\n    \n    # List of the new one-hot encoded feature columns\n    one_hot_features = [col for col in df_one_hot.columns if any(cat in col for cat in cat_features)]\n    \n    # Group by customer_ID and compute the aggregated features\n    aggregated_df = df_one_hot.groupby(\"customer_ID\")[one_hot_features].agg(\n        [\"mean\", \"std\", \"sum\", \"last\"]\n    )\n    \n    # Flatten the MultiIndex columns\n    aggregated_df.columns = [\"_\".join(x) for x in aggregated_df.columns]\n    aggregated_df = aggregated_df.reset_index()\n    \n    return aggregated_df","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:12.741321Z","iopub.execute_input":"2024-06-09T10:58:12.741713Z","iopub.status.idle":"2024-06-09T10:58:12.749438Z","shell.execute_reply.started":"2024-06-09T10:58:12.741678Z","shell.execute_reply":"2024-06-09T10:58:12.748339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"aggregated_one_hot_df = one_hot_and_aggregate(train_df, cat_features)","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:12.750543Z","iopub.execute_input":"2024-06-09T10:58:12.750861Z","iopub.status.idle":"2024-06-09T10:58:32.881966Z","shell.execute_reply.started":"2024-06-09T10:58:12.750834Z","shell.execute_reply":"2024-06-09T10:58:32.880400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"aggregated_one_hot_df","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:32.883508Z","iopub.execute_input":"2024-06-09T10:58:32.883866Z","iopub.status.idle":"2024-06-09T10:58:33.013978Z","shell.execute_reply.started":"2024-06-09T10:58:32.883833Z","shell.execute_reply":"2024-06-09T10:58:33.012648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_agg_df =  train_df.groupby(\"customer_ID\")[cat_features].agg([\"count\", \"last\", \"nunique\"])\ncat_agg_df.columns = [\"_\".join(x) for x in cat_agg_df.columns]\ncat_agg_df = cat_agg_df.reset_index()\ncat_agg_df","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:33.015551Z","iopub.execute_input":"2024-06-09T10:58:33.016027Z","iopub.status.idle":"2024-06-09T10:58:37.609425Z","shell.execute_reply.started":"2024-06-09T10:58:33.015985Z","shell.execute_reply":"2024-06-09T10:58:37.608288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_features","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:37.611228Z","iopub.execute_input":"2024-06-09T10:58:37.611546Z","iopub.status.idle":"2024-06-09T10:58:37.618486Z","shell.execute_reply.started":"2024-06-09T10:58:37.611520Z","shell.execute_reply":"2024-06-09T10:58:37.617265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[\n    train_df[\"customer_ID\"].str.startswith(\"ffff518bb2075e4816ee3fe9f3b152c57fc0e6f01bf7fd\")\n][\"D_117\"].value_counts(dropna=False)","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:37.619932Z","iopub.execute_input":"2024-06-09T10:58:37.620310Z","iopub.status.idle":"2024-06-09T10:58:39.309326Z","shell.execute_reply.started":"2024-06-09T10:58:37.620281Z","shell.execute_reply":"2024-06-09T10:58:39.307978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vis_col = [col for col in aggregated_one_hot_df.columns if col.startswith(\"D_117_3.\")]\n\naggregated_one_hot_df[\n    aggregated_one_hot_df[\"customer_ID\"].str.startswith(\"ffff518bb2075e4816ee3fe9f3b152c57fc0e6f01bf7fd\")\n][vis_col]","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:39.310941Z","iopub.execute_input":"2024-06-09T10:58:39.311988Z","iopub.status.idle":"2024-06-09T10:58:39.479345Z","shell.execute_reply.started":"2024-06-09T10:58:39.311942Z","shell.execute_reply":"2024-06-09T10:58:39.478230Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vis_col = [col for col in cat_agg_df.columns if col.startswith(\"D_117\")]\ncat_agg_df[\n    cat_agg_df[\"customer_ID\"].str.startswith(\"ffff518bb2075e4816ee3fe9f3b152c57fc0e6f01bf7fd\")\n][vis_col]","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:39.480826Z","iopub.execute_input":"2024-06-09T10:58:39.481312Z","iopub.status.idle":"2024-06-09T10:58:39.644854Z","shell.execute_reply.started":"2024-06-09T10:58:39.481271Z","shell.execute_reply":"2024-06-09T10:58:39.643759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"agg_train_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:39.646099Z","iopub.execute_input":"2024-06-09T10:58:39.646433Z","iopub.status.idle":"2024-06-09T10:58:39.654580Z","shell.execute_reply.started":"2024-06-09T10:58:39.646407Z","shell.execute_reply":"2024-06-09T10:58:39.653385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:39.655962Z","iopub.execute_input":"2024-06-09T10:58:39.656355Z","iopub.status.idle":"2024-06-09T10:58:39.664110Z","shell.execute_reply.started":"2024-06-09T10:58:39.656322Z","shell.execute_reply":"2024-06-09T10:58:39.663056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"aggregated_one_hot_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:39.665344Z","iopub.execute_input":"2024-06-09T10:58:39.665687Z","iopub.status.idle":"2024-06-09T10:58:39.675096Z","shell.execute_reply.started":"2024-06-09T10:58:39.665660Z","shell.execute_reply":"2024-06-09T10:58:39.673934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_agg_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:39.676437Z","iopub.execute_input":"2024-06-09T10:58:39.676794Z","iopub.status.idle":"2024-06-09T10:58:39.687292Z","shell.execute_reply.started":"2024-06-09T10:58:39.676765Z","shell.execute_reply":"2024-06-09T10:58:39.686198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_df","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:39.688630Z","iopub.execute_input":"2024-06-09T10:58:39.689025Z","iopub.status.idle":"2024-06-09T10:58:39.704582Z","shell.execute_reply.started":"2024-06-09T10:58:39.688994Z","shell.execute_reply":"2024-06-09T10:58:39.703093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"features_train_df = pd.merge(agg_train_df, result_df, on=\"customer_ID\", how='left')\ndel agg_train_df\ndel result_df\n\nfeatures_train_df = pd.merge(features_train_df, aggregated_one_hot_df, on=\"customer_ID\", how='left')\ndel aggregated_one_hot_df\n\n\nfeatures_train_df = pd.merge(features_train_df, cat_agg_df, on=\"customer_ID\", how='left')\ndel cat_agg_df","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:58:39.706484Z","iopub.execute_input":"2024-06-09T10:58:39.706853Z","iopub.status.idle":"2024-06-09T10:59:04.983401Z","shell.execute_reply.started":"2024-06-09T10:58:39.706823Z","shell.execute_reply":"2024-06-09T10:59:04.982215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"features_train_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-06-09T10:59:04.984962Z","iopub.execute_input":"2024-06-09T10:59:04.986004Z","iopub.status.idle":"2024-06-09T10:59:04.992805Z","shell.execute_reply.started":"2024-06-09T10:59:04.985959Z","shell.execute_reply":"2024-06-09T10:59:04.991723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Decision Trees</div>\n\n## <span style=\"color: #016FD0;\">Introduction</span>\n\nDecision trees are arguably the most easily interpretable ML algorithms you can find and, if used in combination with the right techniques, can be quite powerful.\n\nA decision tree has this name because of its visual shape, which looks like a tree, with a root and many nodes and leaves. Imagine you take a list of the Titanic’s survivors, with some information such as their age and gender, and a binary variable telling who survived the disaster and who didn’t. You now want to create a classification model, to predict who will survive, based on this data. A very simple one would look like this:\n\nAs you can see, decision trees are just a sequence of simple decision rules that, combined, produce a prediction of the desired variable.\n\n<img width=\"909\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/1309ee0f-daea-4766-830b-da414dbbe797\">\n\nDecision trees create a model that predicts the label by evaluating a tree of if-then-else true/false feature questions, and estimating the minimum number of questions needed to assess the probability of making a correct decision. Decision trees can be used for classification to predict a category, or regression to predict a continuous numeric value. In the simple example below, a decision tree is used to estimate a house price (the label) based on the size and number of bedrooms (the features).\n\n<img width=\"934\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/ecae0814-542c-4a95-a56f-da937b628607\">\n\n## <span style=\"color: #016FD0;\">What is a decision tree?</span>\n\n**A decision tree is a non-parametric supervised learning algorithm, which is utilized for both classification and regression tasks.** It has a hierarchical, tree structure, which consists of a root node, branches, internal nodes and leaf nodes.\n\nAs you can see from the diagram below, a decision tree starts with a root node, which does not have any incoming branches. The outgoing branches from the root node then feed into the internal nodes, also known as decision nodes. Based on the available features, both node types conduct evaluations to **form homogenous subsets**, which are **denoted by leaf nodes**, or **terminal nodes**. **The leaf nodes represent all the possible outcomes within the dataset.**\n\n<img width=\"786\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/05dde1ea-035e-478d-89f4-a7ff00a1c4ac\">\n\nAs an example, let’s imagine that you were trying to assess whether or not you should go surf, you may use the following decision rules to make a choice:\n\n<img width=\"903\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/1ff3f519-5a25-48cc-af59-a1859b114c4c\">\n\nThis type of flowchart structure also **creates an easy to digest representation of decision-making,** allowing different groups across an organization to **better understand why a decision was made.**\n\n\nDecision tree learning employs a divide and conquer strategy by conducting a greedy search to identify the optimal split points within a tree. This process of splitting is then repeated in a top-down, recursive manner until all, or the majority of records have been classified under specific class labels. Whether or not all data points are classified as homogenous sets is largely dependent on the complexity of the decision tree. **Smaller trees are more easily able to attain pure leaf nodes—i.e. data points in a single class.** However, as a tree grows in size, it becomes increasingly difficult to maintain this purity, and it usually results in too little data falling within a given subtree. When this occurs, it is known as data fragmentation, and it can often lead to overfitting. As a result, decision trees have preference for small trees, which is consistent with the principle of parsimony in Occam’s Razor; that is, “entities should not be multiplied beyond necessity.” Said differently, decision trees should add complexity only if necessary, as the simplest explanation is often the best. **To reduce complexity and prevent overfitting, pruning is usually employed;** this is a process, which removes branches that split on features with low importance. The model’s fit can then be evaluated through the process of **cross-validation.** **Another way that decision trees can maintain their accuracy is by forming an ensemble via a random forest algorithm;** this classifier predicts more accurate results, particularly when the individual trees are uncorrelated with each other.\n\n\n## <span style=\"color: #016FD0;\">Types of Decision Trees</span>\n\nHunt’s algorithm, which was developed in the 1960s to model human learning in Psychology, forms the foundation of many popular decision tree algorithms, such as the following: \n\n- ID3: Ross Quinlan is credited within the development of ID3, which is shorthand for “Iterative Dichotomiser 3.” This algorithm leverages **entropy** and **information gain** as metrics to **evaluate candidate splits.**\n\n- C4.5: This algorithm is considered a later iteration of ID3, which was also developed by Quinlan. It can use information gain or gain ratios to evaluate split points within the decision trees. \n\n- **CART**: The term, CART, is an abbreviation for **“classification and regression trees”** and was introduced by Leo Breiman. This algorithm typically utilizes Gini impurity to **identify the ideal attribute to split on.** **Gini** impurity measures how often a randomly chosen attribute is misclassified. When evaluating using Gini impurity, a lower value is more ideal. \n\n\n## <span style=\"color: #016FD0;\">Classification And Regression Trees (CART)</span>\n\nA **decision tree** is a supervised machine learning algorithm used for predictive modeling of a dependent variable (target) based on the input of several independent variables. It has a tree-like structure with the root at the top. **CART** which stands for **C**lassification **A**nd **R**egression **T**rees is used as **an umbrella term to refer to the following types of decision trees:**\n\n- **Classification Trees:** where the target variable is fixed or categorical, this algorithm is used to identify the class/category within which the target would most likely fall.\n\n- **Regression Trees:** where the target variable is continuous and the tree/algorithm is used to predict its value, e.g. predicting the weather.\n\n## <span style=\"color: #016FD0;\">How to choose the best attribute at each node</span>\n\nWhile there are multiple ways to select the best attribute at each node, two methods, \n\n- **information gain** and\n- **Gini impurity,**\n\nact as popular splitting criterion for decision tree models. ***They help to evaluate the quality of each test condition and how well it will be able to classify samples into a class.***  \n\n\n### Entropy\n\nIt’s difficult to explain information gain without first discussing entropy. **Entropy** is a concept that stems from information theory, which **measures the impurity of the sample values.** It is defined with by the following formula, where: \n\n<img width=\"903\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/3f2c42d9-aa22-4f65-83f4-9d4821848396\">\n\n\n- S represents the data set that entropy is calculated \n- c represents the classes in set, S\n- p(c) represents the proportion of data points that belong to class c to the number of total data points in set, S\n\n- **Entropy values can fall between 0 and 1.** \n\n- *If all samples in data set, S, belong to one class, then entropy will equal zero. If half of the samples are classified as one class and the other half are in another class, entropy will be at its highest at 1.* \n\n- In order to select the best feature to split on and find the optimal decision tree, **the attribute with the smallest amount of entropy should be used.**\n\n\n### Information gain\n\nInformation gain represents the difference in entropy before and after a split on a given attribute. **The attribute with the highest information gain will produce the best split as it’s doing the best job at classifying the training data according to its target classification.** Information gain is usually represented with the following formula, where: \n\n- α represents a specific attribute or class label\n- Entropy(S) is the entropy of dataset, S\n- |Sv|/ |S| represents the proportion of the values in Sv to the number of values in dataset, S\n- Entropy(Sv) is the entropy of dataset, Sv\n\nLet’s walk through an example to solidify these concepts. Imagine that we have the following arbitrary dataset:\n\n<img width=\"896\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/b88193ec-6390-4d3c-ab4e-75b2b2c7a37f\">\n\nFor this dataset, the entropy is 0.94. This can be calculated by finding the proportion of days where “Play Tennis” is “Yes”, which is 9/14, and the proportion of days where “Play Tennis” is “No”, which is 5/14. Then, these values can be plugged into the entropy formula above.\n\nEntropy (Tennis) = -(9/14) log2(9/14) – (5/14) log2 (5/14) = 0.94\n\n**We can then compute the information gain for each of the attributes individually.** For example, the information gain for the attribute, “Humidity” would be the following:\n\nGain (Tennis, Humidity) = (0.94)-(7/14)*(0.985) – (7/14)*(0.592) = 0.151\n\n\nAs a recap,\n\n- 7/14 represents the proportion of values where humidity equals “high” to the total number of humidity values. In this case, the number of values where humidity equals “high” is the same as the number of values where humidity equals “normal”.\n\n- 0.985 is the entropy when Humidity = “high”\n\n- 0.59 is the entropy when Humidity = “normal”\n\n**Then, repeat the calculation for information gain for each attribute in the table above, and select the attribute with the highest information gain to be the first split point in the decision tree.** In this case, outlook produces the highest information gain. **From there, the process is repeated for each subtree.**\n\n### Gini Impurity \n\nGini impurity **is the probability of incorrectly classifying random data point in the dataset if it were labeled based on the class distribution of the dataset.** ***Similar to entropy, if set, S, is pure—i.e. belonging to one class) then, its impurity is zero.*** This is denoted by the following formula: \n\n<img width=\"998\" alt=\"image\" src=\"https://github.com/eraikakou/hackathon_23/assets/28102493/ee0fa39c-bdc8-405e-989f-275f6e248188\">\n\n## <span style=\"color: #016FD0;\">Advantages and disadvantages of Decision Trees</span>\n\n### Advantages\n\n- **Easy to interpret:** The Boolean logic and visual representations of decision trees make them easier to understand and consume. The hierarchical nature of a decision tree also makes it easy to see which attributes are most important, which isn’t always clear with other algorithms, like neural networks.\n\n- **Little to no data preparation required:** Decision trees have a number of characteristics, which make it more flexible than other classifiers. It can handle various data types—i.e. discrete or continuous values, and continuous values can be converted into categorical values through the use of thresholds. Additionally, **it can also handle values with missing values, which can be problematic for other classifiers, like Naïve Bayes.**  \n\n- **More flexible:** Decision trees can be leveraged for both classification and regression tasks, making it more flexible than some other algorithms. **It’s also insensitive to underlying relationships between attributes;** this means that if **two variables are highly correlated, the algorithm will only choose one of the features to split on.** \n\n### Disadvantages\n\n- **Prone to overfitting:** Complex decision trees tend to overfit and do not generalize well to new data. This scenario can be avoided through the processes of pre-pruning or post-pruning. Pre-pruning halts tree growth when there is insufficient data while post-pruning removes subtrees with inadequate data after tree construction. \n\n- **High variance estimators:** Small variations within data can produce a very different decision tree. Bagging, or the averaging of estimates, can be a method of reducing variance of decision trees. However, this approach is limited as it can lead to highly correlated predictors.  \n\n- **More costly:** Given that decision trees take a greedy search approach during construction, they can be more expensive to train compared to other algorithms. \n\n## <span style=\"color: #016FD0;\">Resources</span>\n\n1. [What is a decision tree?\n](https://www.ibm.com/topics/decision-trees)","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#016FD0;overflow:hidden\">Concepts Recap and FAQ</div>\n\nWeak rules are generated at each iteration by base learning algorithms which in our case can be of two types:\n\n- **Tree** as base learner\n- **Linear** base learner\n\nGenerally, **decision trees** are **default** base learners **for boosting**.\n\n## Summary\n\n- Decision trees\n\n- **Ensemble learning:**\n\n    - **Bagging (+Boostrapping):**\n\n        - Random forest\n\n    - **Boosting:**\n\n        - Adaptative boosting (adaboost), LogitBoost (classification), L2Boost (regression) \n\n        - **Gradient boosting**\n            - Gradient Boosting Trees\n            - XGBoost\n            - LightGBM\n            - CatBoost\n\n    - **Stacking:**\n    \n\n- Most of the errors from a model’s learning are from three main factors: variance, noise, and bias.\n\n- The need for ensemble learning arises in several problematic situations that can be both data-centric and algorithm-centric, like a scarcity/excess of data, the complexity of the problem, constraint in computational resources, etc. High Problem Complexity: Sometimes, a problem can have a complex decision boundary, and it might become impossible for a single classifier to generate the appropriate boundary. For example, if we have a linear classifier and we try to tackle a problem with a parabolic (polynomial) decision boundary. One linear classifier obviously cannot do the job well. However, an ensemble of multiple linear classifiers can generate any polynomial decision boundary.\n\n- a gradient descent approach and can more easily be adapted to large number of loss functions. Thus, gradient boosting can be considered as a generalization of adaboost to arbitrary differentiable loss functions.\n\n### - Decision Trees Disadvantages\n\n- high-variance machine learning algorithm\n\n### + Ensemble Learning Advantages\n\n- improved predictive performance, \n- reduced overfitting,  \n- increased robustness and generalization\n- capturing different aspects of the data and reducing the impact of individual model biases\n\n### - Ensemble Learning Disadvantages\n\n-  it is often not preferred in the industries where interpretability is more important.\n\n### + Bagging Advantages\n\n- It reduces over-fitting of the model. Using random subsets of data, the risk of overfitting is reduced and flattened by averaging the results of the sub-models.\n- It can handle higher dimensionality data very well.\n- Maintains accuracy even for missing data. (observations (from the training dataset or not) with missing data can still be regressed or classified based on the trees that take into account only features where data are not missing)\n- **Random Forests** are robust against overfitting and perform well on many datasets. Compared to individual decision trees, they are also less sensitive to hyperparameters.\n\n### + Boosting Advantages\n\n- Boosting supports different loss functions.\n- It can work well with interactions.\n- The key idea here is clearly to create models that are also able to predict the more difficult data entries. This can then lead to a better fit of the model and reduces the bias.\n\n### - Boosting Disadvantages\n\n- It prone to over-fitting. In comparison to Bagging, this technique uses weighted voting or weighted averaging based on the coefficients of the models that are considered together with their predictions. Therefore, this model can reduce underfitting, but might also tend to overfit sometimes.\n- It requires careful tuning of different hyper-parameters.\n\n### + CatBoost Advantages\n\n- The ability to handle categorical features natively,\n- Models can be trained on several GPUs,\n- It reduces parameter tuning time by providing great results with default parameters,\n- Models can be exported to Core ML for on-device inference (iOS),\n- It handles missing values internally,\n- It can be used for both regression and classification problems.\n\n\n### FAQ and Questions\n\n- depth-wise growth VS Leaf-wise growth\n","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:#016FD0;margin:0;font-size:32px;font-family:Georgia;text-align:center;display:fill;border-radius:5px;overflow:hidden;font-weight:600;\">Work in Progress...</div>\n","metadata":{}}]}