{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:24px;text-align:left;display:fill;border-radius:5px;background-color:#3A578A;overflow:hidden\">Things to Know</div>\n\n## <span style=\"color: #3A578A;\">Convert a column from `float32` to `float16`. What does happen?</span>\n\n\nWhen we convert a column from `float32 `to `float16`, we **decrease the number of unique values** because `float16` **uses fewer bits to represent a number** compared to `float32`. Specifically, `float32` uses 32 bits to represent a number, while `float16` uses only 16 bits.\n\nIn a `float32` column, the number of unique values that can be represented is `2^32`, which is approximately 4.3 billion. This is a very large number, and it allows for a high degree of precision in representing decimal values.\n\nHowever, in a `float16` column, the number of unique values that can be represented is `2^16`, which is approximately 65,000. This is a **much smaller number compared to `float32`**, and it means that **the precision of the decimal values that can be represented is lower.**\n\nWhen we convert a `float32` column to a `float16` column, some of the unique values in the original column will not be representable in the new column. These values will be rounded to the nearest representable value in the `float16` format, resulting in a decrease in the number of unique values. **This decrease in the number of unique values can have an impact on the accuracy of the data, especially if the data has a wide range of values or if the values are very close to each other.**\n\n--------------------------------------------------------------------------------------------------------------------\n\n**Why when we convert a column from `float32` to `float16` do we lose information?**\n\nIn `float32`, 32 bits are used to represent a number, with 23 bits used for the mantissa (fractional part) and 8 bits used for the exponent (the power of 2 that the mantissa is multiplied by). This allows for a wide range of values to be represented with high precision, **making `float32 `suitable for scientific and engineering applications where accuracy is important.**\n\nHowever, in `float16`, only 16 bits are used to represent a number, with 10 bits used for the mantissa and 5 bits used for the exponent. **This means that `float16` can represent a smaller range of values with less precision compared to `float32`.**\n\nWhen we convert a `float32` column to a `float16` column, **some of the information may be lost because the number of bits used to represent each value has been reduced.** Specifically, the mantissa and exponent **values in `float32` may need to be rounded or truncated to fit into the smaller number of bits available in `float16`.** This can result in a loss of precision or a rounding error, which may affect the accuracy of the data.\n\nIn addition, **some operations on `float16` numbers may result in overflow or underflow errors because the range of values that can be represented is smaller.** For example, if a `float16 `number is multiplied by a large value, the result may exceed the maximum value that can be represented, resulting in an overflow error. **These types of errors can also contribute to the loss of information when converting from `float32` to `float16`.**\n\n## <span style=\"color: #3A578A;\">Money has this unique property that values although highly cardinal it is still discrete with an increment of 0.01 USD.</span>\n\nThe term “high cardinality” means that there can be many possible values for a single attribute.\n\nMoney has the unique property that its values are highly cardinal, meaning that there are a large number of distinct possible values, but it is still discrete with an increment of 0.01USD.\n\nWhat this means is that there are a finite number of possible values that can be represented in the currency system, and **each value is separated from the next by a fixed increment of 0.01USD**. For example, in the United States currency system, the smallest unit is the penny, which is worth 0.01USD. This means that any amount of money can be represented as a combination of a whole number of dollars and a fractional amount in cents, with each cent representing a discrete unit of value.\n\nThis discrete nature of money with a fixed increment of 0.01USD is important for many financial calculations and transactions, such as calculating interest, rounding amounts, and making change. It also means that when working with money, it is often necessary to round values to the nearest cent, which can introduce rounding errors.\n\n## <span style=\"color: #3A578A;\">What is a uniform distribution ?</span>\n\nA **uniform distribution** is a type of probability distribution in which **all possible values within a given range are equally likely to occur**. In other words, the probability density function (PDF) of a uniform distribution is constant over a specified range of values.\n\nFor example, if we consider a uniform distribution between 0 and 1, any value between 0 and 1 is equally likely to occur. The PDF for this distribution is a horizontal line at a height of 1, between 0 and 1, and 0 elsewhere. The cumulative distribution function (CDF) for a uniform distribution is a straight line that increases linearly from 0 to 1 over the specified range of values.\n\nA uniform distribution, **sometimes also known as a rectangular distribution, is a distribution that has constant probability.**\n\n<img width=\"926\" alt=\"image\" src=\"https://user-images.githubusercontent.com/28102493/220114129-7896bb12-1d53-49a7-8144-af4dc1637e61.png\">\n\nUniform distributions are often used as models for situations in which there is no a priori reason to expect one value to be more likely than any other value within the range. For example, if we were randomly selecting a number between 1 and 10, and had no reason to expect any particular value to be more likely than any other, we might use a uniform distribution to model the probability of selecting each value.\n\nUniform distributions are also commonly used in computer simulations and in generating random numbers for various applications, such as in cryptography, gambling, and scientific experiments. They are often generated using pseudo-random number generators, which are algorithms that generate a sequence of random numbers that appear to be random, but are actually determined by a starting value known as a seed.\n\nA **random uniform value** refers to a value that is generated randomly from a uniform distribution.\n\nA random value generated from a uniform distribution has the same probability of being any value within the specified range. For example, if we generate a random uniform value between 0 and 1, any value between 0 and 1 is equally likely to be generated. This means that the probability of generating a value between 0.1 and 0.2 is the same as the probability of generating a value between 0.8 and 0.9.\n\n## <span style=\"color: #3A578A;\">What is a random uniform noise?</span>\nRandom uniform noise is a type of noise that is added to data in a random way, such that each value added is drawn from a uniform distribution. A uniform distribution is a probability distribution where every value in a given range is equally likely to occur.\n\nIn the context of machine learning, random uniform noise can be added to data **as a form of regularization or to introduce randomness into the learning process.** This can help **prevent overfitting and improve the generalization ability of the model.** For example, if we are training a model to recognize images of cats and dogs, we might add random uniform noise to the training images to simulate small variations in the images, such as changes in lighting or minor distortions. This can help the model learn to recognize cats and dogs in a more robust way, rather than simply memorizing specific examples from the training data.\n\nRandom uniform noise **can also be used to generate synthetic data for training models, or to simulate noise in real-world data.** In general, adding random uniform noise can help make machine learning models more flexible and robust to different sources of variation and noise in the data.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:24px;text-align:left;display:fill;border-radius:5px;background-color:#3A578A;overflow:hidden\">Available Datasets</div>\n\n1. [AMEX data - integer dtypes - parquet format](https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format): Post processed dataset to transform float data types to int data types where it can be done, by raddar with the discuassion [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514)\n\n1. [AMEX-Feather-Dataset](https://www.kaggle.com/datasets/munumbutt/amexfeather): In this Feather file, the floating point precision has been reduced from 64 bit to 16 bit, compressed version of the train and test sets.\n\n1. [Amex Competition Data in Parquet Format\n](https://www.kaggle.com/datasets/odins0n/amex-parquet): American Express - Default Prediction Competition Data in Parquet Format","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:24px;text-align:left;display:fill;border-radius:5px;background-color:#3A578A;overflow:hidden\">Data Loading</div>\n\n## <span style=\"color: #3A578A;\">Python Garbage Collection</span>\n\nIn Python, `gc.collect()` is a function provided by the gc module which **forces garbage collection, i.e., it frees up memory that is no longer needed by the program.**\n\nPython's garbage collector automatically deallocates memory that is no longer being used by the program. However, this automatic garbage collection may not always happen immediately, as it can depend on various factors such as the amount of memory available or the workload of the system. In some cases, the automatic garbage collection may not free up memory as quickly as needed, leading to memory leaks and other issues.\n\nCalling `gc.collect()` manually **forces the garbage collector to run immediately, which can help free up memory that is no longer needed by the program.** This can be particularly useful in situations where memory usage is a concern, such as when working with large datasets or when running long-running processes.\n\nIt's worth noting that calling gc.collect() too frequently or at inappropriate times can actually slow down your program's performance, so it should be used judiciously and only when necessary.\n\nIn general, deleting a DataFrame with the `del` command in a Jupyter notebook will remove the variable reference to the DataFrame and free up memory, as long as there are no other references to the DataFrame.\n\nPython's garbage collector will eventually clean up any memory that is no longer being used by the program, including the memory previously used by the DataFrame. **However, calling gc.collect() manually may speed up the garbage collection process and free up memory more quickly.**\n\n**That being said, in most cases, it is not necessary to call gc.collect() manually after deleting a DataFrame with del, as the garbage collector should eventually clean up any unused memory. However, in cases where you are working with very large datasets or running long-running processes, calling gc.collect() manually may be useful to free up memory more quickly.**\n\n\n## <span style=\"color: #3A578A;\">Rapids</span>\n\nRAPIDS is a suite of open-source software libraries and APIs developed by NVIDIA that allows users to **execute end-to-end data science and analytics pipelines entirely on GPUs.** The RAPIDS suite includes libraries for data preparation, machine learning, graph analytics, and visualization, and is designed to enable faster and more efficient data processing and analysis.It relies on **NVIDIA® CUDA®** primitives for low-level compute optimization but exposes that GPU parallelism and high-bandwidth memory speed through **user-friendly Python interfaces.**\n\nOne of the key advantages of RAPIDS is that it leverages the power of GPUs to accelerate data processing and analysis. GPUs are highly parallel processors that can perform many calculations simultaneously, making them well-suited for data-intensive tasks such as machine learning and data visualization. By running data science and analytics pipelines entirely on GPUs, RAPIDS can dramatically reduce the time required to process and analyze large datasets.\n\n**Another advantage of RAPIDS is that it integrates seamlessly with popular data science tools and frameworks, such as Python, Jupyter, and scikit-learn.** This makes it easy for users to incorporate RAPIDS into their existing data science workflows and take advantage of its powerful capabilities without having to learn new tools or programming languages. RAPIDS is designed to look and feel like Python and offers a collection of libraries for running a data science pipeline completely through GPUs. RAPIDS includes a Dataframe API, which integrates with machine learning algorithms.\n\n\nOverall, RAPIDS is a powerful tool for data scientists and analysts who need to process and analyze large datasets quickly and efficiently. By leveraging the power of GPUs, RAPIDS can help accelerate the pace of data-driven discovery and enable new insights that would be difficult or impossible to achieve using traditional CPU-based approaches.\n\nIn order to use RAPIDS, you will need a machine with a compatible NVIDIA GPU. This is because RAPIDS is designed to run entirely on GPUs, leveraging their high parallelism and computational power to accelerate data processing and analysis.\n\nHowever, not all NVIDIA GPUs are compatible with RAPIDS, and the specific requirements will depend on the library or tool you are using within the RAPIDS suite. For example, some libraries may require a GPU with a minimum amount of memory, or a specific version of the CUDA Toolkit (NVIDIA's parallel computing platform and programming model).\n\nBefore using RAPIDS, it's important to check the system requirements for the specific library or tool you plan to use, and ensure that your machine meets these requirements. Additionally, if you plan to use RAPIDS in a distributed computing environment, you may need to set up a cluster of machines with compatible GPUs to run your workflows.\n\n<img width=\"1476\" alt=\"image\" src=\"https://user-images.githubusercontent.com/28102493/220171678-545b471a-5866-47c7-bcf1-300eb1c7678d.png\">\n\nRead more about RAPIDS here. The RAPIDS libraries allow us to perform all our data science on GPUs including reading data, transforming data, modeling, validation, and prediction. The package cuDF provides Pandas functionality and cuML provides Scikit-learn functionality. Other packages provide additional tools.\n\nSince GPUs are faster than CPUs, we save time, save money, and can increase model accuracy by performing additional tasks like hyperparameter searches, feature engineering and selection, data augmentation, and ensembling with bagging and boosting.\n\n\n\n- **`cuDF`** is a GPU DataFrame library that provides a pandas-like API for loading, filtering, and manipulating data.\n\n\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport gc\nimport numpy as np\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport scipy.interpolate\nimport scipy.integrate\n#import cuDF # Normally, it worked out of the box with the GPU P100 on and Environment = Always use latest environment\n# but the last two days it doesn't work and I'm getting the error : \"ModuleNotFoundError: No module named 'cuDF'\"","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:37:38.249991Z","iopub.execute_input":"2023-02-27T10:37:38.250521Z","iopub.status.idle":"2023-02-27T10:37:38.258053Z","shell.execute_reply.started":"2023-02-27T10:37:38.250478Z","shell.execute_reply":"2023-02-27T10:37:38.257039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_data(data_path: str, usecols: [] = None):\n    if usecols is not None: \n        df = pd.read_parquet(data_path, columns=usecols)\n    else: \n        df = pd.read_parquet(data_path)\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2023-02-22T10:07:01.709670Z","iopub.execute_input":"2023-02-22T10:07:01.710078Z","iopub.status.idle":"2023-02-22T10:07:01.717963Z","shell.execute_reply.started":"2023-02-22T10:07:01.710046Z","shell.execute_reply":"2023-02-22T10:07:01.716068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = load_data(data_path=\"../input/amex-data-integer-dtypes-parquet-format/train.parquet\")","metadata":{"execution":{"iopub.status.busy":"2023-02-22T10:07:02.299408Z","iopub.execute_input":"2023-02-22T10:07:02.299783Z","iopub.status.idle":"2023-02-22T10:07:21.685753Z","shell.execute_reply.started":"2023-02-22T10:07:02.299755Z","shell.execute_reply":"2023-02-22T10:07:21.684722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-22T10:07:21.687735Z","iopub.execute_input":"2023-02-22T10:07:21.688312Z","iopub.status.idle":"2023-02-22T10:07:21.699694Z","shell.execute_reply.started":"2023-02-22T10:07:21.688277Z","shell.execute_reply":"2023-02-22T10:07:21.698636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-02-22T10:07:21.701518Z","iopub.execute_input":"2023-02-22T10:07:21.702219Z","iopub.status.idle":"2023-02-22T10:07:21.747107Z","shell.execute_reply.started":"2023-02-22T10:07:21.702177Z","shell.execute_reply":"2023-02-22T10:07:21.745749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.info(memory_usage=\"deep\")","metadata":{"execution":{"iopub.status.busy":"2023-02-22T10:07:21.750051Z","iopub.execute_input":"2023-02-22T10:07:21.750562Z","iopub.status.idle":"2023-02-22T10:07:22.769090Z","shell.execute_reply.started":"2023-02-22T10:07:21.750519Z","shell.execute_reply":"2023-02-22T10:07:22.768119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n# <div style=\"padding:20px;color:white;margin:0;font-size:24px;text-align:left;display:fill;border-radius:5px;background-color:#3A578A;overflow:hidden\">Feature Engineering</div>\n\nWe can create the following groups of features:\n\n## <span style=\"color: #3A578A;\">Features aggregated by time period</span>\n\nmin, max, mean, std for the last 3 and last 6 months\n\n1. Selected **features averaged over all statements** of a customer \n1. The **minimum or maximum of selected features** over all statements of a customer \n1. Selected features taken from the **last statement** of a customer \n1. **First:** Just like the last feature, but first.\n\nOn this competition we get information about clients of AMEX over time. Most high scoring notebooks on this competiion focused on aggregating the information per client and create a single row of extracted features: One for each client.\n\nOne of such agg function is last.\n\nQuick examination revealed that the last feature is extreamly powerful at predicting if the client defaults or not (well.. make sense..). So I took this two steps further ..... Lag Features\n\n\n## <span style=\"color: #3A578A;\">Lag features</span>\n\nIntroduction to feature engineering for time series forecasting\n\nLag features are values at prior timesteps that are considered useful because they are created on the assumption that what happened in the past can influence or contain a sort of intrinsic information about the future. Lag features are features that capture the behavior of a target variable (such as a time series) at different time points in the past. They are commonly used in time series analysis and forecasting, as they can help capture patterns in the data over time.\n\nTo create a lag feature, you simply **shift** the target variable forward in time by a certain number of time steps (referred to as the lag time), and use the resulting values as input features for your predictive model. For example, suppose you are working with daily stock price data and want to predict tomorrow's closing price based on past price data. You could create lag features that capture the closing price from the past 1, 5, and 10 days, and use these values as input features for your predictive model.\n\nIn addition to capturing patterns in the data over time, lag features can also help account for autocorrelation in the data. **Autocorrelation refers to the tendency for a time series to be correlated with its own past values. By incorporating lag features into your predictive model, you can help capture this autocorrelation and improve the accuracy of your predictions.**\n\nIt's important to note that when creating lag features, you should be careful not to include future information in your model. For example, if you are trying to predict tomorrow's stock price, you cannot use information from tomorrow or beyond in your model. Additionally, the appropriate lag time period will depend on the specific problem you are working on and the nature of the data.\n\n\"Lag\" fearures can capture the change over time about each client I calculated two features for every first, last pair:\n\n\n1. **First / Last interactions:** (Last - First) and (Last / First)\n    1. Last - First: The change since we first see the client to the last time we see the client.\n    1. Last / First: The fractional difference since we first see the client to the last time we see the client.\n1. The **difference between last value and the average** (this features gives a nice boost)  \n1. The **difference between last value and the lag1**\n\n## <span style=\"color: #3A578A;\">Null count features</span>\n","metadata":{}},{"cell_type":"code","source":"def feature_engineer(df: pd.DataFrame, labels: pd.DataFrame = None):\n    \n    # We keep all the features except for customer_ID and the date\n    features = df.drop([\"customer_ID\", \"S_2\"], axis = 1).columns.to_list()\n    # As described in the data tab in the competition description\n    cat_features = [\n        \"B_30\",\n        \"B_38\",\n        \"D_114\",\n        \"D_116\",\n        \"D_117\",\n        \"D_120\",\n        \"D_126\",\n        \"D_63\",\n        \"D_64\",\n        \"D_66\",\n        \"D_68\",\n    ]\n    # The rest of the features - columns are numerical features\n    num_features = [col for col in features if col not in cat_features]\n    \n    print(\"Starting feature engineer on numerical features...\")\n    # Numerical features aggregated by time period\n    df_num_agg = df.groupby(\"customer_ID\")[num_features].agg([\"mean\", \"std\", \"min\", \"max\", \"first\", \"last\"])\n    df_num_agg.columns = [\"_\".join(x) for x in df_num_agg.columns]\n    df_num_agg.reset_index(inplace = True)\n    \n    print(\"Starting feature engineer on lag features...\")\n    # Lag Features: First / Last interactions\n    for col in df_num_agg.columns:\n        if 'last' in col and col.replace('last', 'first') in df_num_agg.columns:\n            \n            df_num_agg[col + \"_lag_sub\"] = df_num_agg[col] - df_num_agg[col.replace(\"last\", \"first\")]\n            df_num_agg[col + \"_lag_div\"] = df_num_agg[col] / df_num_agg[col.replace(\"last\", \"first\")]\n\n    print(\"Starting feature engineer on categorical features...\")\n    # Categorical features aggregated by time period\n    df_cat_agg = df.groupby(\"customer_ID\")[cat_features].agg([\"count\", \"first\", \"last\", \"nunique\"])\n    df_cat_agg.columns = [\"_\".join(x) for x in df_cat_agg.columns]\n    df_cat_agg.reset_index(inplace = True)\n    \n    if labels is not None: \n        df_agg = df_num_agg.merge(df_cat_agg, how = \"inner\", on = \"customer_ID\").merge(labels, how = 'inner', on = 'customer_ID')\n    else:\n        df_agg = df_num_agg.merge(df_cat_agg, how = \"inner\", on = \"customer_ID\")\n    del df_num_agg, df_cat_agg\n\n    print(\"Dimensions after engineering\", df_agg.shape )\n    \n    return df_agg\n","metadata":{"execution":{"iopub.status.busy":"2023-02-22T10:07:22.770375Z","iopub.execute_input":"2023-02-22T10:07:22.770904Z","iopub.status.idle":"2023-02-22T10:07:22.782289Z","shell.execute_reply.started":"2023-02-22T10:07:22.770874Z","shell.execute_reply":"2023-02-22T10:07:22.781266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels = pd.read_csv(\"../input/amex-default-prediction/train_labels.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-02-23T10:05:56.511369Z","iopub.execute_input":"2023-02-23T10:05:56.512780Z","iopub.status.idle":"2023-02-23T10:05:57.647741Z","shell.execute_reply.started":"2023-02-23T10:05:56.512735Z","shell.execute_reply":"2023-02-23T10:05:57.646559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-23T10:05:59.329751Z","iopub.execute_input":"2023-02-23T10:05:59.330148Z","iopub.status.idle":"2023-02-23T10:05:59.339174Z","shell.execute_reply.started":"2023-02-23T10:05:59.330117Z","shell.execute_reply":"2023-02-23T10:05:59.338155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels[\"target\"].value_counts(normalize=True)","metadata":{"execution":{"iopub.status.busy":"2023-02-23T10:06:17.531326Z","iopub.execute_input":"2023-02-23T10:06:17.531758Z","iopub.status.idle":"2023-02-23T10:06:17.555551Z","shell.execute_reply.started":"2023-02-23T10:06:17.531723Z","shell.execute_reply":"2023-02-23T10:06:17.554218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels[\"target\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-02-23T10:06:17.722109Z","iopub.execute_input":"2023-02-23T10:06:17.722583Z","iopub.status.idle":"2023-02-23T10:06:17.737260Z","shell.execute_reply.started":"2023-02-23T10:06:17.722542Z","shell.execute_reply":"2023-02-23T10:06:17.736027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = feature_engineer(train_data, labels)","metadata":{"execution":{"iopub.status.busy":"2023-02-23T10:06:17.949972Z","iopub.execute_input":"2023-02-23T10:06:17.950708Z","iopub.status.idle":"2023-02-23T10:06:17.970661Z","shell.execute_reply.started":"2023-02-23T10:06:17.950654Z","shell.execute_reply":"2023-02-23T10:06:17.968497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-21T17:31:03.862904Z","iopub.execute_input":"2023-02-21T17:31:03.863272Z","iopub.status.idle":"2023-02-21T17:31:03.871479Z","shell.execute_reply.started":"2023-02-21T17:31:03.863240Z","shell.execute_reply":"2023-02-21T17:31:03.870111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.target = train_df.target.astype(\"int8\")","metadata":{"execution":{"iopub.status.busy":"2023-02-21T17:33:22.497270Z","iopub.execute_input":"2023-02-21T17:33:22.497733Z","iopub.status.idle":"2023-02-21T17:33:22.506518Z","shell.execute_reply.started":"2023-02-21T17:33:22.497693Z","shell.execute_reply":"2023-02-21T17:33:22.505248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# CLEAN RAM\ndel train_data\ndel labels\n# callgarbage collection,to manage a performance-intensive section of\n# our program that generates many temporary objects.\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-02-21T17:31:03.873009Z","iopub.execute_input":"2023-02-21T17:31:03.873418Z","iopub.status.idle":"2023-02-21T17:31:04.000857Z","shell.execute_reply.started":"2023-02-21T17:31:03.873375Z","shell.execute_reply":"2023-02-21T17:31:03.999465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:24px;text-align:left;display:fill;border-radius:5px;background-color:#3A578A;overflow:hidden\">Dealing with Imbalanced Dataset</div>\n\n## <span style=\"color: #3A578A;\">Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.</span>\n\n### Problem Description\n\nA classification data set with skewed class proportions is called **imbalanced**. Classes that make up a large proportion of the data set are called **majority classes**. Those that make up a smaller proportion are **minority classes**. Machine learning techniques often fail or give misleadingly optimistic performance on classification datasets with an imbalanced class distribution.\n\nThe reason is that many machine learning algorithms are designed to operate on classification data with an equal number of observations for each class. When this is not the case, algorithms can learn that very few examples are not important and can be ignored in order to achieve good performance. As such, many machine learning algorithms, like decision trees, k-nearest neighbors, and neural networks, will therefore learn that the minority class is not as important as the majority class and put more attention and perform better on the majority class. **This is a problem because the minority class is exactly the class that we care most about in imbalanced classification problems.** The reason for this is because the majority class often reflects a normal case, whereas the minority class represents a positive case for a diagnostic, fault, fraud, or other types of exceptional circumstance.\n\n### Solutions\n\n#### 1. Balance the Class Distribution With Data Sampling\nThe most popular solution to an imbalanced classification problem is to change the composition of the training dataset.\n\n**Techniques designed to change the class distribution in the training dataset are generally referred to as sampling methods** or resampling methods as we are sampling an existing data sample. \n\nThe reason that sampling methods are so common is because they are simple to understand and implement, and because once applied to transform the training dataset, a suite of standard machine learning algorithms can then be used directly.\n\nThis means that any from tens or hundreds of machine learning algorithms developed for balanced (or mostly balanced) classification can then be fit on the training dataset without any modification adapting them for the imbalance in observations.\n\n1. **Undersampling Techniques:** Undersampling methods delete or select a subset of examples from the majority class. Undersampling can be defined as removing some observations of the majority class. Undersampling can be a good choice when you have a ton of data -think millions of rows. But a drawback is that we are removing information that may be valuable. This could lead to underfitting and poor generalization to the test set.\n\n    Of 458913 customer_IDs in the training data, 340085 (~ **74%**) have a label of 0 (good customer, no default) and 118828 (~ **26%**) have a label of 1 (bad customer, default). In reality, however, American Express has subsampled the negative class for this dataset at 5%. This means that :\n    \n    `Total number of negative samples x 5/100 = 340085 => Total number of negative samples = (340085 x 100) / 5 = ~6.8 million non-defaulting customers` \n    \n    Out of 6920528 total training samples the ~ 6.8 million are non-defaulting customers (**98%**) and only **2%** are the bad customers. \n    \n    Downsampling means you sample from the majority class (the 98.5%) to reduce the imbalance between majority and minority class. If you keep the ratio constant you simply reduce your number of trainings examples. This doesn't make sense. However, you don't have to sample down to a ratio of 50:50. If you have a ratio of 98:2, you can sample to 80:2 instead of 2:2.\n    \n    The main goal of downsampling (and upsampling) is to increase the discriminative power between the two classes. Ideally, you would have a classifier that outputs a decision surface that is not simply binary (e.g. logistic regression (where you don't have to select a cut-off point of 0.5)) but gives you a continuous decision value. You can then order the data and set a decision threshold that gives you the best outcome.\n    \n    Since downsampling (or upsampling) changes your training distribution from your true distribution, you only want to downsample (or upsample) so much that your classifier can start discriminating between the two classes. **You then fine-tune the results by selecting an appropriate decision threshold.**\n    \n    **Downsampling and Upweighting**\n    \n    An effective way to handle imbalanced data is to downsample and upweight the majority class. Let's start by defining those two new terms:\n    \n    1. **Downsampling (in this context)** means training on a disproportionately low subset of the majority class examples.\n\n    1. **Upweighting** means adding an example weight to the downsampled class equal to the factor by which you downsampled.\n    \n    Step 1: Downsample the majority class. Consider again our example of the fraud data set, with 1 positive to 200 negatives. Downsampling by a factor of 20 improves the balance to 1 positive to 10 negatives (10%). Although the resulting training set is still moderately imbalanced, the proportion of positives to negatives is much better than the original extremely imbalanced proportion (0.5%).\n    \n    Step 2: Upweight the downsampled class: The last step is to add example weights to the downsampled class. **Since we downsampled by a factor of 20, the example weight should be 20.** You may be used to hearing the term weight when it refers to model parameters, like connections in a neural network. Here we're talking about example weights, which means counting an individual example more importantly during training. An example weight of 10 means the model treats the example as 10 times as important (when computing loss) as it would an example of weight 1. **The weight should be equal to the factor you used to downsample:**\n    \n    <img width=\"782\" alt=\"image\" src=\"https://user-images.githubusercontent.com/28102493/220610379-e5485e1e-bc29-4166-a0de-62a6fc70bde6.png\">\n    \n    **Why Downsample and Upweight?**\n\n    It may seem odd to add example weights after downsampling. We were trying to make our model improve on the minority class -- why would we upweight the majority? These are the resulting changes:\n\n    1. **Faster convergence:** During training, we see the minority class more often, which will help the model converge faster.\n\n    1. **Disk space:** By consolidating the majority class into fewer examples with larger weights, we spend less disk space storing them. This savings allows more disk space for the minority class, so we can collect a greater number and a wider range of examples from that class.\n\n    1. **Calibration:** Upweighting ensures our model is still calibrated; the outputs can still be interpreted as probabilities.\n\n1. **Oversampling Techniques**\n1. **Combination of both of them**\n\n\n### 2. Weighted Loss Functions\n\nIn undersampling, you randomly remove samples from the majority class so that the ratio of the minority class to the majority class is more balanced. However, this can result in loss of valuable information from the majority class, which may lead to a model that does not generalize well.\n\nTo avoid this issue, some practitioners use \"weighted loss functions\", **which assign higher weight to the underrepresented class in the loss function to give it more importance during training.** However, if the loss function is not properly weighted, the model may be biased towards the majority class.\n\n\n### 3. Choose an appropriate Evaluation Metric\n\n\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:24px;text-align:left;display:fill;border-radius:5px;background-color:#3A578A;overflow:hidden\">Evaluation Metric Explained</div>\n\n\n## <span style=\"color: #3A578A;\">Recall: True Positive Rate</span>\n\nRecall attempts to answer the following question:\n\nWhat proportion of actual positives was identified correctly? => **How many actual positives the model predicted correctly.**\n\n![image](https://user-images.githubusercontent.com/28102493/220618679-2c62cf78-b5ca-4e1f-ab30-edbeb9284f12.png)\n\n## <span style=\"color: #3A578A;\">False Positive Rate</span>\n\nWhat is False Positive Rate?\n\nIn data science, the false positive rate measures the percentage of false positives against all negative predictions (the sum of false positives and true negatives) in a binary classification problem. The false positive rate is based on **how many actual negatives the model predicted incorrectly.** This metric is complementary to the true positive rate, or recall, which shows how many actual positives the model predicted correctly.\nThe false positive rate is calculated as FP/FP+TN, where FP is the number of false positives and TN is the number of true negatives (FP+TN being the total number of negatives). It's the probability that a false alarm will be raised\n\n![image](https://user-images.githubusercontent.com/28102493/220620304-6ea4d3e6-f4ff-46ee-8ed6-636be77a0ff2.png)\n\n## <span style=\"color: #3A578A;\">ROC curve</span>\n\nAn ROC curve (receiver operating characteristic curve) **is a graph showing the performance of a classification model at all classification thresholds.** This curve plots two parameters:\n\n- True Positive Rate\n- False Positive Rate\n\nIn a binary classification problem, the output of a classifier can be represented as a probability score between 0 and 1, where values close to 0 indicate that the sample belongs to the negative class and values close to 1 indicate that the sample belongs to the positive class. **The decision classification threshold is the probability value above which a sample is classified as belonging to the positive class and below which it is classified as belonging to the negative class.**\n\nROC (Receiver Operating Characteristic) is a plot that visualizes the performance of a binary classifier at different decision classification thresholds. The plot shows the true positive rate (TPR) on the y-axis, and the false positive rate (FPR) on the x-axis, at different threshold values. The **TPR** is the proportion of true positive predictions (i.e., **correctly classified positive samples) among all positive samples**, while the FPR is the proportion of false positive predictions (i.e., **incorrectly classified negative samples) among all negative samples.**\n\n**As the decision classification threshold is varied from 0 to 1, the TPR and FPR change, and this change is represented as a curve in the ROC plot.** The curve starts at the point (0,0) (corresponding to a threshold of 1, which classifies all samples as negative) and moves towards the point (1,1) (corresponding to a threshold of 0, which classifies all samples as positive). **The area under this curve (AUC) is a measure of the classifier's overall performance, where an AUC of 1 indicates perfect performance and an AUC of 0.5 indicates random performance.**\n\nLowering the classification threshold classifies more items as positive, thus increasing both False Positives and True Positives. The following figure shows a typical ROC curve.\n\n![image](https://user-images.githubusercontent.com/28102493/220627394-9fe48f3b-2740-4bfb-a9fc-86c47fb665a3.png)\n\n\nTP vs. FP rate at different classification thresholds.\n\nTo compute the points in an ROC curve, we could evaluate a logistic regression model many times with different classification thresholds, but this would be inefficient. Fortunately, there's an efficient, sorting-based algorithm that can provide this information for us, called AUC.\n\n<img width=\"810\" alt=\"image\" src=\"https://user-images.githubusercontent.com/28102493/221170676-43f81237-46e7-4582-93d3-6865eefc175e.png\">\n\n\nAUC stands for \"Area under the ROC Curve.\" That is, AUC measures the entire two-dimensional area underneath the entire ROC curve (think integral calculus) from (0,0) to (1,1).\nAUC provides an aggregate measure of performance across all possible classification thresholds. One way of interpreting AUC is as the probability that the model ranks a random positive example more highly than a random negative example. \n\nAUC ranges in value from 0 to 1. A model whose predictions are 100% wrong has an AUC of 0.0; one whose predictions are 100% correct has an AUC of 1.0.\n\nAUC is desirable for the following two reasons:\n\n1. AUC is scale-invariant. It measures how well predictions are ranked, rather than their absolute values.\n\n1. AUC is classification-threshold-invariant. It measures the quality of the model's predictions irrespective of what classification threshold is chosen.\n\n### Purpose of ROC Curve: Purpose 1 — Determining optimal threshold\nHowever, both these reasons come with caveats, which may limit the usefulness of AUC in certain use cases:\n\n1. **Scale invariance** is not always desirable. For example, sometimes we really do need well calibrated probability outputs, and AUC won’t tell us about that.\n\n1. **Classification-threshold invariance is not always desirable.** In cases where there are wide disparities in the cost of false negatives vs. false positives, it may be critical to minimize one type of classification error. For example, when doing email spam detection, you likely want to prioritize minimizing false positives (even if that results in a significant increase of false negatives). AUC isn't a useful metric for this type of optimization. **The performance of our system would vary as we change this threshold. It can be adjusted to tune the behavior of the model to reduce more of one or another type of error (FP/FN).**\n\nWe have a classification model that predicts the probability that an observation belongs to Class YES. This probability can range from 0 to 1. The default threshold of 0.5 that is used to determine the class of this observation is not always the best threshold. The ROC curve helps us find the threshold where the TPR is high and FPR is low i.e. misclassifications are low. Therefore, ROC curves should be used to determine the optimal probability threshold for a classification model.\n\n\n**Example:**\n\n<img width=\"645\" alt=\"image\" src=\"https://user-images.githubusercontent.com/28102493/220629275-179436bd-fb5c-4a34-a2a4-383a674c85f7.png\">\n\n\n1. In the above curve if you wanted a model with a very low false positive rate, you might pick 0.8 as your threshold of choice. If you favour a low FPR, but you don’t want an abysmal TPR, you might go for 0.5, the point where the curve starts turning hard to the right.\n\n2. If you prefer a low false negative rate/high TPR, you might decide that somewhere between 0.2 and 0.1 is the region where you start getting severely diminishing returns for improving the TPR any further.\n\n3. Notice the graph at threshold 0.5 and 0.4. The TPR at both the thresholds is ~0.6, but FPR is higher at threshold 0.4. It’s clear that if we are happy with TPR = 0.6 we should choose threshold = 0.5.\n\n## <span style=\"color: #3A578A;\">Lorenz Curve & Gini Coefficient</span>\n\n### Definition\nA Lorenz curve, developed by American economist Max Lorenz in 1905, is a graphical representation of income inequality or wealth inequality. **The graph plots percentiles of the population on the horizontal axis according to income or wealth and plots cumulative income or wealth on the vertical axis.**\n\n<img width=\"690\" alt=\"image\" src=\"https://user-images.githubusercontent.com/28102493/220911221-f3384f45-e884-45fe-836c-1a20059b549e.png\">\n\nIn practice, a Lorenz curve is usually a mathematical function estimated from an incomplete set of observations of income or wealth. **The Lorenz curve is often accompanied by a straight diagonal line with a slope of 1, which represents perfect equality in income or wealth distribution; the Lorenz curve lies beneath it, showing the observed or estimated distribution.**\n\n**While the Lorenz curve is most often used to represent economic inequality, it can also demonstrate unequal distribution in any system. The farther the curve is from the baseline, represented by the straight diagonal line, the higher the level of inequality.**\n\nIn economics, the Lorenz curve denotes inequality in the distribution of either wealth or income. These are not synonymous since it is possible to have either high earnings but zero or negative net worth, or low earnings but a large net worth. \n\n- The x-axis is often denoted as the percentile. \n\n- The y-axis is often denoted as the cumulative percentage of occurrences.\n\n- The line of equality is demonstrated by a 45-degree, upward-sloping line. In the graph above, it is denoted as the dashed line.\n\n- The Lorenz curve is demonstrated often by an upward-sloping but often exponentially rising curve. In the graph above, it is denoted as a solid line.\n\n- The Gini coefficient (discussed below) is the gap between the line of equality and the Lorenz curve.\n\n### Example\n\nThe curve above shows a continuous Lorenz curve that has been fitted to data that describes income distribution in Brazil as of 2020. The data set is also compared to a straight diagonal line representing perfect equality.\n\nAt the 55th income percentile, the value of the Lorenz curve is 22.39%. This means that the Lorenz curve estimates that the bottom 55% of the population takes in 22.39% of the nation’s total income. If Brazil were a perfectly equal society, the bottom 55% would earn 55% of the total.\n\nElsewhere, we can see that the 99th percentile corresponds to 89.32% in cumulative income. This means that the top 1% takes in 11.68% of Brazil’s income.\n\n\n\n### KEY TAKEAWAYS\n\n- A Lorenz curve is a graphical representation of the distribution of income or wealth within a population.\n\n- Lorenz curves graph percentiles of the population against cumulative income or wealth of people at or below that percentile. \n\n- Lorenz curves, along with their derivative statistics, are widely used to measure inequality across a population. \n\n- The Lorenz curve is a central piece in calculating **the Gini coefficient, a mathematical representation of inequality levels.**\n\n- Because Lorenz curves are mathematical estimates based on fitting a continuous curve to incomplete and discontinuous data, they may be imperfect measures of true inequality.\n\n### Gini Coefficient\n\n**The Gini coefficient is used to express the extent of inequality in a single figure.** The Gini coefficient measures the inequality in a distribution, and **its range depends on the context in which it is used.** It most often ranges from 0 (or 0%) to 1 (or 100%). Complete equality, in which every individual has the exact same income or wealth, corresponds to a coefficient of 0. Plotted as a Lorenz curve, complete equality would be a straight diagonal line with a slope of 1 (the area between this curve and itself is 0, so the Gini coefficient is 0).\n\nA coefficient of 1 means that one person earns all of the income or holds all of the wealth. A Gini coefficient of 0.5 would represent a situation where half the population has half the income or wealth, and the other half has the other half.\n\nIn theory, the Gini coefficient can exceed 100% in extreme situations. For example, when handling negative wealth or income, the figure can theoretically be higher than 1; in that case, the Lorenz curve would dip below the horizontal axis.\n\n**The Gini coefficient is equal to the area below the line of perfect equality minus the area below the Lorenz curve, divided by the area below the line of perfect equality.** In the graph above, the Gini coefficient is the area below the dashed line but above the solid line. The Gini coefficient is used to measure the extent of inequality. It can also be used to compare two different nations or countries to see which has more inequality.\n\n<img width=\"624\" alt=\"image\" src=\"https://user-images.githubusercontent.com/28102493/220958224-c0080cf7-1518-487a-b7e6-2b2aa1444d3e.png\">\n\n\n### Gini Coefficient in Binary Classification Problems\nIn this section I will explain the Gini coefficient’s usage and relevance for the data science professionals.\n\n\n**The Gini coefficient is a popular metric on Kaggle, especially for imbalanced class values.** But googling \"Gini coefficient\" gives you mostly economic explanations. Here is a descriptive explanation with regard to using it as an evaluation metric in classification.\n\nIn the context of evaluating a binary classifier, such as a credit default model, the Gini coefficient is typically scaled to the range [0, 1]. A perfect classifier that correctly separates all positive and negative examples would have a Gini coefficient of 1, while a completely random classifier with no predictive power would have a Gini coefficient of 0. A classifier that performs worse than random would have a negative Gini coefficient. The Gini Coefficient is `2 * AUC – 1`, and its purpose is to normalize the AUC so that a random classifier scores 0, and a perfect classifier scores 1. \n\n***Why the Gini coefficient metric is preferred especially for imbalanced class classification problems?***\n\n**Gini or the Gini coefficient is one of the most popular metrics used by the financial industry for evaluating the performance of credit score models** because it has several desirable properties for assessing binary classification models:\n\n1. Firstly, the Gini coefficient **is sensitive to the correct ranking of examples, which is particularly important in credit scoring where the focus is on identifying high-risk customers. This means that the metric is able to effectively capture the relative ordering of the predictions made by the model.** The Gini coefficient metric is preferred for imbalanced classification problems because it is sensitive to the ranking of the predicted probabilities of the positive class. In imbalanced problems, the minority class is of particular interest, but it is also more difficult to predict accurately due to its low prevalence. Therefore, metrics that rely on absolute counts, such as accuracy or precision, can be misleading and may not reflect the true performance of a model. Moreover, the Gini coefficient is particularly useful in evaluating models that are intended for ranking or scoring purposes, such as credit scoring or fraud detection, where the objective is to rank the instances according to their likelihood of belonging to the positive class, rather than to make binary predictions.\n\n1. Secondly, the Gini coefficient is relatively **robust to class imbalance**, which is often a problem in credit scoring where the number of bad (i.e., high-risk) customers is typically much smaller than the number of good (i.e., low-risk) customers.\n\n1. Lastly, the Gini coefficient is easy to understand and interpret, which makes it a popular choice for both model developers and end-users. It provides a simple and **intuitive way to measure the predictive power of a credit score model, which can be used to compare different models and to make decisions about which model to use in practice.** The Gini coefficient is a metric that indicates the model’s discriminatory power, namely, the effectiveness of the model in differentiating between “bad” borrowers, who will default in the future, and “good” borrowers, who won’t default in the future. This metric is often used to compare the quality of different models and evaluate their prediction power.\n\n\n**In summary, the Gini coefficient is preferred for imbalanced classification problems because it is a ranking-based metric that is sensitive to the performance of a model in distinguishing between the positive and negative classes, regardless of their relative sizes, and is particularly useful for evaluating models that are intended for ranking or scoring purposes.**\n\n### Intuitive Explanation of the Gini coefficient calculation for binary classification problems\n\nIt is calculated based on the distribution of the predicted probabilities for the positive class (the class of interest). \n\nThe Gini coefficient compares the cumulative distribution of the predicted probabilities of the positive class to the cumulative distribution of a uniform random variable. This comparison provides a measure of how well the model is able to distinguish between the positive and negative classes, regardless of their relative sizes.\n\nThe calculation involves the following steps:\n\n#### Step 1: Sort the predicted probabilities in descending order\nIn Gini coefficient they go through the population from poorest to richest and plot the running total / cumulative share of income, which gives them the Lorenz Curve. The Gini coefficient is then defined as the blue area divided by the total area of the triangle.\n\n<img width=\"644\" alt=\"image\" src=\"https://user-images.githubusercontent.com/28102493/221203026-c19eefb3-3707-43ed-bf47-e122d79f0124.png\">\n\nInstead of going through the population from poorest to richest, we go through our predictions from highest to lowest.\n\nSort the predicted probabilities in descending order, from highest to lowest. The Lorenz curve is the inverse of the CAP curve; it is constructed using the same mechanism of sampling observations and aggregating the cumulative default rate, but the sampling is done in reverse order (from highest to lowest score). \n\n<img width=\"717\" alt=\"image\" src=\"https://user-images.githubusercontent.com/28102493/221240235-4c7083de-a7c3-44c5-a434-ac27aa0ac0ac.png\">\n\nWhen using the Lorenz curve to evaluate an imbalance binary classification problem, we should sort the probabilities by descending order.\n\nThe Lorenz curve is a graphical representation of the cumulative distribution function (CDF) of a variable. In the context of binary classification, the variable of interest is the predicted probability of belonging to the positive class. Sorting the probabilities in descending order means that we are placing the highest predicted probabilities at the top of the curve, and the lowest predicted probabilities at the bottom.\n\nThe Lorenz curve is used to assess the degree of imbalance in the binary classification problem by comparing the actual distribution of positive samples to the ideal distribution. The ideal distribution assumes that the positive samples are ranked in descending order of their predicted probabilities. By sorting the probabilities in descending order, we are aligning the actual distribution of positive samples with the ideal distribution.\n\nIn summary, sorting the probabilities by descending order is recommended when using the Lorenz curve to evaluate an imbalance binary classification problem because it aligns the actual distribution of positive samples with the ideal distribution.\n\nThe Lorenz curve is typically sorted in ascending order. This is because in economics, the variable of interest is typically a measure of economic inequality, such as income or wealth. In this context, the Lorenz curve is used to show the distribution of income or wealth across a population, and the ideal distribution assumes that everyone has an equal share.\n\nWhen the Lorenz curve is sorted in ascending order, the x-axis represents the cumulative proportion of the population, starting from the poorest individuals and increasing towards the richest. The y-axis represents the cumulative proportion of the total income or wealth held by that proportion of the population.\n\nThe difference in sorting order between the use of the Lorenz curve in economics and in binary classification problems arises from the different variables of interest. In economics, the Lorenz curve is used to show the distribution of income or wealth, and it is sorted in ascending order to compare the actual distribution to the ideal distribution of equal shares. In binary classification problems, the Lorenz curve is used to assess the degree of imbalance in the predicted probabilities of belonging to the positive class, and it is sorted in descending order to align the actual distribution of positive samples with the ideal distribution.\n\nIn summary, the sorting order of the Lorenz curve depends on the variable of interest and the specific context in which it is being used. In economics, the Lorenz curve is typically sorted in ascending order to show the distribution of income or wealth, while in binary classification problems, it is sorted in descending order to assess the degree of imbalance in predicted probabilities.\n\n\n**Other explanation:**\n\n\nThe Gini coefficient is a measure of the inequality of a distribution, often used in the context of evaluating imbalanced binary classification problems. When calculating the Gini coefficient for a binary classification problem, the true class labels are typically represented as 0 (for the negative or majority class) and 1 (for the positive or minority class).\n\nTo calculate the Gini coefficient for a given binary classifier, we need to sort the predicted probabilities for the positive class in descending order. This is because the Gini coefficient measures the extent to which the positive class is overrepresented at the top of the sorted list of predictions. The more the positive class is concentrated at the top of the list, the higher the Gini coefficient, indicating better performance.\n\nIn contrast, the Lorenz curve is used to visualize the degree of inequality in a distribution by plotting the cumulative proportion of the population against the cumulative proportion of the resource being measured. The Lorenz curve is used to illustrate the degree of inequality in a distribution, but it does not measure the performance of a binary classification model.\n\n\nWhen we sort the predicted probabilities in descending order for a binary classification problem, we are effectively creating a ranking of the instances according to their likelihood of belonging to the positive class. This ranking can be used to construct the Lorenz curve by plotting the cumulative proportion of the positive class instances on the y-axis against the cumulative proportion of the instances on the x-axis.\n\nIf the predicted probabilities are well-calibrated, meaning that they accurately reflect the true likelihood of the instances belonging to the positive class, then the Lorenz curve will generally lie above the diagonal line. This is because the instances that are more likely to belong to the positive class will be concentrated towards the top of the ranking, and as we move up the ranking, the proportion of positive class instances will increase more quickly than the proportion of instances overall.\n\nThis results in a curve that is above the diagonal line, indicating that the distribution of the positive class is more concentrated towards the top of the ranking than would be expected under perfect balance. On the other hand, if the predicted probabilities are poorly calibrated, then the Lorenz curve may lie below the diagonal line or be close to it, indicating that the positive class is not well-separated from the negative class and that the binary classifier is not performing well.\n\nIn summary, when we sort the probabilities in descending order, the Lorenz curve is shaped above the diagonal line if the predicted probabilities are well-calibrated and accurately reflect the true likelihood of the instances belonging to the positive class.\n","metadata":{}},{"cell_type":"code","source":"predictions = [0.9, 0.3, 0.8, 0.75, 0.65, 0.6, 0.78, 0.7, 0.05, 0.4, 0.4, 0.05, 0.5, 0.1, 0.1]\nactual = [1, 0, 1, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0]\ndata = {\"prediction\":predictions, \"target\": actual}  \ndata_df = pd.DataFrame(data)  \ndata_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:44:59.711455Z","iopub.execute_input":"2023-02-27T10:44:59.711845Z","iopub.status.idle":"2023-02-27T10:44:59.727230Z","shell.execute_reply.started":"2023-02-27T10:44:59.711812Z","shell.execute_reply":"2023-02-27T10:44:59.725699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:44:59.953290Z","iopub.execute_input":"2023-02-27T10:44:59.953687Z","iopub.status.idle":"2023-02-27T10:44:59.960780Z","shell.execute_reply.started":"2023-02-27T10:44:59.953651Z","shell.execute_reply":"2023-02-27T10:44:59.959799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = (pd.concat([data_df[\"target\"], data_df[\"prediction\"]], axis='columns').sort_values('prediction', ascending=False))\ndf","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:45:00.170772Z","iopub.execute_input":"2023-02-27T10:45:00.171316Z","iopub.status.idle":"2023-02-27T10:45:00.190037Z","shell.execute_reply.started":"2023-02-27T10:45:00.171264Z","shell.execute_reply":"2023-02-27T10:45:00.188534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = zip(actual, predictions)\nsorted_data = sorted(data, key=lambda d: d[1], reverse=True)\nsorted_actual = [d[0] for d in sorted_data]\nprint(sorted_actual)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:45:00.354183Z","iopub.execute_input":"2023-02-27T10:45:00.354616Z","iopub.status.idle":"2023-02-27T10:45:00.361436Z","shell.execute_reply.started":"2023-02-27T10:45:00.354580Z","shell.execute_reply":"2023-02-27T10:45:00.359778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Step 2: Sum up the actual values - Calculate Lorenz Curve\n\nInstead of summing up the income, we sum up the actual values of our predictions by taking into account their weight as well:\n","metadata":{}},{"cell_type":"code","source":"# We assign the weight to the majority class (0) that we subsample\ndf['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\ndf","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:45:00.706893Z","iopub.execute_input":"2023-02-27T10:45:00.707662Z","iopub.status.idle":"2023-02-27T10:45:00.721632Z","shell.execute_reply.started":"2023-02-27T10:45:00.707619Z","shell.execute_reply":"2023-02-27T10:45:00.720354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sum up the targets\ntotal_pos = (df['target'] * df['weight']).sum()\ntotal_pos","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:45:00.862015Z","iopub.execute_input":"2023-02-27T10:45:00.862823Z","iopub.status.idle":"2023-02-27T10:45:00.872652Z","shell.execute_reply.started":"2023-02-27T10:45:00.862759Z","shell.execute_reply":"2023-02-27T10:45:00.871507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Instead of summing up the income, we sum up the actual values of our predictions\n# This corresponds to the Lorenz Curve in the diagram above. See the following chart\ndf['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\ndf","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:45:01.047722Z","iopub.execute_input":"2023-02-27T10:45:01.048459Z","iopub.status.idle":"2023-02-27T10:45:01.061363Z","shell.execute_reply.started":"2023-02-27T10:45:01.048417Z","shell.execute_reply":"2023-02-27T10:45:01.060309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sum up the actual values\ncumulative_actual = np.cumsum(sorted_actual)\ncumulative_index = np.arange(1, len(cumulative_actual)+1)\ncumulative_actual","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:45:01.229824Z","iopub.execute_input":"2023-02-27T10:45:01.230253Z","iopub.status.idle":"2023-02-27T10:45:01.238263Z","shell.execute_reply.started":"2023-02-27T10:45:01.230197Z","shell.execute_reply":"2023-02-27T10:45:01.237186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cumulative_index","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:45:01.412669Z","iopub.execute_input":"2023-02-27T10:45:01.413230Z","iopub.status.idle":"2023-02-27T10:45:01.420201Z","shell.execute_reply.started":"2023-02-27T10:45:01.413197Z","shell.execute_reply":"2023-02-27T10:45:01.419287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nplt.plot(cumulative_index, cumulative_actual)\nplt.xlabel('Cumulative Number of Predictions')\nplt.ylabel('Cumulative Actual Values')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:45:01.625087Z","iopub.execute_input":"2023-02-27T10:45:01.625575Z","iopub.status.idle":"2023-02-27T10:45:01.786365Z","shell.execute_reply.started":"2023-02-27T10:45:01.625504Z","shell.execute_reply":"2023-02-27T10:45:01.784590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This corresponds to the Lorenz Curve in the diagram above. \n**We normalize both axes so that they go from 0 to 100% like in the economic figure** and display the 45° line to illustrate random guessing:\n\n\n\n","metadata":{}},{"cell_type":"code","source":"# We normalize both axes so that they go from 0 to 100% like in the economic figure and display the 45° line to illustrate random guessing:\ndf['lorentz'] = df['cum_pos_found'] / total_pos\ndf","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:45:02.018758Z","iopub.execute_input":"2023-02-27T10:45:02.020044Z","iopub.status.idle":"2023-02-27T10:45:02.036101Z","shell.execute_reply.started":"2023-02-27T10:45:02.020000Z","shell.execute_reply":"2023-02-27T10:45:02.034430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cumulative_actual_shares = cumulative_actual / sum(actual)\ncumulative_index_shares = cumulative_index / len(predictions)\n\n# Add (0, 0) to the plot\nx_values = [0] + list(cumulative_index_shares)\ny_values = [0] + list(cumulative_actual_shares)\n\n# Display the 45° line stacked on top of the y values\ndiagonal = [(x - y) for (x, y) in zip(x_values, y_values)]\n#diagonal.reverse()\nplt.stackplot(x_values, y_values, diagonal)\nplt.xlabel('Cumulative Share of Predictions')\nplt.ylabel('Cumulative Share of Actual Values')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T10:45:02.452524Z","iopub.execute_input":"2023-02-27T10:45:02.452875Z","iopub.status.idle":"2023-02-27T10:45:02.616825Z","shell.execute_reply.started":"2023-02-27T10:45:02.452846Z","shell.execute_reply":"2023-02-27T10:45:02.616024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Step 3: Weighted Gini Coefficient:  Calculate the reference line. Is it always a straight diagonal line?\n\nIf the distribution of positive and negative instances in the dataset is not random, then the diagonal line that represents perfect equality in the Lorenz curve may not be the appropriate reference line to use for calculating the Gini coefficient. In such cases, a different reference line may be used that represents the expected distribution of positive and negative instances based on prior knowledge or assumptions.\n\nFor example, if it is known that the proportion of positive instances in the population is 10%, then a reference line that represents this distribution can be used instead of the diagonal line. Similarly, if the dataset has been sampled from a stratified population with known proportions of positive and negative instances in each stratum, then a reference line that reflects these proportions can be used.\n\nThe Gini coefficient can still be calculated using the same formula, where A is the area under the reference line and B is the area between the Lorenz curve and the reference line. The choice of reference line depends on the specific context and assumptions of the problem, and may require some prior knowledge or domain expertise. It's important to note that the Gini coefficient is not necessarily sensitive to the choice of reference line, as long as the same reference line is used consistently across different models or datasets for comparability.\n\n***Why do we create a reference line using instance weights instead of the typical random line in the Gini coefficient calculation?***\n\nWe create a reference line using instance weights in the weighted Gini coefficient calculation because it allows us to take into account the imbalance in the dataset and to compare the performance of the model against a baseline that reflects this imbalance.\n\nIn a binary classification problem with an imbalanced dataset, the random line in the Lorenz curve would not be an appropriate baseline because it assumes that the positive and negative classes are equally important and have equal weights. This is not the case in an imbalanced dataset where the positive class is often more rare and more important than the negative class. For example, in a medical diagnosis problem, the positive class may represent a rare disease that requires urgent treatment, while the negative class represents the absence of the disease, which may not be as critical.\n\nBy creating a reference line using the instance weights, we can define a baseline that reflects the importance of each instance in the population and takes into account the imbalance in the dataset. This allows us to compare the performance of the model against a baseline that is more appropriate for the problem at hand and to quantify how much better the model is doing than random chance given the imbalance in the dataset.\n\nThe reference line is calculated as the cumulative sum of the weights, normalized by the total weight of the dataset. Keep in mind that the instance weights should be normalized so that they sum to 1, and that the reference line should be created by calculating the cumulative proportion of class weight for each instance rank based on the instance weights. ","metadata":{}},{"cell_type":"code","source":"df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\ndf['cumulative_index'] = df.index + 1\ndf","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:43:50.288221Z","iopub.execute_input":"2023-02-27T11:43:50.289374Z","iopub.status.idle":"2023-02-27T11:43:50.307525Z","shell.execute_reply.started":"2023-02-27T11:43:50.289327Z","shell.execute_reply":"2023-02-27T11:43:50.306218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(cumulative_index /len(predictions), df['random'] )\nplt.xlabel('Cumulative Number of Predictions')\nplt.ylabel('Cumulative Actual Values')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:43:51.018734Z","iopub.execute_input":"2023-02-27T11:43:51.019179Z","iopub.status.idle":"2023-02-27T11:43:51.181980Z","shell.execute_reply.started":"2023-02-27T11:43:51.019139Z","shell.execute_reply":"2023-02-27T11:43:51.179725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Step 4: Calculate a Weighted Gini coefficient\n\n\nCalculate the area between the Lorenz curve (which represents the cumulative distribution of the predicted probabilities) and the line of perfect equality (which represents a random classifier)\n\n\n**BUT** A weighted Gini coefficient is a variation of the Gini coefficient that takes into account the different weights or importance of different observations or groups. It is often used in situations where the data is not equally distributed, and some groups or observations are more important than others.\n\nIn the context of credit risk modeling, for example, the weighted Gini coefficient can be used to evaluate the performance of a model that predicts the likelihood of default for a portfolio of loans, where some loans are more valuable or riskier than others. In this case, the Gini coefficient is calculated by weighting the contribution of each loan to the overall Gini coefficient by its value or risk weight.\n\nThe formula for the weighted Gini coefficient is similar to the regular Gini coefficient, but includes a weight factor:\n\n**The weighted Gini coefficient takes into account both the proportion of positive and negative instances in the dataset and the prediction performance of the model. It is calculated as the area between the ROC curve (Receiver Operating Characteristic curve) of the model and the diagonal line that represents random guessing, multiplied by a weight factor that reflects the relative importance of the positive and negative classes.**\n\nThe weighted Gini coefficient is calculated by first constructing the Lorenz curve using the instance weights and the predicted class probabilities or scores from the model. Then, the Gini coefficient is calculated based on the area between the Lorenz curve and the reference line, where the reference line represents the cumulative proportion of positive class weight under the assumption of random classification.\n\nThe calculation of the weighted Gini coefficient involves integrating the area between the Lorenz curve and the reference line, with the weights of each instance factored in. This means that the weighted Gini coefficient reflects the performance of the model while taking into account the importance of each instance in the population.\n\n\nThe weighted Gini coefficient ranges from 0 to 1, where a value of 0 indicates perfect equality (all observations have the same weight) and a value of 1 indicates perfect inequality (one observation has all the weight).","metadata":{}},{"cell_type":"code","source":"df['gini'] = (df['lorentz'] - df['random']) * df['weight']\ndf","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:55:33.237467Z","iopub.execute_input":"2023-02-27T11:55:33.238373Z","iopub.status.idle":"2023-02-27T11:55:33.256507Z","shell.execute_reply.started":"2023-02-27T11:55:33.238322Z","shell.execute_reply":"2023-02-27T11:55:33.255206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['gini'].sum()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T11:56:25.364461Z","iopub.execute_input":"2023-02-27T11:56:25.364900Z","iopub.status.idle":"2023-02-27T11:56:25.374917Z","shell.execute_reply.started":"2023-02-27T11:56:25.364855Z","shell.execute_reply":"2023-02-27T11:56:25.373331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Step 5: Normalize\n\nFinally, normalize the actual Gini coefficient by dividing it by the maximum possible Gini coefficient:\n\nnormalized Gini coefficient = G / G_max\n\nA higher normalized Gini coefficient indicates a better performance of the model in discriminating between the positive and negative classes, taking into account the relative importance of each class.\n\n\n\n","metadata":{}},{"cell_type":"code","source":"# Code Provided by the competition hosts\n# This is a python version of the metric for the Amex competition.\ndef amex_metric(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n\n    def top_four_percent_captured(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n        \n    def weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        \n        # Instead of going through the population from poorest to richest, we go through our predictions from lowest to highest.?? (other way round)\n        # Sort the actual values by the predictions\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        \n        # We assign the weight to the majority class (0) that we subsample\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        \n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        \n        # Sum up the targets\n        total_pos = (df['target'] * df['weight']).sum()\n        \n         # Instead of summing up the income, we sum up the actual values of our predictions:\n        # Sum up the actual values\n        # This corresponds to the Lorenz Curve\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        \n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n\n    def normalized_weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        y_true_pred = y_true.rename(columns={'target': 'prediction'})\n        return weighted_gini(y_true, y_pred) / weighted_gini(y_true, y_true_pred)\n\n    g = normalized_weighted_gini(y_true, y_pred)\n    d = top_four_percent_captured(y_true, y_pred)\n\n    return 0.5 * (g + d)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nThe resulting value is the Gini coefficient, which ranges from 0 to 1. A Gini coefficient of 0 indicates a random classifier, while a Gini coefficient of 1 indicates a perfect classifier.\n\nThe Gini coefficient is similar to the area under the receiver operating characteristic curve (ROC AUC), which is another commonly used evaluation metric for binary classification models. However, the Gini coefficient is often preferred in cases where the positive class is rare, as it is more sensitive to the performance of the model on the positive class.\n\nThe major disadvantage of the Gini coefficient comes from the fact that it is an ordinal metric, i.e., it captures the order of values while ignoring the distance between them. This characteristic of the Gini coefficient can sometimes mask poor model performances.\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:24px;text-align:left;display:fill;border-radius:5px;background-color:#3A578A;overflow:hidden\">Work In Progress</div>","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}