{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport matplotlib.pyplot as plt\nimport random \nfrom numpy import dtype\nimport matplotlib.pyplot as plt\nfrom pandas import DataFrame\nimport seaborn as sns\nfrom tqdm import tqdm\nfrom pathlib import Path\n\n#ignore warning messages \nimport warnings\nwarnings.filterwarnings('ignore') \nimport gc\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n        \n# Seed all\nseed = 12\nrandom.seed(seed)\nos.environ[\"PYTHONHASHSEED\"] = str(seed)\nnp.random.seed(seed)\ngc.collect()\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-25T05:55:05.593411Z","iopub.execute_input":"2023-01-25T05:55:05.593939Z","iopub.status.idle":"2023-01-25T05:55:07.054794Z","shell.execute_reply.started":"2023-01-25T05:55:05.593837Z","shell.execute_reply":"2023-01-25T05:55:07.053632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# AutoML on Playground Series Season 3, Episode 4\nThe Playground Series Season 3, Episode 4 is the chance to practice our skills on classification problems. The dataset for this competition (both train and test) was generated from a deep learning model trained on a [Credit Card Fraud Detection](https://www.kaggle.com/datasets/mlg-ulb/creditcardfraud). Feature distributions are close to, but not exactly the same, as the original. \n\nWe are going to use AutoViz and AutoML to explore and choose the best ML algorithms in order to use in this competition. Let's go!\n\n<img src=\"https://storage.googleapis.com/kaggle-competitions/kaggle/44632/logos/header.png?t=2022-12-21-20-51-50\" style=\"height:180px\"/>","metadata":{}},{"cell_type":"markdown","source":"### Premise\n*In this notebook, I would like to experiment with some of the most famous AutoML libraries: AutoViz and PyCaret\nIn particular, I would like to understand what results can be obtained by exploiting fully automatic libraries for EDA analysis and for the automatic training even of complex ensembling models.*\n\nEnjoy the reading!","metadata":{}},{"cell_type":"markdown","source":"\n# AutoML\n\n<img src=\"https://i.postimg.cc/RCDnLrpJ/automl-autoviz.png\" style=\"width: 200px;float: right;margin: 20px;\"/>\n<img src=\"https://i.postimg.cc/2SpbgQDz/automl-pycaret.png\" style=\"width: 200px;float: right;margin: 20px;\"/>\n\nAutomated Machine learning (AutoML) is the process of automating the tasks of applying machine learning to real-world problems. AutoML potentially includes every stage, from starting with auto exploration data analysis to building a machine learning model ready for deployment.\nIn this notebook we are going to see two of this tools:\n- **AutoViz** for automated exploration data analysis\n- **PyCaret** for compare, build and tune ML models\n\n","metadata":{}},{"cell_type":"markdown","source":"# AutoViz\nAutoViz performs automatic visualization of any dataset with one line of code. Give it any input file (CSV, txt or json format) or a dataframe and AutoViz will visualize it","metadata":{}},{"cell_type":"markdown","source":"# AutoViz - Installation","metadata":{}},{"cell_type":"code","source":"!pip install autoviz\n!pip install autoviz --upgrade ","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-25T05:55:07.057037Z","iopub.execute_input":"2023-01-25T05:55:07.058222Z","iopub.status.idle":"2023-01-25T05:55:46.888889Z","shell.execute_reply.started":"2023-01-25T05:55:07.058162Z","shell.execute_reply":"2023-01-25T05:55:46.887425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA - Exploration Data Analysis with AutoViz","metadata":{}},{"cell_type":"code","source":"path = Path('/kaggle/input/playground-series-s3e4/')\nsubmission = pd.read_csv(path / 'sample_submission.csv', index_col='id')\n\n# adding Credit Cart Fraud Detection datasource (extra)\nccf = pd.read_csv(\"/kaggle/input/creditcardfraud/creditcard.csv\")\nccf = ccf.query('Class==1')  # we include only positive examples\n\n# playground dataset\nX_train = pd.read_csv(path / 'train.csv', index_col='id')\nX_train = pd.concat([X_train, ccf],axis=0)\n\nX_test = pd.read_csv(path / 'test.csv', index_col='id')\nTARGET = 'Class'\ny_train = X_train[TARGET]\nX_train\n","metadata":{"execution":{"iopub.status.busy":"2023-01-25T05:56:06.901545Z","iopub.execute_input":"2023-01-25T05:56:06.901967Z","iopub.status.idle":"2023-01-25T05:56:18.632792Z","shell.execute_reply.started":"2023-01-25T05:56:06.901934Z","shell.execute_reply":"2023-01-25T05:56:18.630836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train","metadata":{"execution":{"iopub.status.busy":"2023-01-22T06:05:30.169207Z","iopub.execute_input":"2023-01-22T06:05:30.169635Z","iopub.status.idle":"2023-01-22T06:05:30.202005Z","shell.execute_reply.started":"2023-01-22T06:05:30.169601Z","shell.execute_reply":"2023-01-22T06:05:30.200562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from autoviz.AutoViz_Class import AutoViz_Class\n%matplotlib inline\n\nAV = AutoViz_Class()\nAV.AutoViz(filename='',\n          dfte=X_train,\n          depVar=TARGET,\n          verbose=1,\n          lowess = False,\n          chart_format='png')\n","metadata":{"execution":{"iopub.status.busy":"2023-01-24T06:24:04.219839Z","iopub.execute_input":"2023-01-24T06:24:04.221167Z","iopub.status.idle":"2023-01-24T06:31:12.513905Z","shell.execute_reply.started":"2023-01-24T06:24:04.221109Z","shell.execute_reply":"2023-01-24T06:31:12.512265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note:\ndepVar - target variable in your dataset. You can leave it as empty string if you don't have a target variable in your data.","metadata":{}},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-01-20T06:51:02.434703Z","iopub.execute_input":"2023-01-20T06:51:02.435191Z","iopub.status.idle":"2023-01-20T06:51:03.236255Z","shell.execute_reply.started":"2023-01-20T06:51:02.435146Z","shell.execute_reply":"2023-01-20T06:51:03.235352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# PyCaret\n\nPyCaret is an open-source, low-code machine learning library in Python that automates machine learning workflows.\nIn this notebook we are going to compare several models in order to see who is the best performer.","metadata":{}},{"cell_type":"markdown","source":"# PyCaret - Installation","metadata":{}},{"cell_type":"code","source":"! pip install --ignore-installed --pre pycaret==3.0.0rc4","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-22T06:05:45.050046Z","iopub.execute_input":"2023-01-22T06:05:45.050467Z","iopub.status.idle":"2023-01-22T06:08:20.217826Z","shell.execute_reply.started":"2023-01-22T06:05:45.050434Z","shell.execute_reply":"2023-01-22T06:08:20.216388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pycaret.classification import *\nfrom sklearn.metrics import log_loss","metadata":{"execution":{"iopub.status.busy":"2023-01-22T06:08:24.551298Z","iopub.execute_input":"2023-01-22T06:08:24.551751Z","iopub.status.idle":"2023-01-22T06:08:29.865420Z","shell.execute_reply.started":"2023-01-22T06:08:24.551712Z","shell.execute_reply":"2023-01-22T06:08:29.863999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# PyCaret - Understand the Workflow\n\nThe following schema describe a generic machine learning workflow usic pycaret\n\n![https://i.postimg.cc/dt34kVwL/pycaret-flow-drawio.png](https://i.postimg.cc/dt34kVwL/pycaret-flow-drawio.png)","metadata":{}},{"cell_type":"markdown","source":"# PyCaret - Configuration\nPyCaret trains and evaluates the performance of many models using many metrics and cross-validation methods\n\nThe setup() function initializes the environment in pycaret and creates the transformation pipeline to prepare the data for modeling and deployment. setup() must be called before executing any other function in pycaret. It takes two mandatory parameters: a pandas dataframe and the name of the target column. All other parameters are optional and are used to customize the pre-processing pipeline (we will see them in later tutorials).\n\n\n\n## Why we need to try many models?\n### => Because there isn't a Golden approach that works for all the problems in data science, so let's go! \"We'll Try all that We can try!\"  \n\n","metadata":{}},{"cell_type":"markdown","source":"# Manual Feature Engineering? No Thanks\n\nPyCaret has several modules activable in the setup function (see advance usage below) in order to perform several feature engineering e data preparation tasks\n\n- **Data preparation**: this module impute NaN and blank data; PyCaret by default imputes the missing value in the dataset by mean for numeric features and constant for categorical features. To change the imputation method, numeric_imputation and categorical_imputation parameters can be used within the setup;\n\n- **Scale and Transform**:There are several methods available for normalization: z-score, minmax, maxabs, robust;  by default, PyCaret uses zscore.\n\n- **Feature Engineering**: Creating a new feature through the interaction of existing features is known as *feature interaction*. It can be achieved in PyCaret using feature_interaction and feature_ratio parameters within setup. Feature interaction creates new features by multiplying two variables (a * b), while feature ratios create new features but by calculating the ratios of existing features (a / b)\n\n- **Feature Selection**: Feature Importance is a process used to select features in the dataset that contributes the most in predicting the target variable. Working with selected features instead of all the features reduces the risk of over-fitting, improves accuracy, and decreases the training time. In PyCaret, this can be achieved using feature_selection parameter. It uses a combination of several supervised feature selection techniques to select the subset of features that are most important for modeling.\n","metadata":{}},{"cell_type":"markdown","source":"## Advanced setup using other optional params and modules\n\n```\n    setup(data=train, \n          target=target,\n          session_id=seed, # our seed\n          normalize = True, # When set to True, the feature space is transformed using the method defined under the normalized_method parameter. default is zscore\n          transformation = True, # When set to True, a power transformer is applied to make the data more normal / Gaussian-like.\n          feature_interaction = True, # When set to True, it will create new features by interacting (a * b) for all numeric variables in the dataset including polynomial and trigonometric features\n          polynomial_features = True, # When set to True, new features are created based on all polynomial combinations that exist within the numeric features in a dataset to the degree defined in polynomial_degree param\n          feature_selection = True, # When set to True, a subset of features are selected using a combination of various permutation importance techniques including Random Forest, Adaboost and Linear correlation with target variable.\n          feature_selection_threshold = 0.8 # Threshold used for feature selection (including newly created polynomial features),\n          fix_imbalance = True # fixing imbvalanced dataset\n          remove_outliers = True # remove outlier for training\n         )\n```","metadata":{}},{"cell_type":"code","source":"def pycaret_models_evaluation(train, target,test, n_select, fold, opt):\n    print('Setup Your Data....')\n    # Note: through setup function we could pass many other parameters in order to activare some features like PCA, specific feature engineering and so on\n    setup(data=train, \n          target=target,\n          session_id=seed, # our seed\n          fix_imbalance = True,\n          normalize = True,\n          remove_outliers = True\n         )   \n    ## Adding extra metrics (optional)\n    # **Example**: \"log loss\" metric is not available by default, so we are going to add this metric manually with *add_metric* function\n    add_metric('logloss', 'Log Loss', log_loss, greater_is_better = False)\n    print('Comparing Models....')\n    # to include only a sub-set of availables models use include = [...]\n    # best = compare_models(sort=opt, n_select=n_select, fold=fold, include=['et', 'lightgbm'])\n    best = compare_models(sort=opt, n_select=n_select, fold=fold)","metadata":{"execution":{"iopub.status.busy":"2023-01-22T06:17:14.948225Z","iopub.execute_input":"2023-01-22T06:17:14.949338Z","iopub.status.idle":"2023-01-22T06:17:14.956377Z","shell.execute_reply.started":"2023-01-22T06:17:14.949298Z","shell.execute_reply":"2023-01-22T06:17:14.955323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" X_test.columns","metadata":{"execution":{"iopub.status.busy":"2023-01-22T06:14:06.662321Z","iopub.execute_input":"2023-01-22T06:14:06.662732Z","iopub.status.idle":"2023-01-22T06:14:06.671042Z","shell.execute_reply.started":"2023-01-22T06:14:06.662699Z","shell.execute_reply":"2023-01-22T06:14:06.669845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DROP_FEATURES = []\n# using selected features by autoviz (optional)\nSELECTED_FEATURES_AUTOVIZ = list(X_test.columns)","metadata":{"execution":{"iopub.status.busy":"2023-01-24T06:31:13.280353Z","iopub.execute_input":"2023-01-24T06:31:13.280776Z","iopub.status.idle":"2023-01-24T06:31:13.286300Z","shell.execute_reply.started":"2023-01-24T06:31:13.280740Z","shell.execute_reply":"2023-01-24T06:31:13.285159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def feature_engineering(df):\n    # no feature engineering\n    df = df.drop(DROP_FEATURES, axis=1, errors=\"ignore\")    \n    return df[SELECTED_FEATURES_AUTOVIZ]\n\nX_train = feature_engineering(X_train)\nX_test = feature_engineering(X_test)\nX_train.columns","metadata":{"execution":{"iopub.status.busy":"2023-01-22T06:20:05.828582Z","iopub.execute_input":"2023-01-22T06:20:05.829052Z","iopub.status.idle":"2023-01-22T06:20:05.843038Z","shell.execute_reply.started":"2023-01-22T06:20:05.829016Z","shell.execute_reply":"2023-01-22T06:20:05.840958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# How to use the GPU with PyCaret\n\nSetting **use_gpu** param to True, it will use GPU for training with algorithms that support it and fall back to CPU if they are unavailable. When set to force it will only use GPU-enabled algorithms and raise exceptions when they are unavailable. When False all algorithms are trained using CPU only.\n\n# How to reduce Dimensionality of our dataset with PCA\n**Principal Component Analysis (PCA)** is an unsupervised technique used in machine learning to reduce the dimensionality of data. It does so by compressing the feature space by identifying a subspace that captures most of the information in the complete feature matrix. This can be achieved in PyCaret using pca parameter within setup.\n\n- **pca**: default = False; When set to True, dimensionality reduction is applied to the data using the method defined in pca_method param. Note that not all datasets can be decomposed efficiently using a linear PCA technique and that apply\n- **pca_method**: default = ‘*linear*’; The ‘linear’ method performs Linear dimensionality reduction using Singular Value Decomposition. The other available options are *kernel* and *incremental* (replacement for ‘linear’ pca when the dataset to be decomposed is too large to fit in memory)\n- **pca_components**: default = 0.99; Number of components to keep. if pca_components is a float, it is treated as a target percentage for information retention. When pca_components is an integer it is treated as the number of features to be kept. pca_components must be strictly less than the original number of features in the dataset.\n\n# How to use several cross-validation methods\nIn our setup() function we can specify some model parameters that can be used for setting parameters for model selection process. These are not related to data preprocessing but can influence your model selection process. \n\nThe cross-validation strategy and data-split params ore some of them:\n\n- **train_size**: float, default = 0.7 is the proportion of the dataset to be used for training and validation. \n- **test_data**: pandas.DataFrame, default = None If not None, the test_data is used as a hold-out set and the train_size is ignored. test_data must be labeled and the shape of the data and test_data must match.\n- **data_split_shuffle**: bool, default = True When set to False, prevents shuffling of rows during train_test_split.\n- **data_split_stratify**: bool or list, default = False Controls stratification during the train_test_split. When set to True, it will stratify by target column. To stratify on any other columns, pass a list of column names. Ignored when data_split_shuffle is False.\n- **fold_strategy**: str or scikit-learn CV generator object, default = ‘stratifiedkfold’ Choice of cross-validation strategy. Possible values are: ‘kfold’, ‘stratifiedkfold’, ‘groupkfold’, ‘timeseries’, a custom CV generator object compatible with scikit-learn.\n- **fold**: int, default = 10 The number of folds to be used in cross-validation. Must be at least 2. This is a global setting that can be over-written at the function level by using the fold parameter. Ignored when fold_strategy is a custom object.\n- **fold_shuffle**: bool, default = False Controls the shuffle parameter of CV. Only applicable when fold_strategy is kfold or stratifiedkfold. Ignored when fold_strategy is a custom object.\n- **fold_groups**: str or array-like, with shape (n_samples,), default = None Optional group labels when ‘GroupKFold’ is used for the cross-validation. It takes an array with shape (n_samples, ) where n_samples is the number of rows in the training dataset. When the string is passed, it is interpreted as the column name in the dataset containing group labels.\n\n### Note\nFor the moment we are going to use the default params of PyCaret","metadata":{}},{"cell_type":"markdown","source":"# PyCaret - Models Comparison Results","metadata":{}},{"cell_type":"code","source":"# train, target, test, 10: number of models, 5 number of folds, metric\npycaret_models_evaluation(X_train, y_train, X_test, 10, 5, 'AUC')","metadata":{"execution":{"iopub.status.busy":"2023-01-22T06:20:12.174267Z","iopub.execute_input":"2023-01-22T06:20:12.174707Z","iopub.status.idle":"2023-01-22T06:22:29.333565Z","shell.execute_reply.started":"2023-01-22T06:20:12.174672Z","shell.execute_reply":"2023-01-22T06:22:29.332279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# PyCaret - Single Model Tuning","metadata":{}},{"cell_type":"code","source":"lr_model = create_model('lr', fold = 5)\ngbm_model = create_model('lightgbm', fold = 5)\ntuned_gbm_model = tune_model(gbm_model, optimize = 'AUC')\ntuned_lr_model = tune_model(lr_model, optimize = 'AUC')","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-20T07:25:34.791069Z","iopub.execute_input":"2023-01-20T07:25:34.791510Z","iopub.status.idle":"2023-01-20T07:27:34.890312Z","shell.execute_reply.started":"2023-01-20T07:25:34.791471Z","shell.execute_reply":"2023-01-20T07:27:34.889073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Sklearn model configurations for tuned models\n\n### Note:\nPrinting a PyCaret model give us the ability to see which are the tuned parameter found by PyCaret so we can use them in other project/notebooks. \n\nMoreover, PyCaret uses sklearn under the hood so the transformations or imputations used can be inspected and printed","metadata":{}},{"cell_type":"code","source":"print(tuned_gbm_model)\nprint(tuned_lr_model)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T07:27:34.895215Z","iopub.execute_input":"2023-01-20T07:27:34.895598Z","iopub.status.idle":"2023-01-20T07:27:34.902393Z","shell.execute_reply.started":"2023-01-20T07:27:34.895562Z","shell.execute_reply":"2023-01-20T07:27:34.901277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluating Single Models","metadata":{}},{"cell_type":"code","source":"evaluate_model(tuned_lr_model)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T07:27:34.903994Z","iopub.execute_input":"2023-01-20T07:27:34.904705Z","iopub.status.idle":"2023-01-20T07:27:35.554396Z","shell.execute_reply.started":"2023-01-20T07:27:34.904667Z","shell.execute_reply":"2023-01-20T07:27:35.553198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Note:\nevaluate_model can only be used in Notebook since it uses ipywidget.\n\nYou can also use the **plot_model** function to generate plots individually. \nplot_model(my_model, plot = 'elbow')\n\n### Plots availables for Classification problems:\n- Area Under the Curve: ‘auc’\n- Discrimination Threshold: ‘threshold’\n- Precision Recall Curve: ‘pr’\n- Confusion Matrix: ‘confusion_matrix’\n- Class Prediction Error: ‘error’\n- Classification Report: ‘class_report’\n- Decision Boundary: ‘boundary’\n- Recursive Feature Selection: ‘rfe’\n- Learning Curve: ‘learning’\n- Manifold Learning: ‘manifold’\n- Calibration Curve: ‘calibration’\n- Validation Curve: ‘vc’\n- Dimension Learning: ‘dimension’\n- Feature Importance (Top 10): ‘feature’\n- Feature IImportance (all): 'feature_all'\n- Model Hyperparameter: ‘parameter’\n- Lift Curve: 'lift'\n- Gain Curve: 'gain'\n- KS Statistic Plot: 'ks'\n\n\n### Plots availables for Regression problems:\n- Residuals Plot: ‘residuals’\n- Prediction Error Plot: ‘error’\n- Cooks Distance Plot: ‘cooks’\n- Recursive Feature Selection: ‘rfe’\n- Learning Curve: ‘learning’\n- Validation Curve: ‘vc’\n- Manifold Learning: ‘manifold’\n- Feature Importance (top 10): ‘feature’\n- Feature Importance (all): 'feature_all'\n- Model Hyperparameter: ‘parameter’\n\n### Plots availables for Clusterization Problems\n- Cluster PCA Plot (2d): ‘cluster’\n- Cluster TSnE (3d): ‘tsne’\n- Elbow Plot: ‘elbow’\n- Silhouette Plot: ‘silhouette’\n- Distance Plot: ‘distance’\n- Distribution Plot: ‘distribution’\n\n","metadata":{}},{"cell_type":"code","source":"plot_model(tuned_gbm_model, plot = 'auc')\nplot_model(tuned_gbm_model, plot = 'confusion_matrix')\nplot_model(tuned_gbm_model, plot = 'pr')\nplot_model(tuned_gbm_model, plot = 'error')\nplot_model(tuned_gbm_model, plot = 'class_report')\nplot_model(tuned_gbm_model, plot = 'boundary')\nplot_model(tuned_gbm_model, plot = 'feature')\nplot_model(tuned_gbm_model, plot = 'parameter')","metadata":{"execution":{"iopub.status.busy":"2023-01-20T07:27:35.556132Z","iopub.execute_input":"2023-01-20T07:27:35.556945Z","iopub.status.idle":"2023-01-20T07:27:45.935804Z","shell.execute_reply.started":"2023-01-20T07:27:35.556896Z","shell.execute_reply":"2023-01-20T07:27:45.934426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# LGBM Model Interpretation","metadata":{}},{"cell_type":"code","source":"interpret_model(tuned_gbm_model)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T07:27:45.937331Z","iopub.execute_input":"2023-01-20T07:27:45.937722Z","iopub.status.idle":"2023-01-20T07:27:47.621946Z","shell.execute_reply.started":"2023-01-20T07:27:45.937687Z","shell.execute_reply":"2023-01-20T07:27:47.620788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ensembling Models (Bagging vs Boosting)\n\n## Bagging\nBagging, also known as Bootstrap aggregating, is a machine learning ensemble meta-algorithm designed to improve the stability and accuracy of machine learning algorithms used in statistical classification and regression. It also reduces variance and helps to avoid overfitting\n\n## Boosting\nBoosting is an ensemble meta-algorithm for primarily reducing bias and variance in supervised learning. Boosting is in the family of machine learning algorithms that convert weak learners to strong ones. \n\n<img src=\"https://i.postimg.cc/FR2hCQfD/ensemble-bagging-boosting.png\"/>\n","metadata":{}},{"cell_type":"code","source":"# ensemble each model with bagging\nbagged_lr_model = ensemble_model(tuned_lr_model, method='Bagging', n_estimators=3)\nbagged_gbm_model = ensemble_model(tuned_gbm_model, method='Bagging', n_estimators=3)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-20T07:27:47.623442Z","iopub.execute_input":"2023-01-20T07:27:47.623808Z","iopub.status.idle":"2023-01-20T07:42:22.829435Z","shell.execute_reply.started":"2023-01-20T07:27:47.623775Z","shell.execute_reply":"2023-01-20T07:42:22.827173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Blending previous models together\n\nblend_models function trains a Soft Voting or a Majority Rule classifier (Hard Voting) for select models passed in the estimator_list parameter.\n\nOptionally we can add the weight for each model passed\n\n## Soft Voting\n\n<img src=\"https://i.postimg.cc/fTxK9Hzd/voting-soft-drawio.png\"/>\n\n## Hard Voting\n\n<img src=\"https://i.postimg.cc/jdVcCVbK/voting-hard-drawio.png\"/>\n","metadata":{}},{"cell_type":"code","source":"blender = blend_models(estimator_list=[bagged_gbm_model, bagged_lr_model], method='soft', weights = [0.5, 0.5])","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-20T07:42:22.830473Z","iopub.status.idle":"2023-01-20T07:42:22.830872Z","shell.execute_reply.started":"2023-01-20T07:42:22.830680Z","shell.execute_reply":"2023-01-20T07:42:22.830698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# tune the model\ntuned_blendr = tune_model(blender, optimize = 'AUC')","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-20T07:42:22.832209Z","iopub.status.idle":"2023-01-20T07:42:22.832579Z","shell.execute_reply.started":"2023-01-20T07:42:22.832396Z","shell.execute_reply":"2023-01-20T07:42:22.832414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(tuned_blendr)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T07:42:22.833667Z","iopub.status.idle":"2023-01-20T07:42:22.834400Z","shell.execute_reply.started":"2023-01-20T07:42:22.834191Z","shell.execute_reply":"2023-01-20T07:42:22.834212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Calibrating model","metadata":{}},{"cell_type":"code","source":"#calibrated_dt = calibrate_model(tuned_blendr)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-01-20T07:42:22.835565Z","iopub.status.idle":"2023-01-20T07:42:22.835933Z","shell.execute_reply.started":"2023-01-20T07:42:22.835747Z","shell.execute_reply":"2023-01-20T07:42:22.835764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Final model\n\nIt will be trained with all train data","metadata":{}},{"cell_type":"code","source":"final_model = finalize_model(tuned_blendr)\n#final_model = finalize_model(calibrated_dt)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T07:42:22.836826Z","iopub.status.idle":"2023-01-20T07:42:22.837689Z","shell.execute_reply.started":"2023-01-20T07:42:22.837486Z","shell.execute_reply":"2023-01-20T07:42:22.837507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(final_model)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T07:42:22.838856Z","iopub.status.idle":"2023-01-20T07:42:22.839244Z","shell.execute_reply.started":"2023-01-20T07:42:22.839061Z","shell.execute_reply":"2023-01-20T07:42:22.839079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"result = predict_model(final_model, data=X_test, raw_score=True)\nresult","metadata":{"execution":{"iopub.status.busy":"2023-01-20T07:42:22.840396Z","iopub.status.idle":"2023-01-20T07:42:22.841110Z","shell.execute_reply.started":"2023-01-20T07:42:22.840896Z","shell.execute_reply":"2023-01-20T07:42:22.840916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nsubmission[TARGET] = result['prediction_score_1'].to_numpy()\nsubmission.to_csv(\"submission.csv\")\nsubmission","metadata":{"execution":{"iopub.status.busy":"2023-01-20T07:42:22.842387Z","iopub.status.idle":"2023-01-20T07:42:22.842752Z","shell.execute_reply.started":"2023-01-20T07:42:22.842572Z","shell.execute_reply":"2023-01-20T07:42:22.842589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"-------------------------------------------------------------------------\n<div style=\"text-align: center;\">\n    <h3>Thanks for watching till the end ;)  </h3>\n    <h2>If you liked this notebook, upvote it ! </h2>\n<img src=\"https://i.postimg.cc/SsChkSJv/upvote.png\" width=\"70\"/>\n</div>\n\n\n\n\n---\n","metadata":{}}]}