{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Tabular Playground Series - Aug 2022\n\n**About the data**\nThis data represents the results of a large product testing study. For each product_code you are given a number of product attributes (fixed for the code) as well as a number of measurement values for each individual product, representing various lab testing methods. Each product is used in a simulated real-world environment experiment, and and absorbs a certain amount of fluid (loading) to see whether or not it fails.\n\nYour task is to use the data to predict individual product failures of new codes with their individual lab test results.\n\n**About the notebook**\nIn this notbook, the `Poisson Regressor` model will be used to classify the target value based on the categories and numerical variables. The following are the steps present within this notebook:\n\n1. Installing Peripheral libraries\n    * Installing dataprep package for visualization and EDA\n2. Importing Necessary Libraries\n3. Import the dataset\n4. Exploratory Data Analysis\n    * Visulize the correlation\n    * Plot statistics for the failure column of the train data\n    * Report generated for the train data\n    * Train and test data comparison\n5. Data Preprocessing\n    * Feature Engineering the number code\n    * Encode categorical dimensions using `LabelEncoder()`\n    * Fill missing values using `KNNImputer()`\n    * Scale data using `StandardScaler()`\n6. Finding the best parameters using Grid Search\n7. Submit predictions","metadata":{}},{"cell_type":"markdown","source":"# Installing peripheral libraries","metadata":{}},{"cell_type":"code","source":"# !pip install dataprep","metadata":{"id":"y3fb9LSAZCfB","outputId":"69ea9858-0a92-43b0-8481-51cbb1744550","execution":{"iopub.status.busy":"2022-08-04T10:11:40.870371Z","iopub.execute_input":"2022-08-04T10:11:40.870790Z","iopub.status.idle":"2022-08-04T10:11:40.876036Z","shell.execute_reply.started":"2022-08-04T10:11:40.870758Z","shell.execute_reply":"2022-08-04T10:11:40.874815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing Necessary Libraries","metadata":{"id":"y4FRSsTjZI-5"}},{"cell_type":"code","source":"# Data Wrangling libraries\nimport numpy as np\nimport pandas as pd\n\n# Visualization Libraries\nfrom IPython.display import display,HTML\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Preprocessing libraries\nfrom sklearn.impute import KNNImputer\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.preprocessing import StandardScaler\n\n# Machine Learning Estimators\nfrom sklearn.linear_model import PoissonRegressor\nfrom sklearn.model_selection import GridSearchCV\n\n# Metrics\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.metrics import accuracy_score, precision_recall_fscore_support\n\n# As per this notebook, we disabled dataprep viz due to viewing errors.\n# You could uncomment the code upon forking to use the module\n# from dataprep.eda import plot, plot_correlation, create_report, plot_missing, plot_diff\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"id":"ZJU80J7dybwB","execution":{"iopub.status.busy":"2022-08-04T10:11:40.877821Z","iopub.execute_input":"2022-08-04T10:11:40.878167Z","iopub.status.idle":"2022-08-04T10:11:40.887902Z","shell.execute_reply.started":"2022-08-04T10:11:40.878134Z","shell.execute_reply":"2022-08-04T10:11:40.886751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import the datasets\n\nData wasa taken from Kaggle's Tabular Playground Series - Aug 2022.","metadata":{"id":"gPXTQMqPae2n"}},{"cell_type":"code","source":"train_data = pd.read_csv('../input/tabular-playground-series-aug-2022/train.csv')\ntest_data = pd.read_csv('../input/tabular-playground-series-aug-2022/test.csv') ","metadata":{"id":"oTTDVFRO0tgK","execution":{"iopub.status.busy":"2022-08-04T10:11:40.889850Z","iopub.execute_input":"2022-08-04T10:11:40.890717Z","iopub.status.idle":"2022-08-04T10:11:41.054141Z","shell.execute_reply.started":"2022-08-04T10:11:40.890672Z","shell.execute_reply":"2022-08-04T10:11:41.052922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory Data Analysis\n\nThe following are interactive visualizations as provided by [dataprep.ai](https://dataprep.ai/). You can visit their website for more information. In this Notebook, this would be the primary tool for the exploratory data analysis. The outputs are interactive and clickable. Therefore, you could view different graphs and analysis by hovering on the columns and figures.","metadata":{"id":"oPfA2hh3Zk2R"}},{"cell_type":"code","source":"# Visulize the correlation\n# plot_correlation(train_data)","metadata":{"id":"FAfLFaijaX_f","outputId":"bbd43b21-763f-4be4-c7b3-0380a5d1083a","execution":{"iopub.status.busy":"2022-08-04T10:11:41.057112Z","iopub.execute_input":"2022-08-04T10:11:41.058157Z","iopub.status.idle":"2022-08-04T10:11:41.062746Z","shell.execute_reply.started":"2022-08-04T10:11:41.058115Z","shell.execute_reply":"2022-08-04T10:11:41.061432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot statistics for the failure column of the train data\n# plot(train_data, 'failure')","metadata":{"id":"jJZc2Ozqas5f","outputId":"903b9235-3ab2-4f79-e4f2-58035a061fbd","execution":{"iopub.status.busy":"2022-08-04T10:11:41.064849Z","iopub.execute_input":"2022-08-04T10:11:41.065767Z","iopub.status.idle":"2022-08-04T10:11:41.076373Z","shell.execute_reply.started":"2022-08-04T10:11:41.065721Z","shell.execute_reply":"2022-08-04T10:11:41.075448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Report generated for the train data\n# Note: This code make take a couple of minutes depending on the speed of your device.\n# create_report(train_data)","metadata":{"id":"0qjPM4rVb15n","outputId":"d2432f8c-13f0-4f0c-ecdd-543e5dcaf5bc","execution":{"iopub.status.busy":"2022-08-04T10:11:41.077981Z","iopub.execute_input":"2022-08-04T10:11:41.078481Z","iopub.status.idle":"2022-08-04T10:11:41.085593Z","shell.execute_reply.started":"2022-08-04T10:11:41.078436Z","shell.execute_reply":"2022-08-04T10:11:41.084782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# View missing data\n# plot_missing(train_data)","metadata":{"id":"UXVAgx8nb8yr","outputId":"cfad177d-b563-4f00-9313-5cf98f339906","execution":{"iopub.status.busy":"2022-08-04T10:11:41.087074Z","iopub.execute_input":"2022-08-04T10:11:41.087751Z","iopub.status.idle":"2022-08-04T10:11:41.096042Z","shell.execute_reply.started":"2022-08-04T10:11:41.087709Z","shell.execute_reply":"2022-08-04T10:11:41.095223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train and test data comparison\n# plot_diff([train_data, test_data])","metadata":{"id":"JBX-hUHvcHhN","outputId":"2aa0b8fd-dd27-42c9-b168-562ab78e7f29","execution":{"iopub.status.busy":"2022-08-04T10:11:41.097303Z","iopub.execute_input":"2022-08-04T10:11:41.097827Z","iopub.status.idle":"2022-08-04T10:11:41.105891Z","shell.execute_reply.started":"2022-08-04T10:11:41.097797Z","shell.execute_reply":"2022-08-04T10:11:41.104929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Preprocessing\n\n**Steps:**\n\n1. Feature Engineering the number code\n2. Encode categorical dimensions using `LabelEncoder()`\n3. Fill missing values using `KNNImputer()`\n4. Scale data using `StandardScaler()`","metadata":{"id":"rDzUR2AzcQrl"}},{"cell_type":"markdown","source":"## Feature Engineering the number code","metadata":{"id":"DO9sMJo3dPq6"}},{"cell_type":"code","source":"def extract_num_code(data, features = ['attribute_0', 'attribute_1']):\n    for col in data[features].columns:\n        data[col] = data[col].str.split('_', 1).str[1].astype('int')\n    return data\n\ntrain_data = extract_num_code(train_data, features = ['attribute_0', 'attribute_1'])\ntest_data = extract_num_code(test_data, features = ['attribute_0', 'attribute_1'])","metadata":{"id":"urkX4JeK0y5n","execution":{"iopub.status.busy":"2022-08-04T10:11:41.107177Z","iopub.execute_input":"2022-08-04T10:11:41.107535Z","iopub.status.idle":"2022-08-04T10:11:41.251485Z","shell.execute_reply.started":"2022-08-04T10:11:41.107500Z","shell.execute_reply":"2022-08-04T10:11:41.250568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Encode categorical dimensions using `LabelEncoder()`","metadata":{"id":"WbMAF7kNdVxX"}},{"cell_type":"code","source":"# List of the categorical features.\ncat_feat = ['product_code',\n 'attribute_0',\n 'attribute_1',\n 'attribute_2',\n 'attribute_3',\n 'measurement_0',\n 'measurement_1',\n 'measurement_2']","metadata":{"id":"hHpJz_rW00M1","execution":{"iopub.status.busy":"2022-08-04T10:11:41.254867Z","iopub.execute_input":"2022-08-04T10:11:41.255579Z","iopub.status.idle":"2022-08-04T10:11:41.260028Z","shell.execute_reply.started":"2022-08-04T10:11:41.255541Z","shell.execute_reply":"2022-08-04T10:11:41.259148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def encode(data, cat_fatures=cat_feat):\n    encoder = LabelEncoder()\n    for feat in cat_feat:\n        data[feat] = encoder.fit_transform(data[[feat]])\n    return data\n\ntrain_data = encode(train_data)\ntest_data = encode(test_data)","metadata":{"id":"Y1fUXpwY01Pk","execution":{"iopub.status.busy":"2022-08-04T10:11:41.261693Z","iopub.execute_input":"2022-08-04T10:11:41.262011Z","iopub.status.idle":"2022-08-04T10:11:41.315923Z","shell.execute_reply.started":"2022-08-04T10:11:41.261982Z","shell.execute_reply":"2022-08-04T10:11:41.314784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Fill missing values using `KNNImputer()`","metadata":{"id":"6oQMMeD-dbOQ"}},{"cell_type":"code","source":"# Credits to TheDevastors' notebook for this imputing method\ndef fill_missing(data):\n    imputer = KNNImputer(n_neighbors=3)\n    for col in data.columns:\n        data[col] = imputer.fit_transform(data[[col]])\n    return data\n\ntrain_data = fill_missing(train_data)\ntest_data = fill_missing(test_data)","metadata":{"id":"0GlpjWLW05cS","execution":{"iopub.status.busy":"2022-08-04T10:11:41.317379Z","iopub.execute_input":"2022-08-04T10:11:41.318059Z","iopub.status.idle":"2022-08-04T10:12:17.961359Z","shell.execute_reply.started":"2022-08-04T10:11:41.318015Z","shell.execute_reply":"2022-08-04T10:12:17.960124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.isna().sum()","metadata":{"id":"Y6N9e58w2KxG","outputId":"bab49de0-eb8e-4256-dbc6-06b2aaa7a679","execution":{"iopub.status.busy":"2022-08-04T10:12:17.963125Z","iopub.execute_input":"2022-08-04T10:12:17.963643Z","iopub.status.idle":"2022-08-04T10:12:17.978187Z","shell.execute_reply.started":"2022-08-04T10:12:17.963590Z","shell.execute_reply":"2022-08-04T10:12:17.976892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Scale data using `StandardScaler()`","metadata":{"id":"PgddieeEdobP"}},{"cell_type":"code","source":"X_train = train_data.drop(['failure', 'id'], axis=1)\ny_train = train_data['failure']","metadata":{"id":"lyn33Vmk2M1w","execution":{"iopub.status.busy":"2022-08-04T10:12:17.979748Z","iopub.execute_input":"2022-08-04T10:12:17.980373Z","iopub.status.idle":"2022-08-04T10:12:17.990669Z","shell.execute_reply.started":"2022-08-04T10:12:17.980330Z","shell.execute_reply":"2022-08-04T10:12:17.989472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def scale_data(data):\n    scaler = StandardScaler()\n    data.loc[:] = scaler.fit_transform(data)\n    return data\n\nX_train = scale_data(X_train)\nX_test = scale_data(test_data.drop('id', axis=1))","metadata":{"id":"MCTKFwja3RwM","execution":{"iopub.status.busy":"2022-08-04T10:12:17.992734Z","iopub.execute_input":"2022-08-04T10:12:17.993161Z","iopub.status.idle":"2022-08-04T10:12:18.022022Z","shell.execute_reply.started":"2022-08-04T10:12:17.993123Z","shell.execute_reply":"2022-08-04T10:12:18.020693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Finding the best parameters using Grid Search\n\nNote that since the Poisson Regressor does not have a predict_proba method, it is necessary to set the `fit_intercept` parameter as `True`. More on this can be found on the sci-kit learn documentation [here](https://scikit-learn.org/stable/modules/generated/sklearn.linear_model.PoissonRegressor.html).","metadata":{"id":"tgKUAWTpds3Y"}},{"cell_type":"code","source":"parameters = {\n    \"alpha\":[1, 2, 3, 4],\n    \"fit_intercept\": [True],  # Important to be True since the PoissonRegressor do not have a predict_proba method\n    \"max_iter\": [500, 1000, 1500],\n    \"tol\":[1e-2, 1e-3, 1e-4, 1e-5],\n    \"verbose\":[0],\n    \n}\n\nmodel_poisson = PoissonRegressor()\n\nmodel_poisson = GridSearchCV(\n    model_poisson, \n    parameters, \n    cv=5,\n    scoring='accuracy',\n)\n\nmodel_poisson.fit(X_train, y_train)\n\nprint('-----')\nprint(f'Best parameters {model_poisson.best_params_}')\nprint(\n    f'Mean cross-validated accuracy score of the best_estimator: ' + \n    f'{model_poisson.best_score_:.3f}'\n)\n\ny_preds_poisson = model_poisson.best_estimator_.predict(X_test)","metadata":{"id":"Uc89cYRa3iNH","outputId":"b05c964d-f171-41be-8443-4af0c64e4694","execution":{"iopub.status.busy":"2022-08-04T10:12:18.023545Z","iopub.execute_input":"2022-08-04T10:12:18.023986Z","iopub.status.idle":"2022-08-04T10:12:25.686938Z","shell.execute_reply.started":"2022-08-04T10:12:18.023940Z","shell.execute_reply":"2022-08-04T10:12:25.685732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Optional: \n\nIf the Grid Search is taking too long. You could opt to just run the following line of code. This includes the best parameters in my case.","metadata":{"id":"9K_2jjZ-edab"}},{"cell_type":"code","source":"# poisson = PoissonRegressor(alpha=1, fit_intercept=True, max_iter=500, tol=0.01, verbose=0)\n# poisson.fit(X_train, y_train)\n# preds = poisson.predict(X_test)\n# preds","metadata":{"id":"FlncF6YKHhgD","execution":{"iopub.status.busy":"2022-08-04T10:12:25.688540Z","iopub.execute_input":"2022-08-04T10:12:25.689611Z","iopub.status.idle":"2022-08-04T10:12:25.694590Z","shell.execute_reply.started":"2022-08-04T10:12:25.689567Z","shell.execute_reply":"2022-08-04T10:12:25.693351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submit Predictions","metadata":{"id":"CiM51vE2fvRf"}},{"cell_type":"code","source":"test_data = pd.read_csv('../input/tabular-playground-series-aug-2022/test.csv')\n\nsubmission_dataframe = pd.DataFrame(test_data['id'])\nsubmission_dataframe['failure'] = y_preds_poisson","metadata":{"id":"XtumL7ll6w1Y","execution":{"iopub.status.busy":"2022-08-04T10:12:25.696784Z","iopub.execute_input":"2022-08-04T10:12:25.697884Z","iopub.status.idle":"2022-08-04T10:12:25.800898Z","shell.execute_reply.started":"2022-08-04T10:12:25.697841Z","shell.execute_reply":"2022-08-04T10:12:25.799575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_dataframe.to_csv('submission_test.csv', index=False)","metadata":{"id":"8j-NJGAJ7LWT","execution":{"iopub.status.busy":"2022-08-04T10:12:25.802940Z","iopub.execute_input":"2022-08-04T10:12:25.803788Z","iopub.status.idle":"2022-08-04T10:12:25.860430Z","shell.execute_reply.started":"2022-08-04T10:12:25.803735Z","shell.execute_reply":"2022-08-04T10:12:25.858881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_dataframe","metadata":{"id":"ReUaNl9s9-Gv","outputId":"fe7df9b2-4c04-48cf-c16f-966fc884dcb8","execution":{"iopub.status.busy":"2022-08-04T10:12:25.862017Z","iopub.execute_input":"2022-08-04T10:12:25.862434Z","iopub.status.idle":"2022-08-04T10:12:25.880512Z","shell.execute_reply.started":"2022-08-04T10:12:25.862374Z","shell.execute_reply":"2022-08-04T10:12:25.878967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"id":"D_hIObWTFN0O"},"execution_count":null,"outputs":[]}]}