{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Competition [INGV - Volcanic Eruption Prediction](https://www.kaggle.com/competitions/predict-volcanic-eruptions-ingv-oe)","metadata":{}},{"cell_type":"markdown","source":"**Thanks to [😸 Basic preprocessing with CatBoost 😸](https://www.kaggle.com/code/miracl16/basic-preprocessing-with-catboost)**","metadata":{}},{"cell_type":"markdown","source":"### My upgrade:\n* added desition and methods descriptions;\n* updated and improved layout;","metadata":{}},{"cell_type":"markdown","source":"<a class=\"anchor\" id=\"0.1\"></a>\n# Table of Contents\n\n1. [Import libraries](#1)\n1. [Download data](#2)\n1. [EDA & FE](#3)\n1. [Modeling & Baseline](#4)\n1. [Prediction & Submission](#5)","metadata":{}},{"cell_type":"markdown","source":"## 1. Import libraries<a class=\"anchor\" id=\"1\"></a>\n\n[Back to Table of Contents](#0.1)","metadata":{}},{"cell_type":"code","source":"# Work with Data\nimport numpy as np \nimport pandas as pd \nimport os\n\n# Visualization\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm\n\n# Modeling and Prediction\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import mean_absolute_error, mean_squared_error\nfrom catboost import CatBoostRegressor, Pool","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-04-12T13:08:37.692054Z","iopub.execute_input":"2022-04-12T13:08:37.692392Z","iopub.status.idle":"2022-04-12T13:08:37.698331Z","shell.execute_reply.started":"2022-04-12T13:08:37.692364Z","shell.execute_reply":"2022-04-12T13:08:37.697139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Download data<a class=\"anchor\" id=\"2\"></a>\n\n[Back to Table of Contents](#0.1)","metadata":{}},{"cell_type":"code","source":"# Download training and test data\nDIR_PATH = '../input/predict-volcanic-eruptions-ingv-oe/'\n\ntrain_list = os.listdir(DIR_PATH + 'train')\ntest_list = os.listdir(DIR_PATH + 'test')\n\nprint('Number of train files: {}'.format(len(train_list)))\nprint('Number of test files: {}'.format(len(test_list )))","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:08:37.794421Z","iopub.execute_input":"2022-04-12T13:08:37.794917Z","iopub.status.idle":"2022-04-12T13:08:37.806479Z","shell.execute_reply.started":"2022-04-12T13:08:37.794883Z","shell.execute_reply":"2022-04-12T13:08:37.80553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the first 5 rows of the example training data set\nexample_train_df = pd.read_csv(DIR_PATH + 'train/' + train_list[0])\nexample_train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:08:37.900265Z","iopub.execute_input":"2022-04-12T13:08:37.900767Z","iopub.status.idle":"2022-04-12T13:08:37.99344Z","shell.execute_reply.started":"2022-04-12T13:08:37.900734Z","shell.execute_reply":"2022-04-12T13:08:37.992644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the first 5 rows of the example test data set\nexample_test_df = pd.read_csv(DIR_PATH + 'test/' + test_list[0])\nexample_test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:08:37.994614Z","iopub.execute_input":"2022-04-12T13:08:37.995033Z","iopub.status.idle":"2022-04-12T13:08:38.095205Z","shell.execute_reply.started":"2022-04-12T13:08:37.995002Z","shell.execute_reply":"2022-04-12T13:08:38.094042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the train data set\ntrain_time_to_eruption = pd.read_csv(DIR_PATH + 'train.csv')\ntrain_time_to_eruption","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:08:38.097188Z","iopub.execute_input":"2022-04-12T13:08:38.097662Z","iopub.status.idle":"2022-04-12T13:08:38.113716Z","shell.execute_reply.started":"2022-04-12T13:08:38.097616Z","shell.execute_reply":"2022-04-12T13:08:38.112826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. EDA & FE <a class=\"anchor\" id=\"3\"></a>\n[Back to Table of Contents](#0.1)","metadata":{}},{"cell_type":"markdown","source":"### Task description\n\nWe are faced with the task of regression. What does this imply?<br />\nRegression is the process of finding the correlations between dependent and independent variables.<br />\nThe task of the Regression algorithm is to find the mapping function to map the input variable(x) to the continuous output variable(y).","metadata":{}},{"cell_type":"code","source":"# Visualization sensor signals from the example train data set\nexample_train_df.plot(figsize=(15,15), subplots=True);","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:08:38.256867Z","iopub.execute_input":"2022-04-12T13:08:38.257257Z","iopub.status.idle":"2022-04-12T13:08:41.587602Z","shell.execute_reply.started":"2022-04-12T13:08:38.257225Z","shell.execute_reply":"2022-04-12T13:08:41.586347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**What to do with the matrix data of our sensors?**<br/>\nWhat to pass to the algorithms as params?<br/>\n<br/>\nSuppose we have a time series that has its own characteristics: lows, highs, etc.<br/>\nIn essence, we compare the values of the graphs: the number of peaks, frequency, variance, etc.<br/>\nIn the set of parameters for each of the sensors, we can process the values for each of the indicators and pass them as parameters to the algorithm.<br/>\n<br/>\nLet's see how to do it on an example data set.","metadata":{}},{"cell_type":"code","source":"# Prepare and display basic statistics for the example train data set\n# replace the NULL values with 0, rename a column\nprocess = pd.DataFrame(example_train_df.fillna(0).describe().iloc[1:, :].unstack()).reset_index()\nprocess = process.rename(columns={0: 'value'})\nprocess","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:08:41.590631Z","iopub.execute_input":"2022-04-12T13:08:41.590942Z","iopub.status.idle":"2022-04-12T13:08:41.654563Z","shell.execute_reply.started":"2022-04-12T13:08:41.590911Z","shell.execute_reply":"2022-04-12T13:08:41.653612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Remove unnecessary columns and transpose data set\nprocess['feature'] = process['level_0'] + '_' + process['level_1']\nprocess = process.drop(['level_0', 'level_1'], axis=1).set_index('feature').T\n\n# Set a column with time to eruption by the specific segment sensors data set\nprocess['time'] = train_time_to_eruption[train_time_to_eruption.segment_id == int(train_list[0].split('.')[0])].time_to_eruption.values[0]\nprocess","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:08:41.65625Z","iopub.execute_input":"2022-04-12T13:08:41.656559Z","iopub.status.idle":"2022-04-12T13:08:41.704238Z","shell.execute_reply.started":"2022-04-12T13:08:41.656503Z","shell.execute_reply":"2022-04-12T13:08:41.703193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### What Is Skewness?\nSkewness refers to a distortion or asymmetry that deviates from the symmetrical bell curve, or normal distribution, in a set of data.<br />\nIf the curve is shifted to the left or to the right, it is said to be skewed.<br />\n<br />\nFor normally distributed data, the skewness should be about zero.<br />\nFor unimodal continuous distributions, a skewness value greater than zero means that there is more weight in the right tail of the distribution.","metadata":{}},{"cell_type":"code","source":"# Computing the sample skewness for each sensor of the example train data set\npd.DataFrame(example_train_df.fillna(0).skew()).T","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:08:41.706176Z","iopub.execute_input":"2022-04-12T13:08:41.706606Z","iopub.status.idle":"2022-04-12T13:08:41.73329Z","shell.execute_reply.started":"2022-04-12T13:08:41.706561Z","shell.execute_reply":"2022-04-12T13:08:41.732137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Combine FE into a function so that it is convenient to apply to different data sets\n\ndef create_frame(data, data_time=None, type_data='train'):\n    data = data.fillna(0)\n    \n    # get basic statistics\n    data_transform = data.describe().iloc[1:, :]\n    \n    # Additional params\n    \n    # Asymmetry coefficient\n    data_transform.loc['skew'] = data.skew().tolist()\n    \n    # Mean absolute deviation\n    data_transform.loc['mad'] = data.mad().tolist()\n    \n    # The kurtosis coefficient\n    # is a measure of the sharpness of the peak of the distribution of a random variable\n    data_transform.loc['kurtosis'] = data.kurtosis().tolist()\n    \n    # Adding quantiles\n    # necessary for a better understanding of what is happening with the dataset\n    for i in range(0, 100, 5):\n        if ((i!=25) & (i!=50)):\n                str_col = f\"{i}%\"\n                int_col = float(i)/100\n                data_transform.loc[str_col] = data_transform.quantile(int_col).tolist()\n        else:\n            continue\n    \n    # Combining all features into the one dataset\n    data_transform = pd.DataFrame(data_transform.unstack()).reset_index()\n    data_transform = data_transform.rename(columns={0: 'value'})\n    data_transform['feature'] = data_transform['level_0'] + '_' + data_transform['level_1']\n    data_transform = data_transform.drop(['level_0', 'level_1'], axis=1).set_index('feature').T\n    \n    # For a train data set adding the time to eruption for the specific segment sensors \n    if type_data=='train':\n        data_transform['time'] = data_time\n\n    return data_transform","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:08:41.734809Z","iopub.execute_input":"2022-04-12T13:08:41.73523Z","iopub.status.idle":"2022-04-12T13:08:41.745766Z","shell.execute_reply.started":"2022-04-12T13:08:41.735186Z","shell.execute_reply":"2022-04-12T13:08:41.744473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Apply the function for each train data set\nall_train = pd.DataFrame()\n\nfor file in tqdm(train_list):\n    df = pd.read_csv(DIR_PATH + 'train/' + file)\n    data_time = train_time_to_eruption[train_time_to_eruption.segment_id == int(file.split('.')[0])].time_to_eruption.values[0]\n    df = create_frame(df, data_time, type_data='train')\n    all_train = all_train.append(df)\n\nall_train = all_train.reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:08:41.747484Z","iopub.execute_input":"2022-04-12T13:08:41.747951Z","iopub.status.idle":"2022-04-12T13:25:26.343039Z","shell.execute_reply.started":"2022-04-12T13:08:41.747908Z","shell.execute_reply":"2022-04-12T13:25:26.341702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Apply the function for each test data set\nall_test = pd.DataFrame()\n\nfor file in tqdm(test_list):\n    df = pd.read_csv(DIR_PATH + 'test/' + file)\n    df = create_frame(df, data_time=None, type_data='test')\n    all_test = all_test.append(df)\n\nall_test = all_test.reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:25:26.344612Z","iopub.execute_input":"2022-04-12T13:25:26.345011Z","iopub.status.idle":"2022-04-12T13:42:05.338539Z","shell.execute_reply.started":"2022-04-12T13:25:26.344969Z","shell.execute_reply":"2022-04-12T13:42:05.33743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:42:05.340871Z","iopub.execute_input":"2022-04-12T13:42:05.341149Z","iopub.status.idle":"2022-04-12T13:42:05.372499Z","shell.execute_reply.started":"2022-04-12T13:42:05.341122Z","shell.execute_reply":"2022-04-12T13:42:05.371421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:42:05.374208Z","iopub.execute_input":"2022-04-12T13:42:05.37455Z","iopub.status.idle":"2022-04-12T13:42:05.409985Z","shell.execute_reply.started":"2022-04-12T13:42:05.374479Z","shell.execute_reply":"2022-04-12T13:42:05.408885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Modeling & Baseline<a class=\"anchor\" id=\"4\"></a>\n\n[Back to Table of Contents](#0.1)","metadata":{}},{"cell_type":"markdown","source":"#### What Is Baseline?\nInsofar as we must know whether the predictions for a given algorithm are good or not we need to use a baseline prediction algorithm.<br />\nA baseline prediction algorithm provides a set of predictions that you can evaluate as you would any predictions for your problem, such as classification accuracy or RMSE.<br />\n<br />\nThe scores from these algorithms provide the required point of comparison when evaluating all other machine learning algorithms on your problem.","metadata":{}},{"cell_type":"code","source":"# Copy and split our data\n\ntest = all_test.copy()\n\nX = all_train.drop('time',axis=1)\ny = all_train['time']\n\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, shuffle=True, random_state=10)\n\nX_train, X_val, y_train, y_val = train_test_split(X_train, y_train, test_size=0.25, shuffle=True, random_state=10)","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:42:05.411407Z","iopub.execute_input":"2022-04-12T13:42:05.411791Z","iopub.status.idle":"2022-04-12T13:42:05.4402Z","shell.execute_reply.started":"2022-04-12T13:42:05.411758Z","shell.execute_reply":"2022-04-12T13:42:05.4393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Mean Absolute Percentage Error** (MAPE) is one of the most commonly used KPIs to measure forecast accuracy.<br/>\nMAPE is the sum of the individual absolute errors divided by the demand (each period separately). It is the average of the percentage errors.<br/>","metadata":{}},{"cell_type":"code","source":"# Define a function for computing a MAPE\ndef mape(y_true, y_pred):\n    return np.mean(np.abs((y_pred-y_true)/y_true))","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:42:05.441453Z","iopub.execute_input":"2022-04-12T13:42:05.441736Z","iopub.status.idle":"2022-04-12T13:42:05.446613Z","shell.execute_reply.started":"2022-04-12T13:42:05.441709Z","shell.execute_reply":"2022-04-12T13:42:05.445617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train a model\nclf = CatBoostRegressor(loss_function='MAPE')\n\ntrain_dataset = Pool(data=X_train, label=y_train)\neval_dataset = Pool(data=X_val, label=y_val)\n\nclf.fit(train_dataset,\n            use_best_model=True,\n            verbose = 100,\n            eval_set=eval_dataset)","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:42:05.448098Z","iopub.execute_input":"2022-04-12T13:42:05.448375Z","iopub.status.idle":"2022-04-12T13:42:45.409807Z","shell.execute_reply.started":"2022-04-12T13:42:05.448348Z","shell.execute_reply":"2022-04-12T13:42:45.40895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Apply the model to the test dataset\ny_pred = clf.predict(Pool(data=X_test))\n    \nprint(f\"MAPE: {mape(y_test, y_pred)}\")\nprint(f\"MAE: {mean_absolute_error(y_test, y_pred)}\")\nprint(f\"RMSE: {np.sqrt(mean_squared_error(y_test, y_pred))}\")","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:42:45.411003Z","iopub.execute_input":"2022-04-12T13:42:45.411374Z","iopub.status.idle":"2022-04-12T13:42:45.442077Z","shell.execute_reply.started":"2022-04-12T13:42:45.411328Z","shell.execute_reply":"2022-04-12T13:42:45.441026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5. Prediction & Submission<a class=\"anchor\" id=\"5\"></a>\n\n[Back to Table of Contents](#0.1)","metadata":{}},{"cell_type":"markdown","source":"**We are going to use KFold with some parametrs with CatBoostRegressor.**<br/>\nWe didn't use GridSearch because save time.\n\n[CatBoost](https://catboost.ai/docs) is a machine learning method based on [gradient boosting](https://en.wikipedia.org/wiki/Gradient_boosting) over decision trees.","metadata":{}},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(\n    X, y, test_size=0.25, shuffle=True, random_state=10)","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:42:45.443486Z","iopub.execute_input":"2022-04-12T13:42:45.4439Z","iopub.status.idle":"2022-04-12T13:42:45.457356Z","shell.execute_reply.started":"2022-04-12T13:42:45.443859Z","shell.execute_reply":"2022-04-12T13:42:45.45629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_fold = 5\ncv = KFold(n_splits=n_fold, shuffle=True, random_state=10)\nprediction = np.zeros(len(test))\nmape_, mae, rmse = [], [], []\n\nparams = {\n            'iterations':1000,\n            'learning_rate':0.1,\n            'depth':6,\n            'eval_metric':'MAPE'\n}\n\n# Split our data for KFolds blocs\nfor fold, (train_index, val_index) in enumerate(cv.split(X)):\n    # Move to the next fold block\n    X_train = X.iloc[train_index,:]\n    X_val = X.iloc[val_index,:]\n\n    y_train = y.iloc[train_index]\n    y_val = y.iloc[val_index]\n    \n    # Traind the model with new train data form the current fold \n    clf = CatBoostRegressor(**params)  \n    \n    train_dataset = Pool(data=X_train, label=y_train)\n    \n    eval_dataset = Pool(data=X_val, label=y_val)\n    \n    clf.fit(train_dataset,\n              use_best_model=True,\n              verbose = 0,\n              eval_set=eval_dataset)\n   \n    y_pred = clf.predict(Pool(data=X_test))\n    \n    mape_.append(mape(y_test, y_pred))\n    mae.append(mean_absolute_error(y_test, y_pred))\n    rmse.append(np.sqrt(mean_squared_error(y_test, y_pred)))\n\n    print(f\"fold: {fold}, MAPE: {mape(y_test, y_pred)}\")\n    print(f\"fold: {fold}, MAE: {mean_absolute_error(y_test, y_pred)}\")\n    print(f\"fold: {fold}, RMSE: {np.sqrt(mean_squared_error(y_test, y_pred))}\")\n\n    # test array predictions\n    prediction += clf.predict(Pool(data=test))\n        \nprediction /= n_fold\n\nprint('CV mean MAPE:  {0:.4f}, std: {1:.4f}.'.format(np.mean(mape_), np.std(mape_)))\nprint('CV mean MAE: {0:.4f}, std: {1:.4f}.'.format(np.mean(mae), np.std(mae)))\nprint('CV mean RMSE: {0:.4f}, std: {1:.4f}.'.format(np.mean(rmse), np.std(rmse)))","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:42:45.458719Z","iopub.execute_input":"2022-04-12T13:42:45.458999Z","iopub.status.idle":"2022-04-12T13:46:05.281725Z","shell.execute_reply.started":"2022-04-12T13:42:45.45897Z","shell.execute_reply":"2022-04-12T13:46:05.280654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the submission dataframe\nsub_example = pd.read_csv(DIR_PATH + 'sample_submission.csv')\nsub_example.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:51:23.203387Z","iopub.execute_input":"2022-04-12T13:51:23.203958Z","iopub.status.idle":"2022-04-12T13:51:23.220241Z","shell.execute_reply.started":"2022-04-12T13:51:23.203904Z","shell.execute_reply":"2022-04-12T13:51:23.219222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the 5 first segment ids\ntest_index = [int(i.split('.')[0]) for i in test_list]\ntest_index[:5]","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:50:28.939503Z","iopub.execute_input":"2022-04-12T13:50:28.940077Z","iopub.status.idle":"2022-04-12T13:50:28.950069Z","shell.execute_reply.started":"2022-04-12T13:50:28.940027Z","shell.execute_reply":"2022-04-12T13:50:28.949013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Saving the result into submission file\nsubmission = pd.DataFrame()\nsubmission['segment_id'] = test_index\nsubmission['time_to_eruption'] = prediction\nsubmission.to_csv('submission.csv', header=True, index=False) # Competition rules require that no index number be saved\n\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-04-12T13:50:32.058837Z","iopub.execute_input":"2022-04-12T13:50:32.059312Z","iopub.status.idle":"2022-04-12T13:50:32.099777Z","shell.execute_reply.started":"2022-04-12T13:50:32.059279Z","shell.execute_reply":"2022-04-12T13:50:32.098619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I hope you find this notebook useful and enjoyable.\n\nYour comments and feedback are most welcome.\n\n[Go to Top](#0)","metadata":{}}]}