{"cells":[{"metadata":{},"cell_type":"markdown","source":"### This kernel was created to the <a target=\"_blank\" href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction\">LANL Earthquake Prediction</a> competition to show how to use the Leave One Out Cross Validation.\n#### Leave One Out Cross Validation is a K-fold cross validation with K equal to N, the number of samples in the data set. It means that for each point the model is trained on all the data except for that point and a prediction is made for that point. As we'll see this technique will lead to overfitting in this case."},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true},"cell_type":"code","source":"# We import the libraries\nimport numpy as np\nimport pandas as pd\npd.options.display.precision = 15\nimport seaborn as sns\nimport lightgbm as lgb\nfrom tqdm import tqdm_notebook\nimport matplotlib.pyplot as plt\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.model_selection import LeaveOneOut, train_test_split\nimport warnings\nwarnings.filterwarnings(\"ignore\")","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"0ee349f0d7f482a0cc4827a3b7515a915e4dd37f","trusted":true},"cell_type":"code","source":"# We load the train.csv\nPATH=\"../input/\"\ntrain_df = pd.read_csv(PATH + 'train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"There are {} rows in train_df\".format(len(train_df)))","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"abcf886f07660f12d170a2182675e15efffefde0","trusted":true},"cell_type":"code","source":"# We define the number of rows in each segment as the same number of rows in the real test segments (150000 rows)\nrows = 150000\nsegments = int(np.floor(train_df.shape[0] / rows))\nprint(\"Number of segments: \", segments)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0cec83d75c3498355fb232c33ef43dcd2017f24e"},"cell_type":"markdown","source":"## We process the train file"},{"metadata":{"_kg_hide-input":true,"_uuid":"c65536bf1bf32037175518c3cc8fc7439e217e9d","trusted":true},"cell_type":"code","source":"train_X = pd.DataFrame(index=range(segments), dtype=np.float64)\ntrain_y = pd.DataFrame(index=range(segments), dtype=np.float64, columns=['time_to_failure'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_X.head(3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_y.head(3)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"0745eb336b1313d3c46c957d7b18c19c98076743","trusted":true},"cell_type":"code","source":"def create_features(seg_id, xc, X):\n    \n    x_roll_std_50 = xc.rolling(50).std().dropna().values\n    x_roll_mean_50 = xc.rolling(50).mean().dropna().values\n    x_roll_std_1000 = xc.rolling(1000).std().dropna().values\n    x_roll_mean_1000 = xc.rolling(1000).mean().dropna().values\n    x_roll_std_5000 = xc.rolling(5000).std().dropna().values\n    x_roll_mean_5000 = xc.rolling(5000).mean().dropna().values\n    \n    X.loc[seg_id, 'q05_roll_std_50'] = np.quantile(x_roll_std_50, 0.05)\n    X.loc[seg_id, 'q90_roll_mean_50'] = np.quantile(x_roll_mean_50, 0.90)\n    X.loc[seg_id, 'q05_roll_std_1000'] = np.quantile(x_roll_std_1000, 0.05)\n    X.loc[seg_id, 'q90_roll_mean_1000'] = np.quantile(x_roll_mean_1000, 0.90)\n    X.loc[seg_id, 'q05_roll_std_5000'] = np.quantile(x_roll_std_5000, 0.05)\n    X.loc[seg_id, 'q90_roll_mean_5000'] = np.quantile(x_roll_mean_5000, 0.90)\n    \n    X.loc[seg_id, 'abs_max'] = np.abs(xc).max()\n    X.loc[seg_id, 'abs_q95'] = np.quantile(np.abs(xc), 0.95)\n    X.loc[seg_id, 'abs_q75'] = np.quantile(np.abs(xc), 0.75)\n    X.loc[seg_id, 'abs_q25'] = np.quantile(np.abs(xc), 0.25)\n    X.loc[seg_id, 'abs_q05'] = np.quantile(np.abs(xc), 0.05)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"5a446450bfd25296d6b98dde3e750e11d96250a4","trusted":true},"cell_type":"code","source":"# iterate over all segments\nfor seg_id in tqdm_notebook(range(segments)):\n    seg = train_df.iloc[seg_id*rows:seg_id*rows + rows]\n    create_features(seg_id, seg['acoustic_data'], train_X)\n    train_y.loc[seg_id, 'time_to_failure'] = seg['time_to_failure'].values[-1]","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"ff98b8f65290cfa14a5fed507a5567a589109fc9","trusted":true},"cell_type":"code","source":"train_X.shape, train_y.shape","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"_uuid":"dc6dd182bf1819c4dcfe823416efca51ff15d5bf","scrolled":true,"trusted":true},"cell_type":"code","source":"train_X.head(3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_y.head(3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## We create a testing set"},{"metadata":{"trusted":true},"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(train_X, train_y, test_size=0.25, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2bb7ba0d1ec723be641adf5f925c24237c9198c0"},"cell_type":"markdown","source":"\n## We create the LGB model"},{"metadata":{"trusted":true},"cell_type":"code","source":"seed = 42\n    \nparams_lgb = {'bagging_fraction': 0.9,\n            'bagging_freq': 5,\n            'early_stopping_rounds': 10,\n            'feature_fraction': 0.7,\n            'learning_rate': 0.02,\n            'num_boost_round': 100,\n            'seed': seed,\n            'feature_fraction_seed': seed,\n            'bagging_seed': seed,\n            'drop_seed': seed,\n            'data_random_seed': seed,\n            'objective': 'regression_l1',\n            'boosting': 'gbdt',\n            'verbosity': -1,\n            'metric': 'mae',\n            'num_threads': 8}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Leave One Out cross validation\nloo = LeaveOneOut()\nn_splits = loo.get_n_splits(X_train)\nt = tqdm_notebook(total=n_splits)\noof = np.zeros(len(X_train))\npredictions = np.zeros(len(X_test))\n\nfor trn_idx, val_idx in loo.split(X_train):\n    t.update(1)\n    trn_data = lgb.Dataset(X_train.iloc[trn_idx], label=y_train.iloc[trn_idx])\n    val_data = lgb.Dataset(X_train.iloc[val_idx], label=y_train.iloc[val_idx])\n    \n    clf = lgb.train(params_lgb, trn_data, valid_sets = [trn_data, val_data], verbose_eval=False)\n    oof[val_idx] = clf.predict(X_train.iloc[val_idx], num_iteration=clf.best_iteration)\n    predictions += clf.predict(X_test, num_iteration=clf.best_iteration) / n_splits\n    \nprint('CV MAE: {}'.format(mean_absolute_error(y_train.values, oof)))\nprint('Test MAE: {}'.format(mean_absolute_error(y_test.values, predictions)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Scatter plot the real time to failure vs predicted (Leave One Out Cross Validation)\nplt.figure(figsize=(6, 6))\nplt.scatter(y_train.values.flatten(), oof)\nplt.xlim(0, 20)\nplt.ylim(0, 20)\nplt.title('Leave One Out Cross Validation')\nplt.xlabel('time to failure', fontsize=12)\nplt.ylabel('predicted', fontsize=12)\nplt.plot([(0, 0), (20, 20)], [(0, 0), (20, 20)])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Scatter plot the real time to failure vs predicted (Testing Set)\nplt.figure(figsize=(6, 6))\nplt.scatter(y_test.values.flatten(), predictions)\nplt.xlim(0, 20)\nplt.ylim(0, 20)\nplt.title('Testing Set')\nplt.xlabel('time to failure', fontsize=12)\nplt.ylabel('predicted', fontsize=12)\nplt.plot([(0, 0), (20, 20)], [(0, 0), (20, 20)])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Conclusion\n#### As we can see, the MAE of the Leave One Out Cross Validation is really good (MAE 1.57) but the MAE of the test set is not good (MAE 2.20). We conclude that LOOCV is not a good choice for this problem because it clearly leads to overfitting."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.6.7"}},"nbformat":4,"nbformat_minor":1}