{"cells":[{"metadata":{"_uuid":"200f33e471557252016863dcdc09540e708ae84b"},"cell_type":"markdown","source":"Greatly inspired by: [Earthquakes FE. More features and samples](https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples)."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"_kg_hide-input":false},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\n\nimport matplotlib.pyplot as plt\n%matplotlib inline\nfrom tqdm import tqdm_notebook\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.svm import NuSVR, SVR\nfrom sklearn.metrics import mean_absolute_error\npd.options.display.precision = 15\n\nimport eli5\nfrom eli5.sklearn import PermutationImportance\n\nimport lightgbm as lgb\nimport xgboost as xgb\nimport time\nimport datetime\nfrom catboost import CatBoostRegressor\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.model_selection import StratifiedKFold, KFold, RepeatedKFold\nfrom sklearn.metrics import mean_absolute_error\nfrom sklearn.linear_model import LinearRegression\nimport gc\nimport seaborn as sns\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nfrom scipy.signal import hilbert\nfrom scipy.signal import hann\nfrom scipy.signal import convolve\nfrom scipy import stats\nfrom sklearn.kernel_ridge import KernelRidge","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5f5f640a3842c1c9e107fc16c92f06292f277117"},"cell_type":"markdown","source":"In test set, the data comes in segments and not like a big chunk in training dataset. I want to know how long the segments are in test segments, we are going to chunk the training data into chunk sizes which are equal to the testing set chunk sizes."},{"metadata":{"trusted":true,"_uuid":"2f6c25d238162e10f3fb185f7f177e37ed46227b"},"cell_type":"code","source":"test_seg = pd.read_csv('../input/test/seg_004cd2.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bb7e665c637efb79d188457baebd49ea5a8cea95"},"cell_type":"code","source":"len(test_seg)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2da64bc29bd3781e2eec9278f8e479ce9559f822"},"cell_type":"markdown","source":"I chunked the training data with 150,000 rows per chunk. I also tried what would it be like when data is supersample, that is, taking overlapping rows, where each starting point differs by 10,000 rows. It turns out that the model overfits quite a bit for that kind of scheme. Maybe, I can reduce the overlap and also tune the parameters for thraining method to reduce overfitting."},{"metadata":{"_uuid":"f6eef7770f5d800594cbf606f48176c9f5d5cbd0"},"cell_type":"markdown","source":"But right now, just a simple chunking with 150,000 rows per chunk."},{"metadata":{"trusted":true,"_uuid":"bc314853a64c3ee3b7b14dd2ad99c06e77e16561"},"cell_type":"code","source":"%%time\ntrain = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0b5611ce49163189ec417894211e07d4bc11914c"},"cell_type":"code","source":"chunk_row = 150_000\nrows = 150_000\nsegments = int(np.floor(train.shape[0] / rows))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"90d8e88f8069c39c41e712c41ae43a01b7ef9e15"},"cell_type":"code","source":"segments","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0f9d5da8437557f584e051f0fff4a14e818ac865"},"cell_type":"markdown","source":"4,194 segments are present so there will be 4,194 rows in our training dataset. The main concern I have is that the model will overfit with such a small dataset."},{"metadata":{"_uuid":"16b0d5c190c4e7654bb47644e866d717087209c0"},"cell_type":"markdown","source":"Now, let's create some placeholder dataframes for features (X_tr) and labels (y_tr)"},{"metadata":{"trusted":true,"_uuid":"772989408b8bd2e02faf2f66da40dd169b925f2c"},"cell_type":"code","source":"X_tr = pd.DataFrame(index=range(segments), dtype=np.float64)\n\ny_tr = pd.DataFrame(index=range(segments), dtype=np.float64, columns=['time_to_failure'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7bab46d3d229ef8268fc937dd494c2ed18f31a0a"},"cell_type":"markdown","source":"I increased the number of features during feature engineering process. I added Fast Fourier Transform to extract some features from the data. I think it makes sense since the acoustic emission might probably be broken down into a bunch of sine waves."},{"metadata":{"trusted":true,"_uuid":"7b7bbb973e1fa15fad44a0388f4bbf837adaa118"},"cell_type":"code","source":"def create_features(X, segment_num, chunk):\n    def add_trend_feature(arr, abs_values=False):\n        idx = np.array(range(len(arr)))\n        if abs_values:\n            arr = np.abs(arr)\n        lr = LinearRegression()\n        lr.fit(idx.reshape(-1, 1), arr)\n        return lr.coef_[0]\n    \n    percentiles = [10, 50, 75, 90, 99]\n    moving_avg_window = [1500, 3000, 6000, 9000, 12000]\n    first_segments = [10000, 50000]\n    last_segments = [10000, 50000]\n    rolling_percentiles = [1, 5, 50, 95, 99]\n    \n    x = pd.Series(chunk['acoustic_data'].values)\n\n    X.loc[segment_num, 'mean'] = x.mean()\n    X.loc[segment_num, 'std'] = x.std()\n    X.loc[segment_num, 'var'] = np.var(x)\n    X.loc[segment_num, 'max'] = x.max()\n    X.loc[segment_num, 'min'] = x.min()\n    \n    z = np.fft.fft(x)\n    realFFT = np.real(z)\n    imagFFT = np.imag(z)\n    X.loc[segment_num, 'Real_mean'] = realFFT.mean()\n    X.loc[segment_num, 'Real_std'] = realFFT.std()\n    X.loc[segment_num, 'Imag_mean'] = imagFFT.mean()\n    X.loc[segment_num, 'Imag_std'] = imagFFT.std()\n    \n    for f_seg in first_segments:\n        X.loc[segment_num, 'std_first_{}'.format(f_seg)] = x[:f_seg].std()\n        X.loc[segment_num, 'mean_first_{}'.format(f_seg)] = x[:f_seg].mean()\n        X.loc[segment_num, 'Rstd_last_{}'.format(f_seg)] = realFFT[:f_seg].std()\n        X.loc[segment_num, 'Rmean_last_{}'.format(f_seg)] = realFFT[:f_seg].mean()\n\n    for l_seg in last_segments:\n        X.loc[segment_num, 'std_last_{}'.format(l_seg)] = x[-l_seg:].std()\n        X.loc[segment_num, 'mean_last_{}'.format(l_seg)] = x[-l_seg:].mean()\n        X.loc[segment_num, 'Rstd_last_{}'.format(l_seg)] = realFFT[-l_seg:].std()\n        X.loc[segment_num, 'Rmean_last_{}'.format(l_seg)] = realFFT[-l_seg:].mean()\n        \n    for percent in percentiles:\n        X.loc[segment_num, 'percentile_{}'.format(percent)] = np.percentile(x, percent)\n    \n    X.loc[segment_num, 'mad'] = x.mad()\n    X.loc[segment_num, 'kurt'] = x.kurtosis()\n    X.loc[segment_num, 'skew'] = x.skew()\n    X.loc[segment_num, 'med'] = x.median()\n    \n    X.loc[segment_num, 'Hilbert_mean'] = np.abs(hilbert(x)).mean()\n    X.loc[segment_num, 'Hann_window_mean'] = (convolve(x, hann(150), mode='same') / sum(hann(150))).mean()\n    \n    X.loc[segment_num, 'trend'] = add_trend_feature(x)\n    X.loc[segment_num, 'abs_trend'] = add_trend_feature(x, abs_values=True)\n    \n    for mv_avg_window in moving_avg_window:\n        X.loc[segment_num, 'Moving_average_{}_mean'.format(mv_avg_window)] = x.rolling(window=mv_avg_window).mean().mean(skipna=True)\n    \n    for windows in [10, 100, 1000]:\n        x_roll_std = x.rolling(windows).std().dropna().values\n        x_roll_mean = x.rolling(windows).mean().dropna().values\n\n        X.loc[segment_num, 'ave_roll_std_' + str(windows)] = x_roll_std.mean()\n        X.loc[segment_num, 'std_roll_std_' + str(windows)] = x_roll_std.std()\n        X.loc[segment_num, 'max_roll_std_' + str(windows)] = x_roll_std.max()\n        X.loc[segment_num, 'min_roll_std_' + str(windows)] = x_roll_std.min()\n        \n        X.loc[segment_num, 'ave_roll_mean_' + str(windows)] = x_roll_mean.mean()\n        X.loc[segment_num, 'std_roll_mean_' + str(windows)] = x_roll_mean.std()\n        X.loc[segment_num, 'max_roll_mean_' + str(windows)] = x_roll_mean.max()\n        X.loc[segment_num, 'min_roll_mean_' + str(windows)] = x_roll_mean.min()\n\n        for roll_perc in rolling_percentiles:\n            X.loc[segment_num, 'p_{}_roll_std_{}'.format(roll_perc, windows)] = np.percentile(x_roll_std, roll_perc)\n            X.loc[segment_num, 'p_{}_roll_mean_{}'.format(roll_perc, windows)] = np.percentile(x_roll_mean, roll_perc)\n\n        X.loc[segment_num, 'av_change_abs_roll_std_' + str(windows)] = np.mean(np.diff(x_roll_std))\n        X.loc[segment_num, 'av_change_rate_roll_std_' + str(windows)] = np.mean(np.nonzero((np.diff(x_roll_std) / x_roll_std[:-1]))[0])\n        X.loc[segment_num, 'abs_max_roll_std_' + str(windows)] = np.abs(x_roll_std).max()\n        X.loc[segment_num, 'av_change_abs_roll_mean_' + str(windows)] = np.mean(np.diff(x_roll_mean))\n        X.loc[segment_num, 'av_change_rate_roll_mean_' + str(windows)] = np.mean(np.nonzero((np.diff(x_roll_mean) / x_roll_mean[:-1]))[0])\n        X.loc[segment_num, 'abs_max_roll_mean_' + str(windows)] = np.abs(x_roll_mean).max()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ffed1f248d36c7a7a7b1b8d2a57994a3522e0c3e"},"cell_type":"code","source":"for segment in tqdm_notebook(range(segments)):\n    seg = train.iloc[segment*rows:segment*rows+chunk_row]\n    y = seg['time_to_failure'].values[-1]\n    y_tr.loc[segment, 'time_to_failure'] = y\n    create_features(X_tr, segment, seg)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6ad94de98384e95f8a769ce9670d8b55f636eec3"},"cell_type":"code","source":"print(f'{X_tr.shape[0]} samples in new train data and {X_tr.shape[1]} columns.')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"825016df814c78c1c9ab1f0a96cef28d28d398cd"},"cell_type":"markdown","source":"Checking how well the added features correlate to labels"},{"metadata":{"trusted":true,"_uuid":"b3e87358c38c3dc2a675be51a241d8bfa7171bbd"},"cell_type":"code","source":"np.abs(X_tr.corrwith(y_tr['time_to_failure'])).sort_values(ascending=False).head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"70f4c5f0b708974342fc2223e9323376bff4694a"},"cell_type":"markdown","source":"For all the missing values, create a dictionary with mean values which will be used both for replacing the missing values in training dataset as well as testing dataset."},{"metadata":{"trusted":true,"_uuid":"73716ded6ea8ba202ed2c52f3ce332371be872ec"},"cell_type":"code","source":"means_dict = {}\nfor col in X_tr.columns:\n    if X_tr[col].isnull().any():\n        mean_value = X_tr.loc[X_tr[col] != -np.inf, col].mean()\n        X_tr.loc[X_tr[col] == -np.inf, col] = mean_value\n        X_tr[col] = X_tr[col].fillna(mean_value)\n        means_dict[col] = mean_value","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7d06146ecb507fd701fbe8453c185b8f1325d4d4"},"cell_type":"markdown","source":"Standardize features by removing the mean and scaling to unit variance"},{"metadata":{"trusted":true,"_uuid":"a154abe0bd5d3e16b216d9375f09166aabba7d36"},"cell_type":"code","source":"scaler = StandardScaler()\nscaler.fit(X_tr)\nX_train_scaled = pd.DataFrame(scaler.transform(X_tr), columns=X_tr.columns)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bb0b51806d92a3e4c788e439ac17d5a5685fdfa8"},"cell_type":"markdown","source":"### Reading test data"},{"metadata":{"_uuid":"9efde3cc418e6b9c37a1394144be49a0fba8744a"},"cell_type":"markdown","source":"Perform the same kind of feature engineering to the test data"},{"metadata":{"trusted":true,"_uuid":"a1b46d0a370c3817a1cc98c022f9c62ebd135331","_kg_hide-input":false,"scrolled":true},"cell_type":"code","source":"submission = pd.read_csv('../input/sample_submission.csv', index_col='seg_id')\nX_test = pd.DataFrame(columns=X_tr.columns, dtype=np.float64, index=submission.index)\n\nfor i, seg_id in enumerate(tqdm_notebook(X_test.index)):\n    seg = pd.read_csv('../input/test/' + seg_id + '.csv')\n    create_features(X_test, seg_id, seg)\n    \nfor col in X_test.columns:\n    if X_test[col].isnull().any():\n        X_test.loc[X_test[col] == -np.inf, col] = means_dict[col]\n        X_test[col] = X_test[col].fillna(means_dict[col])\n        \nX_test_scaled = pd.DataFrame(scaler.transform(X_test), columns=X_test.columns)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"759245da1d8ecdcda5f78cf6c866ef9ccb4f3aff"},"cell_type":"markdown","source":"## Building models"},{"metadata":{"trusted":true,"_uuid":"b1e871cec60e445e82260cb6c3d7b553e290417a"},"cell_type":"code","source":"n_fold = 5\nfolds = KFold(n_splits=n_fold, shuffle=True, random_state=11)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b0f5152fc9a9a851838d0d1a9b134f3bcde03d0d","_kg_hide-input":false},"cell_type":"code","source":"def train_model(X=X_train_scaled, X_test=X_test_scaled, y=y_tr, params=None, folds=folds, model_type='lgb', plot_feature_importance=False, model=None):\n\n    oof = np.zeros(len(X))\n    prediction = np.zeros(len(X_test))\n    scores = []\n    feature_importance = pd.DataFrame()\n    for fold_n, (train_index, valid_index) in enumerate(folds.split(X)):\n        print('Fold', fold_n, 'started at', time.ctime())\n        X_train, X_valid = X.iloc[train_index], X.iloc[valid_index]\n        y_train, y_valid = y.iloc[train_index], y.iloc[valid_index]\n        \n        if model_type == 'lgb':\n            model = lgb.LGBMRegressor(**params, n_estimators = 50000, n_jobs = -1)\n            model.fit(X_train, y_train, \n                    eval_set=[(X_train, y_train), (X_valid, y_valid)], eval_metric='mae',\n                    verbose=10000, early_stopping_rounds=200)\n            \n            y_pred_valid = model.predict(X_valid)\n            y_pred = model.predict(X_test, num_iteration=model.best_iteration_)\n            \n        if model_type == 'xgb':\n            train_data = xgb.DMatrix(data=X_train, label=y_train, feature_names=X.columns)\n            valid_data = xgb.DMatrix(data=X_valid, label=y_valid, feature_names=X.columns)\n\n            watchlist = [(train_data, 'train'), (valid_data, 'valid_data')]\n            model = xgb.train(dtrain=train_data, num_boost_round=20000, evals=watchlist, early_stopping_rounds=200, verbose_eval=500, params=params)\n            y_pred_valid = model.predict(xgb.DMatrix(X_valid, feature_names=X.columns), ntree_limit=model.best_ntree_limit)\n            y_pred = model.predict(xgb.DMatrix(X_test, feature_names=X.columns), ntree_limit=model.best_ntree_limit)\n        \n        if model_type == 'sklearn':\n            model = model\n            model.fit(X_train, y_train)\n            \n            y_pred_valid = model.predict(X_valid).reshape(-1,)\n            score = mean_absolute_error(y_valid, y_pred_valid)\n            print(f'Fold {fold_n}. MAE: {score:.4f}.')\n            print('')\n            \n            y_pred = model.predict(X_test).reshape(-1,)\n        \n        if model_type == 'cat':\n            model = CatBoostRegressor(iterations=20000,  eval_metric='MAE', task_type='GPU', **params)\n            model.fit(X_train, y_train, eval_set=(X_valid, y_valid), cat_features=[], use_best_model=True, verbose=False)\n\n            y_pred_valid = model.predict(X_valid)\n            y_pred = model.predict(X_test)\n        \n        oof[valid_index] = y_pred_valid.reshape(-1,)\n        scores.append(mean_absolute_error(y_valid, y_pred_valid))\n\n        prediction += y_pred    \n        \n        if model_type == 'lgb':\n            # feature importance\n            fold_importance = pd.DataFrame()\n            fold_importance[\"feature\"] = X.columns\n            fold_importance[\"importance\"] = model.feature_importances_\n            fold_importance[\"fold\"] = fold_n + 1\n            feature_importance = pd.concat([feature_importance, fold_importance], axis=0)\n\n    prediction /= n_fold\n    \n    print('CV mean score: {0:.4f}, std: {1:.4f}.'.format(np.mean(scores), np.std(scores)))\n    \n    if model_type == 'lgb':\n        feature_importance[\"importance\"] /= n_fold\n        if plot_feature_importance:\n            cols = feature_importance[[\"feature\", \"importance\"]].groupby(\"feature\").mean().sort_values(\n                by=\"importance\", ascending=False)[:50].index\n\n            best_features = feature_importance.loc[feature_importance.feature.isin(cols)]\n\n            plt.figure(figsize=(16, 12));\n            sns.barplot(x=\"importance\", y=\"feature\", data=best_features.sort_values(by=\"importance\", ascending=False));\n            plt.title('LGB Features (avg over folds)');\n        \n            return oof, prediction, feature_importance, model\n        return oof, prediction, model\n    \n    else:\n        return oof, prediction, model","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ee7a472a875a6adaa520619ffca9bbe3068d7d6d"},"cell_type":"markdown","source":"I just want to train with one machine learning method (light gradient boosting) rather than using xgboost, catboot and SVN. I didn't see the results any getting better with multiple models."},{"metadata":{"trusted":true,"_uuid":"f368d0dc3a116451650e878e81da909ef4e942ae","scrolled":false},"cell_type":"code","source":"params = {'num_leaves': 128,\n          'min_data_in_leaf': 79,\n          'objective': 'huber',\n          'max_depth': -1,\n          'learning_rate': 0.01,\n          \"boosting\": \"gbdt\",\n          \"bagging_freq\": 5,\n          \"bagging_fraction\": 0.8126672064208567,\n          \"bagging_seed\": 11,\n          \"metric\": 'mae',\n          \"verbosity\": -1,\n          'reg_alpha': 0.1302650970728192,\n          'reg_lambda': 0.3603427518866501\n         }\noof_lgb, prediction_lgb, feature_importance, model = train_model(params=params,\n                                                                 model_type='lgb',\n                                                                 plot_feature_importance=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7809fee3ef8f9dee398b2d2512bd5f38588ca06d"},"cell_type":"markdown","source":"This is just for the sake of testing. I actually am not planning to use XGB outputs for submission."},{"metadata":{"trusted":true,"_uuid":"f6212be7ada70583786954c5ab4d059883f2e7b4"},"cell_type":"markdown","source":"xgb_params = {'eta': 0.03,\n              'max_depth': 10,\n              'subsample': 0.9,\n              'objective': 'reg:linear',\n              'eval_metric': 'mae',\n              'silent': True,\n              'nthread': 4}\noof_xgb, prediction_xgb = train_model(X=X_train_scaled, X_test=X_test_scaled, params=xgb_params, model_type='xgb')"},{"metadata":{"trusted":true,"_uuid":"09d14631b0eb612cd2c0335262704131a2816469"},"cell_type":"markdown","source":"xgb_preds = oof_xgb\nlgb_preds = oof_lgb\n\nfig, ax1 = plt.subplots(figsize=(16, 8))\nplt.title(\"LGB Vs XGB prediction values on test data\")\nplt.plot(xgb_preds, color='red', alpha=0.9)\nax1.set_ylabel('time to failure', color='red')\nplt.legend(['xgb'])\nax2 = ax1.twinx()\nplt.plot(lgb_preds, color='blue', alpha=0.5)\nax2.set_ylabel('time to failure', color='blue')\nplt.legend(['lgb'], loc=(0.875, 0.9))\nplt.grid(False)"},{"metadata":{"trusted":true,"_uuid":"583aeee168d8605cbfd6856f0c23be51db136df9"},"cell_type":"code","source":"prediction = oof_lgb\ny_values = y_tr\n\nfig, ax1 = plt.subplots(figsize=(16, 8))\nplt.title(\"Predictions Vs Actual time to failure (Training data)\")\nplt.plot(prediction, color='aqua')\nax1.set_ylabel('time to failure', color='b')\nplt.legend(['predictions'])\nax2 = ax1.twinx()\nplt.plot(y_values, color='g')\nax2.set_ylabel('actual value', color='g')\nplt.legend(['actual value'], loc=(0.875, 0.9))\nplt.grid(False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b138f781bea0ba4356f448efd567ac8e90342acd"},"cell_type":"markdown","source":"The predictions are too cyclical and thus missing out to capture the different time_to_failure after the quakes have occured. The predicions pretty much think that the time will be rest the same way (thus, near perfect cycle). I think this is the next step I have to work on without overfitting."},{"metadata":{"trusted":true,"_uuid":"c2bdecded52fd2613afefc6375fd0dd27a9dd5c0"},"cell_type":"code","source":"feature_importance_values = np.abs(X_test_scaled.corrwith(pd.DataFrame(prediction_lgb)[0]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b59bd5b7132b9ee6c5824dd4028069dd061da21f"},"cell_type":"code","source":"feature_importance_values.sort_values(ascending=False).head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"391a7c17d863cba2c9cdf49c222314cbc89ae4a3"},"cell_type":"markdown","source":"Rolling windows are the most important features! I am figuring whether I should reduce training my model with just the important features thresholded by certain value. For, example, only the features where the feature importance is above 0.6"},{"metadata":{"trusted":true,"_uuid":"90f3df665752cd3b8ca7c865d637fb718f53a383"},"cell_type":"code","source":"feature_importance_values[feature_importance_values > 0.6].sort_values(ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5e6695b7139ac7f1d669839fbc7d24447c5d6e4e"},"cell_type":"markdown","source":"# Investigating Feature Importance and parameter fine tuning"},{"metadata":{"_uuid":"695234f0fbe8befc1de83e27871a322a471c2a41"},"cell_type":"markdown","source":"What would happen if we only train the features with feature importance greater than 0.5?"},{"metadata":{"trusted":true,"_uuid":"2913197491f64d5c584e3706a3ca36b1395040f8"},"cell_type":"markdown","source":"selected_cols = feature_importance_values[feature_importance_values > 0].index\nreduced_X_tr = X_tr[selected_cols]\nreduced_X_test = X_test[selected_cols]\n\nscaler = StandardScaler()\nscaler.fit(reduced_X_tr)\nreduced_X_train_scaled = pd.DataFrame(scaler.transform(reduced_X_tr), columns=reduced_X_tr.columns)\nreduced_X_test_scaled = pd.DataFrame(scaler.transform(reduced_X_test), columns=reduced_X_test.columns)"},{"metadata":{"_uuid":"57bd4608f6d35872e733b2353180980df00a6342"},"cell_type":"markdown","source":"I have heavily fine-tuned num_leaves in tandem with max_depth. Reduced learning rate. Increased lambda_l1 (L1 regularization)."},{"metadata":{"trusted":true,"_uuid":"bf9331c0153096cdc84f267616b41a2bdbb43039"},"cell_type":"markdown","source":"n_fold = 5\nfolds = KFold(n_splits=n_fold, shuffle=True, random_state=11)\n\nparams = {'num_leaves': 100,\n          'min_data_in_leaf': 80,\n          'objective': 'huber',\n          'max_depth': 14,\n          'learning_rate': 0.008,\n          \"boosting\": \"dart\",\n          \"bagging_freq\": 2,\n          \"bagging_fraction\": 0.6,\n          \"feature_fraction\": 0.75,\n          \"bagging_seed\": 11,\n          \"metric\": 'mae',\n          \"verbosity\": -1,\n          'lambda_l1': 0.18,\n          'lambda_l2': 0.25\n         }\nreduced_oof_lgb, reduced_prediction_lgb, reduced_feature_importance = train_model(X=reduced_X_train_scaled,\n                                                                                  X_test=reduced_X_test_scaled,\n                                                                                  y=y_tr,\n                                                                                  params=params,\n                                                                                  folds=folds,\n                                                                                  model_type='lgb',\n                                                                                  plot_feature_importance=True)"},{"metadata":{"_uuid":"2636e3c1570f753a324dd0d1ddb2fc2afd99345a"},"cell_type":"markdown","source":"Plots of time_to_failure values (training set) with rest to important features"},{"metadata":{"trusted":true,"_uuid":"65ef76b516d44d1147012a8a07ef610902a9f32d"},"cell_type":"markdown","source":"plt.figure(figsize=(44, 24))\ncols = list(feature_importance_values.sort_values(ascending=False).head(20).index)\nfor i, col in enumerate(cols):\n    plt.subplot(5, 4, i + 1)\n    plt.plot(X_tr[col], color='blue')\n    plt.title(\"{} - {:.3f}\".format(col, feature_importance_values[col]))\n    ax1.set_ylabel(col, color='b')\n\n    ax2 = ax1.twinx()\n    plt.plot(y_tr, color='g')\n    ax2.set_ylabel('time_to_failure', color='g')\n    plt.legend([col, 'time_to_failure'], loc=(0.875, 0.9))\n    plt.grid(False)"},{"metadata":{"_uuid":"ca7e5831b4a11bfeb6f41ea87fe451fd031ea030"},"cell_type":"markdown","source":"Right now I am just going to take the LGB's output predictions. I will think about what to do with XGB outputs. I don't want to average the two results either."},{"metadata":{"trusted":true,"_uuid":"f92181558df3363e07225f372e40cde81278738d"},"cell_type":"code","source":"submission['time_to_failure'] = prediction_lgb","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d3c5db2e9f7539ebd5758bd6c92e187d7b3a1faa"},"cell_type":"code","source":"submission.to_csv('submission_rolling_v6.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cd8e751304d45a5cf916bf0d2f581a1d025f3800"},"cell_type":"markdown","source":"Sometimes, I just play around with stuff and don't want to commit the whole thing just to download the submission file. Here is a great solution to get around that. Feature engineering and hyper parameter tuning takes a long time so I would like to have a download link of submission file when I have played around."},{"metadata":{"trusted":true,"_uuid":"0d082dd601dce53e6d989e0859a46ccac649514d"},"cell_type":"markdown","source":"from IPython.display import HTML\n\ndef create_download_link(title = \"Download CSV file\", filename = \"data.csv\"):  \n    html = '<a href={filename}>{title}</a>'\n    html = html.format(title=title,filename=filename)\n    return HTML(html)\n\n# create a link to download the dataframe\ncreate_download_link(filename='submission_rolling_v5.csv')\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}