{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### In [previous notebook](https://www.kaggle.com/code/alexryzhkov/how-to-kill-all-your-efforts) we figure out how sampling can destroy all your efforts to get the best LogLoss score. But how we can fix that? The answer is simple (and @mateuscco already said about it in the comments) - you just need to use calibration and in this kernel we will figure out how to do it.\n\n### Preparations (we will use only the `train_0.csv` train file to speedup the experiments)","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nimport lightgbm\nfrom sklearn.metrics import log_loss\nimport matplotlib.pyplot as plt","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-25T15:58:34.700984Z","iopub.execute_input":"2022-10-25T15:58:34.701632Z","iopub.status.idle":"2022-10-25T15:58:36.351483Z","shell.execute_reply.started":"2022-10-25T15:58:34.701516Z","shell.execute_reply":"2022-10-25T15:58:36.350124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"INPUT_DIR = '../input/tabular-playground-series-oct-2022/'","metadata":{"execution":{"iopub.status.busy":"2022-10-25T15:58:36.483146Z","iopub.execute_input":"2022-10-25T15:58:36.483657Z","iopub.status.idle":"2022-10-25T15:58:36.488948Z","shell.execute_reply.started":"2022-10-25T15:58:36.483615Z","shell.execute_reply":"2022-10-25T15:58:36.487720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dtypes_df = pd.read_csv(f'{INPUT_DIR}/train_dtypes.csv')\ntrain_dtypes = {k: v for (k, v) in zip(train_dtypes_df.column, train_dtypes_df.dtype)}","metadata":{"execution":{"iopub.status.busy":"2022-10-25T15:58:36.897405Z","iopub.execute_input":"2022-10-25T15:58:36.897902Z","iopub.status.idle":"2022-10-25T15:58:36.924233Z","shell.execute_reply.started":"2022-10-25T15:58:36.897848Z","shell.execute_reply":"2022-10-25T15:58:36.922819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = pd.read_csv(f'{INPUT_DIR}/train_0.csv', dtype=train_dtypes)\ntrain_data.shape","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-10-25T15:58:37.259296Z","iopub.execute_input":"2022-10-25T15:58:37.259747Z","iopub.status.idle":"2022-10-25T15:59:23.872614Z","shell.execute_reply.started":"2022-10-25T15:58:37.259707Z","shell.execute_reply":"2022-10-25T15:59:23.871389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### All the work further we will do on the `team_A_scoring_within_10sec` target (but it will also work for the second target as well)","metadata":{}},{"cell_type":"code","source":"train_data['team_A_scoring_within_10sec'].mean()","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:01:11.149283Z","iopub.execute_input":"2022-10-25T16:01:11.149970Z","iopub.status.idle":"2022-10-25T16:01:11.161387Z","shell.execute_reply.started":"2022-10-25T16:01:11.149831Z","shell.execute_reply":"2022-10-25T16:01:11.160314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### So the mean (the prior probability to be equal to 1) for this target is ~0.0583 and if we want to use the constant prediction - this value will be the best one. Let me show you what I'm talking about:","metadata":{}},{"cell_type":"code","source":"y = train_data['team_A_scoring_within_10sec'].values\npreds_1 = np.zeros(len(y)) + np.mean(y)\npreds_2 = np.zeros(len(y)) + 0.05\npreds_3 = np.zeros(len(y)) + 0.06\npreds_4 = np.zeros(len(y)) + 0.1\npreds_5 = np.zeros(len(y)) + 0.01\nprint('LogLoss for mean prediction: {:.5f}'.format(log_loss(y, preds_1)))\nprint('LogLoss for 0.05 prediction: {:.5f}'.format(log_loss(y, preds_2)))\nprint('LogLoss for 0.06 prediction: {:.5f}'.format(log_loss(y, preds_3)))\nprint('LogLoss for 0.1  prediction: {:.5f}'.format(log_loss(y, preds_4)))\nprint('LogLoss for 0.01 prediction: {:.5f}'.format(log_loss(y, preds_5)))","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:01:18.791245Z","iopub.execute_input":"2022-10-25T16:01:18.791688Z","iopub.status.idle":"2022-10-25T16:01:20.974746Z","shell.execute_reply.started":"2022-10-25T16:01:18.791645Z","shell.execute_reply":"2022-10-25T16:01:20.973137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### As you can see, the mean prediction is really the best one and equal to `0.22229`. So we need to have the predictions near (or exactly equal) our train mean to receive good scores. Now it's time to create our train and validation sets to train the models, but we need them with the equal target mean and non-intersecting `game_num` (according to [this post](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359714)):","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import StratifiedGroupKFold\ngroups = train_data['game_num'].values\ncv = StratifiedGroupKFold(n_splits=3)\nfor train_ids, test_ids in cv.split(y, y, groups):\n    break\ntr_data = train_data.iloc[train_ids, :]\nte_data = train_data.iloc[test_ids, :]\ntr_data.shape, te_data.shape","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:01:23.110134Z","iopub.execute_input":"2022-10-25T16:01:23.110629Z","iopub.status.idle":"2022-10-25T16:01:25.950713Z","shell.execute_reply.started":"2022-10-25T16:01:23.110590Z","shell.execute_reply":"2022-10-25T16:01:25.949321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Let's check how the data is splitted:","metadata":{}},{"cell_type":"code","source":"tr_data['team_A_scoring_within_10sec'].mean(), te_data['team_A_scoring_within_10sec'].mean()","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:01:25.953072Z","iopub.execute_input":"2022-10-25T16:01:25.953936Z","iopub.status.idle":"2022-10-25T16:01:25.965412Z","shell.execute_reply.started":"2022-10-25T16:01:25.953884Z","shell.execute_reply":"2022-10-25T16:01:25.964047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"set(tr_data['game_num'].values).intersection(set(te_data['game_num'].values))","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:01:25.966595Z","iopub.execute_input":"2022-10-25T16:01:25.967614Z","iopub.status.idle":"2022-10-25T16:01:26.125173Z","shell.execute_reply.started":"2022-10-25T16:01:25.967574Z","shell.execute_reply":"2022-10-25T16:01:26.124117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The target mean between train and test is almost equal and the `game_nums` are non-intersected. Let's split the train data into train and validation parts:","metadata":{}},{"cell_type":"code","source":"y = tr_data['team_A_scoring_within_10sec'].values\ngroups = tr_data['game_num'].values\ncv = StratifiedGroupKFold(n_splits=5)\nfor train_ids, valid_ids in cv.split(y, y, groups):\n    break\nva_data = tr_data.iloc[valid_ids, :]\ntr_data = tr_data.iloc[train_ids, :]\n\ntr_data.shape, va_data.shape, te_data.shape","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:01:31.321311Z","iopub.execute_input":"2022-10-25T16:01:31.321721Z","iopub.status.idle":"2022-10-25T16:01:33.406060Z","shell.execute_reply.started":"2022-10-25T16:01:31.321685Z","shell.execute_reply":"2022-10-25T16:01:33.404851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(tr_data['team_A_scoring_within_10sec'].mean(), \n      va_data['team_A_scoring_within_10sec'].mean(), \n      te_data['team_A_scoring_within_10sec'].mean())","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:01:33.407582Z","iopub.execute_input":"2022-10-25T16:01:33.407914Z","iopub.status.idle":"2022-10-25T16:01:33.417431Z","shell.execute_reply.started":"2022-10-25T16:01:33.407883Z","shell.execute_reply":"2022-10-25T16:01:33.415998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Mean targets are a little bit different but still almost the same. Now we are ready to train the models:","metadata":{}},{"cell_type":"code","source":"y_train = tr_data['team_A_scoring_within_10sec'].values\ny_valid = va_data['team_A_scoring_within_10sec'].values\ny_test  = te_data['team_A_scoring_within_10sec'].values\nprint(y_train.shape, y_valid.shape, y_test.shape)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:01:34.251103Z","iopub.execute_input":"2022-10-25T16:01:34.251571Z","iopub.status.idle":"2022-10-25T16:01:34.259129Z","shell.execute_reply.started":"2022-10-25T16:01:34.251513Z","shell.execute_reply":"2022-10-25T16:01:34.257712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_to_drop = ['game_num', 'event_id', 'event_time', 'player_scoring_next', 'team_scoring_next', \n                'team_A_scoring_within_10sec','team_B_scoring_within_10sec']\ntr_data.drop(cols_to_drop, axis = 1, inplace = True)\nva_data.drop(cols_to_drop, axis = 1, inplace = True)\nte_data.drop(cols_to_drop, axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:01:35.574602Z","iopub.execute_input":"2022-10-25T16:01:35.575006Z","iopub.status.idle":"2022-10-25T16:01:35.793968Z","shell.execute_reply.started":"2022-10-25T16:01:35.574971Z","shell.execute_reply":"2022-10-25T16:01:35.792703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def add_dist_cols(df):\n    for i in range(6):\n        df[f'dist_p_ball_{i}'] = np.sqrt((df.ball_pos_x - df[f'p{i}_pos_x'])**2 + (df.ball_pos_y - df[f'p{i}_pos_y'])**2 + (df.ball_pos_z - df[f'p{i}_pos_z'])**2)\n    return df\n\ntr_data = add_dist_cols(tr_data)\nva_data = add_dist_cols(va_data)\nte_data = add_dist_cols(te_data)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:01:36.958693Z","iopub.execute_input":"2022-10-25T16:01:36.959391Z","iopub.status.idle":"2022-10-25T16:01:37.117723Z","shell.execute_reply.started":"2022-10-25T16:01:36.959335Z","shell.execute_reply":"2022-10-25T16:01:37.116841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tr_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:01:39.067827Z","iopub.execute_input":"2022-10-25T16:01:39.068543Z","iopub.status.idle":"2022-10-25T16:01:39.107812Z","shell.execute_reply.started":"2022-10-25T16:01:39.068503Z","shell.execute_reply":"2022-10-25T16:01:39.106841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nMODEL_PARAMS = {    \n    'objective': 'binary',\n    'seed': 42,\n    'num_leaves': 120,\n    'n_estimators': 500,\n    'max_depth': 8,\n    'learning_rate': 0.01,\n    'feature_fraction': 0.75,\n    'subsample': 0.7,\n    'subsample_freq': 8,\n    'n_jobs': -1,\n    'reg_alpha': 1,\n    'reg_lambda': 2,\n    'min_child_samples': 80}\n\nmodel = lightgbm.LGBMClassifier(**MODEL_PARAMS)\n\nmodel.fit(X=tr_data, y=y_train,\n          eval_set=[(va_data, y_valid)],\n          eval_names=['valid'],\n          callbacks=[lightgbm.log_evaluation(period=50, show_stdv=True),\n                     lightgbm.early_stopping(stopping_rounds=50)])","metadata":{"execution":{"iopub.status.busy":"2022-10-25T17:31:48.392127Z","iopub.execute_input":"2022-10-25T17:31:48.392565Z","iopub.status.idle":"2022-10-25T17:34:04.086770Z","shell.execute_reply.started":"2022-10-25T17:31:48.392527Z","shell.execute_reply":"2022-10-25T17:34:04.085516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid_pred = model.predict_proba(va_data)[:, 1]\ntest_pred = model.predict_proba(te_data)[:, 1]","metadata":{"execution":{"iopub.status.busy":"2022-10-25T17:34:04.088821Z","iopub.execute_input":"2022-10-25T17:34:04.089992Z","iopub.status.idle":"2022-10-25T17:34:23.417634Z","shell.execute_reply.started":"2022-10-25T17:34:04.089949Z","shell.execute_reply":"2022-10-25T17:34:23.416290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('LogLoss for valid_data: {:.5f}'.format(log_loss(y_valid, valid_pred)))\nprint('LogLoss for  test_data: {:.5f}'.format(log_loss(y_test, test_pred)))","metadata":{"execution":{"iopub.status.busy":"2022-10-25T17:34:23.419314Z","iopub.execute_input":"2022-10-25T17:34:23.420469Z","iopub.status.idle":"2022-10-25T17:34:23.617557Z","shell.execute_reply.started":"2022-10-25T17:34:23.420420Z","shell.execute_reply":"2022-10-25T17:34:23.616165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The LogLoss for this variant of LGBM model looks fine. Let's check the prediction histograms:","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (20, 10))\nplt.hist(valid_pred, bins = 100, alpha=0.5, label='Valid preds')\nplt.hist(test_pred, bins = 100, alpha=0.5, label='Test preds')\nplt.legend(loc='upper right')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:05:14.869854Z","iopub.execute_input":"2022-10-25T16:05:14.870244Z","iopub.status.idle":"2022-10-25T16:05:15.542197Z","shell.execute_reply.started":"2022-10-25T16:05:14.870207Z","shell.execute_reply":"2022-10-25T16:05:15.540188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### So the peaks of both predictions are near the prior mean:","metadata":{}},{"cell_type":"code","source":"print('Mean prediction for valid_data: {:.5f}'.format(np.mean(valid_pred)))\nprint('Mean prediction for  test_data: {:.5f}'.format(np.mean(test_pred)))","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:05:15.543997Z","iopub.execute_input":"2022-10-25T16:05:15.544807Z","iopub.status.idle":"2022-10-25T16:05:15.552951Z","shell.execute_reply.started":"2022-10-25T16:05:15.544764Z","shell.execute_reply":"2022-10-25T16:05:15.551484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Now it's time to change the mean for the train and valid data (using sampling) to see how it will change the situation:","metadata":{}},{"cell_type":"code","source":"# Here we create new DataFrames with target mean equal to 33%\ndef create_sampled_df(data, target):\n    data['target'] = target\n    pos_part = data[data['target'] == 1]\n    neg_part = data[data['target'] == 0]\n    df = pd.concat([pos_part, \n                   neg_part.sample(len(pos_part) * 2)])\n    return df.drop('target', axis = 1), df['target'].values\n\ntr_data_sampled, y_train_sampled = create_sampled_df(tr_data, y_train)\nva_data_sampled, y_valid_sampled = create_sampled_df(va_data, y_valid)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:05:15.555567Z","iopub.execute_input":"2022-10-25T16:05:15.556126Z","iopub.status.idle":"2022-10-25T16:05:16.447399Z","shell.execute_reply.started":"2022-10-25T16:05:15.556071Z","shell.execute_reply":"2022-10-25T16:05:16.446277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tr_data.drop('target', axis = 1, inplace = True)\nva_data.drop('target', axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:05:16.448913Z","iopub.execute_input":"2022-10-25T16:05:16.449365Z","iopub.status.idle":"2022-10-25T16:05:16.596233Z","shell.execute_reply.started":"2022-10-25T16:05:16.449319Z","shell.execute_reply":"2022-10-25T16:05:16.595070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(np.mean(y_train_sampled), np.mean(y_valid_sampled), np.mean(y_test))","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:05:16.597613Z","iopub.execute_input":"2022-10-25T16:05:16.597977Z","iopub.status.idle":"2022-10-25T16:05:16.606208Z","shell.execute_reply.started":"2022-10-25T16:05:16.597945Z","shell.execute_reply":"2022-10-25T16:05:16.604321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### So now we have train and valid target rate at 33% and test one is still ~6%. Training LGBM model once again:","metadata":{}},{"cell_type":"code","source":"%%time\n\nmodel = lightgbm.LGBMClassifier(**MODEL_PARAMS)\n\nmodel.fit(X=tr_data_sampled, y=y_train_sampled,\n          eval_set=[(va_data_sampled, y_valid_sampled)],\n          eval_names=['valid'],\n          callbacks=[lightgbm.log_evaluation(period=50, show_stdv=True),\n                     lightgbm.early_stopping(stopping_rounds=50)])","metadata":{"execution":{"iopub.status.busy":"2022-10-25T17:34:23.620095Z","iopub.execute_input":"2022-10-25T17:34:23.620590Z","iopub.status.idle":"2022-10-25T17:34:50.327599Z","shell.execute_reply.started":"2022-10-25T17:34:23.620540Z","shell.execute_reply":"2022-10-25T17:34:50.326448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Below we are predicting for the initial validation and test datasets to check the final LogLoss scores:","metadata":{}},{"cell_type":"code","source":"valid_pred_sampled = model.predict_proba(va_data)[:, 1]\ntest_pred_sampled = model.predict_proba(te_data)[:, 1]","metadata":{"execution":{"iopub.status.busy":"2022-10-25T17:34:50.331602Z","iopub.execute_input":"2022-10-25T17:34:50.332288Z","iopub.status.idle":"2022-10-25T17:34:59.992233Z","shell.execute_reply.started":"2022-10-25T17:34:50.332244Z","shell.execute_reply":"2022-10-25T17:34:59.991260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('LogLoss for valid_data: {:.5f}'.format(log_loss(y_valid, valid_pred_sampled)))\nprint('LogLoss for  test_data: {:.5f}'.format(log_loss(y_test, test_pred_sampled)))","metadata":{"execution":{"iopub.status.busy":"2022-10-25T17:35:13.065679Z","iopub.execute_input":"2022-10-25T17:35:13.066067Z","iopub.status.idle":"2022-10-25T17:35:13.271659Z","shell.execute_reply.started":"2022-10-25T17:35:13.066036Z","shell.execute_reply":"2022-10-25T17:35:13.270466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The scores become dramatically worse in comparison with the initial ones (and they will be even worse if we increase the sampled dataset target rate). Now we can check how the predictions look like in comparison with previous model:","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (20, 10))\nplt.hist(valid_pred, bins = 100, alpha=0.5, label='Valid preds')\nplt.hist(test_pred, bins = 100, alpha=0.5, label='Test preds')\nplt.hist(valid_pred_sampled, bins = 100, alpha=0.5, label='Valid preds (sampled)')\nplt.hist(test_pred_sampled, bins = 100, alpha=0.5, label='Test preds (sampled)')\nplt.legend(loc='upper right')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:05:53.299054Z","iopub.execute_input":"2022-10-25T16:05:53.299463Z","iopub.status.idle":"2022-10-25T16:05:54.436645Z","shell.execute_reply.started":"2022-10-25T16:05:53.299423Z","shell.execute_reply":"2022-10-25T16:05:54.434945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Mean prediction for         valid_data: {:.5f}'.format(np.mean(valid_pred)))\nprint('Mean prediction for          test_data: {:.5f}'.format(np.mean(test_pred)))\nprint('Mean prediction for sampled valid_data: {:.5f}'.format(np.mean(valid_pred_sampled)))\nprint('Mean prediction for sampled  test_data: {:.5f}'.format(np.mean(test_pred_sampled)))","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:05:54.438259Z","iopub.execute_input":"2022-10-25T16:05:54.438694Z","iopub.status.idle":"2022-10-25T16:05:54.449249Z","shell.execute_reply.started":"2022-10-25T16:05:54.438658Z","shell.execute_reply":"2022-10-25T16:05:54.447774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Fix time - let's use calibration\n\n### To fix the problem with changed mean we will apply simple calibration technique using linear model like `A*p+B` (where `p` is the prediction) on prediction for the full validation set","metadata":{}},{"cell_type":"code","source":"best_A = None\nbest_B = None\nbest_sc = 1\nfor A in np.arange(0.1, 0.5, 0.01):\n    for B in np.arange(-0.1, 0.1, 0.02):\n        sc = log_loss(y_valid, A * valid_pred_sampled + B)\n        if sc < best_sc:\n            best_sc = sc\n            best_A = A\n            best_B = B\n            print('Score = {:.5f}, best_A = {:.2f}, best_B = {:.2f}'.format(best_sc, best_A, best_B))\n            \nprint('='*30)\nprint('Best_score = {:.5f}, best_A = {:.2f}, best_B = {:.2f}'.format(best_sc, best_A, best_B))","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:48:37.841203Z","iopub.execute_input":"2022-10-25T16:48:37.842007Z","iopub.status.idle":"2022-10-25T16:48:59.900335Z","shell.execute_reply.started":"2022-10-25T16:48:37.841968Z","shell.execute_reply":"2022-10-25T16:48:59.898679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('LogLoss for valid_data (after calibration): {:.5f}'.format(log_loss(y_valid, best_A*valid_pred_sampled + best_B)))\nprint('LogLoss for  test_data (after calibration): {:.5f}'.format(log_loss(y_test, best_A*test_pred_sampled + best_B)))","metadata":{"execution":{"iopub.status.busy":"2022-10-25T16:50:23.421152Z","iopub.execute_input":"2022-10-25T16:50:23.421781Z","iopub.status.idle":"2022-10-25T16:50:23.630548Z","shell.execute_reply.started":"2022-10-25T16:50:23.421744Z","shell.execute_reply":"2022-10-25T16:50:23.629204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Amazing - after the calibration we have 0.21004 and 0.20689 LogLoss metrics for validation and test sets respectively instead of 0.38150 and 0.37728 beforehand. Now the scores are comparable with the usual ones and are a little bit lower because during sampling we have created a much smaller dataset than the initial one.\n\n### We can also do the same trick using the `LogisticRegression` model from sklearn library:","metadata":{}},{"cell_type":"code","source":"%%time\n\nfrom sklearn.linear_model import LogisticRegression\nclf = LogisticRegression(random_state=0).fit(valid_pred_sampled.reshape(-1, 1), y_valid)\nvalid_pred_calibrated = clf.predict_proba(valid_pred_sampled.reshape(-1, 1))[:, 1]\ntest_pred_calibrated = clf.predict_proba(test_pred_sampled.reshape(-1, 1))[:, 1]","metadata":{"execution":{"iopub.status.busy":"2022-10-25T17:36:41.058730Z","iopub.execute_input":"2022-10-25T17:36:41.059191Z","iopub.status.idle":"2022-10-25T17:36:41.704170Z","shell.execute_reply.started":"2022-10-25T17:36:41.059145Z","shell.execute_reply":"2022-10-25T17:36:41.702363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('LogLoss for valid_data (after calibration): {:.5f}'.format(log_loss(y_valid, valid_pred_calibrated)))\nprint('LogLoss for  test_data (after calibration): {:.5f}'.format(log_loss(y_test, test_pred_calibrated)))","metadata":{"execution":{"iopub.status.busy":"2022-10-25T17:36:44.885613Z","iopub.execute_input":"2022-10-25T17:36:44.887021Z","iopub.status.idle":"2022-10-25T17:36:45.121028Z","shell.execute_reply.started":"2022-10-25T17:36:44.886966Z","shell.execute_reply":"2022-10-25T17:36:45.119326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### And `LogisticRegression` performs here even better - now we have 0.20717 and 0.20370 LogLoss metrics for validation and test sets respectively. They are almost equal to the usual ones 0.20583 and 0.20236 LogLoss scores.\n\n### It's time to check how they look like on the histogram:","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (20, 10))\nplt.hist(valid_pred, bins = 100, alpha=0.5, label='Valid preds')\nplt.hist(test_pred, bins = 100, alpha=0.5, label='Test preds')\nplt.hist(valid_pred_calibrated, bins = 100, alpha=0.5, label='Valid preds (calibrated)')\nplt.hist(test_pred_calibrated, bins = 100, alpha=0.5, label='Test preds (calibrated)')\nplt.legend(loc='upper right')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-25T17:08:50.051778Z","iopub.execute_input":"2022-10-25T17:08:50.052168Z","iopub.status.idle":"2022-10-25T17:08:51.245662Z","shell.execute_reply.started":"2022-10-25T17:08:50.052135Z","shell.execute_reply":"2022-10-25T17:08:51.244332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Mean prediction for         valid_data                    : {:.5f}'.format(np.mean(valid_pred)))\nprint('Mean prediction for          test_data                    : {:.5f}'.format(np.mean(test_pred)))\nprint('Mean prediction for sampled valid_data (after calibration): {:.5f}'.format(np.mean(valid_pred_calibrated)))\nprint('Mean prediction for sampled  test_data (after calibration): {:.5f}'.format(np.mean(test_pred_calibrated)))","metadata":{"execution":{"iopub.status.busy":"2022-10-25T17:10:03.399734Z","iopub.execute_input":"2022-10-25T17:10:03.400159Z","iopub.status.idle":"2022-10-25T17:10:03.410713Z","shell.execute_reply.started":"2022-10-25T17:10:03.400123Z","shell.execute_reply":"2022-10-25T17:10:03.409357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### On the histogram we can see that all of the prediction has almost the same distribution and the means for them are almost equal. Now we can say that we fixed the sampling problem.","metadata":{}},{"cell_type":"markdown","source":"# Conclusion: if you are working with LogLoss metric, be careful with you validation and check the target rates to get the best model. Sampling and changing the target rate will kill all your work!!! But if the work is killed - you can still revive it using calibration ☺️","metadata":{}}]}