{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Public and Private LB simulations\n\npreprocessing inspired from notebook : https://www.kaggle.com/code/takanashihumbert/tps-aug22-lb-0-59013","metadata":{}},{"cell_type":"markdown","source":"Obviously it is impossible to obtain an example of the PLB but we can try to approach it even if the result remains limited and open to criticism.\nWe are going to consider only 4 groups among the 5 of the train file and take the fifth as PLB. Indeed, some have noted that the PLB was made up of a single group and therefore its result could be biased.\nWhen we do a GroupKold on 5 groups, the oof test is done on the fifth group, in our case we simply consider the fifth group as the PLB and the oof test is done on the 4th group. Of course, this will never replace the real PLB. but it allows to simulate the behavior of the oof result versus the result on the PLB\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nfrom matplotlib.ticker import MaxNLocator\nimport seaborn as sns\nfrom cycler import cycler\nfrom IPython.display import display\nimport math\nimport os\nimport random\nimport gc\nimport sys\nimport warnings\nwarnings.filterwarnings('ignore')\nfrom tqdm import tqdm\nimport optuna\nfrom colorama import Fore, Back, Style\n\nfrom sklearn.model_selection import KFold, StratifiedGroupKFold, StratifiedKFold, train_test_split, GroupKFold, RepeatedKFold, RepeatedStratifiedKFold\nfrom sklearn.metrics import roc_auc_score\nfrom sklearn.calibration import CalibrationDisplay\nfrom sklearn.preprocessing import StandardScaler,RobustScaler,LabelEncoder\nfrom sklearn.impute import KNNImputer\nfrom sklearn import linear_model\nfrom sklearn.linear_model import HuberRegressor\nfrom sklearn.decomposition import PCA\nfrom sklearn.naive_bayes import BernoulliNB\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.calibration import CalibratedClassifierCV\n\nimport matplotlib.pyplot as plt\n","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:20.252651Z","iopub.execute_input":"2022-08-13T08:24:20.253153Z","iopub.status.idle":"2022-08-13T08:24:20.265960Z","shell.execute_reply.started":"2022-08-13T08:24:20.253111Z","shell.execute_reply":"2022-08-13T08:24:20.264714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('../input/tabular-playground-series-aug-2022/train.csv')\ntest = pd.read_csv('../input/tabular-playground-series-aug-2022/test.csv')\nsubmission = pd.read_csv('../input/tabular-playground-series-aug-2022/sample_submission.csv')\ntarget = train['failure']\ntrain.drop('failure',axis=1, inplace = True)\ndata = pd.concat([train, test])\ntrain.shape,test.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:20.322468Z","iopub.execute_input":"2022-08-13T08:24:20.323387Z","iopub.status.idle":"2022-08-13T08:24:20.572644Z","shell.execute_reply.started":"2022-08-13T08:24:20.323337Z","shell.execute_reply":"2022-08-13T08:24:20.571497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# library for coding string values :\n! pip install feature_engine\nfrom feature_engine.encoding import WoEEncoder","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:20.574953Z","iopub.execute_input":"2022-08-13T08:24:20.575660Z","iopub.status.idle":"2022-08-13T08:24:31.768250Z","shell.execute_reply.started":"2022-08-13T08:24:20.575611Z","shell.execute_reply":"2022-08-13T08:24:31.766567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1> Preprocessing (nothing new)","metadata":{}},{"cell_type":"code","source":"# From Ambrosm's discussion :\ndata['m3_missing'] = data['measurement_3'].isnull().astype(np.int8)\ndata['m5_missing'] = data['measurement_5'].isnull().astype(np.int8)\n\ndata['area'] = data['attribute_2'] * data['attribute_3']\n\nfeature = [f for f in test.columns if f.startswith('measurement') or f=='loading']\n\n# dictionnary of dictionnaries (for the 11 best correlated measurement columns), \n# we will use the dictionnaries below to select the best correlated columns according to the product code)\n# Only for 'measurement_17' we make a 'manual' selection :\nfull_fill_dict ={}\nfull_fill_dict['measurement_17'] = {\n    'A': ['measurement_5','measurement_6','measurement_8'],\n    'B': ['measurement_4','measurement_5','measurement_7'],\n    'C': ['measurement_5','measurement_7','measurement_8','measurement_9'],\n    'D': ['measurement_5','measurement_6','measurement_7','measurement_8'],\n    'E': ['measurement_4','measurement_5','measurement_6','measurement_8'],\n    'F': ['measurement_4','measurement_5','measurement_6','measurement_7'],\n    'G': ['measurement_4','measurement_6','measurement_8','measurement_9'],\n    'H': ['measurement_4','measurement_5','measurement_7','measurement_8','measurement_9'],\n    'I': ['measurement_3','measurement_7','measurement_8']\n}\n\n# collect the name of the next 10 best measurement columns sorted by correlation (except 17 already done above):\ncol = [col for col in test.columns if 'measurement' not in col]+ ['loading','m3_missing','m5_missing']\na = []\nb =[]\nfor x in range(3,17):\n    corr = np.absolute(data.drop(col, axis=1).corr()[f'measurement_{x}']).sort_values(ascending=False)\n    a.append(np.round(np.sum(corr[1:4]),3)) # we add the 3 first lines of the correlation values to get the \"most correlated\"\n    b.append(f'measurement_{x}')\nc = pd.DataFrame()\nc['Selected columns'] = b\nc['correlation total'] = a\nc = c.sort_values(by = 'correlation total',ascending=False).reset_index(drop = True)\nprint(f'Columns selected by correlation sum of the 3 first rows : ')\ndisplay(c.head(10))\n\nfor i in range(10):\n    measurement_col = 'measurement_' + c.iloc[i,0][12:] # we select the next best correlated column \n    fill_dict ={}\n    for x in data.product_code.unique() : \n        corr = np.absolute(data[data.product_code == x].drop(col, axis=1).corr()[measurement_col]).sort_values(ascending=False)\n        measurement_col_dic = {}\n        measurement_col_dic[measurement_col] = corr[1:5].index.tolist()\n        fill_dict[x] = measurement_col_dic[measurement_col]\n    full_fill_dict[measurement_col] =fill_dict\n    \nfeature = [f for f in data.columns if f.startswith('measurement') or f=='loading']\nnullValue_cols = [col for col in train.columns if train[col].isnull().sum()!=0]\n    \nfor code in data.product_code.unique():\n    total_na_filled_by_linear_model = 0\n    print(f'\\n-------- Product code {code} ----------\\n')\n    print(f'filled by linear model :')\n    for measurement_col in list(full_fill_dict.keys()):\n        tmp = data[data.product_code==code]\n        column = full_fill_dict[measurement_col][code]\n        tmp_train = tmp[column+[measurement_col]].dropna(how='any')\n        tmp_test = tmp[(tmp[column].isnull().sum(axis=1)==0)&(tmp[measurement_col].isnull())]\n\n        model = HuberRegressor(epsilon=1.9)\n        model.fit(tmp_train[column], tmp_train[measurement_col])\n        data.loc[(data.product_code==code)&(data[column].isnull().sum(axis=1)==0)&(data[measurement_col].isnull()),measurement_col] = model.predict(tmp_test[column])\n        print(f'{measurement_col} : {len(tmp_test)}')\n        total_na_filled_by_linear_model += len(tmp_test)\n        \n    # others NA columns:\n    NA = data.loc[data[\"product_code\"] == code,nullValue_cols ].isnull().sum().sum()\n    model1 = KNNImputer(n_neighbors=3)\n    data.loc[data.product_code==code, feature] = model1.fit_transform(data.loc[data.product_code==code, feature])\n    print(f'\\n{total_na_filled_by_linear_model} filled by linear model ') \n    print(f'{NA} filled by KNN ')\n    \n#data['measurement_avg'] = data[[f'measurement_{i}' for i in range(3, 17)]].mean(axis=1)\nmeas_gr1_cols = [f\"measurement_{i:d}\" for i in list(range(3, 5)) + list(range(9, 17))]\ndata['meas_gr1_avg'] = np.mean(data[meas_gr1_cols], axis = 1)\ndata['meas_gr1_std'] = np.std(data[meas_gr1_cols], axis= 1) \n\nmeas_gr2_cols = [f\"measurement_{i:d}\" for i in list(range(5, 9))]\ndata['meas_gr2_avg'] = np.mean(data[meas_gr2_cols], axis = 1)\n\ndata['measurement_avg'] = data[[f'measurement_{i}' for i in range(3, 17)]].mean(axis=1)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:31.772540Z","iopub.execute_input":"2022-08-13T08:24:31.773045Z","iopub.status.idle":"2022-08-13T08:24:48.782431Z","shell.execute_reply.started":"2022-08-13T08:24:31.772991Z","shell.execute_reply":"2022-08-13T08:24:48.781213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def scale(train_data, val_data, test_data, feats):\n    scaler = StandardScaler()\n    # scaler = PowerTransformer()\n    \n    scaled_train = scaler.fit_transform(train_data[feats])\n    scaled_val = scaler.transform(val_data[feats])\n    scaled_test = scaler.transform(test_data[feats])\n    \n    #back to dataframe\n    new_train = train_data.copy()\n    new_val = val_data.copy()\n    new_test = test_data.copy()\n    \n    new_train[feats] = scaled_train\n    new_val[feats] = scaled_val\n    new_test[feats] = scaled_test\n    \n    assert len(train_data) == len(new_train)\n    assert len(val_data) == len(new_val)\n    assert len(test_data) == len(new_test)\n    \n    return new_train, new_val, new_test","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:48.785705Z","iopub.execute_input":"2022-08-13T08:24:48.786312Z","iopub.status.idle":"2022-08-13T08:24:48.795182Z","shell.execute_reply.started":"2022-08-13T08:24:48.786265Z","shell.execute_reply":"2022-08-13T08:24:48.794107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"le = LabelEncoder()\ndata['attribute_1'] = le.fit_transform(data['attribute_1'])\nle = LabelEncoder()\ndata['product_code'] = le.fit_transform(data['product_code'])","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:48.796426Z","iopub.execute_input":"2022-08-13T08:24:48.797286Z","iopub.status.idle":"2022-08-13T08:24:48.835539Z","shell.execute_reply.started":"2022-08-13T08:24:48.797247Z","shell.execute_reply":"2022-08-13T08:24:48.834457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = data.iloc[:train.shape[0],:].copy()\ntest = data.iloc[train.shape[0]:,:].copy()\nprint(train.shape, test.shape)\n\ngroups = train.product_code\nX = train\ny = target","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:48.836891Z","iopub.execute_input":"2022-08-13T08:24:48.837690Z","iopub.status.idle":"2022-08-13T08:24:48.855608Z","shell.execute_reply.started":"2022-08-13T08:24:48.837654Z","shell.execute_reply":"2022-08-13T08:24:48.854670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Thanks to @MAXSARMENTO \nwoe_encoder = WoEEncoder(variables=['attribute_0'])\nwoe_encoder.fit(X, y)\nX = woe_encoder.transform(X)\ntest = woe_encoder.transform(test)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:48.856894Z","iopub.execute_input":"2022-08-13T08:24:48.858005Z","iopub.status.idle":"2022-08-13T08:24:48.915306Z","shell.execute_reply.started":"2022-08-13T08:24:48.857942Z","shell.execute_reply":"2022-08-13T08:24:48.914131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1> Charts for LB score progression versus oof score by gradually adding features","metadata":{}},{"cell_type":"code","source":"def PLB_evolution(result_dict):\n\n    fig, (axs1, axs2, axs3, axs4, axs5) = plt.subplots(1, 5, figsize=(15,5))\n    if strategy == True :\n        plt.suptitle(f'No cross validation',fontsize=16)\n    elif strategy == skf :\n        plt.suptitle(f'Public LB score progression by gradually adding features, versus oof\\n\\n{strategy}',fontsize=16)\n    else :\n        plt.suptitle(f'{strategy}',fontsize=16)\n    \n    fig.tight_layout(pad=2)\n\n    axs1.plot(result_dict['A'].values,label = ['oof','Public LB'])\n    axs1.set_title('PLB = A')\n    axs1.set_ylabel('Auc score',fontsize=12)\n\n    axs2.plot(result_dict['B'].values,label = ['oof','Public LB'])\n    axs2.set_title('PLB = B')\n\n    axs3.plot(result_dict['C'].values,label = ['oof','Public LB'])\n    axs3.set_title('PLB = C')\n\n    axs4.plot(result_dict['D'].values,label = ['oof','Public LB'])\n    axs4.set_title('PLB = D')\n\n    axs5.plot(result_dict['E'].values,label = ['oof','Public LB'])\n    axs5.set_title('PLB = E')\n\n    for ax in fig.get_axes():\n        ax.set_xlabel('Number of features')\n        ax.legend(loc=\"center\")\n        #ax.legend(loc=\"lower right\")\n\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:48.917131Z","iopub.execute_input":"2022-08-13T08:24:48.917845Z","iopub.status.idle":"2022-08-13T08:24:48.930139Z","shell.execute_reply.started":"2022-08-13T08:24:48.917798Z","shell.execute_reply":"2022-08-13T08:24:48.928916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1>Charts for Simulated Private LB score","metadata":{}},{"cell_type":"markdown","source":"Subject to criticism...","metadata":{}},{"cell_type":"code","source":"def simulated_private_lb(pred_PLB_dict,result_dict) :\n    \n    private_LB = X[['id','product_code']].copy()\n    private_LB['pred'] = np.hstack((pred_PLB_dict['A'],pred_PLB_dict['B'],pred_PLB_dict['C'],pred_PLB_dict['D'],pred_PLB_dict['E']))\n    score_private_LB = np.round(roc_auc_score(y, private_LB['pred']), 5)\n\n    df = pd.DataFrame(index = ['A','B','C','D','E'])\n    score_oof = []\n    score_plb = []\n    for i in ['A','B','C','D','E'] :\n        score_oof.append(result_dict[i].iloc[-1,0])\n        score_plb.append(result_dict[i].iloc[-1,1])\n    \n    barWidth = 0.3\n    y1 = list((np.array(score_oof)))\n    y2 = list((np.array(score_plb)))\n    y3 = list((np.array([score_private_LB, score_private_LB, score_private_LB, score_private_LB,score_private_LB])))\n    r1 = range(len(y1))\n    r2 = [x + barWidth for x in r1]\n    #r3 = [x + barWidth for x in r2]\n\n    plt.figure(figsize =(10,7))\n    \n    plt.bar(r1, y1, width = barWidth, color = ['blue' for i in y1],label = 'oof')\n    plt.bar(r2, y2, width = barWidth, color = ['orange' for i in y1], label = 'Public LB with only one product code')\n    #plt.bar(r3, y3, width = barWidth, color = ['red' for i in y1], label = 'Private LB (Public LB per product concatenated)')\n    plt.yticks(np.arange(0.5,0.6,0.001))\n    plt.ylim((0.58,0.6))\n    plt.grid(axis='y',ls ='dashed')\n    plt.hlines(y=score_private_LB, xmin=-0.5, xmax=4.5, linewidth=2, color='r',label = 'Private LB score')\n    plt.hlines(y=np.mean(np.array(score_oof)), xmin=-0.5, xmax=4.5, linewidth=2, color='b',label = 'oof mean')\n\n    plt.legend()\n    if strategy == True :\n        plt.title(label =f'\\nSimulated Private LB score :{score_private_LB}\\n\\nNo cross validation\\n ',fontsize=16)\n    else :\n        plt.title(label =f'\\nSimulated Private LB score :{score_private_LB}\\n\\n{strategy}\\n',fontsize=16)\n\n    plt.xticks([r + barWidth / 2 for r in range(len(y1))], ['PLB with A', 'PLB with B', 'PLB with C', 'PLB with D','PLB with E'])\n\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:48.931755Z","iopub.execute_input":"2022-08-13T08:24:48.932167Z","iopub.status.idle":"2022-08-13T08:24:48.950450Z","shell.execute_reply.started":"2022-08-13T08:24:48.932130Z","shell.execute_reply":"2022-08-13T08:24:48.949238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1> Cross validation strategy and training","metadata":{}},{"cell_type":"code","source":"def scoring_feature(features,cross_val) :\n   \n    lr_oof_scoring_feature = np.zeros(len(X_new))\n\n    if cross_val == True:\n        scaler = StandardScaler()\n        scaled_train = scaler.fit_transform(X_new[features])\n        scaled_test = scaler.transform(X_PLB[features])\n        model = linear_model.LogisticRegression(max_iter=200, C=0.0001, penalty='l2', solver='newton-cg')\n        model.fit(scaled_train, y_new)\n        val_preds = model.predict_proba(scaled_train)[:, 1]\n        lr_oof_scoring_feature = val_preds\n        \n        score_scoring_feature = np.round(roc_auc_score(y_new, lr_oof_scoring_feature), 5)\n        pred_PLB_scoring_feature = model.predict_proba(scaled_test)[:, 1]\n        score_PLB_scoring_feature = np.round(roc_auc_score(y_PLB,pred_PLB_scoring_feature), 5)\n \n    else :\n        for fold_idx, (train_idx, val_idx) in enumerate(cross_val.split(X_new, y_new,groups = X_new.product_code )):\n            x_train, x_val = X_new.iloc[train_idx], X_new.iloc[val_idx]\n            y_train, y_val = y_new.iloc[train_idx], y_new.iloc[val_idx]\n            x_train, x_val, x_test = scale(x_train, x_val, X_PLB,features)\n\n            model = linear_model.LogisticRegression(max_iter=200, C=0.0001, penalty='l2', solver='newton-cg')\n            model.fit(x_train[features], y_train)\n            val_preds = model.predict_proba(x_val[features])[:, 1]\n            lr_oof_scoring_feature[val_idx] = val_preds\n\n        score_scoring_feature = np.round(roc_auc_score(y_new, lr_oof_scoring_feature), 5)\n        pred_PLB_scoring_feature = model.predict_proba(x_test[features])[:, 1]\n        score_PLB_scoring_feature = np.round(roc_auc_score(y_PLB,pred_PLB_scoring_feature), 5)\n        \n    return score_scoring_feature, score_PLB_scoring_feature, pred_PLB_scoring_feature","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:48.954148Z","iopub.execute_input":"2022-08-13T08:24:48.954593Z","iopub.status.idle":"2022-08-13T08:24:48.970028Z","shell.execute_reply.started":"2022-08-13T08:24:48.954555Z","shell.execute_reply":"2022-08-13T08:24:48.969045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1> Gradual addition of features","metadata":{}},{"cell_type":"code","source":"def features_selected(features,cross_val):\n    \n    feat_list_features_selected = []\n    score_list_features_selected = []\n    score_PLB_list_features_selected = []\n    baseline = ['loading']\n    \n    for run,feat in enumerate(features) : \n\n        select_feature = baseline + [feat]\n        score_features_selected, score_PLB_features_selected, pred_PLB_features_selected = scoring_feature(select_feature,cross_val)\n\n        baseline = baseline + [feat]\n        feat_list_features_selected.append(feat)\n        score_list_features_selected.append(score_features_selected)\n        score_PLB_list_features_selected.append(score_PLB_features_selected)\n\n    return feat_list_features_selected, score_list_features_selected, score_PLB_list_features_selected, pred_PLB_features_selected","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:48.971801Z","iopub.execute_input":"2022-08-13T08:24:48.972668Z","iopub.status.idle":"2022-08-13T08:24:48.985139Z","shell.execute_reply.started":"2022-08-13T08:24:48.972632Z","shell.execute_reply":"2022-08-13T08:24:48.983970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1> Launch of calculations","metadata":{}},{"cell_type":"code","source":"features = [\n        'measurement_17',\n        'attribute_0',\n        'measurement_0',\n        'measurement_1',\n        'measurement_2',\n        'm3_missing',\n        'm5_missing',\n        'attribute_1',\n        'measurement_4'\n        ]\n\nskf = StratifiedKFold(n_splits=4,shuffle = True,random_state=1)\ngkf = GroupKFold(n_splits=4)\nrkf = KFold(n_splits=4,shuffle = True, random_state=1)\nno_cross_val = True\n\nfor PLB in [True,False] :\n    \n    for strategy in [skf,gkf,rkf,no_cross_val] : \n\n        if strategy == no_cross_val:\n            strategy = True\n\n        result_dict = {}\n        pred_PLB_dict = {}\n\n        for product in X.product_code.unique() :\n\n            X_PLB = X[X.product_code == product]\n            y_PLB = y.loc[X_PLB.index]\n            X_new = X[X.product_code != product]\n            y_new = y.loc[X_new.index]\n            product_name = le.classes_[product]\n            feat_list, score_list, score_PLB_list, pred_PLB = features_selected(features, cross_val = strategy)\n            result = pd.DataFrame(index = feat_list)\n            result[f'score_oof_without_{product_name}'] = score_list\n            result[f'score_PLB_simulated_with_{product_name}'] = score_PLB_list\n\n            result_dict[product_name] = result\n            pred_PLB_dict[product_name] = pred_PLB\n\n        if PLB == True :\n            PLB_evolution(result_dict)\n        else :\n            simulated_private_lb(pred_PLB_dict,result_dict)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T08:24:48.986433Z","iopub.execute_input":"2022-08-13T08:24:48.987216Z","iopub.status.idle":"2022-08-13T08:26:33.195271Z","shell.execute_reply.started":"2022-08-13T08:24:48.987177Z","shell.execute_reply":"2022-08-13T08:26:33.193834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1> Conclusions","metadata":{}},{"cell_type":"markdown","source":"**Preliminary remarks :**\n\nBehavior and results depend on features chosen, so this could be a major bias\n\n**From PLB evoution :**\n1. The trend of the PLB score evolution (better or worse) compared to the oof score depends little on the cross validation strategy.\n\n2. We don't see a huge difference between Stratified and GroupKfold for the PLB results on the other hand the result oof is quickly superior for stratified than for groupkfold with an overfitting risk, visible for group C and E.\n\n3. As we can anticipate, GroupKfold better manage the complicated groups (C and E) with less overfitting but not unfortunately with a huge better PLB score...\n\n4. No cross validation, Stratified and Kfold have similar behavior\n\n**From simulated private LB**\n\n1. Without cross validation seems to provide a better sumulated private score\n2. For the complicated groups (C and E) GroupKfold gives Private LB worse performance than others cross validation\n3. GroupKfold oof score (mean) is the only one that is worse than the simulated private score.\n\n**Synthesis**\n\n1. 2 groups on 5 provides difficulties for prediction. So if the Public LB is based upon a difficult group, GroupKold can give worse result than others cross validation.\n2. GroupKold provides the better aligned oof score with the simulated private score (smaller gap)\n3. Kfold provides the worse result for Private LB and no cross validation the best but there is maybe a biais for no cross validation. Obviously No cross validation has a best oof score because not oof...\n4. If we choose GroupKfold, we can not easily overfit the PLB and the Private LB score could be better than the oof score but it does not mean that it will be the best score we can achieve.\n\n**My appologies, in case of stupid conclusions :-)**\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}