{
  "id": 372030,
  "title": "What is wrong with my GBT Ranker Pipeline? [LB = CV -> lower than expected]",
  "url": "/competitions/otto-recommender-system/discussion/372030",
  "author_name": "Simon Veitner",
  "post_date": "2022-12-13T18:45:57.007000",
  "votes": 19,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Hello everyone.<br>\nI face the problem that the LB of my Ranker model is much lower than what I would it expect to be from the CV.<br>\nThe CV is 0.565 and the LB aswell. From what I had before I expeted the LB in the region of 0.575.<br>\nI describe my pipeline (which is essentially the one described by Chris) in greater detail in the hope that somebody can help me to fix the problem.<br>\n 1) Step: I use the validation dataset given by chris <a href=\"https://www.kaggle.com/datasets/cdeotte/otto-validation\" target=\"_blank\">here</a> (I call it in following dataset A) to generate co visitation matrices. I use these co visitation matrices to generate candidates. For example the clicks candidates are generated as follows:</p>\n<pre><code>def suggest_clicks(df):\n    # USE USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    result = list(dict.fromkeys(unique_aids + aids2))\n    # USE TOP20 TEST CLICKS\n    return list(dict.fromkeys(result + list(top_clicks)))[:30]\n</code></pre>\n<p>I apply this function for each user and explode then to get an N * 30 long dataframe where N is number of users.</p>\n<p>2) Step: Using the same dataset as in 1) i generate features for items and users and also interaction features.<br>\nFor example for item features i use both train and test set from dataset A and generate them in the following way:</p>\n<pre><code>T = (data.ts.max()-data.ts.min())/(24*60*60)\nTmin = data.ts.min()\ndata.ts -= Tmin\nitem_features = data.groupby('aid').agg({'aid':'count','session':'nunique','type':'mean', 'ts':['min', 'max']})\nitem_features.columns = ['item_item_count','item_user_count','item_buy_ratio', 'item_first_time_seen', 'item_last_time_seen']\n# CONVERT COLUMNS TO INT32 and FLOAT32 HERE\nitem_features.item_item_count /= T\nitem_features.item_user_count /= T\nitem_features.to_pandas().to_parquet('item_features.pqt')\n</code></pre>\n<p>I scale by the total time to be sure that I can generate same features for the \"real\" training and test set. data above is simply given by <code>data = cudf.concat([train, test])</code>.</p>\n<p>User features i generate just with the test data from dataset A:</p>\n<pre><code>T = (test.ts.max()-test.ts.min())/(24*60*60)\nTmin = test.ts.min()\ntest.ts -= Tmin\nuser_features = test.groupby('session').agg({'session':'count','aid':'nunique','type':'mean', 'ts':['min', 'max']})\nuser_features.columns = ['user_user_count','user_item_count','user_buy_ratio', 'user_first_time_seen', 'user_last_time_seen']\n# CONVERT COLUMNS TO INT32 and FLOAT32 HERE\nuser_features.user_user_count /= T\nuser_features.user_item_count /= T\nuser_features.to_parquet('user_features.pqt')\n</code></pre>\n<p>And an example of interaction dataframe is:</p>\n<pre><code>items_clicked = test[[\"session\", \"aid\"]].loc[test.type == 0].drop_duplicates(subset=[\"session\",\"aid\"]).rename({'session':'user', 'aid':'item'},axis=1)\nitems_clicked[\"item_clicked\"] = 1\nitems_clicked.to_parquet('items_clicked.pqt')\n</code></pre>\n<p>.</p>\n<p>3) Step:<br>\nI train the reranker after generating candidates dataframe. The steps taken to generate the dataframe for candidates are as follows (for example for clicks)</p>\n<pre><code>candidates = pd.read_parquet(\"cand_clicks.pqt\")\ncandidates.columns = [\"session\", \"aid\"]\n\nitem_features = pd.read_parquet('item_features.pqt')\ncandidates = candidates.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\nuser_features = pd.read_parquet('user_features.pqt')\ncandidates = candidates.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\ndel item_features, user_features\n_ = gc.collect()\n\ntar = cudf.read_parquet('/notebooks/pipeline_cv_score/otto/test_labels.parquet')\ntar = tar.loc[ tar['type']=='clicks' ]\naids = tar.ground_truth.explode().astype('int32').rename('item')\ntar = tar[['session']].astype('int32').rename({'session':'user'},axis=1)\ntar = tar.merge(aids, left_index=True, right_index=True, how='left')\ntar['click'] = 1\n\nitems_interacted = cudf.read_parquet(\"items_interacted.pqt\")\nitems_clicked = cudf.read_parquet(\"items_clicked.pqt\")\nitems_carted = cudf.read_parquet(\"items_carts.pqt\")\nitems_ordered = cudf.read_parquet(\"items_orders.pqt\")\ncandidates_ = cudf.DataFrame(candidates)\n\ncandidates_ = candidates_.rename({'session':'user', 'aid':'item'},axis=1)\ncandidates_ = candidates_.merge(items_interacted,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(items_clicked,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(items_carted,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(items_ordered,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(tar,on=['user','item'],how='left').fillna(0)\n</code></pre>\n<p>Then I train as suggested by Chris:</p>\n<pre><code>import xgboost as xgb\nfrom sklearn.model_selection import GroupKFold\n\nif TRAIN:\n    FEATURES = candidates.columns[2 : -1]\n\n    skf = GroupKFold(n_splits=5)\n    for fold,(train_idx, valid_idx) in enumerate(skf.split(candidates, candidates['click'], groups=candidates['user'] )):\n        print(f\"Train fold {fold}\")\n\n        X_train = candidates.loc[train_idx, FEATURES]\n        y_train = candidates.loc[train_idx, 'click']\n        X_valid = candidates.loc[valid_idx, FEATURES]\n        y_valid = candidates.loc[valid_idx, 'click']\n\n        # IF YOU HAVE 50 CANDIDATE WE USE 50 BELOW\n        dtrain = xgb.DMatrix(X_train, y_train, group=[30] * (len(train_idx)//30) ) \n        dvalid = xgb.DMatrix(X_valid, y_valid, group=[30] * (len(valid_idx)//30) ) \n\n        xgb_parms = {'objective':'rank:pairwise', 'tree_method':'gpu_hist'}\n        model = xgb.train(xgb_parms, \n            dtrain=dtrain,\n            evals=[(dtrain,'train'),(dvalid,'valid')],\n            num_boost_round=30,\n            verbose_eval=100)\n        model.save_model(f'XGB_fold{fold}_click.xgb')\n</code></pre>\n<p>I infer in the following way to calculate CV score:</p>\n<pre><code>FEATURES = candidates.columns[2 : -1]\n\ndef pred(df):\n    preds = np.zeros(len(df))\n    for fold in range(5):\n        print(f\"Inferece for fold {fold}\")\n        model = xgb.Booster()\n        model.load_model(f'XGB_fold{fold}_click.xgb')\n        model.set_param({'predictor': 'gpu_predictor'})\n        dtest = xgb.DMatrix(data=df[FEATURES])\n        preds += model.predict(dtest)/5\n    predictions = df[['user','item']].copy()\n    predictions['pred'] = preds\n    predictions = predictions.sort_values(['user','pred'], ascending=[True,False]).reset_index(drop=True)\n    predictions['n'] = predictions.groupby('user').item.cumcount().astype('int8')\n    predictions = predictions.loc[predictions.n&lt;20]\n    pred_df_clicks = predictions.groupby('user').item.apply(list)\n    pred_df_clicks = pred_df_clicks.to_frame().reset_index()\n    pred_df_clicks.item = pred_df_clicks.item.apply(lambda x: \" \".join(map(str,x)))\n    pred_df_clicks.columns = ['session_type','labels']\n    pred_df_clicks.session_type = pred_df_clicks.session_type.astype('str')+ '_clicks'\n    return pred_df_clicks\n\n# SPLIT FOR MEMORY MANAGMENT:\ntest_candidates1 = test_candidates.loc[test_candidates.user.isin(grp_idx[:len(grp_idx)//2])]\ntest_candidates2 = test_candidates.loc[test_candidates.user.isin(grp_idx[len(grp_idx)//2:])]\nprint(\"start with prediction 1.\")\npred_df_clicks1 = pred(test_candidates1)\nprint(\"start with prediction 2.\")\npred_df_clicks2 = pred(test_candidates2)\nprint(\"concat...\")\npred_df_clicks = pd.concat([pred_df_clicks1, pred_df_clicks2])\ndel pred_df_clicks1, pred_df_clicks2\n_ = gc.collect\n</code></pre>\n<p>The same I do for carts and buy and then i concat my dataframes and calculcate CV as in co visitation notebook.<br>\nI get CV 0.565.</p>\n<p>---LB PREP---<br>\nI repeat steps 1) and 2) but I use the following dataset <a href=\"https://www.kaggle.com/datasets/columbia2131/otto-chunk-data-inparquet-format\" target=\"_blank\">here</a>.<br>\nThen I prepare the test candidates in step 3 exactly same way BUT i dont have test labels. So i leave out the merge of tar.<br>\nThen I infer and get resulting dataframe.<br>\nWhen I submit the resulting dataframe I only get LB 0.565. (exactly the same as CV) In fact I did multiple experiments and the LB always closely resembles the CV score but I expect to be higher (like for co visitation approach!)<br>\nCan anyone point me to the mistake I made? I cant see it.</p>",
  "messages": [
    {
      "id": 2064435,
      "postDate": "2022-12-13T18:45:57.007Z",
      "content": "<p>Hello everyone.<br>\nI face the problem that the LB of my Ranker model is much lower than what I would it expect to be from the CV.<br>\nThe CV is 0.565 and the LB aswell. From what I had before I expeted the LB in the region of 0.575.<br>\nI describe my pipeline (which is essentially the one described by Chris) in greater detail in the hope that somebody can help me to fix the problem.<br>\n 1) Step: I use the validation dataset given by chris <a href=\"https://www.kaggle.com/datasets/cdeotte/otto-validation\" target=\"_blank\">here</a> (I call it in following dataset A) to generate co visitation matrices. I use these co visitation matrices to generate candidates. For example the clicks candidates are generated as follows:</p>\n<pre><code>def suggest_clicks(df):\n    # USE USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    result = list(dict.fromkeys(unique_aids + aids2))\n    # USE TOP20 TEST CLICKS\n    return list(dict.fromkeys(result + list(top_clicks)))[:30]\n</code></pre>\n<p>I apply this function for each user and explode then to get an N * 30 long dataframe where N is number of users.</p>\n<p>2) Step: Using the same dataset as in 1) i generate features for items and users and also interaction features.<br>\nFor example for item features i use both train and test set from dataset A and generate them in the following way:</p>\n<pre><code>T = (data.ts.max()-data.ts.min())/(24*60*60)\nTmin = data.ts.min()\ndata.ts -= Tmin\nitem_features = data.groupby('aid').agg({'aid':'count','session':'nunique','type':'mean', 'ts':['min', 'max']})\nitem_features.columns = ['item_item_count','item_user_count','item_buy_ratio', 'item_first_time_seen', 'item_last_time_seen']\n# CONVERT COLUMNS TO INT32 and FLOAT32 HERE\nitem_features.item_item_count /= T\nitem_features.item_user_count /= T\nitem_features.to_pandas().to_parquet('item_features.pqt')\n</code></pre>\n<p>I scale by the total time to be sure that I can generate same features for the \"real\" training and test set. data above is simply given by <code>data = cudf.concat([train, test])</code>.</p>\n<p>User features i generate just with the test data from dataset A:</p>\n<pre><code>T = (test.ts.max()-test.ts.min())/(24*60*60)\nTmin = test.ts.min()\ntest.ts -= Tmin\nuser_features = test.groupby('session').agg({'session':'count','aid':'nunique','type':'mean', 'ts':['min', 'max']})\nuser_features.columns = ['user_user_count','user_item_count','user_buy_ratio', 'user_first_time_seen', 'user_last_time_seen']\n# CONVERT COLUMNS TO INT32 and FLOAT32 HERE\nuser_features.user_user_count /= T\nuser_features.user_item_count /= T\nuser_features.to_parquet('user_features.pqt')\n</code></pre>\n<p>And an example of interaction dataframe is:</p>\n<pre><code>items_clicked = test[[\"session\", \"aid\"]].loc[test.type == 0].drop_duplicates(subset=[\"session\",\"aid\"]).rename({'session':'user', 'aid':'item'},axis=1)\nitems_clicked[\"item_clicked\"] = 1\nitems_clicked.to_parquet('items_clicked.pqt')\n</code></pre>\n<p>.</p>\n<p>3) Step:<br>\nI train the reranker after generating candidates dataframe. The steps taken to generate the dataframe for candidates are as follows (for example for clicks)</p>\n<pre><code>candidates = pd.read_parquet(\"cand_clicks.pqt\")\ncandidates.columns = [\"session\", \"aid\"]\n\nitem_features = pd.read_parquet('item_features.pqt')\ncandidates = candidates.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\nuser_features = pd.read_parquet('user_features.pqt')\ncandidates = candidates.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\ndel item_features, user_features\n_ = gc.collect()\n\ntar = cudf.read_parquet('/notebooks/pipeline_cv_score/otto/test_labels.parquet')\ntar = tar.loc[ tar['type']=='clicks' ]\naids = tar.ground_truth.explode().astype('int32').rename('item')\ntar = tar[['session']].astype('int32').rename({'session':'user'},axis=1)\ntar = tar.merge(aids, left_index=True, right_index=True, how='left')\ntar['click'] = 1\n\nitems_interacted = cudf.read_parquet(\"items_interacted.pqt\")\nitems_clicked = cudf.read_parquet(\"items_clicked.pqt\")\nitems_carted = cudf.read_parquet(\"items_carts.pqt\")\nitems_ordered = cudf.read_parquet(\"items_orders.pqt\")\ncandidates_ = cudf.DataFrame(candidates)\n\ncandidates_ = candidates_.rename({'session':'user', 'aid':'item'},axis=1)\ncandidates_ = candidates_.merge(items_interacted,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(items_clicked,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(items_carted,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(items_ordered,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(tar,on=['user','item'],how='left').fillna(0)\n</code></pre>\n<p>Then I train as suggested by Chris:</p>\n<pre><code>import xgboost as xgb\nfrom sklearn.model_selection import GroupKFold\n\nif TRAIN:\n    FEATURES = candidates.columns[2 : -1]\n\n    skf = GroupKFold(n_splits=5)\n    for fold,(train_idx, valid_idx) in enumerate(skf.split(candidates, candidates['click'], groups=candidates['user'] )):\n        print(f\"Train fold {fold}\")\n\n        X_train = candidates.loc[train_idx, FEATURES]\n        y_train = candidates.loc[train_idx, 'click']\n        X_valid = candidates.loc[valid_idx, FEATURES]\n        y_valid = candidates.loc[valid_idx, 'click']\n\n        # IF YOU HAVE 50 CANDIDATE WE USE 50 BELOW\n        dtrain = xgb.DMatrix(X_train, y_train, group=[30] * (len(train_idx)//30) ) \n        dvalid = xgb.DMatrix(X_valid, y_valid, group=[30] * (len(valid_idx)//30) ) \n\n        xgb_parms = {'objective':'rank:pairwise', 'tree_method':'gpu_hist'}\n        model = xgb.train(xgb_parms, \n            dtrain=dtrain,\n            evals=[(dtrain,'train'),(dvalid,'valid')],\n            num_boost_round=30,\n            verbose_eval=100)\n        model.save_model(f'XGB_fold{fold}_click.xgb')\n</code></pre>\n<p>I infer in the following way to calculate CV score:</p>\n<pre><code>FEATURES = candidates.columns[2 : -1]\n\ndef pred(df):\n    preds = np.zeros(len(df))\n    for fold in range(5):\n        print(f\"Inferece for fold {fold}\")\n        model = xgb.Booster()\n        model.load_model(f'XGB_fold{fold}_click.xgb')\n        model.set_param({'predictor': 'gpu_predictor'})\n        dtest = xgb.DMatrix(data=df[FEATURES])\n        preds += model.predict(dtest)/5\n    predictions = df[['user','item']].copy()\n    predictions['pred'] = preds\n    predictions = predictions.sort_values(['user','pred'], ascending=[True,False]).reset_index(drop=True)\n    predictions['n'] = predictions.groupby('user').item.cumcount().astype('int8')\n    predictions = predictions.loc[predictions.n&lt;20]\n    pred_df_clicks = predictions.groupby('user').item.apply(list)\n    pred_df_clicks = pred_df_clicks.to_frame().reset_index()\n    pred_df_clicks.item = pred_df_clicks.item.apply(lambda x: \" \".join(map(str,x)))\n    pred_df_clicks.columns = ['session_type','labels']\n    pred_df_clicks.session_type = pred_df_clicks.session_type.astype('str')+ '_clicks'\n    return pred_df_clicks\n\n# SPLIT FOR MEMORY MANAGMENT:\ntest_candidates1 = test_candidates.loc[test_candidates.user.isin(grp_idx[:len(grp_idx)//2])]\ntest_candidates2 = test_candidates.loc[test_candidates.user.isin(grp_idx[len(grp_idx)//2:])]\nprint(\"start with prediction 1.\")\npred_df_clicks1 = pred(test_candidates1)\nprint(\"start with prediction 2.\")\npred_df_clicks2 = pred(test_candidates2)\nprint(\"concat...\")\npred_df_clicks = pd.concat([pred_df_clicks1, pred_df_clicks2])\ndel pred_df_clicks1, pred_df_clicks2\n_ = gc.collect\n</code></pre>\n<p>The same I do for carts and buy and then i concat my dataframes and calculcate CV as in co visitation notebook.<br>\nI get CV 0.565.</p>\n<p>---LB PREP---<br>\nI repeat steps 1) and 2) but I use the following dataset <a href=\"https://www.kaggle.com/datasets/columbia2131/otto-chunk-data-inparquet-format\" target=\"_blank\">here</a>.<br>\nThen I prepare the test candidates in step 3 exactly same way BUT i dont have test labels. So i leave out the merge of tar.<br>\nThen I infer and get resulting dataframe.<br>\nWhen I submit the resulting dataframe I only get LB 0.565. (exactly the same as CV) In fact I did multiple experiments and the LB always closely resembles the CV score but I expect to be higher (like for co visitation approach!)<br>\nCan anyone point me to the mistake I made? I cant see it.</p>",
      "rawMarkdown": "Hello everyone.\nI face the problem that the LB of my Ranker model is much lower than what I would it expect to be from the CV.\nThe CV is 0.565 and the LB aswell. From what I had before I expeted the LB in the region of 0.575.\nI describe my pipeline (which is essentially the one described by Chris) in greater detail in the hope that somebody can help me to fix the problem.\n 1) Step: I use the validation dataset given by chris [here](https://www.kaggle.com/datasets/cdeotte/otto-validation) (I call it in following dataset A) to generate co visitation matrices. I use these co visitation matrices to generate candidates. For example the clicks candidates are generated as follows:\n```\ndef suggest_clicks(df):\n    # USE USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    result = list(dict.fromkeys(unique_aids + aids2))\n    # USE TOP20 TEST CLICKS\n    return list(dict.fromkeys(result + list(top_clicks)))[:30]\n```\nI apply this function for each user and explode then to get an N * 30 long dataframe where N is number of users.\n\n2) Step: Using the same dataset as in 1) i generate features for items and users and also interaction features.\nFor example for item features i use both train and test set from dataset A and generate them in the following way:\n```\nT = (data.ts.max()-data.ts.min())/(24*60*60)\nTmin = data.ts.min()\ndata.ts -= Tmin\nitem_features = data.groupby('aid').agg({'aid':'count','session':'nunique','type':'mean', 'ts':['min', 'max']})\nitem_features.columns = ['item_item_count','item_user_count','item_buy_ratio', 'item_first_time_seen', 'item_last_time_seen']\n# CONVERT COLUMNS TO INT32 and FLOAT32 HERE\nitem_features.item_item_count /= T\nitem_features.item_user_count /= T\nitem_features.to_pandas().to_parquet('item_features.pqt')\n```\nI scale by the total time to be sure that I can generate same features for the \"real\" training and test set. data above is simply given by `data = cudf.concat([train, test])`.\n\nUser features i generate just with the test data from dataset A:\n```\nT = (test.ts.max()-test.ts.min())/(24*60*60)\nTmin = test.ts.min()\ntest.ts -= Tmin\nuser_features = test.groupby('session').agg({'session':'count','aid':'nunique','type':'mean', 'ts':['min', 'max']})\nuser_features.columns = ['user_user_count','user_item_count','user_buy_ratio', 'user_first_time_seen', 'user_last_time_seen']\n# CONVERT COLUMNS TO INT32 and FLOAT32 HERE\nuser_features.user_user_count /= T\nuser_features.user_item_count /= T\nuser_features.to_parquet('user_features.pqt')\n```\nAnd an example of interaction dataframe is:\n```\nitems_clicked = test[[\"session\", \"aid\"]].loc[test.type == 0].drop_duplicates(subset=[\"session\",\"aid\"]).rename({'session':'user', 'aid':'item'},axis=1)\nitems_clicked[\"item_clicked\"] = 1\nitems_clicked.to_parquet('items_clicked.pqt')\n```.\n\n3) Step:\nI train the reranker after generating candidates dataframe. The steps taken to generate the dataframe for candidates are as follows (for example for clicks)\n\n```\ncandidates = pd.read_parquet(\"cand_clicks.pqt\")\ncandidates.columns = [\"session\", \"aid\"]\n\nitem_features = pd.read_parquet('item_features.pqt')\ncandidates = candidates.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\nuser_features = pd.read_parquet('user_features.pqt')\ncandidates = candidates.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\ndel item_features, user_features\n_ = gc.collect()\n\ntar = cudf.read_parquet('/notebooks/pipeline_cv_score/otto/test_labels.parquet')\ntar = tar.loc[ tar['type']=='clicks' ]\naids = tar.ground_truth.explode().astype('int32').rename('item')\ntar = tar[['session']].astype('int32').rename({'session':'user'},axis=1)\ntar = tar.merge(aids, left_index=True, right_index=True, how='left')\ntar['click'] = 1\n\nitems_interacted = cudf.read_parquet(\"items_interacted.pqt\")\nitems_clicked = cudf.read_parquet(\"items_clicked.pqt\")\nitems_carted = cudf.read_parquet(\"items_carts.pqt\")\nitems_ordered = cudf.read_parquet(\"items_orders.pqt\")\ncandidates_ = cudf.DataFrame(candidates)\n\ncandidates_ = candidates_.rename({'session':'user', 'aid':'item'},axis=1)\ncandidates_ = candidates_.merge(items_interacted,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(items_clicked,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(items_carted,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(items_ordered,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(tar,on=['user','item'],how='left').fillna(0)\n```\nThen I train as suggested by Chris:\n```\nimport xgboost as xgb\nfrom sklearn.model_selection import GroupKFold\n\nif TRAIN:\n    FEATURES = candidates.columns[2 : -1]\n\n    skf = GroupKFold(n_splits=5)\n    for fold,(train_idx, valid_idx) in enumerate(skf.split(candidates, candidates['click'], groups=candidates['user'] )):\n        print(f\"Train fold {fold}\")\n\n        X_train = candidates.loc[train_idx, FEATURES]\n        y_train = candidates.loc[train_idx, 'click']\n        X_valid = candidates.loc[valid_idx, FEATURES]\n        y_valid = candidates.loc[valid_idx, 'click']\n\n        # IF YOU HAVE 50 CANDIDATE WE USE 50 BELOW\n        dtrain = xgb.DMatrix(X_train, y_train, group=[30] * (len(train_idx)//30) ) \n        dvalid = xgb.DMatrix(X_valid, y_valid, group=[30] * (len(valid_idx)//30) ) \n\n        xgb_parms = {'objective':'rank:pairwise', 'tree_method':'gpu_hist'}\n        model = xgb.train(xgb_parms, \n            dtrain=dtrain,\n            evals=[(dtrain,'train'),(dvalid,'valid')],\n            num_boost_round=30,\n            verbose_eval=100)\n        model.save_model(f'XGB_fold{fold}_click.xgb')\n```\n\nI infer in the following way to calculate CV score:\n```\nFEATURES = candidates.columns[2 : -1]\n\ndef pred(df):\n    preds = np.zeros(len(df))\n    for fold in range(5):\n        print(f\"Inferece for fold {fold}\")\n        model = xgb.Booster()\n        model.load_model(f'XGB_fold{fold}_click.xgb')\n        model.set_param({'predictor': 'gpu_predictor'})\n        dtest = xgb.DMatrix(data=df[FEATURES])\n        preds += model.predict(dtest)/5\n    predictions = df[['user','item']].copy()\n    predictions['pred'] = preds\n    predictions = predictions.sort_values(['user','pred'], ascending=[True,False]).reset_index(drop=True)\n    predictions['n'] = predictions.groupby('user').item.cumcount().astype('int8')\n    predictions = predictions.loc[predictions.n<20]\n    pred_df_clicks = predictions.groupby('user').item.apply(list)\n    pred_df_clicks = pred_df_clicks.to_frame().reset_index()\n    pred_df_clicks.item = pred_df_clicks.item.apply(lambda x: \" \".join(map(str,x)))\n    pred_df_clicks.columns = ['session_type','labels']\n    pred_df_clicks.session_type = pred_df_clicks.session_type.astype('str')+ '_clicks'\n    return pred_df_clicks\n\n# SPLIT FOR MEMORY MANAGMENT:\ntest_candidates1 = test_candidates.loc[test_candidates.user.isin(grp_idx[:len(grp_idx)//2])]\ntest_candidates2 = test_candidates.loc[test_candidates.user.isin(grp_idx[len(grp_idx)//2:])]\nprint(\"start with prediction 1.\")\npred_df_clicks1 = pred(test_candidates1)\nprint(\"start with prediction 2.\")\npred_df_clicks2 = pred(test_candidates2)\nprint(\"concat...\")\npred_df_clicks = pd.concat([pred_df_clicks1, pred_df_clicks2])\ndel pred_df_clicks1, pred_df_clicks2\n_ = gc.collect\n```\n\nThe same I do for carts and buy and then i concat my dataframes and calculcate CV as in co visitation notebook.\nI get CV 0.565.\n\n\n---LB PREP---\nI repeat steps 1) and 2) but I use the following dataset [here](https://www.kaggle.com/datasets/columbia2131/otto-chunk-data-inparquet-format).\nThen I prepare the test candidates in step 3 exactly same way BUT i dont have test labels. So i leave out the merge of tar.\nThen I infer and get resulting dataframe.\nWhen I submit the resulting dataframe I only get LB 0.565. (exactly the same as CV) In fact I did multiple experiments and the LB always closely resembles the CV score but I expect to be higher (like for co visitation approach!)\nCan anyone point me to the mistake I made? I cant see it.",
      "votes": 19
    },
    {
      "id": 2064699,
      "postDate": "2022-12-14T03:57:26.387Z",
      "content": "<p>I think as Chris said, there are some 'ts' features involved in the features. The 'ts' distribution of the data for local validation B (second half of week 4) is inconsistent with the test set (second half of week 5). This can lead to poor generalization of the model. You can change these 'ts' characteristics from absolute numbers to relative numbers, such as calculating sequence duration.Finally, these absolute 'ts' features are dropped.</p>",
      "rawMarkdown": "I think as Chris said, there are some 'ts' features involved in the features. The 'ts' distribution of the data for local validation B (second half of week 4) is inconsistent with the test set (second half of week 5). This can lead to poor generalization of the model. You can change these 'ts' characteristics from absolute numbers to relative numbers, such as calculating sequence duration.Finally, these absolute 'ts' features are dropped.",
      "votes": 5,
      "replies": [
        {
          "id": 2064757,
          "postDate": "2022-12-14T05:13:35.713Z",
          "content": "<p>Will try this Jensen!</p>",
          "rawMarkdown": "Will try this Jensen!"
        },
        {
          "id": 2064944,
          "postDate": "2022-12-14T07:26:50.233Z",
          "content": "<p>Also, we should note the difference between local validation and LB testing. </p>\n<p>First, we use the co-visitation matrix to generate candidates. When we generate local validation candidates, our co-visit matrix uses information from validation B. But when we generate candidates for test data, we can't use the ground-truth of test data. I think this would lead to a certain data leak for local validation. </p>\n<p>Second, in local validation, our user sequence is continuous. This is equivalent to us using the historical information of some users to predict the behavior of the same users. The correlation between these messages is high. But for online LB, the users of the test data are a new group of users.So the test data users' information we have is not as rich as local validation. </p>\n<p>In the above two points, I feel that the score of local validation should be greater than the score of LB. The phenomenon you have in this article may be the cause.This may be a normal phenomenon</p>",
          "rawMarkdown": "Also, we should note the difference between local validation and LB testing. \n\nFirst, we use the co-visitation matrix to generate candidates. When we generate local validation candidates, our co-visit matrix uses information from validation B. But when we generate candidates for test data, we can't use the ground-truth of test data. I think this would lead to a certain data leak for local validation. \n\nSecond, in local validation, our user sequence is continuous. This is equivalent to us using the historical information of some users to predict the behavior of the same users. The correlation between these messages is high. But for online LB, the users of the test data are a new group of users.So the test data users' information we have is not as rich as local validation. \n\nIn the above two points, I feel that the score of local validation should be greater than the score of LB. The phenomenon you have in this article may be the cause.This may be a normal phenomenon",
          "votes": 1
        },
        {
          "id": 2064977,
          "postDate": "2022-12-14T08:07:09.790Z",
          "content": "<p>I understand your points but we have one more week to learn co visitation for our LB candidates. That should give us more robust candidates. I guess this is also the reason why we get increase in LB compared to CV for pure co visitation approach. Not sure which one has higher influence of the score. Tonight I will submit my newest results to LB and see if I could fix the problem with LB score. I tried to take into account with the points mentioned by u and Chris in this model.</p>",
          "rawMarkdown": "I understand your points but we have one more week to learn co visitation for our LB candidates. That should give us more robust candidates. I guess this is also the reason why we get increase in LB compared to CV for pure co visitation approach. Not sure which one has higher influence of the score. Tonight I will submit my newest results to LB and see if I could fix the problem with LB score. I tried to take into account with the points mentioned by u and Chris in this model."
        }
      ]
    },
    {
      "id": 2106719,
      "postDate": "2023-01-19T10:20:51.153Z",
      "content": "<p>I added a feature that is the order of aid in rerank model and so gbt model works better</p>",
      "rawMarkdown": "I added a feature that is the order of aid in rerank model and so gbt model works better",
      "votes": 3,
      "replies": [
        {
          "id": 2106724,
          "postDate": "2023-01-19T10:24:38.950Z",
          "content": "<p>I will try that! Thanks!</p>",
          "rawMarkdown": "I will try that! Thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2064522,
      "postDate": "2022-12-13T20:45:42.667Z",
      "content": "<p>Hi Simon, how do you generate candidates for your submission. You must make different co-visitation matrices for training and inference. </p>",
      "rawMarkdown": "Hi Simon, how do you generate candidates for your submission. You must make different co-visitation matrices for training and inference. ",
      "votes": 1,
      "replies": [
        {
          "id": 2064524,
          "postDate": "2022-12-13T20:49:00.497Z",
          "content": "<p>Hallo Chris. I create the candidates in the same way.<br>\nBut I take the dataset B (<a href=\"https://www.kaggle.com/datasets/columbia2131/otto-chunk-data-inparquet-format\" target=\"_blank\">here</a>) to create the Co visitation matrices I use for candidate generation.<br>\nThe same dataset i use to generate features that I merge Onto the candidates before infering my submission from the models I trained before </p>",
          "rawMarkdown": "Hallo Chris. I create the candidates in the same way.\nBut I take the dataset B ([here](https://www.kaggle.com/datasets/columbia2131/otto-chunk-data-inparquet-format)) to create the Co visitation matrices I use for candidate generation.\nThe same dataset i use to generate features that I merge Onto the candidates before infering my submission from the models I trained before \n"
        },
        {
          "id": 2064530,
          "postDate": "2022-12-13T20:57:42.933Z",
          "content": "<p></p>\n<p>oh nevermind, i see that you use my validation Kaggle dataset (which uses milliseconds just like colum2131)</p>",
          "rawMarkdown": "~~Ok and do you account for the fact that that dataset uses milliseconds and Radek's dataset uses seconds?~~\n\noh nevermind, i see that you use my validation Kaggle dataset (which uses milliseconds just like colum2131)"
        },
        {
          "id": 2064537,
          "postDate": "2022-12-13T21:23:12.547Z",
          "content": "<p>yes. but one strange thing i noted:<br>\nin the other thread you said<br>\n<code>When training models and computing CV, we use our \"new train\" dataset (3 weeks of data) and our \"validation A\" dataset (roughly half week of data). Then when we build our submission we use \"Kaggle train\" (4 weeks of data) and \"Kaggle test\" (roughly half week of data).</code><br>\nThe first point i see.<br>\nBut the second point i dont see! the difference between minimum ts of test and maximum ts of test for validation data is also 7 days. <br>\nAlso the length of the corresponding test dataframe is roughly the same as for my real test set.</p>",
          "rawMarkdown": "yes. but one strange thing i noted:\nin the other thread you said\n`When training models and computing CV, we use our \"new train\" dataset (3 weeks of data) and our \"validation A\" dataset (roughly half week of data). Then when we build our submission we use \"Kaggle train\" (4 weeks of data) and \"Kaggle test\" (roughly half week of data).`\nThe first point i see.\nBut the second point i dont see! the difference between minimum ts of test and maximum ts of test for validation data is also 7 days. \nAlso the length of the corresponding test dataframe is roughly the same as for my real test set."
        },
        {
          "id": 2064545,
          "postDate": "2022-12-13T21:39:46.837Z",
          "content": "<p>The <code>ts</code> column in Kaggle's test ranges from Aug 29th thru Sept 4th. And the <code>ts</code> column in our \"new test\" ranges from Aug 22nd thru Aug 28th. We cannot use any absolute features from <code>ts</code> we must create relative features from <code>ts</code>. I recommend you plot a histogram for every one of your features. For each feature plot the distribution from your train dataframe and the distribution from your inference dataframe. If it looks like below, then your XGB will not work. The histograms need to be the same for each of your features when comparing train to infer.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Dec-2022/test.png\" alt=\"\"></p>",
          "rawMarkdown": "The `ts` column in Kaggle's test ranges from Aug 29th thru Sept 4th. And the `ts` column in our \"new test\" ranges from Aug 22nd thru Aug 28th. We cannot use any absolute features from `ts` we must create relative features from `ts`. I recommend you plot a histogram for every one of your features. For each feature plot the distribution from your train dataframe and the distribution from your inference dataframe. If it looks like below, then your XGB will not work. The histograms need to be the same for each of your features when comparing train to infer.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Dec-2022/test.png)",
          "votes": 8
        },
        {
          "id": 2064551,
          "postDate": "2022-12-13T21:50:31.807Z",
          "content": "<p>i will try harder and try to spot flaws in the features i create, chris. i guess that will be the solution to a working model :-) tnx already for the help!!!</p>",
          "rawMarkdown": "i will try harder and try to spot flaws in the features i create, chris. i guess that will be the solution to a working model :-) tnx already for the help!!!",
          "votes": 1,
          "replies": [
            {
              "id": 2071684,
              "postDate": "2022-12-21T09:20:21.427Z",
              "content": "<p>Have you been able to find the source of the problem? I had a similar issue when I downsampled negative samples too much (to save memory). </p>",
              "rawMarkdown": "Have you been able to find the source of the problem? I had a similar issue when I downsampled negative samples too much (to save memory). "
            },
            {
              "id": 2071704,
              "postDate": "2022-12-21T09:33:40.777Z",
              "content": "<p>For now im working on improvement of my candidate generation (Co visitation matrices). Then I will continue. For me was not relaxed to sampling process. I think we need to look into features. I will Explorer next with eda to try to find what I Messed up.</p>",
              "rawMarkdown": "For now im working on improvement of my candidate generation (Co visitation matrices). Then I will continue. For me was not relaxed to sampling process. I think we need to look into features. I will Explorer next with eda to try to find what I Messed up."
            },
            {
              "id": 2072231,
              "postDate": "2022-12-21T21:59:01.527Z",
              "content": "<p>I posted comment <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210#2072224\" target=\"_blank\">here</a> with a screenshot of code that shows that feature distribution should be the reason for our problem as <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> guessed. Until now I was not able to fix this problem unfortunately..</p>",
              "rawMarkdown": "I posted comment [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210#2072224) with a screenshot of code that shows that feature distribution should be the reason for our problem as @cdeotte guessed. Until now I was not able to fix this problem unfortunately.."
            }
          ]
        }
      ]
    },
    {
      "id": 2072327,
      "postDate": "2022-12-22T00:45:57.620Z",
      "content": "<p>Did you use all the test_labels.parquet to label your training data？ And after training and infering, your CV score was also calculated on the test_labels.parquet？</p>",
      "rawMarkdown": "Did you use all the test_labels.parquet to label your training data？ And after training and infering, your CV score was also calculated on the test_labels.parquet？",
      "replies": [
        {
          "id": 2072613,
          "postDate": "2022-12-22T09:15:12.427Z",
          "content": "<p>Yes. Is something wrong about that?<br>\nFurthermore the features seem to have different distribution although i cut out the first week of the training data. See picture below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11646918%2Fc45758ba7233a064e4807c5a8a99949c%2FBildschirmfoto%20vom%202022-12-21%2023-04-59.png?generation=1671700497921582&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Yes. Is something wrong about that?\nFurthermore the features seem to have different distribution although i cut out the first week of the training data. See picture below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11646918%2Fc45758ba7233a064e4807c5a8a99949c%2FBildschirmfoto%20vom%202022-12-21%2023-04-59.png?generation=1671700497921582&alt=media)",
          "replies": [
            {
              "id": 2072964,
              "postDate": "2022-12-22T14:51:22.987Z",
              "content": "<p>I think you should split the test_labels into two parts. One part is used to label the training samples and the other one is to calculate the score, otherwise there will be leaks in training and evaluating.</p>",
              "rawMarkdown": "I think you should split the test_labels into two parts. One part is used to label the training samples and the other one is to calculate the score, otherwise there will be leaks in training and evaluating.",
              "votes": 1
            },
            {
              "id": 2072975,
              "postDate": "2022-12-22T14:58:00.533Z",
              "content": "<p>Ok, that I didn't do. <br>\nBut how should I split test_labels? In the dataframe I have there is no timestamp I could split on..<br>\nAnd what I do about the problem with the features I mention above?<br>\nTnx already im beginner sorry for so many questions :D</p>",
              "rawMarkdown": "Ok, that I didn't do. \nBut how should I split test_labels? In the dataframe I have there is no timestamp I could split on..\nAnd what I do about the problem with the features I mention above?\nTnx already im beginner sorry for so many questions :D"
            },
            {
              "id": 2072978,
              "postDate": "2022-12-22T15:00:16.197Z",
              "content": "<p>For clarification: I use the otto_validation dataset Provided by Chris which is based on the validation dataset radek created.</p>",
              "rawMarkdown": "For clarification: I use the otto_validation dataset Provided by Chris which is based on the validation dataset radek created."
            },
            {
              "id": 2073001,
              "postDate": "2022-12-22T15:26:02.287Z",
              "content": "<p>Split test_labels by user, e.g. 50％ users for labeling and the other 50％ for evaluating. <br>\nI am using <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's dataset too and above is my setting. The LB is higher than CV by ~0.01 in my case, which is similar to those shown in chris's notebook.</p>",
              "rawMarkdown": "Split test_labels by user, e.g. 50％ users for labeling and the other 50％ for evaluating. \nI am using @cdeotte 's dataset too and above is my setting. The LB is higher than CV by ~0.01 in my case, which is similar to those shown in chris's notebook.",
              "votes": 1
            },
            {
              "id": 2073011,
              "postDate": "2022-12-22T15:35:26.747Z",
              "content": "<p>fantastic sirius! i hope that helps. <br>\nthis split i also have to use for feature generation, correct? and only take last three weeks of original train set?</p>",
              "rawMarkdown": "fantastic sirius! i hope that helps. \nthis split i also have to use for feature generation, correct? and only take last three weeks of original train set?"
            },
            {
              "id": 2075048,
              "postDate": "2022-12-25T01:27:55.527Z",
              "content": "<p>Hello, sirius.I would like to ask a question. How is CV score calculated?<br>\nI used Chris' version 9 notebook(0.562 in LB) <a href=\"https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\" target=\"_blank\">here</a> to calculate the local CV. I suspect something went wrong with my idea of calculating local CV scores. In the notebook, I generated candidates for val A's validation set and then calculated the local CV score, I only made the following modifications, but the CV score looks so outrageous. What went wrong with me? I am looking forward to hearing from you.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11356689%2F4790bea02e2cbba61be13b9532ad4411%2F1.png?generation=1671931642626666&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11356689%2F06d86aa264d1a4e9b986751b4542cb08%2F2.png?generation=1671931655265379&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Hello, sirius.I would like to ask a question. How is CV score calculated?\nI used Chris' version 9 notebook(0.562 in LB) [here](https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost) to calculate the local CV. I suspect something went wrong with my idea of calculating local CV scores. In the notebook, I generated candidates for val A's validation set and then calculated the local CV score, I only made the following modifications, but the CV score looks so outrageous. What went wrong with me? I am looking forward to hearing from you.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11356689%2F4790bea02e2cbba61be13b9532ad4411%2F1.png?generation=1671931642626666&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11356689%2F06d86aa264d1a4e9b986751b4542cb08%2F2.png?generation=1671931655265379&alt=media)\n"
            },
            {
              "id": 2075364,
              "postDate": "2022-12-25T09:34:58.390Z",
              "content": "<p>The CV calculation seems no error. Maybe you have some data leaks in the training leading to high CV.</p>",
              "rawMarkdown": "The CV calculation seems no error. Maybe you have some data leaks in the training leading to high CV."
            },
            {
              "id": 2077840,
              "postDate": "2022-12-27T22:48:12.813Z",
              "content": "<p>hello sirius. thanks for your useful advice. using this my CV behaves as expected compared to LB</p>",
              "rawMarkdown": "hello sirius. thanks for your useful advice. using this my CV behaves as expected compared to LB"
            },
            {
              "id": 2082429,
              "postDate": "2023-01-01T15:16:47.180Z",
              "content": "<p>Thank you, I found the data leak and it's normal now.</p>",
              "rawMarkdown": "Thank you, I found the data leak and it's normal now."
            },
            {
              "id": 2082449,
              "postDate": "2023-01-01T15:35:23.953Z",
              "content": "<p>Can you say where was it?</p>",
              "rawMarkdown": "Can you say where was it?"
            },
            {
              "id": 2082463,
              "postDate": "2023-01-01T15:38:44.557Z",
              "content": "<p>For me the problem was solve by following advice of sirius</p>",
              "rawMarkdown": "For me the problem was solve by following advice of sirius"
            },
            {
              "id": 2082470,
              "postDate": "2023-01-01T15:46:55.753Z",
              "content": "<p>One additional thing to keep in mind if you want to use count features: sample down (by user) to get same len for the corresponding train and test Sets!</p>",
              "rawMarkdown": "One additional thing to keep in mind if you want to use count features: sample down (by user) to get same len for the corresponding train and test Sets!"
            }
          ]
        }
      ]
    },
    {
      "id": 2082432,
      "postDate": "2023-01-01T15:25:27.370Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2082436,
          "postDate": "2023-01-01T15:29:55.650Z",
          "content": "<p>Because we only need information about the users in test Set. But from the items we can Aggregate also information from training set. Note: it's important to follow the advice from Sirius below to get a meaningful CV score!!!</p>",
          "rawMarkdown": "Because we only need information about the users in test Set. But from the items we can Aggregate also information from training set. Note: it's important to follow the advice from Sirius below to get a meaningful CV score!!!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2064699,
      "author_name": "Jensen",
      "author_url": "",
      "post_date": "2022-12-14T03:57:26.387000",
      "content": "<p>I think as Chris said, there are some 'ts' features involved in the features. The 'ts' distribution of the data for local validation B (second half of week 4) is inconsistent with the test set (second half of week 5). This can lead to poor generalization of the model. You can change these 'ts' characteristics from absolute numbers to relative numbers, such as calculating sequence duration.Finally, these absolute 'ts' features are dropped.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2064757,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2022-12-14T05:13:35.713000",
          "content": "<p>Will try this Jensen!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2064944,
          "author_name": "Jensen",
          "author_url": "",
          "post_date": "2022-12-14T07:26:50.233000",
          "content": "<p>Also, we should note the difference between local validation and LB testing. </p>\n<p>First, we use the co-visitation matrix to generate candidates. When we generate local validation candidates, our co-visit matrix uses information from validation B. But when we generate candidates for test data, we can't use the ground-truth of test data. I think this would lead to a certain data leak for local validation. </p>\n<p>Second, in local validation, our user sequence is continuous. This is equivalent to us using the historical information of some users to predict the behavior of the same users. The correlation between these messages is high. But for online LB, the users of the test data are a new group of users.So the test data users' information we have is not as rich as local validation. </p>\n<p>In the above two points, I feel that the score of local validation should be greater than the score of LB. The phenomenon you have in this article may be the cause.This may be a normal phenomenon</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2064977,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2022-12-14T08:07:09.790000",
          "content": "<p>I understand your points but we have one more week to learn co visitation for our LB candidates. That should give us more robust candidates. I guess this is also the reason why we get increase in LB compared to CV for pure co visitation approach. Not sure which one has higher influence of the score. Tonight I will submit my newest results to LB and see if I could fix the problem with LB score. I tried to take into account with the points mentioned by u and Chris in this model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2106719,
      "author_name": "Anh Bui",
      "author_url": "",
      "post_date": "2023-01-19T10:20:51.153000",
      "content": "<p>I added a feature that is the order of aid in rerank model and so gbt model works better</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2106724,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2023-01-19T10:24:38.950000",
          "content": "<p>I will try that! Thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2064522,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-12-13T20:45:42.667000",
      "content": "<p>Hi Simon, how do you generate candidates for your submission. You must make different co-visitation matrices for training and inference. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2064524,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2022-12-13T20:49:00.497000",
          "content": "<p>Hallo Chris. I create the candidates in the same way.<br>\nBut I take the dataset B (<a href=\"https://www.kaggle.com/datasets/columbia2131/otto-chunk-data-inparquet-format\" target=\"_blank\">here</a>) to create the Co visitation matrices I use for candidate generation.<br>\nThe same dataset i use to generate features that I merge Onto the candidates before infering my submission from the models I trained before </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2064530,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-13T20:57:42.933000",
          "content": "<p></p>\n<p>oh nevermind, i see that you use my validation Kaggle dataset (which uses milliseconds just like colum2131)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2064537,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2022-12-13T21:23:12.547000",
          "content": "<p>yes. but one strange thing i noted:<br>\nin the other thread you said<br>\n<code>When training models and computing CV, we use our \"new train\" dataset (3 weeks of data) and our \"validation A\" dataset (roughly half week of data). Then when we build our submission we use \"Kaggle train\" (4 weeks of data) and \"Kaggle test\" (roughly half week of data).</code><br>\nThe first point i see.<br>\nBut the second point i dont see! the difference between minimum ts of test and maximum ts of test for validation data is also 7 days. <br>\nAlso the length of the corresponding test dataframe is roughly the same as for my real test set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2064545,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-13T21:39:46.837000",
          "content": "<p>The <code>ts</code> column in Kaggle's test ranges from Aug 29th thru Sept 4th. And the <code>ts</code> column in our \"new test\" ranges from Aug 22nd thru Aug 28th. We cannot use any absolute features from <code>ts</code> we must create relative features from <code>ts</code>. I recommend you plot a histogram for every one of your features. For each feature plot the distribution from your train dataframe and the distribution from your inference dataframe. If it looks like below, then your XGB will not work. The histograms need to be the same for each of your features when comparing train to infer.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Dec-2022/test.png\" alt=\"\"></p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 2064551,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2022-12-13T21:50:31.807000",
          "content": "<p>i will try harder and try to spot flaws in the features i create, chris. i guess that will be the solution to a working model :-) tnx already for the help!!!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2071684,
              "author_name": "Oleg Kokorin",
              "author_url": "",
              "post_date": "2022-12-21T09:20:21.427000",
              "content": "<p>Have you been able to find the source of the problem? I had a similar issue when I downsampled negative samples too much (to save memory). </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2071704,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2022-12-21T09:33:40.777000",
              "content": "<p>For now im working on improvement of my candidate generation (Co visitation matrices). Then I will continue. For me was not relaxed to sampling process. I think we need to look into features. I will Explorer next with eda to try to find what I Messed up.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2072231,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2022-12-21T21:59:01.527000",
              "content": "<p>I posted comment <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210#2072224\" target=\"_blank\">here</a> with a screenshot of code that shows that feature distribution should be the reason for our problem as <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> guessed. Until now I was not able to fix this problem unfortunately..</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2072327,
      "author_name": "sirius",
      "author_url": "",
      "post_date": "2022-12-22T00:45:57.620000",
      "content": "<p>Did you use all the test_labels.parquet to label your training data？ And after training and infering, your CV score was also calculated on the test_labels.parquet？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2072613,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2022-12-22T09:15:12.427000",
          "content": "<p>Yes. Is something wrong about that?<br>\nFurthermore the features seem to have different distribution although i cut out the first week of the training data. See picture below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11646918%2Fc45758ba7233a064e4807c5a8a99949c%2FBildschirmfoto%20vom%202022-12-21%2023-04-59.png?generation=1671700497921582&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2072964,
              "author_name": "sirius",
              "author_url": "",
              "post_date": "2022-12-22T14:51:22.987000",
              "content": "<p>I think you should split the test_labels into two parts. One part is used to label the training samples and the other one is to calculate the score, otherwise there will be leaks in training and evaluating.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2072975,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2022-12-22T14:58:00.533000",
              "content": "<p>Ok, that I didn't do. <br>\nBut how should I split test_labels? In the dataframe I have there is no timestamp I could split on..<br>\nAnd what I do about the problem with the features I mention above?<br>\nTnx already im beginner sorry for so many questions :D</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2072978,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2022-12-22T15:00:16.197000",
              "content": "<p>For clarification: I use the otto_validation dataset Provided by Chris which is based on the validation dataset radek created.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2073001,
              "author_name": "sirius",
              "author_url": "",
              "post_date": "2022-12-22T15:26:02.287000",
              "content": "<p>Split test_labels by user, e.g. 50％ users for labeling and the other 50％ for evaluating. <br>\nI am using <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's dataset too and above is my setting. The LB is higher than CV by ~0.01 in my case, which is similar to those shown in chris's notebook.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2073011,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2022-12-22T15:35:26.747000",
              "content": "<p>fantastic sirius! i hope that helps. <br>\nthis split i also have to use for feature generation, correct? and only take last three weeks of original train set?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2075048,
              "author_name": "Jensen",
              "author_url": "",
              "post_date": "2022-12-25T01:27:55.527000",
              "content": "<p>Hello, sirius.I would like to ask a question. How is CV score calculated?<br>\nI used Chris' version 9 notebook(0.562 in LB) <a href=\"https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\" target=\"_blank\">here</a> to calculate the local CV. I suspect something went wrong with my idea of calculating local CV scores. In the notebook, I generated candidates for val A's validation set and then calculated the local CV score, I only made the following modifications, but the CV score looks so outrageous. What went wrong with me? I am looking forward to hearing from you.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11356689%2F4790bea02e2cbba61be13b9532ad4411%2F1.png?generation=1671931642626666&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11356689%2F06d86aa264d1a4e9b986751b4542cb08%2F2.png?generation=1671931655265379&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2075364,
              "author_name": "sirius",
              "author_url": "",
              "post_date": "2022-12-25T09:34:58.390000",
              "content": "<p>The CV calculation seems no error. Maybe you have some data leaks in the training leading to high CV.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2077840,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2022-12-27T22:48:12.813000",
              "content": "<p>hello sirius. thanks for your useful advice. using this my CV behaves as expected compared to LB</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082429,
              "author_name": "Jensen",
              "author_url": "",
              "post_date": "2023-01-01T15:16:47.180000",
              "content": "<p>Thank you, I found the data leak and it's normal now.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082449,
              "author_name": "Anil Ozturk",
              "author_url": "",
              "post_date": "2023-01-01T15:35:23.953000",
              "content": "<p>Can you say where was it?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082463,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2023-01-01T15:38:44.557000",
              "content": "<p>For me the problem was solve by following advice of sirius</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2082470,
              "author_name": "Simon Veitner",
              "author_url": "",
              "post_date": "2023-01-01T15:46:55.753000",
              "content": "<p>One additional thing to keep in mind if you want to use count features: sample down (by user) to get same len for the corresponding train and test Sets!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2082432,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-01T15:25:27.370000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2082436,
          "author_name": "Simon Veitner",
          "author_url": "",
          "post_date": "2023-01-01T15:29:55.650000",
          "content": "<p>Because we only need information about the users in test Set. But from the items we can Aggregate also information from training set. Note: it's important to follow the advice from Sirius below to get a meaningful CV score!!!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2064435": "Hello everyone.\nI face the problem that the LB of my Ranker model is much lower than what I would it expect to be from the CV.\nThe CV is 0.565 and the LB aswell. From what I had before I expeted the LB in the region of 0.575.\nI describe my pipeline (which is essentially the one described by Chris) in greater detail in the hope that somebody can help me to fix the problem.\n 1) Step: I use the validation dataset given by chris [here](https://www.kaggle.com/datasets/cdeotte/otto-validation) (I call it in following dataset A) to generate co visitation matrices. I use these co visitation matrices to generate candidates. For example the clicks candidates are generated as follows:\n```\ndef suggest_clicks(df):\n    # USE USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    result = list(dict.fromkeys(unique_aids + aids2))\n    # USE TOP20 TEST CLICKS\n    return list(dict.fromkeys(result + list(top_clicks)))[:30]\n```\nI apply this function for each user and explode then to get an N * 30 long dataframe where N is number of users.\n\n2) Step: Using the same dataset as in 1) i generate features for items and users and also interaction features.\nFor example for item features i use both train and test set from dataset A and generate them in the following way:\n```\nT = (data.ts.max()-data.ts.min())/(24*60*60)\nTmin = data.ts.min()\ndata.ts -= Tmin\nitem_features = data.groupby('aid').agg({'aid':'count','session':'nunique','type':'mean', 'ts':['min', 'max']})\nitem_features.columns = ['item_item_count','item_user_count','item_buy_ratio', 'item_first_time_seen', 'item_last_time_seen']\n# CONVERT COLUMNS TO INT32 and FLOAT32 HERE\nitem_features.item_item_count /= T\nitem_features.item_user_count /= T\nitem_features.to_pandas().to_parquet('item_features.pqt')\n```\nI scale by the total time to be sure that I can generate same features for the \"real\" training and test set. data above is simply given by `data = cudf.concat([train, test])`.\n\nUser features i generate just with the test data from dataset A:\n```\nT = (test.ts.max()-test.ts.min())/(24*60*60)\nTmin = test.ts.min()\ntest.ts -= Tmin\nuser_features = test.groupby('session').agg({'session':'count','aid':'nunique','type':'mean', 'ts':['min', 'max']})\nuser_features.columns = ['user_user_count','user_item_count','user_buy_ratio', 'user_first_time_seen', 'user_last_time_seen']\n# CONVERT COLUMNS TO INT32 and FLOAT32 HERE\nuser_features.user_user_count /= T\nuser_features.user_item_count /= T\nuser_features.to_parquet('user_features.pqt')\n```\nAnd an example of interaction dataframe is:\n```\nitems_clicked = test[[\"session\", \"aid\"]].loc[test.type == 0].drop_duplicates(subset=[\"session\",\"aid\"]).rename({'session':'user', 'aid':'item'},axis=1)\nitems_clicked[\"item_clicked\"] = 1\nitems_clicked.to_parquet('items_clicked.pqt')\n```.\n\n3) Step:\nI train the reranker after generating candidates dataframe. The steps taken to generate the dataframe for candidates are as follows (for example for clicks)\n\n```\ncandidates = pd.read_parquet(\"cand_clicks.pqt\")\ncandidates.columns = [\"session\", \"aid\"]\n\nitem_features = pd.read_parquet('item_features.pqt')\ncandidates = candidates.merge(item_features, left_on='aid', right_index=True, how='left').fillna(-1)\nuser_features = pd.read_parquet('user_features.pqt')\ncandidates = candidates.merge(user_features, left_on='session', right_index=True, how='left').fillna(-1)\ndel item_features, user_features\n_ = gc.collect()\n\ntar = cudf.read_parquet('/notebooks/pipeline_cv_score/otto/test_labels.parquet')\ntar = tar.loc[ tar['type']=='clicks' ]\naids = tar.ground_truth.explode().astype('int32').rename('item')\ntar = tar[['session']].astype('int32').rename({'session':'user'},axis=1)\ntar = tar.merge(aids, left_index=True, right_index=True, how='left')\ntar['click'] = 1\n\nitems_interacted = cudf.read_parquet(\"items_interacted.pqt\")\nitems_clicked = cudf.read_parquet(\"items_clicked.pqt\")\nitems_carted = cudf.read_parquet(\"items_carts.pqt\")\nitems_ordered = cudf.read_parquet(\"items_orders.pqt\")\ncandidates_ = cudf.DataFrame(candidates)\n\ncandidates_ = candidates_.rename({'session':'user', 'aid':'item'},axis=1)\ncandidates_ = candidates_.merge(items_interacted,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(items_clicked,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(items_carted,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(items_ordered,on=['user','item'],how='left').fillna(0)\ncandidates_ = candidates_.merge(tar,on=['user','item'],how='left').fillna(0)\n```\nThen I train as suggested by Chris:\n```\nimport xgboost as xgb\nfrom sklearn.model_selection import GroupKFold\n\nif TRAIN:\n    FEATURES = candidates.columns[2 : -1]\n\n    skf = GroupKFold(n_splits=5)\n    for fold,(train_idx, valid_idx) in enumerate(skf.split(candidates, candidates['click'], groups=candidates['user'] )):\n        print(f\"Train fold {fold}\")\n\n        X_train = candidates.loc[train_idx, FEATURES]\n        y_train = candidates.loc[train_idx, 'click']\n        X_valid = candidates.loc[valid_idx, FEATURES]\n        y_valid = candidates.loc[valid_idx, 'click']\n\n        # IF YOU HAVE 50 CANDIDATE WE USE 50 BELOW\n        dtrain = xgb.DMatrix(X_train, y_train, group=[30] * (len(train_idx)//30) ) \n        dvalid = xgb.DMatrix(X_valid, y_valid, group=[30] * (len(valid_idx)//30) ) \n\n        xgb_parms = {'objective':'rank:pairwise', 'tree_method':'gpu_hist'}\n        model = xgb.train(xgb_parms, \n            dtrain=dtrain,\n            evals=[(dtrain,'train'),(dvalid,'valid')],\n            num_boost_round=30,\n            verbose_eval=100)\n        model.save_model(f'XGB_fold{fold}_click.xgb')\n```\n\nI infer in the following way to calculate CV score:\n```\nFEATURES = candidates.columns[2 : -1]\n\ndef pred(df):\n    preds = np.zeros(len(df))\n    for fold in range(5):\n        print(f\"Inferece for fold {fold}\")\n        model = xgb.Booster()\n        model.load_model(f'XGB_fold{fold}_click.xgb')\n        model.set_param({'predictor': 'gpu_predictor'})\n        dtest = xgb.DMatrix(data=df[FEATURES])\n        preds += model.predict(dtest)/5\n    predictions = df[['user','item']].copy()\n    predictions['pred'] = preds\n    predictions = predictions.sort_values(['user','pred'], ascending=[True,False]).reset_index(drop=True)\n    predictions['n'] = predictions.groupby('user').item.cumcount().astype('int8')\n    predictions = predictions.loc[predictions.n<20]\n    pred_df_clicks = predictions.groupby('user').item.apply(list)\n    pred_df_clicks = pred_df_clicks.to_frame().reset_index()\n    pred_df_clicks.item = pred_df_clicks.item.apply(lambda x: \" \".join(map(str,x)))\n    pred_df_clicks.columns = ['session_type','labels']\n    pred_df_clicks.session_type = pred_df_clicks.session_type.astype('str')+ '_clicks'\n    return pred_df_clicks\n\n# SPLIT FOR MEMORY MANAGMENT:\ntest_candidates1 = test_candidates.loc[test_candidates.user.isin(grp_idx[:len(grp_idx)//2])]\ntest_candidates2 = test_candidates.loc[test_candidates.user.isin(grp_idx[len(grp_idx)//2:])]\nprint(\"start with prediction 1.\")\npred_df_clicks1 = pred(test_candidates1)\nprint(\"start with prediction 2.\")\npred_df_clicks2 = pred(test_candidates2)\nprint(\"concat...\")\npred_df_clicks = pd.concat([pred_df_clicks1, pred_df_clicks2])\ndel pred_df_clicks1, pred_df_clicks2\n_ = gc.collect\n```\n\nThe same I do for carts and buy and then i concat my dataframes and calculcate CV as in co visitation notebook.\nI get CV 0.565.\n\n\n---LB PREP---\nI repeat steps 1) and 2) but I use the following dataset [here](https://www.kaggle.com/datasets/columbia2131/otto-chunk-data-inparquet-format).\nThen I prepare the test candidates in step 3 exactly same way BUT i dont have test labels. So i leave out the merge of tar.\nThen I infer and get resulting dataframe.\nWhen I submit the resulting dataframe I only get LB 0.565. (exactly the same as CV) In fact I did multiple experiments and the LB always closely resembles the CV score but I expect to be higher (like for co visitation approach!)\nCan anyone point me to the mistake I made? I cant see it.",
    "2064699": "I think as Chris said, there are some 'ts' features involved in the features. The 'ts' distribution of the data for local validation B (second half of week 4) is inconsistent with the test set (second half of week 5). This can lead to poor generalization of the model. You can change these 'ts' characteristics from absolute numbers to relative numbers, such as calculating sequence duration.Finally, these absolute 'ts' features are dropped.",
    "2106719": "I added a feature that is the order of aid in rerank model and so gbt model works better",
    "2064522": "Hi Simon, how do you generate candidates for your submission. You must make different co-visitation matrices for training and inference. ",
    "2072327": "Did you use all the test_labels.parquet to label your training data？ And after training and infering, your CV score was also calculated on the test_labels.parquet？",
    "2082432": ""
  }
}