{
  "id": 381773,
  "title": "Why LGBM ndcg nearly 0.9 while LB score is only 0.4?",
  "url": "/competitions/otto-recommender-system/discussion/381773",
  "author_name": "",
  "post_date": "2023-01-28T03:24:28.527940700Z",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi kagglers,</p>\n<p>I trained a LGBM Ranker following this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">link</a> by <a href=\"https://www.kaggle.com/CHRIS\" target=\"_blank\">@CHRIS</a> DEOTTE, the eval score looks good, yet the lb score is surprisely low.</p>\n<p>Here are the key steps I did:</p>\n<ol>\n<li>recall: 100 candidate per session</li>\n<li><strong>neg_sample dealing: filter sessions whose candidate have no carts/orders action label</strong></li>\n<li>feature: perpare some features</li>\n<li>Ranker: train a LGBM ranker, and got≈0.6 ndcg@20</li>\n<li>Eval: prediction on valid set, and calc by OTTO metric and get 0.35</li>\n</ol>\n<p>Neg sample </p>\n<pre><code>def filter_out_neg_sessions(train_set,target,frac=0.9):\n    session_target_sum = train_set.groupby('session')[target].sum().to_frame()\n    filter_out_session = session_target_sum[session_target_sum[target]==0].sample(frac=frac)\n    train_set = train_set[~train_set.session.isin(filter_out_session.index)]\n    return train_set\n</code></pre>\n<p>LGBM</p>\n<pre><code>params = {\n    'learning_rate': 0.005,\n    'max_depth': 8,\n    'objective':\"lambdarank\",\n    'metric':\"ndcg\",\n    'colsample_bytree': 0.9,\n    'subsample':0.9,\n    'boosting_type':\"gbdt\",\n    'n_estimators':200,\n    'importance_type':'gain'\n}\nskf = GroupKFold(n_splits=KFOLDS)\nfor fold,(train_idx, valid_idx) in tqdm(enumerate(skf.split(train_set, train_set[TARGET], groups=train_set['session']))):\n\n    X_train = train_set.iloc[train_idx][FEATURES]\n    y_train = train_set.iloc[train_idx][TARGET]\n    X_valid = train_set.iloc[valid_idx][FEATURES]\n    y_valid = train_set.iloc[valid_idx][TARGET]\n    group_train = train_set.iloc[train_idx].groupby('session')['session'].count()\n    group_valid = train_set.iloc[valid_idx].groupby('session')['session'].count()\n    print(train_idx.shape,X_train.shape,y_train.shape)\n    ranker = LGBMRanker(\n        **params\n    )\n    ranker = ranker.fit(\n        X_train,\n        y_train,\n        group=group_train,\n        eval_set=[(X_train,y_train),(X_valid, y_valid)],\n        eval_group=[group_train,group_valid],\n        eval_at=(1,5,10,20,30)\n    )\n    lgb.plot_metric(ranker)\n    lgb.plot_importance(ranker,max_num_features=20)\n    plt.show()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1194567%2F9199ea823b62538d5ba08d02c4ad2eb5%2Fkaggle.png?generation=1674876055881572&amp;alt=media\" alt=\"\"></p>\n<p>I had noticed that the more I keep neg_samples, the higer ndcg score I have. <em>(if keep all, I got 0.91)</em></p>\n<p>Also, I submitted recall result separatly, LB score is 0.575, so the candidate is fine.<br>\nPersonaly I believe <strong>there is something wrong on the [neg_sample dealing] step</strong>, but still can't get a clue</p>\n<p>Does someone have any idea? Really appreciate it!</p>",
  "messages": [
    {
      "id": "2118476",
      "postDate": "01/28/2023 03:24:28",
      "content": "<p>Hi kagglers,</p>\n<p>I trained a LGBM Ranker following this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">link</a> by <a href=\"https://www.kaggle.com/CHRIS\" target=\"_blank\">@CHRIS</a> DEOTTE, the eval score looks good, yet the lb score is surprisely low.</p>\n<p>Here are the key steps I did:</p>\n<ol>\n<li>recall: 100 candidate per session</li>\n<li><strong>neg_sample dealing: filter sessions whose candidate have no carts/orders action label</strong></li>\n<li>feature: perpare some features</li>\n<li>Ranker: train a LGBM ranker, and got≈0.6 ndcg@20</li>\n<li>Eval: prediction on valid set, and calc by OTTO metric and get 0.35</li>\n</ol>\n<p>Neg sample </p>\n<pre><code>def filter_out_neg_sessions(train_set,target,frac=0.9):\n    session_target_sum = train_set.groupby('session')[target].sum().to_frame()\n    filter_out_session = session_target_sum[session_target_sum[target]==0].sample(frac=frac)\n    train_set = train_set[~train_set.session.isin(filter_out_session.index)]\n    return train_set\n</code></pre>\n<p>LGBM</p>\n<pre><code>params = {\n    'learning_rate': 0.005,\n    'max_depth': 8,\n    'objective':\"lambdarank\",\n    'metric':\"ndcg\",\n    'colsample_bytree': 0.9,\n    'subsample':0.9,\n    'boosting_type':\"gbdt\",\n    'n_estimators':200,\n    'importance_type':'gain'\n}\nskf = GroupKFold(n_splits=KFOLDS)\nfor fold,(train_idx, valid_idx) in tqdm(enumerate(skf.split(train_set, train_set[TARGET], groups=train_set['session']))):\n\n    X_train = train_set.iloc[train_idx][FEATURES]\n    y_train = train_set.iloc[train_idx][TARGET]\n    X_valid = train_set.iloc[valid_idx][FEATURES]\n    y_valid = train_set.iloc[valid_idx][TARGET]\n    group_train = train_set.iloc[train_idx].groupby('session')['session'].count()\n    group_valid = train_set.iloc[valid_idx].groupby('session')['session'].count()\n    print(train_idx.shape,X_train.shape,y_train.shape)\n    ranker = LGBMRanker(\n        **params\n    )\n    ranker = ranker.fit(\n        X_train,\n        y_train,\n        group=group_train,\n        eval_set=[(X_train,y_train),(X_valid, y_valid)],\n        eval_group=[group_train,group_valid],\n        eval_at=(1,5,10,20,30)\n    )\n    lgb.plot_metric(ranker)\n    lgb.plot_importance(ranker,max_num_features=20)\n    plt.show()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1194567%2F9199ea823b62538d5ba08d02c4ad2eb5%2Fkaggle.png?generation=1674876055881572&amp;alt=media\" alt=\"\"></p>\n<p>I had noticed that the more I keep neg_samples, the higer ndcg score I have. <em>(if keep all, I got 0.91)</em></p>\n<p>Also, I submitted recall result separatly, LB score is 0.575, so the candidate is fine.<br>\nPersonaly I believe <strong>there is something wrong on the [neg_sample dealing] step</strong>, but still can't get a clue</p>\n<p>Does someone have any idea? Really appreciate it!</p>",
      "rawMarkdown": "Hi kagglers,\n\nI trained a LGBM Ranker following this [link](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210) by @CHRIS DEOTTE, the eval score looks good, yet the lb score is surprisely low.\n\nHere are the key steps I did:\n1. recall: 100 candidate per session\n2. **neg_sample dealing: filter sessions whose candidate have no carts/orders action label**\n3. feature: perpare some features\n4. Ranker: train a LGBM ranker, and got≈0.6 ndcg@20\n5. Eval: prediction on valid set, and calc by OTTO metric and get 0.35\n\nNeg sample \n```\ndef filter_out_neg_sessions(train_set,target,frac=0.9):\n    session_target_sum = train_set.groupby('session')[target].sum().to_frame()\n    filter_out_session = session_target_sum[session_target_sum[target]==0].sample(frac=frac)\n    train_set = train_set[~train_set.session.isin(filter_out_session.index)]\n    return train_set\n```\nLGBM\n```\nparams = {\n    'learning_rate': 0.005,\n    'max_depth': 8,\n    'objective':\"lambdarank\",\n    'metric':\"ndcg\",\n    'colsample_bytree': 0.9,\n    'subsample':0.9,\n    'boosting_type':\"gbdt\",\n    'n_estimators':200,\n    'importance_type':'gain'\n}\nskf = GroupKFold(n_splits=KFOLDS)\nfor fold,(train_idx, valid_idx) in tqdm(enumerate(skf.split(train_set, train_set[TARGET], groups=train_set['session']))):\n\n    X_train = train_set.iloc[train_idx][FEATURES]\n    y_train = train_set.iloc[train_idx][TARGET]\n    X_valid = train_set.iloc[valid_idx][FEATURES]\n    y_valid = train_set.iloc[valid_idx][TARGET]\n    group_train = train_set.iloc[train_idx].groupby('session')['session'].count()\n    group_valid = train_set.iloc[valid_idx].groupby('session')['session'].count()\n    print(train_idx.shape,X_train.shape,y_train.shape)\n    ranker = LGBMRanker(\n        **params\n    )\n    ranker = ranker.fit(\n        X_train,\n        y_train,\n        group=group_train,\n        eval_set=[(X_train,y_train),(X_valid, y_valid)],\n        eval_group=[group_train,group_valid],\n        eval_at=(1,5,10,20,30)\n    )\n    lgb.plot_metric(ranker)\n    lgb.plot_importance(ranker,max_num_features=20)\n    plt.show()\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1194567%2F9199ea823b62538d5ba08d02c4ad2eb5%2Fkaggle.png?generation=1674876055881572&alt=media)\n\nI had noticed that the more I keep neg_samples, the higer ndcg score I have. *(if keep all, I got 0.91)*\n\nAlso, I submitted recall result separatly, LB score is 0.575, so the candidate is fine.\nPersonaly I believe **there is something wrong on the [neg_sample dealing] step**, but still can't get a clue\n\n\nDoes someone have any idea? Really appreciate it!",
      "votes": null
    },
    {
      "id": "2118510",
      "postDate": "01/28/2023 04:12:24",
      "content": "<p>I think one possible reason is that your feature is not strong enough, so your rank model get overfitted. </p>",
      "rawMarkdown": "I think one possible reason is that your feature is not strong enough, so your rank model get overfitted.",
      "votes": null
    },
    {
      "id": "2118515",
      "postDate": "01/28/2023 04:15:51",
      "content": "<p>Thx for reply, in most case I will consider overfiting first, but in this case the metric plot shows the valid/train scores are quite near, it doesn't seem like overfiting?</p>",
      "rawMarkdown": "Thx for reply, in most case I will consider overfiting first, but in this case the metric plot shows the valid/train scores are quite near, it doesn't seem like overfiting?",
      "votes": null
    },
    {
      "id": "2118525",
      "postDate": "01/28/2023 04:31:51",
      "content": "<p>You are probably dropping negatives from your validation set too. You shouldn't do that.</p>",
      "rawMarkdown": "You are probably dropping negatives from your validation set too. You shouldn't do that.",
      "votes": null
    },
    {
      "id": "2118526",
      "postDate": "01/28/2023 04:31:57",
      "content": "<p>In short, your average session length is too short after downsampling. details discussed here:  <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/377442\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/377442</a></p>",
      "rawMarkdown": "In short, your average session length is too short after downsampling. details discussed here:  https://www.kaggle.com/competitions/otto-recommender-system/discussion/377442",
      "votes": null
    },
    {
      "id": "2118569",
      "postDate": "01/28/2023 05:50:41",
      "content": "<p>Thx🙏. I read about that discussion earlier, but in my case I sample on sessions not candidates, each session have 100 candidates after neg sample. (before sample there are <strong>1.8M</strong> sessions, after neg sample remains <strong>370K</strong> sessions)<br>\nMeanwhile, I got ndcg@20≈0.9 when I kept all negative samples, the more I drop neg samples, the lower ndcg@20 score I have. Is it reasonable?</p>",
      "rawMarkdown": "Thx🙏. I read about that discussion earlier, but in my case I sample on sessions not candidates, each session have 100 candidates after neg sample. (before sample there are **1.8M** sessions, after neg sample remains **370K** sessions)\nMeanwhile, I got ndcg@20≈0.9 when I kept all negative samples, the more I drop neg samples, the lower ndcg@20 score I have. Is it reasonable?",
      "votes": null
    },
    {
      "id": "2118576",
      "postDate": "01/28/2023 05:54:10",
      "content": "<p>Yes. I just update more detail of my pipeline. I used same dataframe for train&amp;eval (under K-folding) , so the neg sample method is same</p>",
      "rawMarkdown": "Yes. I just update more detail of my pipeline. I used same dataframe for train&eval (under K-folding) , so the neg sample method is same",
      "votes": null
    },
    {
      "id": "2118677",
      "postDate": "01/28/2023 07:21:53",
      "content": "<p>I just read your code, and found it is quite a different story. <br>\nIn short, in your case, more negative sample have a higher ndcg is an expected outcome.<br>\nIn details, because your downsampling method did not reduce the length of each session, so the ndcg will not increase due to less noise in each session like in other people's cases (they reduce the session length 👉 less noise each session 👉 higher ndcg).<br>\nInstead, you got higher ndcg while keeping all negative samples, simply because the model sees more data while having the same session length with after downsampling.</p>",
      "rawMarkdown": "I just read your code, and found it is quite a different story. \nIn short, in your case, more negative sample have a higher ndcg is an expected outcome.\nIn details, because your downsampling method did not reduce the length of each session, so the ndcg will not increase due to less noise in each session like in other people's cases (they reduce the session length 👉 less noise each session 👉 higher ndcg).\nInstead, you got higher ndcg while keeping all negative samples, simply because the model sees more data while having the same session length with after downsampling.",
      "votes": null
    },
    {
      "id": "2118942",
      "postDate": "01/28/2023 11:58:59",
      "content": "<p>Indeed, I noticed the difference too. Just to make sure: do you keep the sessions who have no carts/orders label in the train_set?</p>",
      "rawMarkdown": "Indeed, I noticed the difference too. Just to make sure: do you keep the sessions who have no carts/orders label in the train_set?",
      "votes": null
    },
    {
      "id": "2118945",
      "postDate": "01/28/2023 12:01:39",
      "content": "<p>No, I dropped all sessions without positive targets.</p>",
      "rawMarkdown": "No, I dropped all sessions without positive targets.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2118510,
      "author_name": "kimoyami",
      "author_url": "",
      "post_date": "01/28/2023 04:12:24",
      "content": "<p>I think one possible reason is that your feature is not strong enough, so your rank model get overfitted. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2118515,
          "author_name": "yuzhang0422",
          "author_url": "",
          "post_date": "01/28/2023 04:15:51",
          "content": "<p>Thx for reply, in most case I will consider overfiting first, but in this case the metric plot shows the valid/train scores are quite near, it doesn't seem like overfiting?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2118525,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "01/28/2023 04:31:51",
      "content": "<p>You are probably dropping negatives from your validation set too. You shouldn't do that.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2118576,
          "author_name": "yuzhang0422",
          "author_url": "",
          "post_date": "01/28/2023 05:54:10",
          "content": "<p>Yes. I just update more detail of my pipeline. I used same dataframe for train&amp;eval (under K-folding) , so the neg sample method is same</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2118526,
      "author_name": "buumoo",
      "author_url": "",
      "post_date": "01/28/2023 04:31:57",
      "content": "<p>In short, your average session length is too short after downsampling. details discussed here:  <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/377442\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/377442</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2118569,
          "author_name": "yuzhang0422",
          "author_url": "",
          "post_date": "01/28/2023 05:50:41",
          "content": "<p>Thx🙏. I read about that discussion earlier, but in my case I sample on sessions not candidates, each session have 100 candidates after neg sample. (before sample there are <strong>1.8M</strong> sessions, after neg sample remains <strong>370K</strong> sessions)<br>\nMeanwhile, I got ndcg@20≈0.9 when I kept all negative samples, the more I drop neg samples, the lower ndcg@20 score I have. Is it reasonable?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2118677,
              "author_name": "buumoo",
              "author_url": "",
              "post_date": "01/28/2023 07:21:53",
              "content": "<p>I just read your code, and found it is quite a different story. <br>\nIn short, in your case, more negative sample have a higher ndcg is an expected outcome.<br>\nIn details, because your downsampling method did not reduce the length of each session, so the ndcg will not increase due to less noise in each session like in other people's cases (they reduce the session length 👉 less noise each session 👉 higher ndcg).<br>\nInstead, you got higher ndcg while keeping all negative samples, simply because the model sees more data while having the same session length with after downsampling.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2118942,
                  "author_name": "yuzhang0422",
                  "author_url": "",
                  "post_date": "01/28/2023 11:58:59",
                  "content": "<p>Indeed, I noticed the difference too. Just to make sure: do you keep the sessions who have no carts/orders label in the train_set?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2118945,
                      "author_name": "buumoo",
                      "author_url": "",
                      "post_date": "01/28/2023 12:01:39",
                      "content": "<p>No, I dropped all sessions without positive targets.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2118476": "Hi kagglers,\n\nI trained a LGBM Ranker following this [link](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210) by @CHRIS DEOTTE, the eval score looks good, yet the lb score is surprisely low.\n\nHere are the key steps I did:\n1. recall: 100 candidate per session\n2. **neg_sample dealing: filter sessions whose candidate have no carts/orders action label**\n3. feature: perpare some features\n4. Ranker: train a LGBM ranker, and got≈0.6 ndcg@20\n5. Eval: prediction on valid set, and calc by OTTO metric and get 0.35\n\nNeg sample \n```\ndef filter_out_neg_sessions(train_set,target,frac=0.9):\n    session_target_sum = train_set.groupby('session')[target].sum().to_frame()\n    filter_out_session = session_target_sum[session_target_sum[target]==0].sample(frac=frac)\n    train_set = train_set[~train_set.session.isin(filter_out_session.index)]\n    return train_set\n```\nLGBM\n```\nparams = {\n    'learning_rate': 0.005,\n    'max_depth': 8,\n    'objective':\"lambdarank\",\n    'metric':\"ndcg\",\n    'colsample_bytree': 0.9,\n    'subsample':0.9,\n    'boosting_type':\"gbdt\",\n    'n_estimators':200,\n    'importance_type':'gain'\n}\nskf = GroupKFold(n_splits=KFOLDS)\nfor fold,(train_idx, valid_idx) in tqdm(enumerate(skf.split(train_set, train_set[TARGET], groups=train_set['session']))):\n\n    X_train = train_set.iloc[train_idx][FEATURES]\n    y_train = train_set.iloc[train_idx][TARGET]\n    X_valid = train_set.iloc[valid_idx][FEATURES]\n    y_valid = train_set.iloc[valid_idx][TARGET]\n    group_train = train_set.iloc[train_idx].groupby('session')['session'].count()\n    group_valid = train_set.iloc[valid_idx].groupby('session')['session'].count()\n    print(train_idx.shape,X_train.shape,y_train.shape)\n    ranker = LGBMRanker(\n        **params\n    )\n    ranker = ranker.fit(\n        X_train,\n        y_train,\n        group=group_train,\n        eval_set=[(X_train,y_train),(X_valid, y_valid)],\n        eval_group=[group_train,group_valid],\n        eval_at=(1,5,10,20,30)\n    )\n    lgb.plot_metric(ranker)\n    lgb.plot_importance(ranker,max_num_features=20)\n    plt.show()\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1194567%2F9199ea823b62538d5ba08d02c4ad2eb5%2Fkaggle.png?generation=1674876055881572&alt=media)\n\nI had noticed that the more I keep neg_samples, the higer ndcg score I have. *(if keep all, I got 0.91)*\n\nAlso, I submitted recall result separatly, LB score is 0.575, so the candidate is fine.\nPersonaly I believe **there is something wrong on the [neg_sample dealing] step**, but still can't get a clue\n\n\nDoes someone have any idea? Really appreciate it!",
    "2118510": "I think one possible reason is that your feature is not strong enough, so your rank model get overfitted.",
    "2118515": "Thx for reply, in most case I will consider overfiting first, but in this case the metric plot shows the valid/train scores are quite near, it doesn't seem like overfiting?",
    "2118525": "You are probably dropping negatives from your validation set too. You shouldn't do that.",
    "2118526": "In short, your average session length is too short after downsampling. details discussed here:  https://www.kaggle.com/competitions/otto-recommender-system/discussion/377442",
    "2118569": "Thx🙏. I read about that discussion earlier, but in my case I sample on sessions not candidates, each session have 100 candidates after neg sample. (before sample there are **1.8M** sessions, after neg sample remains **370K** sessions)\nMeanwhile, I got ndcg@20≈0.9 when I kept all negative samples, the more I drop neg samples, the lower ndcg@20 score I have. Is it reasonable?",
    "2118576": "Yes. I just update more detail of my pipeline. I used same dataframe for train&eval (under K-folding) , so the neg sample method is same",
    "2118677": "I just read your code, and found it is quite a different story. \nIn short, in your case, more negative sample have a higher ndcg is an expected outcome.\nIn details, because your downsampling method did not reduce the length of each session, so the ndcg will not increase due to less noise in each session like in other people's cases (they reduce the session length 👉 less noise each session 👉 higher ndcg).\nInstead, you got higher ndcg while keeping all negative samples, simply because the model sees more data while having the same session length with after downsampling.",
    "2118942": "Indeed, I noticed the difference too. Just to make sure: do you keep the sessions who have no carts/orders label in the train_set?",
    "2118945": "No, I dropped all sessions without positive targets."
  },
  "source": "meta"
}