{
  "id": 374824,
  "title": "Why my rank model doesn't perform better than the public notebook",
  "url": "/competitions/otto-recommender-system/discussion/374824",
  "author_name": "",
  "post_date": "2022-12-29T02:28:19.918321200Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Recently I have been confused by a problem. <br>\nI recall 200 candidate items for each session, and the recall score for each behavior is as follows：<br>\n| clicks recall | 0.73326395 |<br>\n| carts recall  | 0.55524809 |<br>\n| orders recall | 0.74062868 |<br>\nThe recall score looks good. Then I generate 100+ features for the rank model, which include session features, item features and session/item features etc. I train a LGBMRanker for each behavior , and choose the AUC and NDCG@20 as the eval metrics.  To avoid the memory consumption of training, I downsample the negative samples,like this :<br>\n<code>\nrecall_df_neg = recall_df[recall_df['label'] == 0]\nrecall_df_neg.sample(frac = 0.2,random_state=2022)\n</code><br>\nHowever, the score I got in LB is not good, not even as good as the score before rank.<br>\nI don't know the reason for the poor performance of my rank model,  is it because of over-fitting?</p>",
  "messages": [
    {
      "id": "2079110",
      "postDate": "12/29/2022 02:28:19",
      "content": "<p>Recently I have been confused by a problem. <br>\nI recall 200 candidate items for each session, and the recall score for each behavior is as follows：<br>\n| clicks recall | 0.73326395 |<br>\n| carts recall  | 0.55524809 |<br>\n| orders recall | 0.74062868 |<br>\nThe recall score looks good. Then I generate 100+ features for the rank model, which include session features, item features and session/item features etc. I train a LGBMRanker for each behavior , and choose the AUC and NDCG@20 as the eval metrics.  To avoid the memory consumption of training, I downsample the negative samples,like this :<br>\n<code>\nrecall_df_neg = recall_df[recall_df['label'] == 0]\nrecall_df_neg.sample(frac = 0.2,random_state=2022)\n</code><br>\nHowever, the score I got in LB is not good, not even as good as the score before rank.<br>\nI don't know the reason for the poor performance of my rank model,  is it because of over-fitting?</p>",
      "rawMarkdown": "Recently I have been confused by a problem. \nI recall 200 candidate items for each session, and the recall score for each behavior is as follows：\n| clicks recall | 0.73326395 |\n| carts recall  | 0.55524809 |\n| orders recall | 0.74062868 |\nThe recall score looks good. Then I generate 100+ features for the rank model, which include session features, item features and session/item features etc. I train a LGBMRanker for each behavior , and choose the AUC and NDCG@20 as the eval metrics.  To avoid the memory consumption of training, I downsample the negative samples,like this :\n``\nrecall_df_neg = recall_df[recall_df['label'] == 0]\nrecall_df_neg.sample(frac = 0.2,random_state=2022)\n``\nHowever, the score I got in LB is not good, not even as good as the score before rank.\nI don't know the reason for the poor performance of my rank model,  is it because of over-fitting?",
      "votes": null
    },
    {
      "id": "2079857",
      "postDate": "12/29/2022 16:53:46",
      "content": "<p>Question: are those recall values from local validation? </p>\n<p>I am encountering a similar issue because the ranker model does worse than manual ranking. So far my hunch is - probably I don't have good enough features and the model is overfitting quickly. I am also trying to revise features and see if I am leaking targets in any way. Let me know if you come across what was causing the issue. </p>",
      "rawMarkdown": "Question: are those recall values from local validation? \n\nI am encountering a similar issue because the ranker model does worse than manual ranking. So far my hunch is - probably I don't have good enough features and the model is overfitting quickly. I am also trying to revise features and see if I am leaking targets in any way. Let me know if you come across what was causing the issue.",
      "votes": null
    },
    {
      "id": "2081117",
      "postDate": "12/30/2022 20:23:25",
      "content": "<p>I have the same problem. The public score of my ranker is much lower than the score of manually crafted candidates. </p>\n<p>I trained my model on features from training set (3 weeks as training set, 1 week as validation set) and then I created features from 4 weeks of training data and 1 week of test data and applied model trained on 3 weeks of training data (as Chris Deotte described here: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210)\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210)</a>. </p>\n<p>May this be a reason of such a bad performance? My features had raw counts and distribution of these counts may be quite different on 4+1 weeks than on 3+1 weeks.</p>",
      "rawMarkdown": "I have the same problem. The public score of my ranker is much lower than the score of manually crafted candidates. \n\nI trained my model on features from training set (3 weeks as training set, 1 week as validation set) and then I created features from 4 weeks of training data and 1 week of test data and applied model trained on 3 weeks of training data (as Chris Deotte described here: https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210). \n\nMay this be a reason of such a bad performance? My features had raw counts and distribution of these counts may be quite different on 4+1 weeks than on 3+1 weeks.",
      "votes": null
    },
    {
      "id": "2081246",
      "postDate": "12/31/2022 01:36:08",
      "content": "<p>Yes, I also use 3 weeks of historical data + 1 week of validation set to generate item features, and then train the ranking model. For inference, I use 4 weeks of historical data + 1 week of test set to generate item features. Indeed, this may be the reason for the poor performance of the ranking model. I will try to use 3 weeks of historical data + 1 week of session dataset to generate features to ensure that the distribution of features during training and inference is consistent.</p>",
      "rawMarkdown": "Yes, I also use 3 weeks of historical data + 1 week of validation set to generate item features, and then train the ranking model. For inference, I use 4 weeks of historical data + 1 week of test set to generate item features. Indeed, this may be the reason for the poor performance of the ranking model. I will try to use 3 weeks of historical data + 1 week of session dataset to generate features to ensure that the distribution of features during training and inference is consistent.",
      "votes": null
    },
    {
      "id": "2091116",
      "postDate": "01/08/2023 03:14:01",
      "content": "<p>What is your ndcg@20 score?</p>",
      "rawMarkdown": "What is your ndcg@20 score?",
      "votes": null
    },
    {
      "id": "2120875",
      "postDate": "01/29/2023 21:48:21",
      "content": "<p>Hi. Is this problem fixed?  It's a bug, right?</p>",
      "rawMarkdown": "Hi. Is this problem fixed?  It's a bug, right?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2079857,
      "author_name": "parthpankajtiwary",
      "author_url": "",
      "post_date": "12/29/2022 16:53:46",
      "content": "<p>Question: are those recall values from local validation? </p>\n<p>I am encountering a similar issue because the ranker model does worse than manual ranking. So far my hunch is - probably I don't have good enough features and the model is overfitting quickly. I am also trying to revise features and see if I am leaking targets in any way. Let me know if you come across what was causing the issue. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2081117,
          "author_name": "jamnik99",
          "author_url": "",
          "post_date": "12/30/2022 20:23:25",
          "content": "<p>I have the same problem. The public score of my ranker is much lower than the score of manually crafted candidates. </p>\n<p>I trained my model on features from training set (3 weeks as training set, 1 week as validation set) and then I created features from 4 weeks of training data and 1 week of test data and applied model trained on 3 weeks of training data (as Chris Deotte described here: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210)\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210)</a>. </p>\n<p>May this be a reason of such a bad performance? My features had raw counts and distribution of these counts may be quite different on 4+1 weeks than on 3+1 weeks.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2081246,
              "author_name": "chaihadyy",
              "author_url": "",
              "post_date": "12/31/2022 01:36:08",
              "content": "<p>Yes, I also use 3 weeks of historical data + 1 week of validation set to generate item features, and then train the ranking model. For inference, I use 4 weeks of historical data + 1 week of test set to generate item features. Indeed, this may be the reason for the poor performance of the ranking model. I will try to use 3 weeks of historical data + 1 week of session dataset to generate features to ensure that the distribution of features during training and inference is consistent.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2091116,
      "author_name": "wangqihanginthesky",
      "author_url": "",
      "post_date": "01/08/2023 03:14:01",
      "content": "<p>What is your ndcg@20 score?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2120875,
      "author_name": "huaguo",
      "author_url": "",
      "post_date": "01/29/2023 21:48:21",
      "content": "<p>Hi. Is this problem fixed?  It's a bug, right?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2079110": "Recently I have been confused by a problem. \nI recall 200 candidate items for each session, and the recall score for each behavior is as follows：\n| clicks recall | 0.73326395 |\n| carts recall  | 0.55524809 |\n| orders recall | 0.74062868 |\nThe recall score looks good. Then I generate 100+ features for the rank model, which include session features, item features and session/item features etc. I train a LGBMRanker for each behavior , and choose the AUC and NDCG@20 as the eval metrics.  To avoid the memory consumption of training, I downsample the negative samples,like this :\n``\nrecall_df_neg = recall_df[recall_df['label'] == 0]\nrecall_df_neg.sample(frac = 0.2,random_state=2022)\n``\nHowever, the score I got in LB is not good, not even as good as the score before rank.\nI don't know the reason for the poor performance of my rank model,  is it because of over-fitting?",
    "2079857": "Question: are those recall values from local validation? \n\nI am encountering a similar issue because the ranker model does worse than manual ranking. So far my hunch is - probably I don't have good enough features and the model is overfitting quickly. I am also trying to revise features and see if I am leaking targets in any way. Let me know if you come across what was causing the issue.",
    "2081117": "I have the same problem. The public score of my ranker is much lower than the score of manually crafted candidates. \n\nI trained my model on features from training set (3 weeks as training set, 1 week as validation set) and then I created features from 4 weeks of training data and 1 week of test data and applied model trained on 3 weeks of training data (as Chris Deotte described here: https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210). \n\nMay this be a reason of such a bad performance? My features had raw counts and distribution of these counts may be quite different on 4+1 weeks than on 3+1 weeks.",
    "2081246": "Yes, I also use 3 weeks of historical data + 1 week of validation set to generate item features, and then train the ranking model. For inference, I use 4 weeks of historical data + 1 week of test set to generate item features. Indeed, this may be the reason for the poor performance of the ranking model. I will try to use 3 weeks of historical data + 1 week of session dataset to generate features to ensure that the distribution of features during training and inference is consistent.",
    "2091116": "What is your ndcg@20 score?",
    "2120875": "Hi. Is this problem fixed?  It's a bug, right?"
  },
  "source": "meta"
}