{
  "id": 377442,
  "title": "NDCG@20 = 0.99 after re-ranking orders candidates",
  "url": "/competitions/otto-recommender-system/discussion/377442",
  "author_name": "",
  "post_date": "2023-01-11T08:05:18.038431800Z",
  "votes": 2,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hello community, I got 0.99 on ndcg@20 for orders re-ranking, which I find very suspicious. Is it because the number of orders are relatively lower than the number of clicks ?</p>\n<p>Also, my click reranker performed well ( 0.54 vs 0.5255 benchmark ). However, when using the same gbdt with the same variables for the orders re-ranker, the performance got worse ( 0.61 vs 0.64 benchmark ).</p>\n<p>Do you guys use the same parameters and variables for clicks / orders / carts ? <br>\nDoes the model overfit more quickly when trying to predict orders regarding  the low proportion of real candidates vs false candidates ?</p>",
  "messages": [
    {
      "id": "2095120",
      "postDate": "01/11/2023 08:05:18",
      "content": "<p>Hello community, I got 0.99 on ndcg@20 for orders re-ranking, which I find very suspicious. Is it because the number of orders are relatively lower than the number of clicks ?</p>\n<p>Also, my click reranker performed well ( 0.54 vs 0.5255 benchmark ). However, when using the same gbdt with the same variables for the orders re-ranker, the performance got worse ( 0.61 vs 0.64 benchmark ).</p>\n<p>Do you guys use the same parameters and variables for clicks / orders / carts ? <br>\nDoes the model overfit more quickly when trying to predict orders regarding  the low proportion of real candidates vs false candidates ?</p>",
      "rawMarkdown": "Hello community, I got 0.99 on ndcg@20 for orders re-ranking, which I find very suspicious. Is it because the number of orders are relatively lower than the number of clicks ?\n\nAlso, my click reranker performed well ( 0.54 vs 0.5255 benchmark ). However, when using the same gbdt with the same variables for the orders re-ranker, the performance got worse ( 0.61 vs 0.64 benchmark ).\n\nDo you guys use the same parameters and variables for clicks / orders / carts ? \nDoes the model overfit more quickly when trying to predict orders regarding  the low proportion of real candidates vs false candidates ?",
      "votes": null
    },
    {
      "id": "2097833",
      "postDate": "01/13/2023 01:04:41",
      "content": "<p>My ndcg@20 for clicks and carts are near 0.86 using 20 trees, for orders it's 0.91. I haven't gotten chance to make a submission with reranking yet. Probably submit one today. Quite busy with the work this year…</p>",
      "rawMarkdown": "My ndcg@20 for clicks and carts are near 0.86 using 20 trees, for orders it's 0.91. I haven't gotten chance to make a submission with reranking yet. Probably submit one today. Quite busy with the work this year...",
      "votes": null
    },
    {
      "id": "2099887",
      "postDate": "01/14/2023 18:37:32",
      "content": "<p>The ultra high ndcg comes from the negtive sample downsampling. Assuming your sampling ratio positive:negtive = 1:20, then recall at 20 means collect 20 samples from 21 samples for each session, therefore even a random guess can reach 90%+ hit rate. For the sake of a fair evalution, put the 'noise' (all negtive samples in your valid fold) back and inference. You will find your model cannot performs that good anymore under such large amount of nosie. You will get the true performance of your ranker model.</p>",
      "rawMarkdown": "The ultra high ndcg comes from the negtive sample downsampling. Assuming your sampling ratio positive:negtive = 1:20, then recall at 20 means collect 20 samples from 21 samples for each session, therefore even a random guess can reach 90%+ hit rate. For the sake of a fair evalution, put the 'noise' (all negtive samples in your valid fold) back and inference. You will find your model cannot performs that good anymore under such large amount of nosie. You will get the true performance of your ranker model.",
      "votes": null
    },
    {
      "id": "2099923",
      "postDate": "01/14/2023 19:12:52",
      "content": "<p>Insightful ! If I well understood you suggest to avoid downsampling negative sample in validation, in order to get a fair evaluation.</p>",
      "rawMarkdown": "Insightful ! If I well understood you suggest to avoid downsampling negative sample in validation, in order to get a fair evaluation.",
      "votes": null
    },
    {
      "id": "2099963",
      "postDate": "01/14/2023 19:47:25",
      "content": "<p>Yes, especially if you use early stopping, ndcg@20 in sessions with only average 21 samples is not a good choice.</p>",
      "rawMarkdown": "Yes, especially if you use early stopping, ndcg@20 in sessions with only average 21 samples is not a good choice.",
      "votes": null
    },
    {
      "id": "2100513",
      "postDate": "01/15/2023 08:35:01",
      "content": "<p>I also face the same problem (ncdg around .95 for carts and orders) even though I didn't downsample. Any ideas from where this comes from?</p>",
      "rawMarkdown": "I also face the same problem (ncdg around .95 for carts and orders) even though I didn't downsample. Any ideas from where this comes from?",
      "votes": null
    },
    {
      "id": "2100522",
      "postDate": "01/15/2023 08:39:31",
      "content": "<p>How does your model generalize performance to test set, namely LB? If it is much lower than your validation, then one of some most posible reasons might be there is some target leakage during your feature generation.</p>",
      "rawMarkdown": "How does your model generalize performance to test set, namely LB? If it is much lower than your validation, then one of some most posible reasons might be there is some target leakage during your feature generation.",
      "votes": null
    },
    {
      "id": "2100525",
      "postDate": "01/15/2023 08:41:26",
      "content": "<p>My CV is 0.552. My LB is 0.568</p>",
      "rawMarkdown": "My CV is 0.552. My LB is 0.568",
      "votes": null
    },
    {
      "id": "2100528",
      "postDate": "01/15/2023 08:44:41",
      "content": "<p>seems correct. Then how about your average session length? If your retrival items per session is very low, ndcg@20 would be high.</p>",
      "rawMarkdown": "seems correct. Then how about your average session length? If your retrival items per session is very low, ndcg@20 would be high.",
      "votes": null
    },
    {
      "id": "2100530",
      "postDate": "01/15/2023 08:44:59",
      "content": "<p>The CV scores for clicks are comparable to my Co visitation matrix approach which scores 0.577. But the carts and orders recall much lower..<br>\nIf you could explain that to me I would be grateful..</p>",
      "rawMarkdown": "The CV scores for clicks are comparable to my Co visitation matrix approach which scores 0.577. But the carts and orders recall much lower..\nIf you could explain that to me I would be grateful..",
      "votes": null
    },
    {
      "id": "2100533",
      "postDate": "01/15/2023 08:46:26",
      "content": "<p>Currently I use 20 candidates per session. I will try more now. But in my Intuition that doesn't explain the extremly high scores for carts and orders.  </p>",
      "rawMarkdown": "Currently I use 20 candidates per session. I will try more now. But in my Intuition that doesn't explain the extremly high scores for carts and orders.",
      "votes": null
    },
    {
      "id": "2100541",
      "postDate": "01/15/2023 08:50:24",
      "content": "<p>So that's the reason… You did not downsample, but your session length is so short which equals to a downsampling effect from whatever total number to only 20.</p>",
      "rawMarkdown": "So that's the reason... You did not downsample, but your session length is so short which equals to a downsampling effect from whatever total number to only 20.",
      "votes": null
    },
    {
      "id": "2100558",
      "postDate": "01/15/2023 08:59:24",
      "content": "<p>Thank you so much! I will fix that and then report how I improve (hopefully)</p>",
      "rawMarkdown": "Thank you so much! I will fix that and then report how I improve (hopefully)",
      "votes": null
    },
    {
      "id": "2102120",
      "postDate": "01/16/2023 12:05:22",
      "content": "<p>Unfortunately with 100 candidates CV score got worse and the issue didn't resolve (high score for order and carts.. )</p>",
      "rawMarkdown": "Unfortunately with 100 candidates CV score got worse and the issue didn't resolve (high score for order and carts.. )",
      "votes": null
    },
    {
      "id": "2118596",
      "postDate": "01/28/2023 06:15:51",
      "content": "<p>May I ask how do you fix this issue later? And how do you do neg sample?</p>",
      "rawMarkdown": "May I ask how do you fix this issue later? And how do you do neg sample?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2097833,
      "author_name": "wuwenmin",
      "author_url": "",
      "post_date": "01/13/2023 01:04:41",
      "content": "<p>My ndcg@20 for clicks and carts are near 0.86 using 20 trees, for orders it's 0.91. I haven't gotten chance to make a submission with reranking yet. Probably submit one today. Quite busy with the work this year…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2099887,
      "author_name": "buumoo",
      "author_url": "",
      "post_date": "01/14/2023 18:37:32",
      "content": "<p>The ultra high ndcg comes from the negtive sample downsampling. Assuming your sampling ratio positive:negtive = 1:20, then recall at 20 means collect 20 samples from 21 samples for each session, therefore even a random guess can reach 90%+ hit rate. For the sake of a fair evalution, put the 'noise' (all negtive samples in your valid fold) back and inference. You will find your model cannot performs that good anymore under such large amount of nosie. You will get the true performance of your ranker model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2099923,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/14/2023 19:12:52",
          "content": "<p>Insightful ! If I well understood you suggest to avoid downsampling negative sample in validation, in order to get a fair evaluation.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2099963,
              "author_name": "buumoo",
              "author_url": "",
              "post_date": "01/14/2023 19:47:25",
              "content": "<p>Yes, especially if you use early stopping, ndcg@20 in sessions with only average 21 samples is not a good choice.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2100513,
                  "author_name": "simonveitner",
                  "author_url": "",
                  "post_date": "01/15/2023 08:35:01",
                  "content": "<p>I also face the same problem (ncdg around .95 for carts and orders) even though I didn't downsample. Any ideas from where this comes from?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2100522,
                      "author_name": "buumoo",
                      "author_url": "",
                      "post_date": "01/15/2023 08:39:31",
                      "content": "<p>How does your model generalize performance to test set, namely LB? If it is much lower than your validation, then one of some most posible reasons might be there is some target leakage during your feature generation.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2100525,
                          "author_name": "simonveitner",
                          "author_url": "",
                          "post_date": "01/15/2023 08:41:26",
                          "content": "<p>My CV is 0.552. My LB is 0.568</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2100528,
                              "author_name": "buumoo",
                              "author_url": "",
                              "post_date": "01/15/2023 08:44:41",
                              "content": "<p>seems correct. Then how about your average session length? If your retrival items per session is very low, ndcg@20 would be high.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2100533,
                                  "author_name": "simonveitner",
                                  "author_url": "",
                                  "post_date": "01/15/2023 08:46:26",
                                  "content": "<p>Currently I use 20 candidates per session. I will try more now. But in my Intuition that doesn't explain the extremly high scores for carts and orders.  </p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2100541,
                                      "author_name": "buumoo",
                                      "author_url": "",
                                      "post_date": "01/15/2023 08:50:24",
                                      "content": "<p>So that's the reason… You did not downsample, but your session length is so short which equals to a downsampling effect from whatever total number to only 20.</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 2100558,
                                          "author_name": "simonveitner",
                                          "author_url": "",
                                          "post_date": "01/15/2023 08:59:24",
                                          "content": "<p>Thank you so much! I will fix that and then report how I improve (hopefully)</p>",
                                          "votes": null,
                                          "replies": [
                                            {
                                              "id": 2102120,
                                              "author_name": "simonveitner",
                                              "author_url": "",
                                              "post_date": "01/16/2023 12:05:22",
                                              "content": "<p>Unfortunately with 100 candidates CV score got worse and the issue didn't resolve (high score for order and carts.. )</p>",
                                              "votes": null,
                                              "replies": [
                                                {
                                                  "id": 2118596,
                                                  "author_name": "yuzhang0422",
                                                  "author_url": "",
                                                  "post_date": "01/28/2023 06:15:51",
                                                  "content": "<p>May I ask how do you fix this issue later? And how do you do neg sample?</p>",
                                                  "votes": null,
                                                  "replies": []
                                                }
                                              ]
                                            }
                                          ]
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            },
                            {
                              "id": 2100530,
                              "author_name": "simonveitner",
                              "author_url": "",
                              "post_date": "01/15/2023 08:44:59",
                              "content": "<p>The CV scores for clicks are comparable to my Co visitation matrix approach which scores 0.577. But the carts and orders recall much lower..<br>\nIf you could explain that to me I would be grateful..</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2095120": "Hello community, I got 0.99 on ndcg@20 for orders re-ranking, which I find very suspicious. Is it because the number of orders are relatively lower than the number of clicks ?\n\nAlso, my click reranker performed well ( 0.54 vs 0.5255 benchmark ). However, when using the same gbdt with the same variables for the orders re-ranker, the performance got worse ( 0.61 vs 0.64 benchmark ).\n\nDo you guys use the same parameters and variables for clicks / orders / carts ? \nDoes the model overfit more quickly when trying to predict orders regarding  the low proportion of real candidates vs false candidates ?",
    "2097833": "My ndcg@20 for clicks and carts are near 0.86 using 20 trees, for orders it's 0.91. I haven't gotten chance to make a submission with reranking yet. Probably submit one today. Quite busy with the work this year...",
    "2099887": "The ultra high ndcg comes from the negtive sample downsampling. Assuming your sampling ratio positive:negtive = 1:20, then recall at 20 means collect 20 samples from 21 samples for each session, therefore even a random guess can reach 90%+ hit rate. For the sake of a fair evalution, put the 'noise' (all negtive samples in your valid fold) back and inference. You will find your model cannot performs that good anymore under such large amount of nosie. You will get the true performance of your ranker model.",
    "2099923": "Insightful ! If I well understood you suggest to avoid downsampling negative sample in validation, in order to get a fair evaluation.",
    "2099963": "Yes, especially if you use early stopping, ndcg@20 in sessions with only average 21 samples is not a good choice.",
    "2100513": "I also face the same problem (ncdg around .95 for carts and orders) even though I didn't downsample. Any ideas from where this comes from?",
    "2100522": "How does your model generalize performance to test set, namely LB? If it is much lower than your validation, then one of some most posible reasons might be there is some target leakage during your feature generation.",
    "2100525": "My CV is 0.552. My LB is 0.568",
    "2100528": "seems correct. Then how about your average session length? If your retrival items per session is very low, ndcg@20 would be high.",
    "2100530": "The CV scores for clicks are comparable to my Co visitation matrix approach which scores 0.577. But the carts and orders recall much lower..\nIf you could explain that to me I would be grateful..",
    "2100533": "Currently I use 20 candidates per session. I will try more now. But in my Intuition that doesn't explain the extremly high scores for carts and orders.",
    "2100541": "So that's the reason... You did not downsample, but your session length is so short which equals to a downsampling effect from whatever total number to only 20.",
    "2100558": "Thank you so much! I will fix that and then report how I improve (hopefully)",
    "2102120": "Unfortunately with 100 candidates CV score got worse and the issue didn't resolve (high score for order and carts.. )",
    "2118596": "May I ask how do you fix this issue later? And how do you do neg sample?"
  },
  "source": "meta"
}