{
  "id": 376994,
  "title": "NDCG@20 increases while loss increases",
  "url": "/competitions/otto-recommender-system/discussion/376994",
  "author_name": "",
  "post_date": "2023-01-09T12:47:40.438356800Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello community, I'm training a lightgbmRanker I've got those results:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2666506%2F1489668b50466a28a904dbc3ef1dc0af%2FCapture%20decran%202023-01-09%20a%201.36.28%20PM.png?generation=1673268188825032&amp;alt=media\" alt=\"\"></p>\n<p>As you can see, my NDCG@20 increases while the binary loss is increasing. I took care of:</p>\n<ul>\n<li>Using DART booster.</li>\n<li>Using lambda_l1 / lambda_l2.</li>\n<li>Reduce the learning rate.</li>\n<li>Reduce the max_depth.</li>\n<li>Increase the min_sum_hessian_in_leaf.</li>\n</ul>\n<p>I also noticed that I have a feature that has an excessive amount of gain, which is the \"<strong>already_see</strong>\" feature. </p>\n<p>This feature equals 1 if the user has already interacts with the candidate or not.</p>\n<p>am I the only one who got this ? </p>",
  "messages": [
    {
      "id": "2092598",
      "postDate": "01/09/2023 12:47:40",
      "content": "<p>Hello community, I'm training a lightgbmRanker I've got those results:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2666506%2F1489668b50466a28a904dbc3ef1dc0af%2FCapture%20decran%202023-01-09%20a%201.36.28%20PM.png?generation=1673268188825032&amp;alt=media\" alt=\"\"></p>\n<p>As you can see, my NDCG@20 increases while the binary loss is increasing. I took care of:</p>\n<ul>\n<li>Using DART booster.</li>\n<li>Using lambda_l1 / lambda_l2.</li>\n<li>Reduce the learning rate.</li>\n<li>Reduce the max_depth.</li>\n<li>Increase the min_sum_hessian_in_leaf.</li>\n</ul>\n<p>I also noticed that I have a feature that has an excessive amount of gain, which is the \"<strong>already_see</strong>\" feature. </p>\n<p>This feature equals 1 if the user has already interacts with the candidate or not.</p>\n<p>am I the only one who got this ? </p>",
      "rawMarkdown": "Hello community, I'm training a lightgbmRanker I've got those results:![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2666506%2F1489668b50466a28a904dbc3ef1dc0af%2FCapture%20decran%202023-01-09%20a%201.36.28%20PM.png?generation=1673268188825032&alt=media)\n\nAs you can see, my NDCG@20 increases while the binary loss is increasing. I took care of:\n\n- Using DART booster.\n- Using lambda_l1 / lambda_l2.\n- Reduce the learning rate.\n- Reduce the max_depth.\n- Increase the min_sum_hessian_in_leaf.\n\n\nI also noticed that I have a feature that has an excessive amount of gain, which is the \"**already_see**\" feature. \n\nThis feature equals 1 if the user has already interacts with the candidate or not.\n\nam I the only one who got this ?",
      "votes": null
    },
    {
      "id": "2093113",
      "postDate": "01/09/2023 20:09:41",
      "content": "<p>What is the objective function (lambdarank etc)? Since binary log loss is calculated point-wise while NDCG is calculated list-wise, it can be happen where one increase while others decrease.<br>\nAlso, did you shuffle the order of train/validation dataset before training? If you didn’t, and the trainin/validation set has pattern in order (e.g. positive set appears before negative set) the model got may get extremely high score when it always outputs constant score. That is because evaluation metrics is sensitive to the order of the dataset when the models output is constant.</p>",
      "rawMarkdown": "What is the objective function (lambdarank etc)? Since binary log loss is calculated point-wise while NDCG is calculated list-wise, it can be happen where one increase while others decrease.\nAlso, did you shuffle the order of train/validation dataset before training? If you didn’t, and the trainin/validation set has pattern in order (e.g. positive set appears before negative set) the model got may get extremely high score when it always outputs constant score. That is because evaluation metrics is sensitive to the order of the dataset when the models output is constant.",
      "votes": null
    },
    {
      "id": "2093145",
      "postDate": "01/09/2023 20:32:55",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>  for all those useful informations.</p>\n<ol>\n<li><p>Yes I use lambdarank as objective function.</p></li>\n<li><p>Yes I shuffled the data, and I use a subsample of the negative labels. However,  I noticed that &gt; 40% of my sessions has no candidates ( all its targets == 0 ). Can this cause a problem ?</p></li>\n</ol>",
      "rawMarkdown": "Thank you @tatamikenn  for all those useful informations.\n1. Yes I use lambdarank as objective function.\n\n2. Yes I shuffled the data, and I use a subsample of the negative labels. However,  I noticed that > 40% of my sessions has no candidates ( all its targets == 0 ). Can this cause a problem ?",
      "votes": null
    },
    {
      "id": "2093160",
      "postDate": "01/09/2023 20:47:45",
      "content": "<blockquote>\n  <p>Yes I use lambdarank as objective function.</p>\n</blockquote>\n<p>Since lambdarank objective is calculated pair-wise, NDCG should be more correlated with it comparing to the binary logloss. So, I think the shared situation (binary logloss increase while NDCG increase) could be happened.</p>\n<p>Consider the following two cases (case 1 and 2). While binary logloss is lower in case 1. However, the NDCG is also loser in case 1 because the rank of the prediction is more accurate in case 2. To summarize, binary logloss only evaluates the sample point(event)-wise, while NDCG evaluates the sample list(session)-wise.</p>\n<p>Case1: binary logloss is lower (0.6596) while NDCG is lower (0.888)</p>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>label</th>\n<th>pred</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>1</td>\n<td>0.60</td>\n</tr>\n<tr>\n<td>0</td>\n<td>1</td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>0</td>\n<td>0</td>\n<td>0.52</td>\n</tr>\n</tbody>\n</table>\n<p>Case2: binary logloss is higher (0.9263) while NDCG is higher (1.0)</p>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>label</th>\n<th>pred</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>1</td>\n<td>0.30</td>\n</tr>\n<tr>\n<td>0</td>\n<td>1</td>\n<td>0.23</td>\n</tr>\n<tr>\n<td>0</td>\n<td>0</td>\n<td>0.1</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "> Yes I use lambdarank as objective function.\n\nSince lambdarank objective is calculated pair-wise, NDCG should be more correlated with it comparing to the binary logloss. So, I think the shared situation (binary logloss increase while NDCG increase) could be happened.\n\nConsider the following two cases (case 1 and 2). While binary logloss is lower in case 1. However, the NDCG is also loser in case 1 because the rank of the prediction is more accurate in case 2. To summarize, binary logloss only evaluates the sample point(event)-wise, while NDCG evaluates the sample list(session)-wise.\n\nCase1: binary logloss is lower (0.6596) while NDCG is lower (0.888)\n\n| session | label | pred |\n|--------:|-------|------|\n|       0 | 1     | 0.60 |\n|       0 | 1     | 0.48 |\n| 0       | 0     | 0.52  |\n\nCase2: binary logloss is higher (0.9263) while NDCG is higher (1.0)\n\n| session | label | pred |\n|--------:|-------|------|\n|       0 | 1     | 0.30 |\n|       0 | 1     | 0.23  |\n| 0       | 0     | 0.1  |",
      "votes": null
    },
    {
      "id": "2093165",
      "postDate": "01/09/2023 20:50:40",
      "content": "<p>I think binary logloss is not appropriate evaluation metrics for ranking models. That is because when evaluating ranking models, the <strong>rank</strong> of the predicted score is more important than the predicted scores itself.</p>",
      "rawMarkdown": "I think binary logloss is not appropriate evaluation metrics for ranking models. That is because when evaluating ranking models, the **rank** of the predicted score is more important than the predicted scores itself.",
      "votes": null
    },
    {
      "id": "2093197",
      "postDate": "01/09/2023 21:19:27",
      "content": "<blockquote>\n  <p>However, I noticed that &gt; 40% of my sessions has no candidates ( all its targets == 0 ). Can this cause a problem ?</p>\n</blockquote>\n<p>I think this is irrelevant. Although NDCG is not properly defined when all candidates are zero (since the maximum DCG is 0), it should be constant if it is defined.</p>",
      "rawMarkdown": "> However, I noticed that > 40% of my sessions has no candidates ( all its targets == 0 ). Can this cause a problem ?\n\nI think this is irrelevant. Although NDCG is not properly defined when all candidates are zero (since the maximum DCG is 0), it should be constant if it is defined.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2093113,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "01/09/2023 20:09:41",
      "content": "<p>What is the objective function (lambdarank etc)? Since binary log loss is calculated point-wise while NDCG is calculated list-wise, it can be happen where one increase while others decrease.<br>\nAlso, did you shuffle the order of train/validation dataset before training? If you didn’t, and the trainin/validation set has pattern in order (e.g. positive set appears before negative set) the model got may get extremely high score when it always outputs constant score. That is because evaluation metrics is sensitive to the order of the dataset when the models output is constant.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2093145,
          "author_name": "rayanaay",
          "author_url": "",
          "post_date": "01/09/2023 20:32:55",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>  for all those useful informations.</p>\n<ol>\n<li><p>Yes I use lambdarank as objective function.</p></li>\n<li><p>Yes I shuffled the data, and I use a subsample of the negative labels. However,  I noticed that &gt; 40% of my sessions has no candidates ( all its targets == 0 ). Can this cause a problem ?</p></li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 2093160,
              "author_name": "tatamikenn",
              "author_url": "",
              "post_date": "01/09/2023 20:47:45",
              "content": "<blockquote>\n  <p>Yes I use lambdarank as objective function.</p>\n</blockquote>\n<p>Since lambdarank objective is calculated pair-wise, NDCG should be more correlated with it comparing to the binary logloss. So, I think the shared situation (binary logloss increase while NDCG increase) could be happened.</p>\n<p>Consider the following two cases (case 1 and 2). While binary logloss is lower in case 1. However, the NDCG is also loser in case 1 because the rank of the prediction is more accurate in case 2. To summarize, binary logloss only evaluates the sample point(event)-wise, while NDCG evaluates the sample list(session)-wise.</p>\n<p>Case1: binary logloss is lower (0.6596) while NDCG is lower (0.888)</p>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>label</th>\n<th>pred</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>1</td>\n<td>0.60</td>\n</tr>\n<tr>\n<td>0</td>\n<td>1</td>\n<td>0.48</td>\n</tr>\n<tr>\n<td>0</td>\n<td>0</td>\n<td>0.52</td>\n</tr>\n</tbody>\n</table>\n<p>Case2: binary logloss is higher (0.9263) while NDCG is higher (1.0)</p>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>label</th>\n<th>pred</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>1</td>\n<td>0.30</td>\n</tr>\n<tr>\n<td>0</td>\n<td>1</td>\n<td>0.23</td>\n</tr>\n<tr>\n<td>0</td>\n<td>0</td>\n<td>0.1</td>\n</tr>\n</tbody>\n</table>",
              "votes": null,
              "replies": [
                {
                  "id": 2093165,
                  "author_name": "tatamikenn",
                  "author_url": "",
                  "post_date": "01/09/2023 20:50:40",
                  "content": "<p>I think binary logloss is not appropriate evaluation metrics for ranking models. That is because when evaluating ranking models, the <strong>rank</strong> of the predicted score is more important than the predicted scores itself.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2093197,
                      "author_name": "tatamikenn",
                      "author_url": "",
                      "post_date": "01/09/2023 21:19:27",
                      "content": "<blockquote>\n  <p>However, I noticed that &gt; 40% of my sessions has no candidates ( all its targets == 0 ). Can this cause a problem ?</p>\n</blockquote>\n<p>I think this is irrelevant. Although NDCG is not properly defined when all candidates are zero (since the maximum DCG is 0), it should be constant if it is defined.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2092598": "Hello community, I'm training a lightgbmRanker I've got those results:![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2666506%2F1489668b50466a28a904dbc3ef1dc0af%2FCapture%20decran%202023-01-09%20a%201.36.28%20PM.png?generation=1673268188825032&alt=media)\n\nAs you can see, my NDCG@20 increases while the binary loss is increasing. I took care of:\n\n- Using DART booster.\n- Using lambda_l1 / lambda_l2.\n- Reduce the learning rate.\n- Reduce the max_depth.\n- Increase the min_sum_hessian_in_leaf.\n\n\nI also noticed that I have a feature that has an excessive amount of gain, which is the \"**already_see**\" feature. \n\nThis feature equals 1 if the user has already interacts with the candidate or not.\n\nam I the only one who got this ?",
    "2093113": "What is the objective function (lambdarank etc)? Since binary log loss is calculated point-wise while NDCG is calculated list-wise, it can be happen where one increase while others decrease.\nAlso, did you shuffle the order of train/validation dataset before training? If you didn’t, and the trainin/validation set has pattern in order (e.g. positive set appears before negative set) the model got may get extremely high score when it always outputs constant score. That is because evaluation metrics is sensitive to the order of the dataset when the models output is constant.",
    "2093145": "Thank you @tatamikenn  for all those useful informations.\n1. Yes I use lambdarank as objective function.\n\n2. Yes I shuffled the data, and I use a subsample of the negative labels. However,  I noticed that > 40% of my sessions has no candidates ( all its targets == 0 ). Can this cause a problem ?",
    "2093160": "> Yes I use lambdarank as objective function.\n\nSince lambdarank objective is calculated pair-wise, NDCG should be more correlated with it comparing to the binary logloss. So, I think the shared situation (binary logloss increase while NDCG increase) could be happened.\n\nConsider the following two cases (case 1 and 2). While binary logloss is lower in case 1. However, the NDCG is also loser in case 1 because the rank of the prediction is more accurate in case 2. To summarize, binary logloss only evaluates the sample point(event)-wise, while NDCG evaluates the sample list(session)-wise.\n\nCase1: binary logloss is lower (0.6596) while NDCG is lower (0.888)\n\n| session | label | pred |\n|--------:|-------|------|\n|       0 | 1     | 0.60 |\n|       0 | 1     | 0.48 |\n| 0       | 0     | 0.52  |\n\nCase2: binary logloss is higher (0.9263) while NDCG is higher (1.0)\n\n| session | label | pred |\n|--------:|-------|------|\n|       0 | 1     | 0.30 |\n|       0 | 1     | 0.23  |\n| 0       | 0     | 0.1  |",
    "2093165": "I think binary logloss is not appropriate evaluation metrics for ranking models. That is because when evaluating ranking models, the **rank** of the predicted score is more important than the predicted scores itself.",
    "2093197": "> However, I noticed that > 40% of my sessions has no candidates ( all its targets == 0 ). Can this cause a problem ?\n\nI think this is irrelevant. Although NDCG is not properly defined when all candidates are zero (since the maximum DCG is 0), it should be constant if it is defined."
  },
  "source": "meta"
}