{
  "id": 377094,
  "title": "Tips: leakage in ranking model: the order of dataset",
  "url": "/competitions/otto-recommender-system/discussion/377094",
  "author_name": "Bilzard",
  "post_date": "2023-01-09T20:34:38.717000",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I have experienced a leakage when training ranking model: the order of dataset.</p>\n<h2>How the leakage occurs?</h2>\n<p>Suppose if your dataset has pattern in your dataset (e.g. positive set appears before negative set for each session), the model get extremely high NDCG while the model outputs the constant value for any input features.</p>\n<p>E.g. assume that your evaluating algorithm selects the first 2 items for each sessions, and calculate Recall@2 in the below dataset (Table 1). You will get maximum recall@2 on the dataset even if your model always outputs a constant score (1.0).</p>\n<p>Table 1: example of leaked dataset (the positive set appears before negative set).</p>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>feat1</th>\n<th>feat2</th>\n<th>label</th>\n<th>pred score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0</td>\n<td>x</td>\n<td>1</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>0</td>\n<td>1</td>\n<td>x</td>\n<td>1</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>0</td>\n<td>0</td>\n<td>x</td>\n<td>0</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>1</td>\n<td>1</td>\n<td>x</td>\n<td>1</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0</td>\n<td>x</td>\n<td>0</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0</td>\n<td>x</td>\n<td>0</td>\n<td>1.0</td>\n</tr>\n</tbody>\n</table>\n<h2>How to avoid the leak</h2>\n<p>To avoid these kind of leaks, </p>\n<ol>\n<li>Shuffle the dataset before training/evaluating</li>\n<li>add tiny amount of noise to the model's predicted score</li>\n</ol>\n<p>Example code of shuffling:</p>\n<pre><code> ():\n    df[] = np.random.randn((df))\n    df = df.sort_values([, , ])\n    df = df.drop(, axis=)\n     df\n</code></pre>",
  "messages": [
    {
      "id": 2093149,
      "postDate": "2023-01-09T20:34:38.717Z",
      "content": "<p>I have experienced a leakage when training ranking model: the order of dataset.</p>\n<h2>How the leakage occurs?</h2>\n<p>Suppose if your dataset has pattern in your dataset (e.g. positive set appears before negative set for each session), the model get extremely high NDCG while the model outputs the constant value for any input features.</p>\n<p>E.g. assume that your evaluating algorithm selects the first 2 items for each sessions, and calculate Recall@2 in the below dataset (Table 1). You will get maximum recall@2 on the dataset even if your model always outputs a constant score (1.0).</p>\n<p>Table 1: example of leaked dataset (the positive set appears before negative set).</p>\n<table>\n<thead>\n<tr>\n<th>session</th>\n<th>feat1</th>\n<th>feat2</th>\n<th>label</th>\n<th>pred score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0</td>\n<td>x</td>\n<td>1</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>0</td>\n<td>1</td>\n<td>x</td>\n<td>1</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>0</td>\n<td>0</td>\n<td>x</td>\n<td>0</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>1</td>\n<td>1</td>\n<td>x</td>\n<td>1</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0</td>\n<td>x</td>\n<td>0</td>\n<td>1.0</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0</td>\n<td>x</td>\n<td>0</td>\n<td>1.0</td>\n</tr>\n</tbody>\n</table>\n<h2>How to avoid the leak</h2>\n<p>To avoid these kind of leaks, </p>\n<ol>\n<li>Shuffle the dataset before training/evaluating</li>\n<li>add tiny amount of noise to the model's predicted score</li>\n</ol>\n<p>Example code of shuffling:</p>\n<pre><code> ():\n    df[] = np.random.randn((df))\n    df = df.sort_values([, , ])\n    df = df.drop(, axis=)\n     df\n</code></pre>",
      "rawMarkdown": "I have experienced a leakage when training ranking model: the order of dataset.\n\n## How the leakage occurs?\n\nSuppose if your dataset has pattern in your dataset (e.g. positive set appears before negative set for each session), the model get extremely high NDCG while the model outputs the constant value for any input features.\n\nE.g. assume that your evaluating algorithm selects the first 2 items for each sessions, and calculate Recall@2 in the below dataset (Table 1). You will get maximum recall@2 on the dataset even if your model always outputs a constant score (1.0).\n\nTable 1: example of leaked dataset (the positive set appears before negative set).\n\n| session | feat1 | feat2 | label | pred score |\n|:-------:|:-----:|:-----:|:-----:|------|\n| 0       | 0     | x     | 1     | 1.0  |\n| 0       | 1     | x     | 1     | 1.0  |\n| 0       | 0     | x     | 0     | 1.0  |\n| 1       | 1     | x     | 1     | 1.0  |\n| 1       | 0     | x     | 0     | 1.0  |\n| 1       | 0     | x     | 0     | 1.0  |\n\n## How to avoid the leak\n\nTo avoid these kind of leaks, \n\n1. Shuffle the dataset before training/evaluating\n2. add tiny amount of noise to the model's predicted score\n\nExample code of shuffling:\n\n```python\ndef shuffle_keep_session(df):\n    df[\"_noise\"] = np.random.randn(len(df))\n    df = df.sort_values([\"type\", \"session\", \"_noise\"])\n    df = df.drop(\"_noise\", axis=1)\n    return df\n```\n",
      "votes": 8
    },
    {
      "id": 2094263,
      "postDate": "2023-01-10T17:19:40.583Z",
      "content": "<p>Thank you mate!!!But will it be working for all the models?</p>",
      "rawMarkdown": "Thank you mate!!!But will it be working for all the models?",
      "replies": [
        {
          "id": 2094265,
          "postDate": "2023-01-10T17:20:57.510Z",
          "content": "<blockquote>\n  <p>But will it be working for all the models?</p>\n</blockquote>\n<p>Sorry, I don't understand what you mean.<br>\nHowever, this thread may be more informative:<br>\n<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/377094#2093481\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/377094#2093481</a></p>",
          "rawMarkdown": "> But will it be working for all the models?\n\nSorry, I don't understand what you mean.\nHowever, this thread may be more informative:\nhttps://www.kaggle.com/competitions/otto-recommender-system/discussion/377094#2093481"
        }
      ]
    },
    {
      "id": 2094240,
      "postDate": "2023-01-10T17:09:36.497Z",
      "content": "<p>Very informative and insightful. Thanks !<br>\nWe should also print the AUC score to help detect this kind of problem earlier ! </p>",
      "rawMarkdown": "Very informative and insightful. Thanks !\nWe should also print the AUC score to help detect this kind of problem earlier ! ",
      "replies": [
        {
          "id": 2094286,
          "postDate": "2023-01-10T17:30:29.537Z",
          "content": "<p>Sorry, I don't understand. How AUC score help detect this kind of leakage?<br>\nI think AUC score also have problem with tie score.</p>",
          "rawMarkdown": "Sorry, I don't understand. How AUC score help detect this kind of leakage?\nI think AUC score also have problem with tie score.",
          "replies": [
            {
              "id": 2094326,
              "postDate": "2023-01-10T17:44:12.863Z",
              "content": "<p>Anyway, I noticed this leakage with extremely hight NDCG (~1.00) with very simple model (always predicts a constant score).</p>",
              "rawMarkdown": "Anyway, I noticed this leakage with extremely hight NDCG (~1.00) with very simple model (always predicts a constant score)."
            }
          ]
        }
      ]
    },
    {
      "id": 2093474,
      "postDate": "2023-01-10T03:19:15.580Z",
      "content": "<p>Does it also affect pairwise-rank model? </p>",
      "rawMarkdown": "Does it also affect pairwise-rank model? ",
      "replies": [
        {
          "id": 2093481,
          "postDate": "2023-01-10T03:24:03.543Z",
          "content": "<p>It could be. Actually it does not depend on the model itself, but how we evaluate the metrics (more precisely, how we select the tie-scored items).</p>\n<p>E.g. assume that your evaluating algorithm selects the first 2 items for each sessions, and calculate Recall@2 in the above example. You will get maximum recall@2 on the dataset.</p>\n<p>If you randomly choose N items instead of the first N items, you can also avoid the leakage.</p>",
          "rawMarkdown": "It could be. Actually it does not depend on the model itself, but how we evaluate the metrics (more precisely, how we select the tie-scored items).\n\nE.g. assume that your evaluating algorithm selects the first 2 items for each sessions, and calculate Recall@2 in the above example. You will get maximum recall@2 on the dataset.\n\nIf you randomly choose N items instead of the first N items, you can also avoid the leakage.",
          "replies": [
            {
              "id": 2093485,
              "postDate": "2023-01-10T03:33:42.527Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2094263,
      "author_name": "Senapati Rajesh",
      "author_url": "",
      "post_date": "2023-01-10T17:19:40.583000",
      "content": "<p>Thank you mate!!!But will it be working for all the models?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2094265,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2023-01-10T17:20:57.510000",
          "content": "<blockquote>\n  <p>But will it be working for all the models?</p>\n</blockquote>\n<p>Sorry, I don't understand what you mean.<br>\nHowever, this thread may be more informative:<br>\n<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/377094#2093481\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/377094#2093481</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2094240,
      "author_name": "Rayan-aay",
      "author_url": "",
      "post_date": "2023-01-10T17:09:36.497000",
      "content": "<p>Very informative and insightful. Thanks !<br>\nWe should also print the AUC score to help detect this kind of problem earlier ! </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2094286,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2023-01-10T17:30:29.537000",
          "content": "<p>Sorry, I don't understand. How AUC score help detect this kind of leakage?<br>\nI think AUC score also have problem with tie score.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2094326,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-01-10T17:44:12.863000",
              "content": "<p>Anyway, I noticed this leakage with extremely hight NDCG (~1.00) with very simple model (always predicts a constant score).</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2093474,
      "author_name": "EeyoreLee",
      "author_url": "",
      "post_date": "2023-01-10T03:19:15.580000",
      "content": "<p>Does it also affect pairwise-rank model? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2093481,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2023-01-10T03:24:03.543000",
          "content": "<p>It could be. Actually it does not depend on the model itself, but how we evaluate the metrics (more precisely, how we select the tie-scored items).</p>\n<p>E.g. assume that your evaluating algorithm selects the first 2 items for each sessions, and calculate Recall@2 in the above example. You will get maximum recall@2 on the dataset.</p>\n<p>If you randomly choose N items instead of the first N items, you can also avoid the leakage.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2093485,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-10T03:33:42.527000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2093149": "I have experienced a leakage when training ranking model: the order of dataset.\n\n## How the leakage occurs?\n\nSuppose if your dataset has pattern in your dataset (e.g. positive set appears before negative set for each session), the model get extremely high NDCG while the model outputs the constant value for any input features.\n\nE.g. assume that your evaluating algorithm selects the first 2 items for each sessions, and calculate Recall@2 in the below dataset (Table 1). You will get maximum recall@2 on the dataset even if your model always outputs a constant score (1.0).\n\nTable 1: example of leaked dataset (the positive set appears before negative set).\n\n| session | feat1 | feat2 | label | pred score |\n|:-------:|:-----:|:-----:|:-----:|------|\n| 0       | 0     | x     | 1     | 1.0  |\n| 0       | 1     | x     | 1     | 1.0  |\n| 0       | 0     | x     | 0     | 1.0  |\n| 1       | 1     | x     | 1     | 1.0  |\n| 1       | 0     | x     | 0     | 1.0  |\n| 1       | 0     | x     | 0     | 1.0  |\n\n## How to avoid the leak\n\nTo avoid these kind of leaks, \n\n1. Shuffle the dataset before training/evaluating\n2. add tiny amount of noise to the model's predicted score\n\nExample code of shuffling:\n\n```python\ndef shuffle_keep_session(df):\n    df[\"_noise\"] = np.random.randn(len(df))\n    df = df.sort_values([\"type\", \"session\", \"_noise\"])\n    df = df.drop(\"_noise\", axis=1)\n    return df\n```\n",
    "2094263": "Thank you mate!!!But will it be working for all the models?",
    "2094240": "Very informative and insightful. Thanks !\nWe should also print the AUC score to help detect this kind of problem earlier ! ",
    "2093474": "Does it also affect pairwise-rank model? "
  }
}