{
  "id": 323081,
  "title": "nasty, silent leakage...",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/323081",
  "author_name": "",
  "post_date": "2022-05-04T18:27:14.190523700Z",
  "votes": 12,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I've had a bizarre and frustrating experience for most of this competition, where the loss would get worse with LGBMRanker with each of the first few rounds. (originally reported <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/317374\" target=\"_blank\">over here</a>)</p>\n<p>I've finally figured it out, and I'm very relieved!</p>\n<p>Wanted to share in case anyone else was having this problem.</p>\n<p>The way I was generating the labels was by merging my cudf <code>candidates_df</code> with a <code>truth_df</code> (containing all the transactions that actually happened in the target week) that had columns: <code>[\"customer_id\", \"article_id\", \"match\"]</code>.</p>\n<p>The end result was an additional column that was 1 (from the merge) if the customer/article had occurred, or NA if it hadn't.</p>\n<p>What I didn't realize is that cudf was putting the rows that had gotten merged before the non-merged rows - so that the positive examples were before the negative examples!<br>\nLater, when I would take the first 12 items for each customer, the positive examples would be taken first.</p>\n<p>By training a LGBMRanker, and using its ranking scores to sort the values, I was losing this leakage, hence the score went down.</p>",
  "messages": [
    {
      "id": "1777713",
      "postDate": "05/04/2022 18:27:14",
      "content": "<p>I've had a bizarre and frustrating experience for most of this competition, where the loss would get worse with LGBMRanker with each of the first few rounds. (originally reported <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/317374\" target=\"_blank\">over here</a>)</p>\n<p>I've finally figured it out, and I'm very relieved!</p>\n<p>Wanted to share in case anyone else was having this problem.</p>\n<p>The way I was generating the labels was by merging my cudf <code>candidates_df</code> with a <code>truth_df</code> (containing all the transactions that actually happened in the target week) that had columns: <code>[\"customer_id\", \"article_id\", \"match\"]</code>.</p>\n<p>The end result was an additional column that was 1 (from the merge) if the customer/article had occurred, or NA if it hadn't.</p>\n<p>What I didn't realize is that cudf was putting the rows that had gotten merged before the non-merged rows - so that the positive examples were before the negative examples!<br>\nLater, when I would take the first 12 items for each customer, the positive examples would be taken first.</p>\n<p>By training a LGBMRanker, and using its ranking scores to sort the values, I was losing this leakage, hence the score went down.</p>",
      "rawMarkdown": "I've had a bizarre and frustrating experience for most of this competition, where the loss would get worse with LGBMRanker with each of the first few rounds. (originally reported [over here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/317374))\n\nI've finally figured it out, and I'm very relieved!\n\nWanted to share in case anyone else was having this problem.\n\nThe way I was generating the labels was by merging my cudf `candidates_df` with a `truth_df` (containing all the transactions that actually happened in the target week) that had columns: `[\"customer_id\", \"article_id\", \"match\"]`.\n\nThe end result was an additional column that was 1 (from the merge) if the customer/article had occurred, or NA if it hadn't.\n\nWhat I didn't realize is that cudf was putting the rows that had gotten merged before the non-merged rows - so that the positive examples were before the negative examples!\nLater, when I would take the first 12 items for each customer, the positive examples would be taken first.\n\nBy training a LGBMRanker, and using its ranking scores to sort the values, I was losing this leakage, hence the score went down.",
      "votes": null
    },
    {
      "id": "1777800",
      "postDate": "05/04/2022 19:41:03",
      "content": "<p>Bugs happen. A week ago it turned out that I was training the models using next 6 days of transactions instead of 7. Typical \"off by one\" error. Luckily my teammate spotted something wrong with my validation table. Funnily enough there was almost no difference between using 6 or 7 days.</p>",
      "rawMarkdown": "Bugs happen. A week ago it turned out that I was training the models using next 6 days of transactions instead of 7. Typical \"off by one\" error. Luckily my teammate spotted something wrong with my validation table. Funnily enough there was almost no difference between using 6 or 7 days.",
      "votes": null
    },
    {
      "id": "1777834",
      "postDate": "05/04/2022 20:14:58",
      "content": "<p>A good practice for cudf is to call sort_values after merge functions. It happens because cudf process everything in parallel and after a dataframe operation it can return rows out of order.</p>",
      "rawMarkdown": "A good practice for cudf is to call sort_values after merge functions. It happens because cudf process everything in parallel and after a dataframe operation it can return rows out of order.",
      "votes": null
    },
    {
      "id": "1777958",
      "postDate": "05/05/2022 00:38:06",
      "content": "<p>Congrats, hard work pays off finally. I also battle with different kinds of bug all the way. It sucks but finally after the bugs are solved, I kind of enjoy the process. Good luck!</p>",
      "rawMarkdown": "Congrats, hard work pays off finally. I also battle with different kinds of bug all the way. It sucks but finally after the bugs are solved, I kind of enjoy the process. Good luck!",
      "votes": null
    },
    {
      "id": "1777964",
      "postDate": "05/05/2022 00:51:16",
      "content": "<p>Thanks, I'll make that a standard practice.</p>",
      "rawMarkdown": "Thanks, I'll make that a standard practice.",
      "votes": null
    },
    {
      "id": "1777965",
      "postDate": "05/05/2022 00:51:49",
      "content": "<p>Thanks!<br>\nCan't say I enjoyed the process, but definitely enjoy the relief…</p>",
      "rawMarkdown": "Thanks!\nCan't say I enjoyed the process, but definitely enjoy the relief...",
      "votes": null
    },
    {
      "id": "1777974",
      "postDate": "05/05/2022 01:06:43",
      "content": "<p>Congratulations👍🏻Every one who has been obsessed with bugs knows the relieve.<br>\nThere is a point I don't understand well. What's the effect of the order of pos and neg on the training? Wrong groups when calculating listwise loss?</p>",
      "rawMarkdown": "Congratulations👍🏻Every one who has been obsessed with bugs knows the relieve.\nThere is a point I don't understand well. What's the effect of the order of pos and neg on the training? Wrong groups when calculating listwise loss?",
      "votes": null
    },
    {
      "id": "1777988",
      "postDate": "05/05/2022 01:54:43",
      "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> Interesting! Thanks for sharing :)</p>",
      "rawMarkdown": "jacob34 Interesting! Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "1777996",
      "postDate": "05/05/2022 02:20:07",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a>!</p>\n<p>In the early stages of training, many of the candidates have the same score, because they end up in the same leaves of the LGBMRanker model.<br>\nHow does LGBM decide how to order candidates with equal scores?<br>\nSeems like they just keep the order within the original DataFrame!<br>\nSince the positives were earlier in the DataFrame, the score in the early stages of training was very high.</p>\n<p>(If I set the num_leaves to 2 (just a single split), then after the first estimator ran, I had a huge score on the cv).</p>\n<p>As training progressed, less and less candidates had the same exact score, so the original ordering in the DataFrame was used less and less, so the score kept going down.</p>\n<p>So it wasn't that training wasn't working, it's just that the metrics being reported as training progressed were polluted by the leakage.</p>",
      "rawMarkdown": "Thanks, @sirius81!\n\nIn the early stages of training, many of the candidates have the same score, because they end up in the same leaves of the LGBMRanker model.\nHow does LGBM decide how to order candidates with equal scores?\nSeems like they just keep the order within the original DataFrame!\nSince the positives were earlier in the DataFrame, the score in the early stages of training was very high.\n\n(If I set the num_leaves to 2 (just a single split), then after the first estimator ran, I had a huge score on the cv).\n\nAs training progressed, less and less candidates had the same exact score, so the original ordering in the DataFrame was used less and less, so the score kept going down.\n\nSo it wasn't that training wasn't working, it's just that the metrics being reported as training progressed were polluted by the leakage.",
      "votes": null
    },
    {
      "id": "1778000",
      "postDate": "05/05/2022 02:30:12",
      "content": "<p>Thanks for sharing. Interesting! </p>",
      "rawMarkdown": "Thanks for sharing. Interesting!",
      "votes": null
    },
    {
      "id": "1778088",
      "postDate": "05/05/2022 05:25:30",
      "content": "<p>You kind of found a new way to do ensembling.😂</p>",
      "rawMarkdown": "You kind of found a new way to do ensembling.😂",
      "votes": null
    },
    {
      "id": "1778327",
      "postDate": "05/05/2022 08:36:09",
      "content": "<p>Yeah. It may be that longer periods lead to less random and sparse target which can actually be positive for the results.</p>",
      "rawMarkdown": "Yeah. It may be that longer periods lead to less random and sparse target which can actually be positive for the results.",
      "votes": null
    },
    {
      "id": "1778509",
      "postDate": "05/05/2022 11:44:16",
      "content": "<p>This is a very helpful comment! Thank you for taking the time to share this.</p>",
      "rawMarkdown": "This is a very helpful comment! Thank you for taking the time to share this.",
      "votes": null
    },
    {
      "id": "1778632",
      "postDate": "05/05/2022 14:30:06",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1777800,
      "author_name": "paweljankiewicz",
      "author_url": "",
      "post_date": "05/04/2022 19:41:03",
      "content": "<p>Bugs happen. A week ago it turned out that I was training the models using next 6 days of transactions instead of 7. Typical \"off by one\" error. Luckily my teammate spotted something wrong with my validation table. Funnily enough there was almost no difference between using 6 or 7 days.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1778088,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "05/05/2022 05:25:30",
          "content": "<p>You kind of found a new way to do ensembling.😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1778327,
          "author_name": "paweljankiewicz",
          "author_url": "",
          "post_date": "05/05/2022 08:36:09",
          "content": "<p>Yeah. It may be that longer periods lead to less random and sparse target which can actually be positive for the results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1777834,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "05/04/2022 20:14:58",
      "content": "<p>A good practice for cudf is to call sort_values after merge functions. It happens because cudf process everything in parallel and after a dataframe operation it can return rows out of order.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1777964,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/05/2022 00:51:16",
          "content": "<p>Thanks, I'll make that a standard practice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1777958,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "05/05/2022 00:38:06",
      "content": "<p>Congrats, hard work pays off finally. I also battle with different kinds of bug all the way. It sucks but finally after the bugs are solved, I kind of enjoy the process. Good luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1777965,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/05/2022 00:51:49",
          "content": "<p>Thanks!<br>\nCan't say I enjoyed the process, but definitely enjoy the relief…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1777974,
      "author_name": "sirius81",
      "author_url": "",
      "post_date": "05/05/2022 01:06:43",
      "content": "<p>Congratulations👍🏻Every one who has been obsessed with bugs knows the relieve.<br>\nThere is a point I don't understand well. What's the effect of the order of pos and neg on the training? Wrong groups when calculating listwise loss?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1777996,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "05/05/2022 02:20:07",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a>!</p>\n<p>In the early stages of training, many of the candidates have the same score, because they end up in the same leaves of the LGBMRanker model.<br>\nHow does LGBM decide how to order candidates with equal scores?<br>\nSeems like they just keep the order within the original DataFrame!<br>\nSince the positives were earlier in the DataFrame, the score in the early stages of training was very high.</p>\n<p>(If I set the num_leaves to 2 (just a single split), then after the first estimator ran, I had a huge score on the cv).</p>\n<p>As training progressed, less and less candidates had the same exact score, so the original ordering in the DataFrame was used less and less, so the score kept going down.</p>\n<p>So it wasn't that training wasn't working, it's just that the metrics being reported as training progressed were polluted by the leakage.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1777988,
      "author_name": "lachlangillian",
      "author_url": "",
      "post_date": "05/05/2022 01:54:43",
      "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> Interesting! Thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1778000,
      "author_name": "biubiug",
      "author_url": "",
      "post_date": "05/05/2022 02:30:12",
      "content": "<p>Thanks for sharing. Interesting! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1778509,
      "author_name": "",
      "author_url": "",
      "post_date": "05/05/2022 11:44:16",
      "content": "<p>This is a very helpful comment! Thank you for taking the time to share this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1778632,
      "author_name": "richardsong0268",
      "author_url": "",
      "post_date": "05/05/2022 14:30:06",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1777713": "I've had a bizarre and frustrating experience for most of this competition, where the loss would get worse with LGBMRanker with each of the first few rounds. (originally reported [over here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/317374))\n\nI've finally figured it out, and I'm very relieved!\n\nWanted to share in case anyone else was having this problem.\n\nThe way I was generating the labels was by merging my cudf `candidates_df` with a `truth_df` (containing all the transactions that actually happened in the target week) that had columns: `[\"customer_id\", \"article_id\", \"match\"]`.\n\nThe end result was an additional column that was 1 (from the merge) if the customer/article had occurred, or NA if it hadn't.\n\nWhat I didn't realize is that cudf was putting the rows that had gotten merged before the non-merged rows - so that the positive examples were before the negative examples!\nLater, when I would take the first 12 items for each customer, the positive examples would be taken first.\n\nBy training a LGBMRanker, and using its ranking scores to sort the values, I was losing this leakage, hence the score went down.",
    "1777800": "Bugs happen. A week ago it turned out that I was training the models using next 6 days of transactions instead of 7. Typical \"off by one\" error. Luckily my teammate spotted something wrong with my validation table. Funnily enough there was almost no difference between using 6 or 7 days.",
    "1777834": "A good practice for cudf is to call sort_values after merge functions. It happens because cudf process everything in parallel and after a dataframe operation it can return rows out of order.",
    "1777958": "Congrats, hard work pays off finally. I also battle with different kinds of bug all the way. It sucks but finally after the bugs are solved, I kind of enjoy the process. Good luck!",
    "1777964": "Thanks, I'll make that a standard practice.",
    "1777965": "Thanks!\nCan't say I enjoyed the process, but definitely enjoy the relief...",
    "1777974": "Congratulations👍🏻Every one who has been obsessed with bugs knows the relieve.\nThere is a point I don't understand well. What's the effect of the order of pos and neg on the training? Wrong groups when calculating listwise loss?",
    "1777988": "jacob34 Interesting! Thanks for sharing :)",
    "1777996": "Thanks, @sirius81!\n\nIn the early stages of training, many of the candidates have the same score, because they end up in the same leaves of the LGBMRanker model.\nHow does LGBM decide how to order candidates with equal scores?\nSeems like they just keep the order within the original DataFrame!\nSince the positives were earlier in the DataFrame, the score in the early stages of training was very high.\n\n(If I set the num_leaves to 2 (just a single split), then after the first estimator ran, I had a huge score on the cv).\n\nAs training progressed, less and less candidates had the same exact score, so the original ordering in the DataFrame was used less and less, so the score kept going down.\n\nSo it wasn't that training wasn't working, it's just that the metrics being reported as training progressed were polluted by the leakage.",
    "1778000": "Thanks for sharing. Interesting!",
    "1778088": "You kind of found a new way to do ensembling.😂",
    "1778327": "Yeah. It may be that longer periods lead to less random and sparse target which can actually be positive for the results.",
    "1778509": "This is a very helpful comment! Thank you for taking the time to share this.",
    "1778632": "Thanks for sharing"
  },
  "source": "meta"
}