{
  "id": 502175,
  "title": "Rerunning previous notebooks: challenges and solutions",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/502175",
  "author_name": "",
  "post_date": "2024-05-12T12:25:15.902920900Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hey, Kagglers, just wanted to share some reflections about scores before and after metric adjustment. What I observed is significant decrease of high previous score (e.g. 0.621 to 0.568) after re-running a notebook with similar LGBM parameters which is scary to some extent. </p>\n<p>Thus, I have some questions and issues to point and discuss, your insights and reflections are important.</p>\n<ol>\n<li><p>As I understood, new metric considerably reduce scores, is it another purpose of it aside from hacking issues?</p></li>\n<li><p>So, does it mean that there is no point of re-running past notebooks with good scores anymore? </p></li>\n<li><p>Instead, is it necessary to develop and test other notebooks, e.g. with different (hyper)parameters? What do you think?</p></li>\n</ol>",
  "messages": [
    {
      "id": "2808867",
      "postDate": "05/12/2024 12:25:15",
      "content": "<p>Hey, Kagglers, just wanted to share some reflections about scores before and after metric adjustment. What I observed is significant decrease of high previous score (e.g. 0.621 to 0.568) after re-running a notebook with similar LGBM parameters which is scary to some extent. </p>\n<p>Thus, I have some questions and issues to point and discuss, your insights and reflections are important.</p>\n<ol>\n<li><p>As I understood, new metric considerably reduce scores, is it another purpose of it aside from hacking issues?</p></li>\n<li><p>So, does it mean that there is no point of re-running past notebooks with good scores anymore? </p></li>\n<li><p>Instead, is it necessary to develop and test other notebooks, e.g. with different (hyper)parameters? What do you think?</p></li>\n</ol>",
      "rawMarkdown": "Hey, Kagglers, just wanted to share some reflections about scores before and after metric adjustment. What I observed is significant decrease of high previous score (e.g. 0.621 to 0.568) after re-running a notebook with similar LGBM parameters which is scary to some extent. \n\nThus, I have some questions and issues to point and discuss, your insights and reflections are important.\n\n1. As I understood, new metric considerably reduce scores, is it another purpose of it aside from hacking issues?\n\n2. So, does it mean that there is no point of re-running past notebooks with good scores anymore? \n\n3. Instead, is it necessary to develop and test other notebooks, e.g. with different (hyper)parameters? What do you think?",
      "votes": null
    },
    {
      "id": "2810770",
      "postDate": "05/13/2024 12:49:36",
      "content": "<ol>\n<li><p>No, the metric did not change. The only change was that <code>WEEK_NUM</code>,  <code>MONTH</code>, <code>date_decision</code> and the other date columns in the <strong>test dataset</strong> have been updated such as to prevent the hack. (However, similar hacking is still possible, see the forums).</p></li>\n<li><p>If your initial notebook didn't implement the hack by artificially lowering performance in the early weeks, then I expect its performance should stay relatively the same. Have you used <code>WEEK_NUM</code> or <code>MONTH</code> explicitly in your model or to alter predictions?</p></li>\n<li><p>It is difficult to answer without knowing your implementation. You can normally ensure exact reproducibility by setting random seeds, etc. In your case the difference is too large to be explained by randomness, so I assume you must have used the date / week / month columns to alter your predictions?</p></li>\n</ol>",
      "rawMarkdown": "1. No, the metric did not change. The only change was that `WEEK_NUM`,  `MONTH`, `date_decision` and the other date columns in the **test dataset** have been updated such as to prevent the hack. (However, similar hacking is still possible, see the forums).\n\n2. If your initial notebook didn't implement the hack by artificially lowering performance in the early weeks, then I expect its performance should stay relatively the same. Have you used `WEEK_NUM` or `MONTH` explicitly in your model or to alter predictions?\n\n3. It is difficult to answer without knowing your implementation. You can normally ensure exact reproducibility by setting random seeds, etc. In your case the difference is too large to be explained by randomness, so I assume you must have used the date / week / month columns to alter your predictions?",
      "votes": null
    },
    {
      "id": "2836832",
      "postDate": "05/26/2024 06:09:52",
      "content": "<p>Thanks for answers, much appreciated. Need to revise and consider your suggestions, hope they would help to improve submission scores.</p>",
      "rawMarkdown": "Thanks for answers, much appreciated. Need to revise and consider your suggestions, hope they would help to improve submission scores.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2810770,
      "author_name": "georgeciobanu",
      "author_url": "",
      "post_date": "05/13/2024 12:49:36",
      "content": "<ol>\n<li><p>No, the metric did not change. The only change was that <code>WEEK_NUM</code>,  <code>MONTH</code>, <code>date_decision</code> and the other date columns in the <strong>test dataset</strong> have been updated such as to prevent the hack. (However, similar hacking is still possible, see the forums).</p></li>\n<li><p>If your initial notebook didn't implement the hack by artificially lowering performance in the early weeks, then I expect its performance should stay relatively the same. Have you used <code>WEEK_NUM</code> or <code>MONTH</code> explicitly in your model or to alter predictions?</p></li>\n<li><p>It is difficult to answer without knowing your implementation. You can normally ensure exact reproducibility by setting random seeds, etc. In your case the difference is too large to be explained by randomness, so I assume you must have used the date / week / month columns to alter your predictions?</p></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2836832,
          "author_name": "ablaydosmaganbetov",
          "author_url": "",
          "post_date": "05/26/2024 06:09:52",
          "content": "<p>Thanks for answers, much appreciated. Need to revise and consider your suggestions, hope they would help to improve submission scores.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2808867": "Hey, Kagglers, just wanted to share some reflections about scores before and after metric adjustment. What I observed is significant decrease of high previous score (e.g. 0.621 to 0.568) after re-running a notebook with similar LGBM parameters which is scary to some extent. \n\nThus, I have some questions and issues to point and discuss, your insights and reflections are important.\n\n1. As I understood, new metric considerably reduce scores, is it another purpose of it aside from hacking issues?\n\n2. So, does it mean that there is no point of re-running past notebooks with good scores anymore? \n\n3. Instead, is it necessary to develop and test other notebooks, e.g. with different (hyper)parameters? What do you think?",
    "2810770": "1. No, the metric did not change. The only change was that `WEEK_NUM`,  `MONTH`, `date_decision` and the other date columns in the **test dataset** have been updated such as to prevent the hack. (However, similar hacking is still possible, see the forums).\n\n2. If your initial notebook didn't implement the hack by artificially lowering performance in the early weeks, then I expect its performance should stay relatively the same. Have you used `WEEK_NUM` or `MONTH` explicitly in your model or to alter predictions?\n\n3. It is difficult to answer without knowing your implementation. You can normally ensure exact reproducibility by setting random seeds, etc. In your case the difference is too large to be explained by randomness, so I assume you must have used the date / week / month columns to alter your predictions?",
    "2836832": "Thanks for answers, much appreciated. Need to revise and consider your suggestions, hope they would help to improve submission scores."
  },
  "source": "meta"
}