{
  "id": 506096,
  "title": "After restoring WEEK_NUM in test dataset, what's next?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/506096",
  "author_name": "",
  "post_date": "2024-05-20T13:32:33.777431700Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>If we succeed in restoring WEEK_NUM in test dataset, how would we improve your score?</p>\n<p>After seeing several discussions and notebooks, the most popular method is to divide the data into the first and second halves by WEEK_NUM, and then subtract a certain value from the predicted value in the first half to improve the apparent \"stability\".<br>\n(<a href=\"https://www.kaggle.com/code/kononenko/metric-trick-home-credit-baseline-inference\" target=\"_blank\">https://www.kaggle.com/code/kononenko/metric-trick-home-credit-baseline-inference</a>)</p>\n<p>Generalizing this approach gives rise to two parameters as follows:<br>\n`DEGRADE_RATIO = 0.5<br>\nREDUCED_SCORE = 0.02</p>\n<p>thresh = (1.0 - DEGRADE_RATIO) * WEEK_NUM.max() + DEGRADE_RATIO * WEEK_NUM.min()<br>\ncondition = list(WEEK_NUM &lt; thresh)<br>\ndf_subm.loc[condition, 'score'] = (df_subm.loc[condition, 'score'] - REDUCED_SCORE).clip(0)`<br>\nNamely, DEGRADE_RATIO for how many predictions to degrade and REDUCED_SCORE for the extent of degradation.</p>\n<p>Then, the question is how can we optimize? Based on what logic? In any case, I think we can't determine the best parameters due to the lack of information about the private dataset distribution along WEEK_NUM and the stabilities of our original model.</p>\n<p>The only thing I could come up with is to pick a parameter that maximizes the public score or roll the dice.</p>\n<p>Share your ideas about how to select parameters or different approaches!</p>",
  "messages": [
    {
      "id": "2825576",
      "postDate": "05/20/2024 13:32:33",
      "content": "<p>If we succeed in restoring WEEK_NUM in test dataset, how would we improve your score?</p>\n<p>After seeing several discussions and notebooks, the most popular method is to divide the data into the first and second halves by WEEK_NUM, and then subtract a certain value from the predicted value in the first half to improve the apparent \"stability\".<br>\n(<a href=\"https://www.kaggle.com/code/kononenko/metric-trick-home-credit-baseline-inference\" target=\"_blank\">https://www.kaggle.com/code/kononenko/metric-trick-home-credit-baseline-inference</a>)</p>\n<p>Generalizing this approach gives rise to two parameters as follows:<br>\n`DEGRADE_RATIO = 0.5<br>\nREDUCED_SCORE = 0.02</p>\n<p>thresh = (1.0 - DEGRADE_RATIO) * WEEK_NUM.max() + DEGRADE_RATIO * WEEK_NUM.min()<br>\ncondition = list(WEEK_NUM &lt; thresh)<br>\ndf_subm.loc[condition, 'score'] = (df_subm.loc[condition, 'score'] - REDUCED_SCORE).clip(0)`<br>\nNamely, DEGRADE_RATIO for how many predictions to degrade and REDUCED_SCORE for the extent of degradation.</p>\n<p>Then, the question is how can we optimize? Based on what logic? In any case, I think we can't determine the best parameters due to the lack of information about the private dataset distribution along WEEK_NUM and the stabilities of our original model.</p>\n<p>The only thing I could come up with is to pick a parameter that maximizes the public score or roll the dice.</p>\n<p>Share your ideas about how to select parameters or different approaches!</p>",
      "rawMarkdown": "If we succeed in restoring WEEK_NUM in test dataset, how would we improve your score?\n\nAfter seeing several discussions and notebooks, the most popular method is to divide the data into the first and second halves by WEEK_NUM, and then subtract a certain value from the predicted value in the first half to improve the apparent \"stability\".\n(https://www.kaggle.com/code/kononenko/metric-trick-home-credit-baseline-inference)\n\nGeneralizing this approach gives rise to two parameters as follows:\n`DEGRADE_RATIO = 0.5\nREDUCED_SCORE = 0.02\n\nthresh = (1.0 - DEGRADE_RATIO) * WEEK_NUM.max() + DEGRADE_RATIO * WEEK_NUM.min()\ncondition = list(WEEK_NUM < thresh)\ndf_subm.loc[condition, 'score'] = (df_subm.loc[condition, 'score'] - REDUCED_SCORE).clip(0)`\nNamely, DEGRADE_RATIO for how many predictions to degrade and REDUCED_SCORE for the extent of degradation.\n\nThen, the question is how can we optimize? Based on what logic? In any case, I think we can't determine the best parameters due to the lack of information about the private dataset distribution along WEEK_NUM and the stabilities of our original model.\n\nThe only thing I could come up with is to pick a parameter that maximizes the public score or roll the dice.\n\nShare your ideas about how to select parameters or different approaches!",
      "votes": null
    },
    {
      "id": "2827392",
      "postDate": "05/21/2024 12:42:45",
      "content": "<p>The adjustment in high scoring notebooks is based on probability rather than WEEK_NUM, that would be the second approach. The first approach as you mentioned would be to reverse engineer WEEK_NUM</p>",
      "rawMarkdown": "The adjustment in high scoring notebooks is based on probability rather than WEEK_NUM, that would be the second approach. The first approach as you mentioned would be to reverse engineer WEEK_NUM",
      "votes": null
    },
    {
      "id": "2827444",
      "postDate": "05/21/2024 13:32:35",
      "content": "<p>Thanks for your comment. I am aware that in very high scoring public notebooks, something like similarity between test and training data is predicted as a probability and used as a threshold. Add that as an option as well. </p>\n<p>Again, there are two parameters: the threshold and the score reduction. It looks like we are over-fitting them to the public score in many public notebooks, but is there a strategy to improve the private score?</p>",
      "rawMarkdown": "Thanks for your comment. I am aware that in very high scoring public notebooks, something like similarity between test and training data is predicted as a probability and used as a threshold. Add that as an option as well. \n\nAgain, there are two parameters: the threshold and the score reduction. It looks like we are over-fitting them to the public score in many public notebooks, but is there a strategy to improve the private score?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2827392,
      "author_name": "fadylabib",
      "author_url": "",
      "post_date": "05/21/2024 12:42:45",
      "content": "<p>The adjustment in high scoring notebooks is based on probability rather than WEEK_NUM, that would be the second approach. The first approach as you mentioned would be to reverse engineer WEEK_NUM</p>",
      "votes": null,
      "replies": [
        {
          "id": 2827444,
          "author_name": "atsuno",
          "author_url": "",
          "post_date": "05/21/2024 13:32:35",
          "content": "<p>Thanks for your comment. I am aware that in very high scoring public notebooks, something like similarity between test and training data is predicted as a probability and used as a threshold. Add that as an option as well. </p>\n<p>Again, there are two parameters: the threshold and the score reduction. It looks like we are over-fitting them to the public score in many public notebooks, but is there a strategy to improve the private score?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2825576": "If we succeed in restoring WEEK_NUM in test dataset, how would we improve your score?\n\nAfter seeing several discussions and notebooks, the most popular method is to divide the data into the first and second halves by WEEK_NUM, and then subtract a certain value from the predicted value in the first half to improve the apparent \"stability\".\n(https://www.kaggle.com/code/kononenko/metric-trick-home-credit-baseline-inference)\n\nGeneralizing this approach gives rise to two parameters as follows:\n`DEGRADE_RATIO = 0.5\nREDUCED_SCORE = 0.02\n\nthresh = (1.0 - DEGRADE_RATIO) * WEEK_NUM.max() + DEGRADE_RATIO * WEEK_NUM.min()\ncondition = list(WEEK_NUM < thresh)\ndf_subm.loc[condition, 'score'] = (df_subm.loc[condition, 'score'] - REDUCED_SCORE).clip(0)`\nNamely, DEGRADE_RATIO for how many predictions to degrade and REDUCED_SCORE for the extent of degradation.\n\nThen, the question is how can we optimize? Based on what logic? In any case, I think we can't determine the best parameters due to the lack of information about the private dataset distribution along WEEK_NUM and the stabilities of our original model.\n\nThe only thing I could come up with is to pick a parameter that maximizes the public score or roll the dice.\n\nShare your ideas about how to select parameters or different approaches!",
    "2827392": "The adjustment in high scoring notebooks is based on probability rather than WEEK_NUM, that would be the second approach. The first approach as you mentioned would be to reverse engineer WEEK_NUM",
    "2827444": "Thanks for your comment. I am aware that in very high scoring public notebooks, something like similarity between test and training data is predicted as a probability and used as a threshold. Add that as an option as well. \n\nAgain, there are two parameters: the threshold and the score reduction. It looks like we are over-fitting them to the public score in many public notebooks, but is there a strategy to improve the private score?"
  },
  "source": "meta"
}