{
  "id": 501817,
  "title": "Is every one who is hacking the metric, overfitting the PLB?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/501817",
  "author_name": "",
  "post_date": "2024-05-10T21:11:08.485082300Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hypothesis: Making the first half less accurate should result in better Gini stability. (empirically)<br>\nAlthough I have figured out the hack but still, I am still hesitant to use it. Why?</p>\n<p>I trained my model on the first 6 months of training data and predicted the rest ~1.5 years in training.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5053803%2F8b6bf1efacf608fde8a51bdea4d355e2%2FScreenshot%202024-05-11%20at%202.31.56AM.png?generation=1715374942790453&amp;alt=media\"></p>\n<p>The above gini stability is without any post-processing.<br>\nAnd this is how the predictions look.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5053803%2F0be54439d57a9d0e01b5f8a0bea24c7e%2FScreenshot%202024-05-11%20at%202.33.22AM.png?generation=1715375025196380&amp;alt=media\"></p>\n<p>Here is the list of ginis if anyone wants to experiment.<br>\n<code>[0.6467616719770706, 0.670325950155453, 0.6703447968128613, 0.6287240612872953, 0.6625411078038856, 0.6634128405368274, 0.6441799294415731, 0.62756413608631, 0.6525058523826972, 0.6490252592098806, 0.6532703905499848, 0.6243405394548445, 0.6358132715062748, 0.6528721569237905, 0.6281497317245581, 0.6256488053374469, 0.6131732699544612, 0.6604694044934822, 0.6045724117720108, 0.5990421314950765, 0.6671857361715852, 0.6067641415159302, 0.6372779483171964, 0.6030645393301526, 0.651696900873965, 0.643071954981006, 0.6594469754467351, 0.6006235455877724, 0.6128542488800133, 0.6331346907274227, 0.6239385392443748, 0.6416887522296484, 0.6233036265359944, 0.6052406258624368, 0.5881237666455368, 0.6143955447365193, 0.6649877271853835, 0.6210255878639059, 0.6487165625500562, 0.5271787960467207, 0.6665716999050331, 0.5793650793650793, 0.6193923723335488, 0.7435293193511645, 0.719949820998885, 0.6615118634186548, 0.6242028936780644, 0.6610524676408609, 0.6838568782345322, 0.6709288152598636, 0.6484316487419803, 0.713358495087467, 0.7126790668705056, 0.6708673413864403, 0.6400214226593313, 0.5916699911244709, 0.6766210163744832, 0.6473571414655055, 0.7063753834515245, 0.7638571090305379, 0.7043342803717219, 0.6981780411342098, 0.6908719163927992, 0.6694295370248824, 0.7311648904677324]</code></p>\n<p>Now if I manipulate the ginis:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5053803%2F95754b8afdb2ba3c519777323c1637e8%2FScreenshot%202024-05-11%20at%202.35.32AM.png?generation=1715375151504939&amp;alt=media\"></p>\n<p>The local score drops. I am still trying to understand more of this scenario, but just thought of sharing this. Also, what is interesting is that my model has better ginis towards the end of training dates and the public leaderboard data might be chosen such that the weeks towards the end have very bad data and thus bad predictions but this might not translate well in private.</p>",
  "messages": [
    {
      "id": "2806072",
      "postDate": "05/10/2024 21:11:08",
      "content": "<p>Hypothesis: Making the first half less accurate should result in better Gini stability. (empirically)<br>\nAlthough I have figured out the hack but still, I am still hesitant to use it. Why?</p>\n<p>I trained my model on the first 6 months of training data and predicted the rest ~1.5 years in training.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5053803%2F8b6bf1efacf608fde8a51bdea4d355e2%2FScreenshot%202024-05-11%20at%202.31.56AM.png?generation=1715374942790453&amp;alt=media\"></p>\n<p>The above gini stability is without any post-processing.<br>\nAnd this is how the predictions look.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5053803%2F0be54439d57a9d0e01b5f8a0bea24c7e%2FScreenshot%202024-05-11%20at%202.33.22AM.png?generation=1715375025196380&amp;alt=media\"></p>\n<p>Here is the list of ginis if anyone wants to experiment.<br>\n<code>[0.6467616719770706, 0.670325950155453, 0.6703447968128613, 0.6287240612872953, 0.6625411078038856, 0.6634128405368274, 0.6441799294415731, 0.62756413608631, 0.6525058523826972, 0.6490252592098806, 0.6532703905499848, 0.6243405394548445, 0.6358132715062748, 0.6528721569237905, 0.6281497317245581, 0.6256488053374469, 0.6131732699544612, 0.6604694044934822, 0.6045724117720108, 0.5990421314950765, 0.6671857361715852, 0.6067641415159302, 0.6372779483171964, 0.6030645393301526, 0.651696900873965, 0.643071954981006, 0.6594469754467351, 0.6006235455877724, 0.6128542488800133, 0.6331346907274227, 0.6239385392443748, 0.6416887522296484, 0.6233036265359944, 0.6052406258624368, 0.5881237666455368, 0.6143955447365193, 0.6649877271853835, 0.6210255878639059, 0.6487165625500562, 0.5271787960467207, 0.6665716999050331, 0.5793650793650793, 0.6193923723335488, 0.7435293193511645, 0.719949820998885, 0.6615118634186548, 0.6242028936780644, 0.6610524676408609, 0.6838568782345322, 0.6709288152598636, 0.6484316487419803, 0.713358495087467, 0.7126790668705056, 0.6708673413864403, 0.6400214226593313, 0.5916699911244709, 0.6766210163744832, 0.6473571414655055, 0.7063753834515245, 0.7638571090305379, 0.7043342803717219, 0.6981780411342098, 0.6908719163927992, 0.6694295370248824, 0.7311648904677324]</code></p>\n<p>Now if I manipulate the ginis:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5053803%2F95754b8afdb2ba3c519777323c1637e8%2FScreenshot%202024-05-11%20at%202.35.32AM.png?generation=1715375151504939&amp;alt=media\"></p>\n<p>The local score drops. I am still trying to understand more of this scenario, but just thought of sharing this. Also, what is interesting is that my model has better ginis towards the end of training dates and the public leaderboard data might be chosen such that the weeks towards the end have very bad data and thus bad predictions but this might not translate well in private.</p>",
      "rawMarkdown": "Hypothesis: Making the first half less accurate should result in better Gini stability. (empirically)\nAlthough I have figured out the hack but still, I am still hesitant to use it. Why?\n\nI trained my model on the first 6 months of training data and predicted the rest ~1.5 years in training.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5053803%2F8b6bf1efacf608fde8a51bdea4d355e2%2FScreenshot%202024-05-11%20at%202.31.56AM.png?generation=1715374942790453&alt=media)\n\nThe above gini stability is without any post-processing.\nAnd this is how the predictions look.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5053803%2F0be54439d57a9d0e01b5f8a0bea24c7e%2FScreenshot%202024-05-11%20at%202.33.22AM.png?generation=1715375025196380&alt=media)\n\nHere is the list of ginis if anyone wants to experiment.\n`[0.6467616719770706, 0.670325950155453, 0.6703447968128613, 0.6287240612872953, 0.6625411078038856, 0.6634128405368274, 0.6441799294415731, 0.62756413608631, 0.6525058523826972, 0.6490252592098806, 0.6532703905499848, 0.6243405394548445, 0.6358132715062748, 0.6528721569237905, 0.6281497317245581, 0.6256488053374469, 0.6131732699544612, 0.6604694044934822, 0.6045724117720108, 0.5990421314950765, 0.6671857361715852, 0.6067641415159302, 0.6372779483171964, 0.6030645393301526, 0.651696900873965, 0.643071954981006, 0.6594469754467351, 0.6006235455877724, 0.6128542488800133, 0.6331346907274227, 0.6239385392443748, 0.6416887522296484, 0.6233036265359944, 0.6052406258624368, 0.5881237666455368, 0.6143955447365193, 0.6649877271853835, 0.6210255878639059, 0.6487165625500562, 0.5271787960467207, 0.6665716999050331, 0.5793650793650793, 0.6193923723335488, 0.7435293193511645, 0.719949820998885, 0.6615118634186548, 0.6242028936780644, 0.6610524676408609, 0.6838568782345322, 0.6709288152598636, 0.6484316487419803, 0.713358495087467, 0.7126790668705056, 0.6708673413864403, 0.6400214226593313, 0.5916699911244709, 0.6766210163744832, 0.6473571414655055, 0.7063753834515245, 0.7638571090305379, 0.7043342803717219, 0.6981780411342098, 0.6908719163927992, 0.6694295370248824, 0.7311648904677324]`\n\nNow if I manipulate the ginis:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5053803%2F95754b8afdb2ba3c519777323c1637e8%2FScreenshot%202024-05-11%20at%202.35.32AM.png?generation=1715375151504939&alt=media)\n\nThe local score drops. I am still trying to understand more of this scenario, but just thought of sharing this. Also, what is interesting is that my model has better ginis towards the end of training dates and the public leaderboard data might be chosen such that the weeks towards the end have very bad data and thus bad predictions but this might not translate well in private.",
      "votes": null
    },
    {
      "id": "2806466",
      "postDate": "05/11/2024 05:41:32",
      "content": "<p>The issue here is that your model is already performing well on the training dataset, and because of that, it already has a positive slope. So, using the metric hack on the training data won't necessarily improve performance because it may just worsen the average gini. This is not the case when predicting on the test dataset. </p>",
      "rawMarkdown": "The issue here is that your model is already performing well on the training dataset, and because of that, it already has a positive slope. So, using the metric hack on the training data won't necessarily improve performance because it may just worsen the average gini. This is not the case when predicting on the test dataset.",
      "votes": null
    },
    {
      "id": "2808143",
      "postDate": "05/12/2024 05:02:07",
      "content": "<p>Optimize false positives !! </p>",
      "rawMarkdown": "Optimize false positives !!",
      "votes": null
    },
    {
      "id": "2811903",
      "postDate": "05/14/2024 01:50:30",
      "content": "<p>why do you clip the predictions? arent they already in [0,1] range?</p>",
      "rawMarkdown": "why do you clip the predictions? arent they already in [0,1] range?",
      "votes": null
    },
    {
      "id": "2812532",
      "postDate": "05/14/2024 09:15:53",
      "content": "<p>By reducing score the probability could go below 0. So clipping is necessary</p>",
      "rawMarkdown": "By reducing score the probability could go below 0. So clipping is necessary",
      "votes": null
    },
    {
      "id": "2812757",
      "postDate": "05/14/2024 11:59:28",
      "content": "<p>yes. exactly</p>",
      "rawMarkdown": "yes. exactly",
      "votes": null
    },
    {
      "id": "2812761",
      "postDate": "05/14/2024 12:05:08",
      "content": "<p>Yes, I agree. But playing with the unknown here if you have built a stable model and not knowing the quality of private test data towards the end can be dangerous. But still the counter is that you can choose two submissions :)</p>",
      "rawMarkdown": "Yes, I agree. But playing with the unknown here if you have built a stable model and not knowing the quality of private test data towards the end can be dangerous. But still the counter is that you can choose two submissions :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2806466,
      "author_name": "mehdi0694",
      "author_url": "",
      "post_date": "05/11/2024 05:41:32",
      "content": "<p>The issue here is that your model is already performing well on the training dataset, and because of that, it already has a positive slope. So, using the metric hack on the training data won't necessarily improve performance because it may just worsen the average gini. This is not the case when predicting on the test dataset. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2812761,
          "author_name": "krishnapriya18",
          "author_url": "",
          "post_date": "05/14/2024 12:05:08",
          "content": "<p>Yes, I agree. But playing with the unknown here if you have built a stable model and not knowing the quality of private test data towards the end can be dangerous. But still the counter is that you can choose two submissions :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2808143,
      "author_name": "kunduruanil",
      "author_url": "",
      "post_date": "05/12/2024 05:02:07",
      "content": "<p>Optimize false positives !! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2811903,
      "author_name": "pustoi",
      "author_url": "",
      "post_date": "05/14/2024 01:50:30",
      "content": "<p>why do you clip the predictions? arent they already in [0,1] range?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2812532,
          "author_name": "lucamtb",
          "author_url": "",
          "post_date": "05/14/2024 09:15:53",
          "content": "<p>By reducing score the probability could go below 0. So clipping is necessary</p>",
          "votes": null,
          "replies": [
            {
              "id": 2812757,
              "author_name": "krishnapriya18",
              "author_url": "",
              "post_date": "05/14/2024 11:59:28",
              "content": "<p>yes. exactly</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2806072": "Hypothesis: Making the first half less accurate should result in better Gini stability. (empirically)\nAlthough I have figured out the hack but still, I am still hesitant to use it. Why?\n\nI trained my model on the first 6 months of training data and predicted the rest ~1.5 years in training.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5053803%2F8b6bf1efacf608fde8a51bdea4d355e2%2FScreenshot%202024-05-11%20at%202.31.56AM.png?generation=1715374942790453&alt=media)\n\nThe above gini stability is without any post-processing.\nAnd this is how the predictions look.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5053803%2F0be54439d57a9d0e01b5f8a0bea24c7e%2FScreenshot%202024-05-11%20at%202.33.22AM.png?generation=1715375025196380&alt=media)\n\nHere is the list of ginis if anyone wants to experiment.\n`[0.6467616719770706, 0.670325950155453, 0.6703447968128613, 0.6287240612872953, 0.6625411078038856, 0.6634128405368274, 0.6441799294415731, 0.62756413608631, 0.6525058523826972, 0.6490252592098806, 0.6532703905499848, 0.6243405394548445, 0.6358132715062748, 0.6528721569237905, 0.6281497317245581, 0.6256488053374469, 0.6131732699544612, 0.6604694044934822, 0.6045724117720108, 0.5990421314950765, 0.6671857361715852, 0.6067641415159302, 0.6372779483171964, 0.6030645393301526, 0.651696900873965, 0.643071954981006, 0.6594469754467351, 0.6006235455877724, 0.6128542488800133, 0.6331346907274227, 0.6239385392443748, 0.6416887522296484, 0.6233036265359944, 0.6052406258624368, 0.5881237666455368, 0.6143955447365193, 0.6649877271853835, 0.6210255878639059, 0.6487165625500562, 0.5271787960467207, 0.6665716999050331, 0.5793650793650793, 0.6193923723335488, 0.7435293193511645, 0.719949820998885, 0.6615118634186548, 0.6242028936780644, 0.6610524676408609, 0.6838568782345322, 0.6709288152598636, 0.6484316487419803, 0.713358495087467, 0.7126790668705056, 0.6708673413864403, 0.6400214226593313, 0.5916699911244709, 0.6766210163744832, 0.6473571414655055, 0.7063753834515245, 0.7638571090305379, 0.7043342803717219, 0.6981780411342098, 0.6908719163927992, 0.6694295370248824, 0.7311648904677324]`\n\nNow if I manipulate the ginis:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5053803%2F95754b8afdb2ba3c519777323c1637e8%2FScreenshot%202024-05-11%20at%202.35.32AM.png?generation=1715375151504939&alt=media)\n\nThe local score drops. I am still trying to understand more of this scenario, but just thought of sharing this. Also, what is interesting is that my model has better ginis towards the end of training dates and the public leaderboard data might be chosen such that the weeks towards the end have very bad data and thus bad predictions but this might not translate well in private.",
    "2806466": "The issue here is that your model is already performing well on the training dataset, and because of that, it already has a positive slope. So, using the metric hack on the training data won't necessarily improve performance because it may just worsen the average gini. This is not the case when predicting on the test dataset.",
    "2808143": "Optimize false positives !!",
    "2811903": "why do you clip the predictions? arent they already in [0,1] range?",
    "2812532": "By reducing score the probability could go below 0. So clipping is necessary",
    "2812757": "yes. exactly",
    "2812761": "Yes, I agree. But playing with the unknown here if you have built a stable model and not knowing the quality of private test data towards the end can be dangerous. But still the counter is that you can choose two submissions :)"
  },
  "source": "meta"
}