{
  "id": 505646,
  "title": "It's funny that the scores were way lower with the original version of the hack",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/505646",
  "author_name": "",
  "post_date": "2024-05-18T12:18:26.605726100Z",
  "votes": 12,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I find it quite hilarious that the scores were way lower when the metric hack was first proposed where we had the original week numbers. After removing the week num, we got more and more creative ideas to hack the metric since then and here we are :D </p>",
  "messages": [
    {
      "id": "2822053",
      "postDate": "05/18/2024 12:18:26",
      "content": "<p>I find it quite hilarious that the scores were way lower when the metric hack was first proposed where we had the original week numbers. After removing the week num, we got more and more creative ideas to hack the metric since then and here we are :D </p>",
      "rawMarkdown": "I find it quite hilarious that the scores were way lower when the metric hack was first proposed where we had the original week numbers. After removing the week num, we got more and more creative ideas to hack the metric since then and here we are :D",
      "votes": null
    },
    {
      "id": "2822076",
      "postDate": "05/18/2024 12:35:43",
      "content": "<p><a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> what a degraded competition it has become!!!<br>\nI think they should simply close it 😀</p>",
      "rawMarkdown": "snnclsr what a degraded competition it has become!!!\nI think they should simply close it 😀",
      "votes": null
    },
    {
      "id": "2822120",
      "postDate": "05/18/2024 13:01:32",
      "content": "<p>I trained a very good cat model locally, but it was terrible on LB. And the better the local, the worse the LB.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20129689%2F9a37950f8ba79b8673ceb13eb8469af1%2Fef65cfa0fbc14770a21989fc1d458dc.png?generation=1716037255279724&amp;alt=media\"></p>",
      "rawMarkdown": "I trained a very good cat model locally, but it was terrible on LB. And the better the local, the worse the LB.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20129689%2F9a37950f8ba79b8673ceb13eb8469af1%2Fef65cfa0fbc14770a21989fc1d458dc.png?generation=1716037255279724&alt=media)",
      "votes": null
    },
    {
      "id": "2822127",
      "postDate": "05/18/2024 13:07:07",
      "content": "<p>Same with us! We also trained some really good models that failed miserably on the leaderboard. Seems like this competition was not worth trying to build good models <a href=\"https://www.kaggle.com/f520hh\" target=\"_blank\">@f520hh</a> </p>",
      "rawMarkdown": "Same with us! We also trained some really good models that failed miserably on the leaderboard. Seems like this competition was not worth trying to build good models @f520hh",
      "votes": null
    },
    {
      "id": "2822140",
      "postDate": "05/18/2024 13:12:54",
      "content": "<p>Who knows, maybe it will turn out good in the private! Well, hopefully.</p>",
      "rawMarkdown": "Who knows, maybe it will turn out good in the private! Well, hopefully.",
      "votes": null
    },
    {
      "id": "2822149",
      "postDate": "05/18/2024 13:15:48",
      "content": "<p>Haha~  Really Funny</p>",
      "rawMarkdown": "Haha~  Really Funny",
      "votes": null
    },
    {
      "id": "2822519",
      "postDate": "05/18/2024 16:37:19",
      "content": "<p>I did the blending based on LGBM and Catboost. It was 0.593 on the public leaderboard. I thought it was okay. But in the morning I was unpleasantly surprised how low I fell on the leaderboard :). Had to accept the rules of the game :).</p>",
      "rawMarkdown": "I did the blending based on LGBM and Catboost. It was 0.593 on the public leaderboard. I thought it was okay. But in the morning I was unpleasantly surprised how low I fell on the leaderboard :). Had to accept the rules of the game :).",
      "votes": null
    },
    {
      "id": "2822559",
      "postDate": "05/18/2024 17:03:47",
      "content": "<p>It's because at the beginning people weren't concentrated on hacking and after the data was modified also not much interest. Cheating became the competition's main goal after the organizer wrote that everything is allowed (without breaking into the servers) - so no fear of disqualification made people to concentrate on the hacking instead of scrubing 0.001 from a fine tune, some feature engineering or other smart ideas.</p>",
      "rawMarkdown": "It's because at the beginning people weren't concentrated on hacking and after the data was modified also not much interest. Cheating became the competition's main goal after the organizer wrote that everything is allowed (without breaking into the servers) - so no fear of disqualification made people to concentrate on the hacking instead of scrubing 0.001 from a fine tune, some feature engineering or other smart ideas.",
      "votes": null
    },
    {
      "id": "2822737",
      "postDate": "05/18/2024 18:25:30",
      "content": "<p>Please!, close this at least to recover some respect!</p>",
      "rawMarkdown": "Please!, close this at least to recover some respect!",
      "votes": null
    },
    {
      "id": "2822978",
      "postDate": "05/18/2024 23:51:12",
      "content": "<p>I observe the same thing!</p>\n<p>Also, the worse the model is at predicting defaults, the better the stability score is. </p>",
      "rawMarkdown": "I observe the same thing!\n\nAlso, the worse the model is at predicting defaults, the better the stability score is.",
      "votes": null
    },
    {
      "id": "2823869",
      "postDate": "05/19/2024 13:11:02",
      "content": "<p>I believe that you are not right. <br>\nIf most of participants used similar way to hack metric, strong models have a good chance of outperforming public scripts. </p>",
      "rawMarkdown": "I believe that you are not right. \nIf most of participants used similar way to hack metric, strong models have a good chance of outperforming public scripts.",
      "votes": null
    },
    {
      "id": "2824307",
      "postDate": "05/19/2024 17:48:44",
      "content": "<p>I trained a very good cat model locally, but it was terrible on LB. And the better the local, the worse the LB!</p>",
      "rawMarkdown": "I trained a very good cat model locally, but it was terrible on LB. And the better the local, the worse the LB!",
      "votes": null
    },
    {
      "id": "2824333",
      "postDate": "05/19/2024 18:10:44",
      "content": "<p>How does this hack do on your ensemble? I had a similar situation, but this hack lowered the LB score compared to the initial ensemble's LB score.</p>",
      "rawMarkdown": "How does this hack do on your ensemble? I had a similar situation, but this hack lowered the LB score compared to the initial ensemble's LB score.",
      "votes": null
    },
    {
      "id": "2824337",
      "postDate": "05/19/2024 18:12:11",
      "content": "<p>That ship has sailed ⛵️ 😄</p>",
      "rawMarkdown": "That ship has sailed ⛵️ 😄",
      "votes": null
    },
    {
      "id": "2824377",
      "postDate": "05/19/2024 18:26:39",
      "content": "<p>Improves the result. But I'll probably keep my original clean version, because I think all these tricks with metrics lead to hard overfitting. In fact, this competition is now about who can reproduce someone else's code the fastest (and there are many such cases, just look at the date of registration of accounts and the number of submissions) or send more than 5 submissions from several accounts and choose the optimal one for the main account. The contest has become a joke/prunk and even such professionals in risk management as <a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a> and <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> are already far behind the top places, which looks surreal. When I started this competition, I thought I would find some fresh ideas for scoring models to apply in my work in bank. But I learned only one lesson - not to use this stable metric in a production environment :). </p>",
      "rawMarkdown": "Improves the result. But I'll probably keep my original clean version, because I think all these tricks with metrics lead to hard overfitting. In fact, this competition is now about who can reproduce someone else's code the fastest (and there are many such cases, just look at the date of registration of accounts and the number of submissions) or send more than 5 submissions from several accounts and choose the optimal one for the main account. The contest has become a joke/prunk and even such professionals in risk management as @johnpateha and @ravi20076 are already far behind the top places, which looks surreal. When I started this competition, I thought I would find some fresh ideas for scoring models to apply in my work in bank. But I learned only one lesson - not to use this stable metric in a production environment :).",
      "votes": null
    },
    {
      "id": "2824532",
      "postDate": "05/19/2024 20:25:14",
      "content": "<p>it's too early to start panic :)  Only final standing count. <br>\nCurrent LB means nothing, because it will change next week</p>",
      "rawMarkdown": "it's too early to start panic :)  Only final standing count. \nCurrent LB means nothing, because it will change next week",
      "votes": null
    },
    {
      "id": "2824568",
      "postDate": "05/19/2024 21:05:48",
      "content": "<p>Yeah, we'll see soon enough. Did you do the same stacking as last time? In general, based on your experience (not only in Uralsib), have you encountered ensembles of scoring models (applicative, transactional) in a production environment or mainly logistic regression because of the requirements for interpretability and ease of creating scorecard?</p>",
      "rawMarkdown": "Yeah, we'll see soon enough. Did you do the same stacking as last time? In general, based on your experience (not only in Uralsib), have you encountered ensembles of scoring models (applicative, transactional) in a production environment or mainly logistic regression because of the requirements for interpretability and ease of creating scorecard?",
      "votes": null
    },
    {
      "id": "2824593",
      "postDate": "05/19/2024 21:56:35",
      "content": "<p>Interesting… I have a multi model ensemble that contains cat, lgb and xgb that doesn't do well with this approach… at least not with the initial parameters. I didn't have the time to experiment with various score shifts  but I may give it a try this upcoming week. I liked my ensemble a lot before the explosion of scores that appeared out of this world. What I thought to be a robust ensemble of several high AUC models couldn't even come close to the scores that I got by running the \"This is the way\" hacked notebook. But I realize that relying on that notebook is probably not a winning strategy, so I'll keep my ensemble as one of my final submissions. </p>",
      "rawMarkdown": "Interesting... I have a multi model ensemble that contains cat, lgb and xgb that doesn't do well with this approach... at least not with the initial parameters. I didn't have the time to experiment with various score shifts  but I may give it a try this upcoming week. I liked my ensemble a lot before the explosion of scores that appeared out of this world. What I thought to be a robust ensemble of several high AUC models couldn't even come close to the scores that I got by running the \"This is the way\" hacked notebook. But I realize that relying on that notebook is probably not a winning strategy, so I'll keep my ensemble as one of my final submissions.",
      "votes": null
    },
    {
      "id": "2824610",
      "postDate": "05/19/2024 23:08:09",
      "content": "<p>I was more on the hacking side so far (hacked ones will be the winners). But with the recent updates, I am now more towards the normal (!) submissions. I think it will include the \"hack\" still but will be less overfitted. Like not making the threshold 0.984583945 or something.</p>",
      "rawMarkdown": "I was more on the hacking side so far (hacked ones will be the winners). But with the recent updates, I am now more towards the normal (!) submissions. I think it will include the \"hack\" still but will be less overfitted. Like not making the threshold 0.984583945 or something.",
      "votes": null
    },
    {
      "id": "2824833",
      "postDate": "05/20/2024 04:45:44",
      "content": "<p>My model ensemble that scores LB=0.592, score much lower i.e. LB=0.515 with the latest one liner trick😀. I hope this useless trick don't fair well in the private LB. We will see next week.</p>",
      "rawMarkdown": "My model ensemble that scores LB=0.592, score much lower i.e. LB=0.515 with the latest one liner trick😀. I hope this useless trick don't fair well in the private LB. We will see next week.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2822076,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "05/18/2024 12:35:43",
      "content": "<p><a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> what a degraded competition it has become!!!<br>\nI think they should simply close it 😀</p>",
      "votes": null,
      "replies": [
        {
          "id": 2822737,
          "author_name": "carloshuertas",
          "author_url": "",
          "post_date": "05/18/2024 18:25:30",
          "content": "<p>Please!, close this at least to recover some respect!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2824337,
              "author_name": "eduardastefanescu",
              "author_url": "",
              "post_date": "05/19/2024 18:12:11",
              "content": "<p>That ship has sailed ⛵️ 😄</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2822120,
      "author_name": "",
      "author_url": "",
      "post_date": "05/18/2024 13:01:32",
      "content": "<p>I trained a very good cat model locally, but it was terrible on LB. And the better the local, the worse the LB.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20129689%2F9a37950f8ba79b8673ceb13eb8469af1%2Fef65cfa0fbc14770a21989fc1d458dc.png?generation=1716037255279724&amp;alt=media\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2822127,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "05/18/2024 13:07:07",
          "content": "<p>Same with us! We also trained some really good models that failed miserably on the leaderboard. Seems like this competition was not worth trying to build good models <a href=\"https://www.kaggle.com/f520hh\" target=\"_blank\">@f520hh</a> </p>",
          "votes": null,
          "replies": [
            {
              "id": 2822140,
              "author_name": "snnclsr",
              "author_url": "",
              "post_date": "05/18/2024 13:12:54",
              "content": "<p>Who knows, maybe it will turn out good in the private! Well, hopefully.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2823869,
              "author_name": "johnpateha",
              "author_url": "",
              "post_date": "05/19/2024 13:11:02",
              "content": "<p>I believe that you are not right. <br>\nIf most of participants used similar way to hack metric, strong models have a good chance of outperforming public scripts. </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2822519,
          "author_name": "bratkovskyevgeny",
          "author_url": "",
          "post_date": "05/18/2024 16:37:19",
          "content": "<p>I did the blending based on LGBM and Catboost. It was 0.593 on the public leaderboard. I thought it was okay. But in the morning I was unpleasantly surprised how low I fell on the leaderboard :). Had to accept the rules of the game :).</p>",
          "votes": null,
          "replies": [
            {
              "id": 2824333,
              "author_name": "eduardastefanescu",
              "author_url": "",
              "post_date": "05/19/2024 18:10:44",
              "content": "<p>How does this hack do on your ensemble? I had a similar situation, but this hack lowered the LB score compared to the initial ensemble's LB score.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2824377,
                  "author_name": "bratkovskyevgeny",
                  "author_url": "",
                  "post_date": "05/19/2024 18:26:39",
                  "content": "<p>Improves the result. But I'll probably keep my original clean version, because I think all these tricks with metrics lead to hard overfitting. In fact, this competition is now about who can reproduce someone else's code the fastest (and there are many such cases, just look at the date of registration of accounts and the number of submissions) or send more than 5 submissions from several accounts and choose the optimal one for the main account. The contest has become a joke/prunk and even such professionals in risk management as <a href=\"https://www.kaggle.com/johnpateha\" target=\"_blank\">@johnpateha</a> and <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> are already far behind the top places, which looks surreal. When I started this competition, I thought I would find some fresh ideas for scoring models to apply in my work in bank. But I learned only one lesson - not to use this stable metric in a production environment :). </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2824532,
                      "author_name": "johnpateha",
                      "author_url": "",
                      "post_date": "05/19/2024 20:25:14",
                      "content": "<p>it's too early to start panic :)  Only final standing count. <br>\nCurrent LB means nothing, because it will change next week</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2824568,
                          "author_name": "bratkovskyevgeny",
                          "author_url": "",
                          "post_date": "05/19/2024 21:05:48",
                          "content": "<p>Yeah, we'll see soon enough. Did you do the same stacking as last time? In general, based on your experience (not only in Uralsib), have you encountered ensembles of scoring models (applicative, transactional) in a production environment or mainly logistic regression because of the requirements for interpretability and ease of creating scorecard?</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    },
                    {
                      "id": 2824593,
                      "author_name": "eduardastefanescu",
                      "author_url": "",
                      "post_date": "05/19/2024 21:56:35",
                      "content": "<p>Interesting… I have a multi model ensemble that contains cat, lgb and xgb that doesn't do well with this approach… at least not with the initial parameters. I didn't have the time to experiment with various score shifts  but I may give it a try this upcoming week. I liked my ensemble a lot before the explosion of scores that appeared out of this world. What I thought to be a robust ensemble of several high AUC models couldn't even come close to the scores that I got by running the \"This is the way\" hacked notebook. But I realize that relying on that notebook is probably not a winning strategy, so I'll keep my ensemble as one of my final submissions. </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2824610,
                          "author_name": "snnclsr",
                          "author_url": "",
                          "post_date": "05/19/2024 23:08:09",
                          "content": "<p>I was more on the hacking side so far (hacked ones will be the winners). But with the recent updates, I am now more towards the normal (!) submissions. I think it will include the \"hack\" still but will be less overfitted. Like not making the threshold 0.984583945 or something.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2822978,
          "author_name": "varuniraothumsi",
          "author_url": "",
          "post_date": "05/18/2024 23:51:12",
          "content": "<p>I observe the same thing!</p>\n<p>Also, the worse the model is at predicting defaults, the better the stability score is. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2822149,
      "author_name": "bruceqdu",
      "author_url": "",
      "post_date": "05/18/2024 13:15:48",
      "content": "<p>Haha~  Really Funny</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2822559,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "05/18/2024 17:03:47",
      "content": "<p>It's because at the beginning people weren't concentrated on hacking and after the data was modified also not much interest. Cheating became the competition's main goal after the organizer wrote that everything is allowed (without breaking into the servers) - so no fear of disqualification made people to concentrate on the hacking instead of scrubing 0.001 from a fine tune, some feature engineering or other smart ideas.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2824307,
      "author_name": "sheemazain",
      "author_url": "",
      "post_date": "05/19/2024 17:48:44",
      "content": "<p>I trained a very good cat model locally, but it was terrible on LB. And the better the local, the worse the LB!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2824833,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "05/20/2024 04:45:44",
      "content": "<p>My model ensemble that scores LB=0.592, score much lower i.e. LB=0.515 with the latest one liner trick😀. I hope this useless trick don't fair well in the private LB. We will see next week.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2822053": "I find it quite hilarious that the scores were way lower when the metric hack was first proposed where we had the original week numbers. After removing the week num, we got more and more creative ideas to hack the metric since then and here we are :D",
    "2822076": "snnclsr what a degraded competition it has become!!!\nI think they should simply close it 😀",
    "2822120": "I trained a very good cat model locally, but it was terrible on LB. And the better the local, the worse the LB.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20129689%2F9a37950f8ba79b8673ceb13eb8469af1%2Fef65cfa0fbc14770a21989fc1d458dc.png?generation=1716037255279724&alt=media)",
    "2822127": "Same with us! We also trained some really good models that failed miserably on the leaderboard. Seems like this competition was not worth trying to build good models @f520hh",
    "2822140": "Who knows, maybe it will turn out good in the private! Well, hopefully.",
    "2822149": "Haha~  Really Funny",
    "2822519": "I did the blending based on LGBM and Catboost. It was 0.593 on the public leaderboard. I thought it was okay. But in the morning I was unpleasantly surprised how low I fell on the leaderboard :). Had to accept the rules of the game :).",
    "2822559": "It's because at the beginning people weren't concentrated on hacking and after the data was modified also not much interest. Cheating became the competition's main goal after the organizer wrote that everything is allowed (without breaking into the servers) - so no fear of disqualification made people to concentrate on the hacking instead of scrubing 0.001 from a fine tune, some feature engineering or other smart ideas.",
    "2822737": "Please!, close this at least to recover some respect!",
    "2822978": "I observe the same thing!\n\nAlso, the worse the model is at predicting defaults, the better the stability score is.",
    "2823869": "I believe that you are not right. \nIf most of participants used similar way to hack metric, strong models have a good chance of outperforming public scripts.",
    "2824307": "I trained a very good cat model locally, but it was terrible on LB. And the better the local, the worse the LB!",
    "2824333": "How does this hack do on your ensemble? I had a similar situation, but this hack lowered the LB score compared to the initial ensemble's LB score.",
    "2824337": "That ship has sailed ⛵️ 😄",
    "2824377": "Improves the result. But I'll probably keep my original clean version, because I think all these tricks with metrics lead to hard overfitting. In fact, this competition is now about who can reproduce someone else's code the fastest (and there are many such cases, just look at the date of registration of accounts and the number of submissions) or send more than 5 submissions from several accounts and choose the optimal one for the main account. The contest has become a joke/prunk and even such professionals in risk management as @johnpateha and @ravi20076 are already far behind the top places, which looks surreal. When I started this competition, I thought I would find some fresh ideas for scoring models to apply in my work in bank. But I learned only one lesson - not to use this stable metric in a production environment :).",
    "2824532": "it's too early to start panic :)  Only final standing count. \nCurrent LB means nothing, because it will change next week",
    "2824568": "Yeah, we'll see soon enough. Did you do the same stacking as last time? In general, based on your experience (not only in Uralsib), have you encountered ensembles of scoring models (applicative, transactional) in a production environment or mainly logistic regression because of the requirements for interpretability and ease of creating scorecard?",
    "2824593": "Interesting... I have a multi model ensemble that contains cat, lgb and xgb that doesn't do well with this approach... at least not with the initial parameters. I didn't have the time to experiment with various score shifts  but I may give it a try this upcoming week. I liked my ensemble a lot before the explosion of scores that appeared out of this world. What I thought to be a robust ensemble of several high AUC models couldn't even come close to the scores that I got by running the \"This is the way\" hacked notebook. But I realize that relying on that notebook is probably not a winning strategy, so I'll keep my ensemble as one of my final submissions.",
    "2824610": "I was more on the hacking side so far (hacked ones will be the winners). But with the recent updates, I am now more towards the normal (!) submissions. I think it will include the \"hack\" still but will be less overfitted. Like not making the threshold 0.984583945 or something.",
    "2824833": "My model ensemble that scores LB=0.592, score much lower i.e. LB=0.515 with the latest one liner trick😀. I hope this useless trick don't fair well in the private LB. We will see next week."
  },
  "source": "meta"
}