{
  "id": 499699,
  "title": "Will the competition hosts disqualify cheaters?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/499699",
  "author_name": "Matous Famera",
  "post_date": "2024-05-02T17:29:38.247000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Dear <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>, <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a>, and <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>,</p>\n<p>First of all, I want to thank the hosts for running this competition. It's been a part of my learning journey. I really enjoy participating and working on the solution with my team. And I hope to climb up the leaderboard before the deadline.</p>\n<p>I know there have been problems with potential metric hacking, and I am aware of the previous competition update regarding the edit of the test dataset. However, there is an opinion that the evaluation metric can still be manipulated.</p>\n<p>I have not tried the evaluation metric hack, so I don't have insight into it like others. The idea of hacking stems from worsening the AUC in the high-score weeks, flattening the slope of the fitted line, and lowering the variance of the score throughout the weeks. It has been possible with the \"WEEK_NUM,\" \"date_decision,\" and \"MONTH\" available in the test dataset. However, it seems it is possible even without these columns available in the test dataset.</p>\n<p>There are a few posts which suggest the hacking method:<br>\n<a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167\" target=\"_blank\">POST 1</a><br>\n<a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/496898\" target=\"_blank\">POST 2</a></p>\n<p>Reading all the discussion posts, it seems the participant who does not utilize this hack might be at a serious disadvantage. I'm sure many of us have put significant effort and time into the solution of this competition. It would be very unfair to get dominated by cheating solutions.</p>\n<p>You can read all of this in the previous discussion posts, so what is my post about?</p>\n<p>The thing I want to ask the competition hosts is: how are you going to handle the cheating solutions. What is considered cheating and what is not? Are all notebooks (even the cheating) eligible for the private leaderboard? Is utilizing the metric hack described above legitimate or not? Are the submissions utilizing the metric hack going to be removed from the private leaderboard?</p>\n<p>Thanks for your answer.</p>\n<p>Good luck, everyone!</p>",
  "messages": [
    {
      "id": 2789485,
      "postDate": "2024-05-02T17:29:38.247Z",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a>, <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a>, and <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>,</p>\n<p>First of all, I want to thank the hosts for running this competition. It's been a part of my learning journey. I really enjoy participating and working on the solution with my team. And I hope to climb up the leaderboard before the deadline.</p>\n<p>I know there have been problems with potential metric hacking, and I am aware of the previous competition update regarding the edit of the test dataset. However, there is an opinion that the evaluation metric can still be manipulated.</p>\n<p>I have not tried the evaluation metric hack, so I don't have insight into it like others. The idea of hacking stems from worsening the AUC in the high-score weeks, flattening the slope of the fitted line, and lowering the variance of the score throughout the weeks. It has been possible with the \"WEEK_NUM,\" \"date_decision,\" and \"MONTH\" available in the test dataset. However, it seems it is possible even without these columns available in the test dataset.</p>\n<p>There are a few posts which suggest the hacking method:<br>\n<a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167\" target=\"_blank\">POST 1</a><br>\n<a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/496898\" target=\"_blank\">POST 2</a></p>\n<p>Reading all the discussion posts, it seems the participant who does not utilize this hack might be at a serious disadvantage. I'm sure many of us have put significant effort and time into the solution of this competition. It would be very unfair to get dominated by cheating solutions.</p>\n<p>You can read all of this in the previous discussion posts, so what is my post about?</p>\n<p>The thing I want to ask the competition hosts is: how are you going to handle the cheating solutions. What is considered cheating and what is not? Are all notebooks (even the cheating) eligible for the private leaderboard? Is utilizing the metric hack described above legitimate or not? Are the submissions utilizing the metric hack going to be removed from the private leaderboard?</p>\n<p>Thanks for your answer.</p>\n<p>Good luck, everyone!</p>",
      "rawMarkdown": "Dear @jetakow, @tomasjeline2, and @inversion,\n\nFirst of all, I want to thank the hosts for running this competition. It's been a part of my learning journey. I really enjoy participating and working on the solution with my team. And I hope to climb up the leaderboard before the deadline.\n\nI know there have been problems with potential metric hacking, and I am aware of the previous competition update regarding the edit of the test dataset. However, there is an opinion that the evaluation metric can still be manipulated.\n\nI have not tried the evaluation metric hack, so I don't have insight into it like others. The idea of hacking stems from worsening the AUC in the high-score weeks, flattening the slope of the fitted line, and lowering the variance of the score throughout the weeks. It has been possible with the \"WEEK_NUM,\" \"date_decision,\" and \"MONTH\" available in the test dataset. However, it seems it is possible even without these columns available in the test dataset.\n\nThere are a few posts which suggest the hacking method:\n[POST 1](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167)\n[POST 2](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/496898)\n\nReading all the discussion posts, it seems the participant who does not utilize this hack might be at a serious disadvantage. I'm sure many of us have put significant effort and time into the solution of this competition. It would be very unfair to get dominated by cheating solutions.\n\nYou can read all of this in the previous discussion posts, so what is my post about?\n\nThe thing I want to ask the competition hosts is: how are you going to handle the cheating solutions. What is considered cheating and what is not? Are all notebooks (even the cheating) eligible for the private leaderboard? Is utilizing the metric hack described above legitimate or not? Are the submissions utilizing the metric hack going to be removed from the private leaderboard?\n\nThanks for your answer.\n\nGood luck, everyone!",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2789485": "Dear @jetakow, @tomasjeline2, and @inversion,\n\nFirst of all, I want to thank the hosts for running this competition. It's been a part of my learning journey. I really enjoy participating and working on the solution with my team. And I hope to climb up the leaderboard before the deadline.\n\nI know there have been problems with potential metric hacking, and I am aware of the previous competition update regarding the edit of the test dataset. However, there is an opinion that the evaluation metric can still be manipulated.\n\nI have not tried the evaluation metric hack, so I don't have insight into it like others. The idea of hacking stems from worsening the AUC in the high-score weeks, flattening the slope of the fitted line, and lowering the variance of the score throughout the weeks. It has been possible with the \"WEEK_NUM,\" \"date_decision,\" and \"MONTH\" available in the test dataset. However, it seems it is possible even without these columns available in the test dataset.\n\nThere are a few posts which suggest the hacking method:\n[POST 1](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/497167)\n[POST 2](https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/496898)\n\nReading all the discussion posts, it seems the participant who does not utilize this hack might be at a serious disadvantage. I'm sure many of us have put significant effort and time into the solution of this competition. It would be very unfair to get dominated by cheating solutions.\n\nYou can read all of this in the previous discussion posts, so what is my post about?\n\nThe thing I want to ask the competition hosts is: how are you going to handle the cheating solutions. What is considered cheating and what is not? Are all notebooks (even the cheating) eligible for the private leaderboard? Is utilizing the metric hack described above legitimate or not? Are the submissions utilizing the metric hack going to be removed from the private leaderboard?\n\nThanks for your answer.\n\nGood luck, everyone!"
  }
}