{
  "id": 508070,
  "title": "The Good, the Bad, and the Hack..",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/508070",
  "author_name": "George Papachristou",
  "post_date": "2024-05-28T06:50:43.463000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>As the competition comes to a close, I want to share my reflections. Many of us were disappointed by the hack that influenced the final results. Initially, I was disappointed that some of the top solutions might a result of metric hack. However, this is part of the Kaggle experience; evaluate the \"hack boost\" and also develop a good baseline model</p>\n<p><strong>Discussion forum, only threads around hack..</strong><br>\nFrom my perspective, the most significant impact of this hack was on the discussion forum. Instead of engaging in meaningful conversations about data aggregation and feature engineering, much of the discourse was dominated by the hack. Personally, I'm here to learn, improve my modeling skills, and apply this knowledge to my daily work, not to simply copy code without understanding, just for the sake of climbing the leaderboard.</p>\n<p><strong>Data size, OOM Issues</strong><br>\nAnother point worth mentioning is the data size. The dataset was nearly at Kaggle's resource limits (memory limit), which made it challenging to explore complex engineering methods within the platform's constraints. While it's possible to work around these limitations with some extra effort it was possible. But a lot of us was struggled with OOM issues in training or in submission. Participants with powerful local machines and RAM of 64GB or more had a distinct advantage. For the sake of fair play, it would be beneficial for future Kaggle competitions to consider these aspects.</p>\n<p>All in all, I'm glad after competition close and the lessons which gave us, finally I see interesting posts with solutions with engineering effort!</p>",
  "messages": [
    {
      "id": 2840593,
      "postDate": "2024-05-28T06:50:43.463Z",
      "content": "<p>As the competition comes to a close, I want to share my reflections. Many of us were disappointed by the hack that influenced the final results. Initially, I was disappointed that some of the top solutions might a result of metric hack. However, this is part of the Kaggle experience; evaluate the \"hack boost\" and also develop a good baseline model</p>\n<p><strong>Discussion forum, only threads around hack..</strong><br>\nFrom my perspective, the most significant impact of this hack was on the discussion forum. Instead of engaging in meaningful conversations about data aggregation and feature engineering, much of the discourse was dominated by the hack. Personally, I'm here to learn, improve my modeling skills, and apply this knowledge to my daily work, not to simply copy code without understanding, just for the sake of climbing the leaderboard.</p>\n<p><strong>Data size, OOM Issues</strong><br>\nAnother point worth mentioning is the data size. The dataset was nearly at Kaggle's resource limits (memory limit), which made it challenging to explore complex engineering methods within the platform's constraints. While it's possible to work around these limitations with some extra effort it was possible. But a lot of us was struggled with OOM issues in training or in submission. Participants with powerful local machines and RAM of 64GB or more had a distinct advantage. For the sake of fair play, it would be beneficial for future Kaggle competitions to consider these aspects.</p>\n<p>All in all, I'm glad after competition close and the lessons which gave us, finally I see interesting posts with solutions with engineering effort!</p>",
      "rawMarkdown": "As the competition comes to a close, I want to share my reflections. Many of us were disappointed by the hack that influenced the final results. Initially, I was disappointed that some of the top solutions might a result of metric hack. However, this is part of the Kaggle experience; evaluate the \"hack boost\" and also develop a good baseline model\n\n\n**Discussion forum, only threads around hack..**\nFrom my perspective, the most significant impact of this hack was on the discussion forum. Instead of engaging in meaningful conversations about data aggregation and feature engineering, much of the discourse was dominated by the hack. Personally, I'm here to learn, improve my modeling skills, and apply this knowledge to my daily work, not to simply copy code without understanding, just for the sake of climbing the leaderboard.\n\n\n**Data size, OOM Issues**\nAnother point worth mentioning is the data size. The dataset was nearly at Kaggle's resource limits (memory limit), which made it challenging to explore complex engineering methods within the platform's constraints. While it's possible to work around these limitations with some extra effort it was possible. But a lot of us was struggled with OOM issues in training or in submission. Participants with powerful local machines and RAM of 64GB or more had a distinct advantage. For the sake of fair play, it would be beneficial for future Kaggle competitions to consider these aspects.\n\nAll in all, I'm glad after competition close and the lessons which gave us, finally I see interesting posts with solutions with engineering effort!",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2840593": "As the competition comes to a close, I want to share my reflections. Many of us were disappointed by the hack that influenced the final results. Initially, I was disappointed that some of the top solutions might a result of metric hack. However, this is part of the Kaggle experience; evaluate the \"hack boost\" and also develop a good baseline model\n\n\n**Discussion forum, only threads around hack..**\nFrom my perspective, the most significant impact of this hack was on the discussion forum. Instead of engaging in meaningful conversations about data aggregation and feature engineering, much of the discourse was dominated by the hack. Personally, I'm here to learn, improve my modeling skills, and apply this knowledge to my daily work, not to simply copy code without understanding, just for the sake of climbing the leaderboard.\n\n\n**Data size, OOM Issues**\nAnother point worth mentioning is the data size. The dataset was nearly at Kaggle's resource limits (memory limit), which made it challenging to explore complex engineering methods within the platform's constraints. While it's possible to work around these limitations with some extra effort it was possible. But a lot of us was struggled with OOM issues in training or in submission. Participants with powerful local machines and RAM of 64GB or more had a distinct advantage. For the sake of fair play, it would be beneficial for future Kaggle competitions to consider these aspects.\n\nAll in all, I'm glad after competition close and the lessons which gave us, finally I see interesting posts with solutions with engineering effort!"
  }
}