{
  "id": 502939,
  "title": "Private LB data",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/502939",
  "author_name": "",
  "post_date": "2024-05-15T12:23:46.334839500Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I was wondering if public test is sequentially continued from train data, why does train case_id ends with ±2_500_000, but contains only ±1_500_000 samples.</p>\n<p>What are the chances private sample consist of missing data from each week, including both train and public test? In this case metric hack may even worse final score.</p>\n<p>Maybe I missed some explanation for this case_id situation. </p>",
  "messages": [
    {
      "id": "2814583",
      "postDate": "05/15/2024 12:23:46",
      "content": "<p>I was wondering if public test is sequentially continued from train data, why does train case_id ends with ±2_500_000, but contains only ±1_500_000 samples.</p>\n<p>What are the chances private sample consist of missing data from each week, including both train and public test? In this case metric hack may even worse final score.</p>\n<p>Maybe I missed some explanation for this case_id situation. </p>",
      "rawMarkdown": "I was wondering if public test is sequentially continued from train data, why does train case_id ends with ±2_500_000, but contains only ±1_500_000 samples.\n\nWhat are the chances private sample consist of missing data from each week, including both train and public test? In this case metric hack may even worse final score.\n\nMaybe I missed some explanation for this case_id situation.",
      "votes": null
    },
    {
      "id": "2814736",
      "postDate": "05/15/2024 14:19:51",
      "content": "<p>Actually I hope hacking won't work. But I believe it will though…</p>",
      "rawMarkdown": "Actually I hope hacking won't work. But I believe it will though...",
      "votes": null
    },
    {
      "id": "2814757",
      "postDate": "05/15/2024 14:28:40",
      "content": "<p>ye, but, as far I know, most of the current top participants also has strong non-hacked models. With 2 submissions one could protect himself from hacking instability</p>",
      "rawMarkdown": "ye, but, as far I know, most of the current top participants also has strong non-hacked models. With 2 submissions one could protect himself from hacking instability",
      "votes": null
    },
    {
      "id": "2815049",
      "postDate": "05/15/2024 16:59:04",
      "content": "<p>That's true. On the other hand I guess the \"strong non-hacked models\" still might remsemble the hacked versions but to a much lesser degree (and done unintentional). That might make them more robust to the \"hacking instability\", but I still wonder if that's what the host really wants. </p>",
      "rawMarkdown": "That's true. On the other hand I guess the \"strong non-hacked models\" still might remsemble the hacked versions but to a much lesser degree (and done unintentional). That might make them more robust to the \"hacking instability\", but I still wonder if that's what the host really wants.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2814736,
      "author_name": "ern711",
      "author_url": "",
      "post_date": "05/15/2024 14:19:51",
      "content": "<p>Actually I hope hacking won't work. But I believe it will though…</p>",
      "votes": null,
      "replies": [
        {
          "id": 2814757,
          "author_name": "bluepill",
          "author_url": "",
          "post_date": "05/15/2024 14:28:40",
          "content": "<p>ye, but, as far I know, most of the current top participants also has strong non-hacked models. With 2 submissions one could protect himself from hacking instability</p>",
          "votes": null,
          "replies": [
            {
              "id": 2815049,
              "author_name": "ern711",
              "author_url": "",
              "post_date": "05/15/2024 16:59:04",
              "content": "<p>That's true. On the other hand I guess the \"strong non-hacked models\" still might remsemble the hacked versions but to a much lesser degree (and done unintentional). That might make them more robust to the \"hacking instability\", but I still wonder if that's what the host really wants. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2814583": "I was wondering if public test is sequentially continued from train data, why does train case_id ends with ±2_500_000, but contains only ±1_500_000 samples.\n\nWhat are the chances private sample consist of missing data from each week, including both train and public test? In this case metric hack may even worse final score.\n\nMaybe I missed some explanation for this case_id situation.",
    "2814736": "Actually I hope hacking won't work. But I believe it will though...",
    "2814757": "ye, but, as far I know, most of the current top participants also has strong non-hacked models. With 2 submissions one could protect himself from hacking instability",
    "2815049": "That's true. On the other hand I guess the \"strong non-hacked models\" still might remsemble the hacked versions but to a much lesser degree (and done unintentional). That might make them more robust to the \"hacking instability\", but I still wonder if that's what the host really wants."
  },
  "source": "meta"
}