{
  "id": 493880,
  "title": "num_group1 index missing in Testset?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/493880",
  "author_name": "",
  "post_date": "2024-04-15T07:42:26.349129900Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Have you noticed that some case ids in the test set are missing the num_group1 index (e.g., 0, 1, 2, 4, 5…), resulting in a non-sequential list? Do you know why they aren't aligned properly? I heard that the test cases are randomly selected from the original test set. Could this be the reason some are missing?</p>",
  "messages": [
    {
      "id": "2752880",
      "postDate": "04/15/2024 07:42:26",
      "content": "<p>Have you noticed that some case ids in the test set are missing the num_group1 index (e.g., 0, 1, 2, 4, 5…), resulting in a non-sequential list? Do you know why they aren't aligned properly? I heard that the test cases are randomly selected from the original test set. Could this be the reason some are missing?</p>",
      "rawMarkdown": "Have you noticed that some case ids in the test set are missing the num_group1 index (e.g., 0, 1, 2, 4, 5...), resulting in a non-sequential list? Do you know why they aren't aligned properly? I heard that the test cases are randomly selected from the original test set. Could this be the reason some are missing?",
      "votes": null
    },
    {
      "id": "2752921",
      "postDate": "04/15/2024 08:14:52",
      "content": "<p>Hi,<br>\nthis is not reason, test cases are selected randomly, but if case_id is selected, then all data for the case_id are included - all num_group1 or num_group2 are in the dataset.</p>",
      "rawMarkdown": "Hi,\nthis is not reason, test cases are selected randomly, but if case_id is selected, then all data for the case_id are included - all num_group1 or num_group2 are in the dataset.",
      "votes": null
    },
    {
      "id": "2754424",
      "postDate": "04/16/2024 03:32:44",
      "content": "<p>Thank you for your explanation. I understand now. <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> </p>",
      "rawMarkdown": "Thank you for your explanation. I understand now. @tomasjeline2",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2752921,
      "author_name": "tomasjeline2",
      "author_url": "",
      "post_date": "04/15/2024 08:14:52",
      "content": "<p>Hi,<br>\nthis is not reason, test cases are selected randomly, but if case_id is selected, then all data for the case_id are included - all num_group1 or num_group2 are in the dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2754424,
          "author_name": "kjihwan",
          "author_url": "",
          "post_date": "04/16/2024 03:32:44",
          "content": "<p>Thank you for your explanation. I understand now. <a href=\"https://www.kaggle.com/tomasjeline2\" target=\"_blank\">@tomasjeline2</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2752880": "Have you noticed that some case ids in the test set are missing the num_group1 index (e.g., 0, 1, 2, 4, 5...), resulting in a non-sequential list? Do you know why they aren't aligned properly? I heard that the test cases are randomly selected from the original test set. Could this be the reason some are missing?",
    "2752921": "Hi,\nthis is not reason, test cases are selected randomly, but if case_id is selected, then all data for the case_id are included - all num_group1 or num_group2 are in the dataset.",
    "2754424": "Thank you for your explanation. I understand now. @tomasjeline2"
  },
  "source": "meta"
}