{
  "id": 505600,
  "title": "Questions About Test Data Selection and Feature Engineering",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/505600",
  "author_name": "",
  "post_date": "2024-05-18T07:33:51.130550300Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I am participating in a Kaggle competition for the first time and I'm still learning the ropes of data science, so I would appreciate your patience and guidance with my questions.</p>\n<ol>\n<li><p>Is the data used for the leaderboard randomly selected from the test data, or is it chosen based on the smallest week_num values? I apologize if this question has been answered before.</p></li>\n<li><p>I've learned that feature engineering is crucial in data science, but I've noticed that there isn't much discussion about it in this competition. Is it common practice not to share feature engineering insights, or is feature engineering not as important in this particular competition?</p></li>\n</ol>\n<p>Thank you in advance for your help!</p>",
  "messages": [
    {
      "id": "2821671",
      "postDate": "05/18/2024 07:33:51",
      "content": "<p>Hello everyone,</p>\n<p>I am participating in a Kaggle competition for the first time and I'm still learning the ropes of data science, so I would appreciate your patience and guidance with my questions.</p>\n<ol>\n<li><p>Is the data used for the leaderboard randomly selected from the test data, or is it chosen based on the smallest week_num values? I apologize if this question has been answered before.</p></li>\n<li><p>I've learned that feature engineering is crucial in data science, but I've noticed that there isn't much discussion about it in this competition. Is it common practice not to share feature engineering insights, or is feature engineering not as important in this particular competition?</p></li>\n</ol>\n<p>Thank you in advance for your help!</p>",
      "rawMarkdown": "Hello everyone,\n\nI am participating in a Kaggle competition for the first time and I'm still learning the ropes of data science, so I would appreciate your patience and guidance with my questions.\n\n1. Is the data used for the leaderboard randomly selected from the test data, or is it chosen based on the smallest week_num values? I apologize if this question has been answered before.\n\n2. I've learned that feature engineering is crucial in data science, but I've noticed that there isn't much discussion about it in this competition. Is it common practice not to share feature engineering insights, or is feature engineering not as important in this particular competition?\n\nThank you in advance for your help!",
      "votes": null
    },
    {
      "id": "2821682",
      "postDate": "05/18/2024 07:38:38",
      "content": "<p><a href=\"https://www.kaggle.com/miulab\" target=\"_blank\">@miulab</a> <br>\nA lot of work on feature engineering is done in the kernels section. Kindly refer to the public code forums to know more </p>\n<p>We have no idea of how the leaderboard is split, but it presume it is the same for all participants to foster comparable results </p>",
      "rawMarkdown": "miulab \nA lot of work on feature engineering is done in the kernels section. Kindly refer to the public code forums to know more \n\nWe have no idea of how the leaderboard is split, but it presume it is the same for all participants to foster comparable results",
      "votes": null
    },
    {
      "id": "2821698",
      "postDate": "05/18/2024 07:47:53",
      "content": "<p>Thank you for the information! </p>\n<p>I was wondering if feature engineering has become less meaningful due to metric hacking. I will focus on feature engineering to avoid being shaken down.</p>\n<p>Thanks again!</p>",
      "rawMarkdown": "Thank you for the information! \n\nI was wondering if feature engineering has become less meaningful due to metric hacking. I will focus on feature engineering to avoid being shaken down.\n\nThanks again!",
      "votes": null
    },
    {
      "id": "2821761",
      "postDate": "05/18/2024 08:25:32",
      "content": "<p>Feature engineering is of importance <a href=\"https://www.kaggle.com/miulab\" target=\"_blank\">@miulab</a> <br>\nMetric hacking will only augment your score from a base score. To get there, you need to do good feature engineering!</p>",
      "rawMarkdown": "Feature engineering is of importance @miulab \nMetric hacking will only augment your score from a base score. To get there, you need to do good feature engineering!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2821682,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "05/18/2024 07:38:38",
      "content": "<p><a href=\"https://www.kaggle.com/miulab\" target=\"_blank\">@miulab</a> <br>\nA lot of work on feature engineering is done in the kernels section. Kindly refer to the public code forums to know more </p>\n<p>We have no idea of how the leaderboard is split, but it presume it is the same for all participants to foster comparable results </p>",
      "votes": null,
      "replies": [
        {
          "id": 2821698,
          "author_name": "miulab",
          "author_url": "",
          "post_date": "05/18/2024 07:47:53",
          "content": "<p>Thank you for the information! </p>\n<p>I was wondering if feature engineering has become less meaningful due to metric hacking. I will focus on feature engineering to avoid being shaken down.</p>\n<p>Thanks again!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2821761,
              "author_name": "ravi20076",
              "author_url": "",
              "post_date": "05/18/2024 08:25:32",
              "content": "<p>Feature engineering is of importance <a href=\"https://www.kaggle.com/miulab\" target=\"_blank\">@miulab</a> <br>\nMetric hacking will only augment your score from a base score. To get there, you need to do good feature engineering!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2821671": "Hello everyone,\n\nI am participating in a Kaggle competition for the first time and I'm still learning the ropes of data science, so I would appreciate your patience and guidance with my questions.\n\n1. Is the data used for the leaderboard randomly selected from the test data, or is it chosen based on the smallest week_num values? I apologize if this question has been answered before.\n\n2. I've learned that feature engineering is crucial in data science, but I've noticed that there isn't much discussion about it in this competition. Is it common practice not to share feature engineering insights, or is feature engineering not as important in this particular competition?\n\nThank you in advance for your help!",
    "2821682": "miulab \nA lot of work on feature engineering is done in the kernels section. Kindly refer to the public code forums to know more \n\nWe have no idea of how the leaderboard is split, but it presume it is the same for all participants to foster comparable results",
    "2821698": "Thank you for the information! \n\nI was wondering if feature engineering has become less meaningful due to metric hacking. I will focus on feature engineering to avoid being shaken down.\n\nThanks again!",
    "2821761": "Feature engineering is of importance @miulab \nMetric hacking will only augment your score from a base score. To get there, you need to do good feature engineering!"
  },
  "source": "meta"
}