{
  "id": 543740,
  "title": "How are you use the questionaire features?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/543740",
  "author_name": "",
  "post_date": "2024-11-01T07:32:21.183640400Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>the test set does not include the questionnaire features <code>PCIAT-PCIAT-*</code>, wonder</p>\n<ol>\n<li>I assume the test set surely does not include the same ids in the training set right?</li>\n<li>based of the above, how are you thinking to use the questionnaire features? e.g. categorical encoding?</li>\n</ol>",
  "messages": [
    {
      "id": "3033481",
      "postDate": "11/01/2024 07:32:21",
      "content": "<p>the test set does not include the questionnaire features <code>PCIAT-PCIAT-*</code>, wonder</p>\n<ol>\n<li>I assume the test set surely does not include the same ids in the training set right?</li>\n<li>based of the above, how are you thinking to use the questionnaire features? e.g. categorical encoding?</li>\n</ol>",
      "rawMarkdown": "the test set does not include the questionnaire features `PCIAT-PCIAT-*`, wonder\n1.  I assume the test set surely does not include the same ids in the training set right?\n2. based of the above, how are you thinking to use the questionnaire features? e.g. categorical encoding?",
      "votes": null
    },
    {
      "id": "3034150",
      "postDate": "11/01/2024 20:25:19",
      "content": "<p>Well, the test set doesn't hold these features because the sii score is based on the total score of the individual scores. Knowing these individual scores will give you a perfect signal for the target value which doesn't make sense. </p>\n<p>To answer your questions: </p>\n<ol>\n<li>No, the test set will have different IDs than the training set. </li>\n<li>Well, you can't use them using inference given the reason above. However, we might be able to use them for imputing missing values. I haven't looked into that myself yet. </li>\n</ol>",
      "rawMarkdown": "Well, the test set doesn't hold these features because the sii score is based on the total score of the individual scores. Knowing these individual scores will give you a perfect signal for the target value which doesn't make sense. \n\nTo answer your questions: \n\n1. No, the test set will have different IDs than the training set. \n2. Well, you can't use them using inference given the reason above. However, we might be able to use them for imputing missing values. I haven't looked into that myself yet.",
      "votes": null
    },
    {
      "id": "3034632",
      "postDate": "11/02/2024 12:20:53",
      "content": "<p>I have dropped all the questionnaire features and I have also noted that almost everyone has dropped those features. We have to make sure data in test and train sets have same columns to ensure that models will run on them without error.</p>",
      "rawMarkdown": "I have dropped all the questionnaire features and I have also noted that almost everyone has dropped those features. We have to make sure data in test and train sets have same columns to ensure that models will run on them without error.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3034150,
      "author_name": "woutermostard",
      "author_url": "",
      "post_date": "11/01/2024 20:25:19",
      "content": "<p>Well, the test set doesn't hold these features because the sii score is based on the total score of the individual scores. Knowing these individual scores will give you a perfect signal for the target value which doesn't make sense. </p>\n<p>To answer your questions: </p>\n<ol>\n<li>No, the test set will have different IDs than the training set. </li>\n<li>Well, you can't use them using inference given the reason above. However, we might be able to use them for imputing missing values. I haven't looked into that myself yet. </li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3034632,
      "author_name": "taimour",
      "author_url": "",
      "post_date": "11/02/2024 12:20:53",
      "content": "<p>I have dropped all the questionnaire features and I have also noted that almost everyone has dropped those features. We have to make sure data in test and train sets have same columns to ensure that models will run on them without error.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3033481": "the test set does not include the questionnaire features `PCIAT-PCIAT-*`, wonder\n1.  I assume the test set surely does not include the same ids in the training set right?\n2. based of the above, how are you thinking to use the questionnaire features? e.g. categorical encoding?",
    "3034150": "Well, the test set doesn't hold these features because the sii score is based on the total score of the individual scores. Knowing these individual scores will give you a perfect signal for the target value which doesn't make sense. \n\nTo answer your questions: \n\n1. No, the test set will have different IDs than the training set. \n2. Well, you can't use them using inference given the reason above. However, we might be able to use them for imputing missing values. I haven't looked into that myself yet.",
    "3034632": "I have dropped all the questionnaire features and I have also noted that almost everyone has dropped those features. We have to make sure data in test and train sets have same columns to ensure that models will run on them without error."
  },
  "source": "meta"
}