{
  "id": 544108,
  "title": "Doubt in how to get started with this problem ?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/544108",
  "author_name": "Freak2209",
  "post_date": "2024-11-03T09:23:41.774000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am stuck with this problem statement and would like to understand the initial approach to tackle such a problem.</p>\n<ul>\n<li><p>There are many labels missing. How can I train my model based on the remaining data?</p></li>\n<li><p>The number of columns or features differs between the Test and Train datasets. How can I incorporate this? How can I predict the Test set without the complete set of features used in the Train set?</p></li>\n<li><p>There is a set of data in “parquet” format, available for only a subset of IDs. I thought of concatenating the parquet data with the CSV data, but this is not possible as some IDs do not have parquet data since the wrist accelerometer was worn by selected participants only. What should I do?</p></li>\n</ul>\n<p>I am new to these types of problems, so please help me learn from this situation.</p>",
  "messages": [
    {
      "id": 3035323,
      "postDate": "2024-11-03T09:23:41.773Z",
      "content": "<p>I am stuck with this problem statement and would like to understand the initial approach to tackle such a problem.</p>\n<ul>\n<li><p>There are many labels missing. How can I train my model based on the remaining data?</p></li>\n<li><p>The number of columns or features differs between the Test and Train datasets. How can I incorporate this? How can I predict the Test set without the complete set of features used in the Train set?</p></li>\n<li><p>There is a set of data in “parquet” format, available for only a subset of IDs. I thought of concatenating the parquet data with the CSV data, but this is not possible as some IDs do not have parquet data since the wrist accelerometer was worn by selected participants only. What should I do?</p></li>\n</ul>\n<p>I am new to these types of problems, so please help me learn from this situation.</p>",
      "rawMarkdown": "I am stuck with this problem statement and would like to understand the initial approach to tackle such a problem.\n\n- There are many labels missing. How can I train my model based on the remaining data?\n\n- The number of columns or features differs between the Test and Train datasets. How can I incorporate this? How can I predict the Test set without the complete set of features used in the Train set?\n\n- There is a set of data in “parquet” format, available for only a subset of IDs. I thought of concatenating the parquet data with the CSV data, but this is not possible as some IDs do not have parquet data since the wrist accelerometer was worn by selected participants only. What should I do?\n\nI am new to these types of problems, so please help me learn from this situation.",
      "votes": 1
    },
    {
      "id": 3037283,
      "postDate": "2024-11-05T15:05:00.220Z",
      "content": "<ol>\n<li>Just train on the data where labels are present (don't concentrate on what is missing).</li>\n<li>Features are the same in train and test, it's just the fact that you interpret as features that many PCIAT 0-20 columns that in fact are labels and in sum result in PCIAT Total that in it's turn is grouped in 4 classes that are represented in 'sii' column. Read carefully the overview and data sections.</li>\n<li>Acelerometer data should be aggregated to turn a long format into a wide format and concatenate with corresponding ID's where present, and where not present then we have NaNs.</li>\n</ol>",
      "rawMarkdown": "1. Just train on the data where labels are present (don't concentrate on what is missing).\n2. Features are the same in train and test, it's just the fact that you interpret as features that many PCIAT 0-20 columns that in fact are labels and in sum result in PCIAT Total that in it's turn is grouped in 4 classes that are represented in 'sii' column. Read carefully the overview and data sections.\n3. Acelerometer data should be aggregated to turn a long format into a wide format and concatenate with corresponding ID's where present, and where not present then we have NaNs.\n\n"
    },
    {
      "id": 3036281,
      "postDate": "2024-11-04T12:37:55.073Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 3036295,
          "postDate": "2024-11-04T12:50:21.910Z",
          "content": "<p>But not every id has its parquet data so how can i concatenate this ?</p>\n<p>There are more training ids but less number of parquet data</p>",
          "rawMarkdown": "But not every id has its parquet data so how can i concatenate this ?\n\nThere are more training ids but less number of parquet data",
          "replies": [
            {
              "id": 3036316,
              "postDate": "2024-11-04T12:58:53.600Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3037283,
      "author_name": "Danu A.",
      "author_url": "",
      "post_date": "2024-11-05T15:05:00.220000",
      "content": "<ol>\n<li>Just train on the data where labels are present (don't concentrate on what is missing).</li>\n<li>Features are the same in train and test, it's just the fact that you interpret as features that many PCIAT 0-20 columns that in fact are labels and in sum result in PCIAT Total that in it's turn is grouped in 4 classes that are represented in 'sii' column. Read carefully the overview and data sections.</li>\n<li>Acelerometer data should be aggregated to turn a long format into a wide format and concatenate with corresponding ID's where present, and where not present then we have NaNs.</li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3036281,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-04T12:37:55.073000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 3036295,
          "author_name": "Freak2209",
          "author_url": "",
          "post_date": "2024-11-04T12:50:21.910000",
          "content": "<p>But not every id has its parquet data so how can i concatenate this ?</p>\n<p>There are more training ids but less number of parquet data</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3036316,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-11-04T12:58:53.600000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3035323": "I am stuck with this problem statement and would like to understand the initial approach to tackle such a problem.\n\n- There are many labels missing. How can I train my model based on the remaining data?\n\n- The number of columns or features differs between the Test and Train datasets. How can I incorporate this? How can I predict the Test set without the complete set of features used in the Train set?\n\n- There is a set of data in “parquet” format, available for only a subset of IDs. I thought of concatenating the parquet data with the CSV data, but this is not possible as some IDs do not have parquet data since the wrist accelerometer was worn by selected participants only. What should I do?\n\nI am new to these types of problems, so please help me learn from this situation.",
    "3037283": "1. Just train on the data where labels are present (don't concentrate on what is missing).\n2. Features are the same in train and test, it's just the fact that you interpret as features that many PCIAT 0-20 columns that in fact are labels and in sum result in PCIAT Total that in it's turn is grouped in 4 classes that are represented in 'sii' column. Read carefully the overview and data sections.\n3. Acelerometer data should be aggregated to turn a long format into a wide format and concatenate with corresponding ID's where present, and where not present then we have NaNs.\n\n",
    "3036281": ""
  }
}