{
  "id": 546966,
  "title": "General Guidance",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/546966",
  "author_name": "",
  "post_date": "2024-11-19T05:10:00.312871900Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello, I wanted to ask about a bit of general guidance for this competition. What is the relationship between the parquet files and the csv files? Additionally, I wanted to get a bit of advice on how to organize the data in such a way that features can be selected to use for models (these features coming from the actigraphy data as well as the csv data files). Thanks.</p>",
  "messages": [
    {
      "id": "3049411",
      "postDate": "11/19/2024 05:10:00",
      "content": "<p>Hello, I wanted to ask about a bit of general guidance for this competition. What is the relationship between the parquet files and the csv files? Additionally, I wanted to get a bit of advice on how to organize the data in such a way that features can be selected to use for models (these features coming from the actigraphy data as well as the csv data files). Thanks.</p>",
      "rawMarkdown": "Hello, I wanted to ask about a bit of general guidance for this competition. What is the relationship between the parquet files and the csv files? Additionally, I wanted to get a bit of advice on how to organize the data in such a way that features can be selected to use for models (these features coming from the actigraphy data as well as the csv data files). Thanks.",
      "votes": null
    },
    {
      "id": "3049563",
      "postDate": "11/19/2024 09:15:48",
      "content": "<p>So for some participants there exists actigraph data, which is stored in the parquet files. You can connect them to the csv by their 'id' because it matches the 'id' in the csv.<br>\nThe parquet files are timeseries. This means it is not adviced to just put the raw data into your csv df, since you would need alot of features to store it and your model will have a hard time to know what's going on. Since the information is how the values changes over time and not how it is compared to other participants at \"step x\".<br>\nThe approach I know is that you try to transform the timeseries information for a feature into a single values, like mean. That's why alot of the public notebooks use the \".describe()\" method and save those values into their csv df.<br>\nI'm also using that but want to improve on it.</p>\n<p>My biggest issue are outliers. As some forum posts suggest, some actigraph data is just nonsense and I want to remove it, but I also dont want to go through 1000 actigraph datasets hand by hand. What's a good approach to do that?<br>\nI tried detecting outliers with the .describe() values, but I don't think that's the best way.</p>",
      "rawMarkdown": "So for some participants there exists actigraph data, which is stored in the parquet files. You can connect them to the csv by their 'id' because it matches the 'id' in the csv.\nThe parquet files are timeseries. This means it is not adviced to just put the raw data into your csv df, since you would need alot of features to store it and your model will have a hard time to know what's going on. Since the information is how the values changes over time and not how it is compared to other participants at \"step x\".\nThe approach I know is that you try to transform the timeseries information for a feature into a single values, like mean. That's why alot of the public notebooks use the \".describe()\" method and save those values into their csv df.\nI'm also using that but want to improve on it.\n\nMy biggest issue are outliers. As some forum posts suggest, some actigraph data is just nonsense and I want to remove it, but I also dont want to go through 1000 actigraph datasets hand by hand. What's a good approach to do that?\nI tried detecting outliers with the .describe() values, but I don't think that's the best way.",
      "votes": null
    },
    {
      "id": "3052135",
      "postDate": "11/22/2024 05:14:37",
      "content": "<p>Thanks for the feedback and the tips. I will check out the data and see if I find anything interesting regarding outlier values. </p>",
      "rawMarkdown": "Thanks for the feedback and the tips. I will check out the data and see if I find anything interesting regarding outlier values.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3049563,
      "author_name": "mariusheuser",
      "author_url": "",
      "post_date": "11/19/2024 09:15:48",
      "content": "<p>So for some participants there exists actigraph data, which is stored in the parquet files. You can connect them to the csv by their 'id' because it matches the 'id' in the csv.<br>\nThe parquet files are timeseries. This means it is not adviced to just put the raw data into your csv df, since you would need alot of features to store it and your model will have a hard time to know what's going on. Since the information is how the values changes over time and not how it is compared to other participants at \"step x\".<br>\nThe approach I know is that you try to transform the timeseries information for a feature into a single values, like mean. That's why alot of the public notebooks use the \".describe()\" method and save those values into their csv df.<br>\nI'm also using that but want to improve on it.</p>\n<p>My biggest issue are outliers. As some forum posts suggest, some actigraph data is just nonsense and I want to remove it, but I also dont want to go through 1000 actigraph datasets hand by hand. What's a good approach to do that?<br>\nI tried detecting outliers with the .describe() values, but I don't think that's the best way.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3052135,
          "author_name": "bransone",
          "author_url": "",
          "post_date": "11/22/2024 05:14:37",
          "content": "<p>Thanks for the feedback and the tips. I will check out the data and see if I find anything interesting regarding outlier values. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3049411": "Hello, I wanted to ask about a bit of general guidance for this competition. What is the relationship between the parquet files and the csv files? Additionally, I wanted to get a bit of advice on how to organize the data in such a way that features can be selected to use for models (these features coming from the actigraphy data as well as the csv data files). Thanks.",
    "3049563": "So for some participants there exists actigraph data, which is stored in the parquet files. You can connect them to the csv by their 'id' because it matches the 'id' in the csv.\nThe parquet files are timeseries. This means it is not adviced to just put the raw data into your csv df, since you would need alot of features to store it and your model will have a hard time to know what's going on. Since the information is how the values changes over time and not how it is compared to other participants at \"step x\".\nThe approach I know is that you try to transform the timeseries information for a feature into a single values, like mean. That's why alot of the public notebooks use the \".describe()\" method and save those values into their csv df.\nI'm also using that but want to improve on it.\n\nMy biggest issue are outliers. As some forum posts suggest, some actigraph data is just nonsense and I want to remove it, but I also dont want to go through 1000 actigraph datasets hand by hand. What's a good approach to do that?\nI tried detecting outliers with the .describe() values, but I don't think that's the best way.",
    "3052135": "Thanks for the feedback and the tips. I will check out the data and see if I find anything interesting regarding outlier values."
  },
  "source": "meta"
}