{
  "id": 535355,
  "title": "Series Data",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/535355",
  "author_name": "",
  "post_date": "2024-09-21T17:15:15.527898500Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Dear colleagues, there are 996 series (train data records are 3960). First thing to note is that series data of 2964 (3960-996) patient is missing. </p>\n<p>On the average, each series comprises of 350000 records. Total records  of all series become  348,600,000 records. It is a kind of big data because when I tried to merge each patient data with the corresponding data, I experienced the message 'out of memory'. </p>\n<p>As mentioned in the data section of the competition page 'series_{train|test}.parquet/id={id} - Series to be used as training data, partitioned by id. Each series is a continuous recording of accelerometer data for a single subject spanning many days.', it should be used as train data.  What is your opinion? Is merging series with corresponding  train data is useful for good performance  of the model. </p>",
  "messages": [
    {
      "id": "2994959",
      "postDate": "09/21/2024 17:15:15",
      "content": "<p>Dear colleagues, there are 996 series (train data records are 3960). First thing to note is that series data of 2964 (3960-996) patient is missing. </p>\n<p>On the average, each series comprises of 350000 records. Total records  of all series become  348,600,000 records. It is a kind of big data because when I tried to merge each patient data with the corresponding data, I experienced the message 'out of memory'. </p>\n<p>As mentioned in the data section of the competition page 'series_{train|test}.parquet/id={id} - Series to be used as training data, partitioned by id. Each series is a continuous recording of accelerometer data for a single subject spanning many days.', it should be used as train data.  What is your opinion? Is merging series with corresponding  train data is useful for good performance  of the model. </p>",
      "rawMarkdown": "Dear colleagues, there are 996 series (train data records are 3960). First thing to note is that series data of 2964 (3960-996) patient is missing. \n\nOn the average, each series comprises of 350000 records. Total records  of all series become  348,600,000 records. It is a kind of big data because when I tried to merge each patient data with the corresponding data, I experienced the message 'out of memory'. \n\nAs mentioned in the data section of the competition page 'series_{train|test}.parquet/id={id} - Series to be used as training data, partitioned by id. Each series is a continuous recording of accelerometer data for a single subject spanning many days.', it should be used as train data.  What is your opinion? Is merging series with corresponding  train data is useful for good performance  of the model.",
      "votes": null
    },
    {
      "id": "2996085",
      "postDate": "09/23/2024 04:25:18",
      "content": "<p>I haven't looked at the data in detail yet, so this is just a guess.</p>\n<p>To merge into tabular data, it may be necessary to process time series data using methods such as aggregation. Otherwise, (train id count: 996) * (each series length: 350000) rows will be created.</p>",
      "rawMarkdown": "I haven't looked at the data in detail yet, so this is just a guess.\n\nTo merge into tabular data, it may be necessary to process time series data using methods such as aggregation. Otherwise, (train id count: 996) * (each series length: 350000) rows will be created.",
      "votes": null
    },
    {
      "id": "2996162",
      "postDate": "09/23/2024 06:30:03",
      "content": "<p>Good point. thanks <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> </p>",
      "rawMarkdown": "Good point. thanks @yuuniekiri",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2996085,
      "author_name": "yuuniekiri",
      "author_url": "",
      "post_date": "09/23/2024 04:25:18",
      "content": "<p>I haven't looked at the data in detail yet, so this is just a guess.</p>\n<p>To merge into tabular data, it may be necessary to process time series data using methods such as aggregation. Otherwise, (train id count: 996) * (each series length: 350000) rows will be created.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2996162,
          "author_name": "tariqcp",
          "author_url": "",
          "post_date": "09/23/2024 06:30:03",
          "content": "<p>Good point. thanks <a href=\"https://www.kaggle.com/yuuniekiri\" target=\"_blank\">@yuuniekiri</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2994959": "Dear colleagues, there are 996 series (train data records are 3960). First thing to note is that series data of 2964 (3960-996) patient is missing. \n\nOn the average, each series comprises of 350000 records. Total records  of all series become  348,600,000 records. It is a kind of big data because when I tried to merge each patient data with the corresponding data, I experienced the message 'out of memory'. \n\nAs mentioned in the data section of the competition page 'series_{train|test}.parquet/id={id} - Series to be used as training data, partitioned by id. Each series is a continuous recording of accelerometer data for a single subject spanning many days.', it should be used as train data.  What is your opinion? Is merging series with corresponding  train data is useful for good performance  of the model.",
    "2996085": "I haven't looked at the data in detail yet, so this is just a guess.\n\nTo merge into tabular data, it may be necessary to process time series data using methods such as aggregation. Otherwise, (train id count: 996) * (each series length: 350000) rows will be created.",
    "2996162": "Good point. thanks @yuuniekiri"
  },
  "source": "meta"
}