{
  "id": 541185,
  "title": " About series_train.parquet",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/541185",
  "author_name": "Arith",
  "post_date": "2024-10-18T04:43:27.138000",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>This is a beginner's question.<br>\np1_train_ts = load_time_series(\"../input/child-mind-institute-problematic-internet-use/series_train.parquet\")<br>\nWhat does \"stat_0~stat_95\" mean when performing</p>",
  "messages": [
    {
      "id": 3021257,
      "postDate": "2024-10-18T10:10:06.897Z",
      "content": "<p>What you're interested in happens in this section:</p>\n<pre><code> ():\n    df = pd.read_parquet(os.path.join(dirname, filename, ))\n    df.drop(, axis=, inplace=)\n     df.describe().values.reshape(-), filename.split()[]\n</code></pre>\n<p>The original time-series file has columns such as \"X\", \"Y\", \"Z\", \"enmo\", \"anglez\" and so on. For more information look at the data section or load one of the parquet files. The function above, specifically df.describe().values.reshape(-1), gives you some summary-statistics for each column of the time-series file and returns them as an array. <br>\nThe names for the columns are then assigned in:</p>\n<p><code>df = pd.DataFrame(stats, columns=[f\"Stat_{i}\" for i in range(len(stats[0]))])</code></p>\n<p>You may change this part to give you more informative column-names.</p>",
      "rawMarkdown": "What you're interested in happens in this section:\n\n```python\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop('step', axis=1, inplace=True)\n    return df.describe().values.reshape(-1), filename.split('=')[1]\n```\n\nThe original time-series file has columns such as \"X\", \"Y\", \"Z\", \"enmo\", \"anglez\" and so on. For more information look at the data section or load one of the parquet files. The function above, specifically df.describe().values.reshape(-1), gives you some summary-statistics for each column of the time-series file and returns them as an array. \nThe names for the columns are then assigned in:\n\n`df = pd.DataFrame(stats, columns=[f\"Stat_{i}\" for i in range(len(stats[0]))])`\n\nYou may change this part to give you more informative column-names.",
      "votes": 1,
      "replies": [
        {
          "id": 3025603,
          "postDate": "2024-10-23T01:36:22.770Z",
          "content": "<p>返信が遅くなり申し訳ありません<br>\n。説明が明確で、データを理解することができました。</p>",
          "rawMarkdown": "返信が遅くなり申し訳ありません\n。説明が明確で、データを理解することができました。"
        }
      ]
    },
    {
      "id": 3020947,
      "postDate": "2024-10-18T04:43:27.140Z",
      "content": "<p>This is a beginner's question.<br>\np1_train_ts = load_time_series(\"../input/child-mind-institute-problematic-internet-use/series_train.parquet\")<br>\nWhat does \"stat_0~stat_95\" mean when performing</p>",
      "rawMarkdown": "This is a beginner's question.\np1_train_ts = load_time_series(\"../input/child-mind-institute-problematic-internet-use/series_train.parquet\")\nWhat does \"stat_0~stat_95\" mean when performing",
      "votes": 1
    },
    {
      "id": 3020963,
      "postDate": "2024-10-18T05:07:04.693Z",
      "content": "<p>stat_0 could be the minimum value (0th percentile), stat_50 the median (50th percentile), and stat_95 the 95th percentile.The notation \"stat_0~stat_95\" likely refers to a set of statistical features or columns in the dataset, possibly representing statistical values like mean, standard deviation, percentiles, etc., for different segments or features in the time series data.</p>",
      "rawMarkdown": "stat_0 could be the minimum value (0th percentile), stat_50 the median (50th percentile), and stat_95 the 95th percentile.The notation \"stat_0~stat_95\" likely refers to a set of statistical features or columns in the dataset, possibly representing statistical values like mean, standard deviation, percentiles, etc., for different segments or features in the time series data.\n\n",
      "replies": [
        {
          "id": 3020988,
          "postDate": "2024-10-18T05:45:56.763Z",
          "content": "<p>Thank you for your answer.<br>\nbut i still don't understand.<br>\nWhat are the minimum and median values ​​based on?<br>\nAlso, in the first line, stat_0~stat_11 are the same numbers, but I don't understand why the numbers change so much from stat12.</p>",
          "rawMarkdown": "Thank you for your answer.\nbut i still don't understand.\nWhat are the minimum and median values ​​based on?\nAlso, in the first line, stat_0~stat_11 are the same numbers, but I don't understand why the numbers change so much from stat12.",
          "replies": [
            {
              "id": 3021299,
              "postDate": "2024-10-18T10:56:24.800Z",
              "content": "<p>In the main data frame we have one row by 'id', but in series we have thousands rows of the same 'id'. To join these dataframes we must do something with the series data to have just one row for each 'id' like the main data frame. That is why we take for example the average of  each column and now we have one single row instead of thousands - this is called aggregation and the resulted columns you can call X_mean, Y_mean and so on, or like in your case stat_1, stat_2… now to squeeze more info we are applying multiple aggregations like min, max, std… and all these go into one long line with many columns of different statistics refering to one single 'id' if you are processing each file one by one</p>",
              "rawMarkdown": "In the main data frame we have one row by 'id', but in series we have thousands rows of the same 'id'. To join these dataframes we must do something with the series data to have just one row for each 'id' like the main data frame. That is why we take for example the average of  each column and now we have one single row instead of thousands - this is called aggregation and the resulted columns you can call X_mean, Y_mean and so on, or like in your case stat_1, stat_2... now to squeeze more info we are applying multiple aggregations like min, max, std... and all these go into one long line with many columns of different statistics refering to one single 'id' if you are processing each file one by one",
              "votes": 2
            },
            {
              "id": 3025599,
              "postDate": "2024-10-23T01:34:08.493Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3025601,
              "postDate": "2024-10-23T01:34:58.183Z",
              "content": "<p>Sorry for the late reply<br>\nYour explanation helped me understand the data.</p>",
              "rawMarkdown": "Sorry for the late reply\nYour explanation helped me understand the data."
            },
            {
              "id": 3025602,
              "postDate": "2024-10-23T01:35:41.470Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3021257,
      "author_name": "Lennart Haupts",
      "author_url": "",
      "post_date": "2024-10-18T10:10:06.897000",
      "content": "<p>What you're interested in happens in this section:</p>\n<pre><code> ():\n    df = pd.read_parquet(os.path.join(dirname, filename, ))\n    df.drop(, axis=, inplace=)\n     df.describe().values.reshape(-), filename.split()[]\n</code></pre>\n<p>The original time-series file has columns such as \"X\", \"Y\", \"Z\", \"enmo\", \"anglez\" and so on. For more information look at the data section or load one of the parquet files. The function above, specifically df.describe().values.reshape(-1), gives you some summary-statistics for each column of the time-series file and returns them as an array. <br>\nThe names for the columns are then assigned in:</p>\n<p><code>df = pd.DataFrame(stats, columns=[f\"Stat_{i}\" for i in range(len(stats[0]))])</code></p>\n<p>You may change this part to give you more informative column-names.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3025603,
          "author_name": "Arith",
          "author_url": "",
          "post_date": "2024-10-23T01:36:22.770000",
          "content": "<p>返信が遅くなり申し訳ありません<br>\n。説明が明確で、データを理解することができました。</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3020963,
      "author_name": "Sumit_08",
      "author_url": "",
      "post_date": "2024-10-18T05:07:04.693000",
      "content": "<p>stat_0 could be the minimum value (0th percentile), stat_50 the median (50th percentile), and stat_95 the 95th percentile.The notation \"stat_0~stat_95\" likely refers to a set of statistical features or columns in the dataset, possibly representing statistical values like mean, standard deviation, percentiles, etc., for different segments or features in the time series data.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3020988,
          "author_name": "Arith",
          "author_url": "",
          "post_date": "2024-10-18T05:45:56.763000",
          "content": "<p>Thank you for your answer.<br>\nbut i still don't understand.<br>\nWhat are the minimum and median values ​​based on?<br>\nAlso, in the first line, stat_0~stat_11 are the same numbers, but I don't understand why the numbers change so much from stat12.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3021299,
              "author_name": "Danu A.",
              "author_url": "",
              "post_date": "2024-10-18T10:56:24.800000",
              "content": "<p>In the main data frame we have one row by 'id', but in series we have thousands rows of the same 'id'. To join these dataframes we must do something with the series data to have just one row for each 'id' like the main data frame. That is why we take for example the average of  each column and now we have one single row instead of thousands - this is called aggregation and the resulted columns you can call X_mean, Y_mean and so on, or like in your case stat_1, stat_2… now to squeeze more info we are applying multiple aggregations like min, max, std… and all these go into one long line with many columns of different statistics refering to one single 'id' if you are processing each file one by one</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3025599,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-10-23T01:34:08.493000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3025601,
              "author_name": "Arith",
              "author_url": "",
              "post_date": "2024-10-23T01:34:58.183000",
              "content": "<p>Sorry for the late reply<br>\nYour explanation helped me understand the data.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3025602,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-10-23T01:35:41.470000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3021257": "What you're interested in happens in this section:\n\n```python\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop('step', axis=1, inplace=True)\n    return df.describe().values.reshape(-1), filename.split('=')[1]\n```\n\nThe original time-series file has columns such as \"X\", \"Y\", \"Z\", \"enmo\", \"anglez\" and so on. For more information look at the data section or load one of the parquet files. The function above, specifically df.describe().values.reshape(-1), gives you some summary-statistics for each column of the time-series file and returns them as an array. \nThe names for the columns are then assigned in:\n\n`df = pd.DataFrame(stats, columns=[f\"Stat_{i}\" for i in range(len(stats[0]))])`\n\nYou may change this part to give you more informative column-names.",
    "3020947": "This is a beginner's question.\np1_train_ts = load_time_series(\"../input/child-mind-institute-problematic-internet-use/series_train.parquet\")\nWhat does \"stat_0~stat_95\" mean when performing",
    "3020963": "stat_0 could be the minimum value (0th percentile), stat_50 the median (50th percentile), and stat_95 the 95th percentile.The notation \"stat_0~stat_95\" likely refers to a set of statistical features or columns in the dataset, possibly representing statistical values like mean, standard deviation, percentiles, etc., for different segments or features in the time series data.\n\n"
  }
}