{
  "id": 486581,
  "title": "What does  eeg_label_offset_seconds show ? ",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/486581",
  "author_name": "",
  "post_date": "2024-03-25T15:24:54.299057400Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I understand that eeg_label_offset_seconds shows from where the data begins that is labeled and plays role on classifying. Now my question is why they did not provided EEG already cut by that offset size ?<br>\nwhy we need the rest in EEG file ? </p>",
  "messages": [
    {
      "id": "2715601",
      "postDate": "03/25/2024 15:24:54",
      "content": "<p>I understand that eeg_label_offset_seconds shows from where the data begins that is labeled and plays role on classifying. Now my question is why they did not provided EEG already cut by that offset size ?<br>\nwhy we need the rest in EEG file ? </p>",
      "rawMarkdown": "I understand that eeg_label_offset_seconds shows from where the data begins that is labeled and plays role on classifying. Now my question is why they did not provided EEG already cut by that offset size ?\nwhy we need the rest in EEG file ?",
      "votes": null
    },
    {
      "id": "2715620",
      "postDate": "03/25/2024 15:32:23",
      "content": "<p>Because many samples overlap. So a parquet for each sample would lead to more repeated data and memmory usage.</p>",
      "rawMarkdown": "Because many samples overlap. So a parquet for each sample would lead to more repeated data and memmory usage.",
      "votes": null
    },
    {
      "id": "2715635",
      "postDate": "03/25/2024 15:37:08",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18779914%2Fdf75ab46e5cacf758edba29c6b90462f%2FUntitled.png?generation=1711381026200934&amp;alt=media\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18779914%2Fdf75ab46e5cacf758edba29c6b90462f%2FUntitled.png?generation=1711381026200934&alt=media)",
      "votes": null
    },
    {
      "id": "2715637",
      "postDate": "03/25/2024 15:37:54",
      "content": "<p>I mean, this red part in that line, is not it unnecessary data ? in any case we take eeg and we cut it based on offset value, and the rest seems remains useless.</p>",
      "rawMarkdown": "I mean, this red part in that line, is not it unnecessary data ? in any case we take eeg and we cut it based on offset value, and the rest seems remains useless.",
      "votes": null
    },
    {
      "id": "2715644",
      "postDate": "03/25/2024 15:42:58",
      "content": "<p>Previous and posterior time of sample will be as unnecessary as your approach consider.<br>\nI was on my phone. You should check<br>\nUnderstanding Competition Data and EfficientNetB2 Starter - LB 0.43<br>\n<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010</a><br>\nAnd now on my computer to be more accurate. That red area can be the in sample window from another window. And even if is not the case it can still be useful. Again, depending on your approach.</p>",
      "rawMarkdown": "Previous and posterior time of sample will be as unnecessary as your approach consider.\n\nI was on my phone. You should check\n\nUnderstanding Competition Data and EfficientNetB2 Starter - LB 0.43\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\n\nAnd now on my computer to be more accurate. That red area can be the in sample window from another window. And even if is not the case it can still be useful. Again, depending on your approach.",
      "votes": null
    },
    {
      "id": "2715741",
      "postDate": "03/25/2024 16:49:06",
      "content": "<p>The variable eeg_label_offset_seconds is the beginning․ The file train_eegs/1628180742.parquet is 90 seconds long. Therefore when eeg_label_offset_seconds = 40, we use the 50 seconds between time 40 and time 90. These are rows 40<em>200 thru 90</em>200 of 1628180742.parquet. What happens the 90 - 50 = 40 seconds ? </p>",
      "rawMarkdown": "The variable eeg_label_offset_seconds is the beginning․ The file train_eegs/1628180742.parquet is 90 seconds long. Therefore when eeg_label_offset_seconds = 40, we use the 50 seconds between time 40 and time 90. These are rows 40*200 thru 90*200 of 1628180742.parquet. What happens the 90 - 50 = 40 seconds ?",
      "votes": null
    },
    {
      "id": "2715760",
      "postDate": "03/25/2024 17:00:58",
      "content": "<p>They are part of the time windows of previous offsets</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fc35e6f75de0be87f6293b8a97004e2a0%2Foffset.png?generation=1711386102690376&amp;alt=media\"></p>\n<p>You need to find and cut every sample. If they provided all cutted samples they would need a lot more memmory and it would contain many repeated data.</p>",
      "rawMarkdown": "They are part of the time windows of previous offsets\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fc35e6f75de0be87f6293b8a97004e2a0%2Foffset.png?generation=1711386102690376&alt=media)\n\nYou need to find and cut every sample. If they provided all cutted samples they would need a lot more memmory and it would contain many repeated data.",
      "votes": null
    },
    {
      "id": "2716893",
      "postDate": "03/26/2024 09:08:27",
      "content": "<p>Now I got it, thanks a lot for clarification</p>",
      "rawMarkdown": "Now I got it, thanks a lot for clarification",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2715620,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "03/25/2024 15:32:23",
      "content": "<p>Because many samples overlap. So a parquet for each sample would lead to more repeated data and memmory usage.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2715635,
          "author_name": "datevik",
          "author_url": "",
          "post_date": "03/25/2024 15:37:08",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18779914%2Fdf75ab46e5cacf758edba29c6b90462f%2FUntitled.png?generation=1711381026200934&amp;alt=media\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 2715637,
              "author_name": "datevik",
              "author_url": "",
              "post_date": "03/25/2024 15:37:54",
              "content": "<p>I mean, this red part in that line, is not it unnecessary data ? in any case we take eeg and we cut it based on offset value, and the rest seems remains useless.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2715644,
                  "author_name": "sacuscreed",
                  "author_url": "",
                  "post_date": "03/25/2024 15:42:58",
                  "content": "<p>Previous and posterior time of sample will be as unnecessary as your approach consider.<br>\nI was on my phone. You should check<br>\nUnderstanding Competition Data and EfficientNetB2 Starter - LB 0.43<br>\n<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010</a><br>\nAnd now on my computer to be more accurate. That red area can be the in sample window from another window. And even if is not the case it can still be useful. Again, depending on your approach.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2715741,
                      "author_name": "datevik",
                      "author_url": "",
                      "post_date": "03/25/2024 16:49:06",
                      "content": "<p>The variable eeg_label_offset_seconds is the beginning․ The file train_eegs/1628180742.parquet is 90 seconds long. Therefore when eeg_label_offset_seconds = 40, we use the 50 seconds between time 40 and time 90. These are rows 40<em>200 thru 90</em>200 of 1628180742.parquet. What happens the 90 - 50 = 40 seconds ? </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2715760,
                          "author_name": "sacuscreed",
                          "author_url": "",
                          "post_date": "03/25/2024 17:00:58",
                          "content": "<p>They are part of the time windows of previous offsets</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fc35e6f75de0be87f6293b8a97004e2a0%2Foffset.png?generation=1711386102690376&amp;alt=media\"></p>\n<p>You need to find and cut every sample. If they provided all cutted samples they would need a lot more memmory and it would contain many repeated data.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2716893,
                              "author_name": "datevik",
                              "author_url": "",
                              "post_date": "03/26/2024 09:08:27",
                              "content": "<p>Now I got it, thanks a lot for clarification</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2715601": "I understand that eeg_label_offset_seconds shows from where the data begins that is labeled and plays role on classifying. Now my question is why they did not provided EEG already cut by that offset size ?\nwhy we need the rest in EEG file ?",
    "2715620": "Because many samples overlap. So a parquet for each sample would lead to more repeated data and memmory usage.",
    "2715635": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18779914%2Fdf75ab46e5cacf758edba29c6b90462f%2FUntitled.png?generation=1711381026200934&alt=media)",
    "2715637": "I mean, this red part in that line, is not it unnecessary data ? in any case we take eeg and we cut it based on offset value, and the rest seems remains useless.",
    "2715644": "Previous and posterior time of sample will be as unnecessary as your approach consider.\n\nI was on my phone. You should check\n\nUnderstanding Competition Data and EfficientNetB2 Starter - LB 0.43\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\n\nAnd now on my computer to be more accurate. That red area can be the in sample window from another window. And even if is not the case it can still be useful. Again, depending on your approach.",
    "2715741": "The variable eeg_label_offset_seconds is the beginning․ The file train_eegs/1628180742.parquet is 90 seconds long. Therefore when eeg_label_offset_seconds = 40, we use the 50 seconds between time 40 and time 90. These are rows 40*200 thru 90*200 of 1628180742.parquet. What happens the 90 - 50 = 40 seconds ?",
    "2715760": "They are part of the time windows of previous offsets\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fc35e6f75de0be87f6293b8a97004e2a0%2Foffset.png?generation=1711386102690376&alt=media)\n\nYou need to find and cut every sample. If they provided all cutted samples they would need a lot more memmory and it would contain many repeated data.",
    "2716893": "Now I got it, thanks a lot for clarification"
  },
  "source": "meta"
}