{
  "id": 468909,
  "title": "【Q】What's the difference between the 'XXX_sub_id's?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/468909",
  "author_name": "",
  "post_date": "2024-01-18T10:15:47.845881800Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>For example, for the same patient_id, maybe have several rows in train.csv, which have different XXX_sub_id but the same expert_concensus, i don't know what's the difference between them? And because of that, when we use all of that to build a dataset for training model, that might cause overfitting during the training process because there are a lot of overlapping in the dataset! Any sharing would be appreciate!</p>",
  "messages": [
    {
      "id": "2607578",
      "postDate": "01/18/2024 10:15:47",
      "content": "<p>For example, for the same patient_id, maybe have several rows in train.csv, which have different XXX_sub_id but the same expert_concensus, i don't know what's the difference between them? And because of that, when we use all of that to build a dataset for training model, that might cause overfitting during the training process because there are a lot of overlapping in the dataset! Any sharing would be appreciate!</p>",
      "rawMarkdown": "For example, for the same patient_id, maybe have several rows in train.csv, which have different XXX_sub_id but the same expert_concensus, i don't know what's the difference between them? And because of that, when we use all of that to build a dataset for training model, that might cause overfitting during the training process because there are a lot of overlapping in the dataset! Any sharing would be appreciate!",
      "votes": null
    },
    {
      "id": "2607856",
      "postDate": "01/18/2024 13:34:42",
      "content": "<p><a href=\"https://www.kaggle.com/suanbaicai\" target=\"_blank\">@suanbaicai</a> go through this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468208\" target=\"_blank\">discussion</a> &amp; <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">discussion</a></p>",
      "rawMarkdown": "suanbaicai go through this [discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468208) & [discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010)",
      "votes": null
    },
    {
      "id": "2609234",
      "postDate": "01/19/2024 10:31:45",
      "content": "<p>First of all, Thanks for the reply!<br>\nYes, I've seen these two discussion, but two point still remain confused:<br>\n1.how does the label_id comes out, for example, for eeg_id=1628180742&amp;eeg_sub_id=0, we can get label_id=127492639, why 127492639? not 3.1415926 or sth else. Is it the time location of the 10-sec signal or sth else?<br>\n2.why just use 10-sec signal in the middle to predict, i learned the fact throught the discussion, but don't know why</p>",
      "rawMarkdown": "First of all, Thanks for the reply!\nYes, I've seen these two discussion, but two point still remain confused:\n1.how does the label_id comes out, for example, for eeg_id=1628180742&eeg_sub_id=0, we can get label_id=127492639, why 127492639? not 3.1415926 or sth else. Is it the time location of the 10-sec signal or sth else?\n2.why just use 10-sec signal in the middle to predict, i learned the fact throught the discussion, but don't know why",
      "votes": null
    },
    {
      "id": "2610216",
      "postDate": "01/20/2024 02:19:32",
      "content": "<p>how does the label_id comes out, for example, for eeg_id=1628180742&amp;eeg_sub_id=0, we can get label_id=127492639, why 127492639? not 3.1415926 or sth else. Is it the time location of the 10-sec signal or sth else?</p>\n<blockquote>\n  <p>label_id not sure, but # of label_id is same as # of rows in train.csv ( mostly my guess is label_id is unique for the 10secs seq pattern )</p>\n</blockquote>\n<p>why just use 10-sec signal in the middle to predict, i learned the fact throught the discussion, but don't know why</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/suanbaicai\" target=\"_blank\">@suanbaicai</a> Experts focus on 10-sec while voting the targets even they see 50secs of eeg</p>\n</blockquote>",
      "rawMarkdown": "how does the label_id comes out, for example, for eeg_id=1628180742&eeg_sub_id=0, we can get label_id=127492639, why 127492639? not 3.1415926 or sth else. Is it the time location of the 10-sec signal or sth else?\n\n> label_id not sure, but # of label_id is same as # of rows in train.csv ( mostly my guess is label_id is unique for the 10secs seq pattern )\n\nwhy just use 10-sec signal in the middle to predict, i learned the fact throught the discussion, but don't know why\n\n> @suanbaicai Experts focus on 10-sec while voting the targets even they see 50secs of eeg",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2607856,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "01/18/2024 13:34:42",
      "content": "<p><a href=\"https://www.kaggle.com/suanbaicai\" target=\"_blank\">@suanbaicai</a> go through this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468208\" target=\"_blank\">discussion</a> &amp; <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">discussion</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2609234,
          "author_name": "suanbaicai",
          "author_url": "",
          "post_date": "01/19/2024 10:31:45",
          "content": "<p>First of all, Thanks for the reply!<br>\nYes, I've seen these two discussion, but two point still remain confused:<br>\n1.how does the label_id comes out, for example, for eeg_id=1628180742&amp;eeg_sub_id=0, we can get label_id=127492639, why 127492639? not 3.1415926 or sth else. Is it the time location of the 10-sec signal or sth else?<br>\n2.why just use 10-sec signal in the middle to predict, i learned the fact throught the discussion, but don't know why</p>",
          "votes": null,
          "replies": [
            {
              "id": 2610216,
              "author_name": "seshurajup",
              "author_url": "",
              "post_date": "01/20/2024 02:19:32",
              "content": "<p>how does the label_id comes out, for example, for eeg_id=1628180742&amp;eeg_sub_id=0, we can get label_id=127492639, why 127492639? not 3.1415926 or sth else. Is it the time location of the 10-sec signal or sth else?</p>\n<blockquote>\n  <p>label_id not sure, but # of label_id is same as # of rows in train.csv ( mostly my guess is label_id is unique for the 10secs seq pattern )</p>\n</blockquote>\n<p>why just use 10-sec signal in the middle to predict, i learned the fact throught the discussion, but don't know why</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/suanbaicai\" target=\"_blank\">@suanbaicai</a> Experts focus on 10-sec while voting the targets even they see 50secs of eeg</p>\n</blockquote>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2607578": "For example, for the same patient_id, maybe have several rows in train.csv, which have different XXX_sub_id but the same expert_concensus, i don't know what's the difference between them? And because of that, when we use all of that to build a dataset for training model, that might cause overfitting during the training process because there are a lot of overlapping in the dataset! Any sharing would be appreciate!",
    "2607856": "suanbaicai go through this [discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468208) & [discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010)",
    "2609234": "First of all, Thanks for the reply!\nYes, I've seen these two discussion, but two point still remain confused:\n1.how does the label_id comes out, for example, for eeg_id=1628180742&eeg_sub_id=0, we can get label_id=127492639, why 127492639? not 3.1415926 or sth else. Is it the time location of the 10-sec signal or sth else?\n2.why just use 10-sec signal in the middle to predict, i learned the fact throught the discussion, but don't know why",
    "2610216": "how does the label_id comes out, for example, for eeg_id=1628180742&eeg_sub_id=0, we can get label_id=127492639, why 127492639? not 3.1415926 or sth else. Is it the time location of the 10-sec signal or sth else?\n\n> label_id not sure, but # of label_id is same as # of rows in train.csv ( mostly my guess is label_id is unique for the 10secs seq pattern )\n\nwhy just use 10-sec signal in the middle to predict, i learned the fact throught the discussion, but don't know why\n\n> @suanbaicai Experts focus on 10-sec while voting the targets even they see 50secs of eeg"
  },
  "source": "meta"
}