{
  "id": 471890,
  "title": "LB probing results in HMS-HBAC",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/471890",
  "author_name": "",
  "post_date": "2024-01-30T01:13:41.079569300Z",
  "votes": 85,
  "comment_count": 9,
  "views": 0,
  "content": "<p>In <a href=\"https://www.kaggle.com/code/tomooinubushi/lb-probing-notebook-for-hms\" target=\"_blank\">this notebook</a>, I tested some basic assumptions about test dataset.<br>\nI demonstrated that following assumptions are all TRUE.</p>\n<ul>\n<li>There are no overlap of EEG IDs in test set.</li>\n<li>There are no overlap of spectrogram IDs in test set.</li>\n<li>There are overlap of patient IDs in test set.</li>\n<li>EEG IDs in test csv are the same as those in test_eeg dir.</li>\n<li>Spectrogram IDs in test csv are the same as those in test_spectrogram dir.</li>\n<li>5 &lt; mean number of spectrogram/eeg per patient &lt; 6 (8.76 spectrogram/patient and 5.71 eeg/patient in train set)</li>\n<li>EEG IDs of train and test set do not overlap.</li>\n<li>Spectrogram IDs of train and test set do not overlap.</li>\n<li>Patient IDs of train and test set do not overlap.</li>\n<li>Shapes of EEGs are all [20,10000].</li>\n<li>Shapes of spectrograms are all [300, 401].</li>\n<li>2.5% &lt; mean ratio of nan data in test spectrogram &lt; 3% (1.82% in train set)</li>\n<li>There are no nan in EEG data (0.005% in train set).</li>\n</ul>\n<p>I am happy if anyone correct me if I am wrong.<br>\nI am also very happy if anyone share us other assumptions/hypothesis about test dataset.</p>\n<p>See also</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471287\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471287</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471044\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471044</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021</a></li>\n</ul>",
  "messages": [
    {
      "id": "2626387",
      "postDate": "01/30/2024 01:13:41",
      "content": "<p>In <a href=\"https://www.kaggle.com/code/tomooinubushi/lb-probing-notebook-for-hms\" target=\"_blank\">this notebook</a>, I tested some basic assumptions about test dataset.<br>\nI demonstrated that following assumptions are all TRUE.</p>\n<ul>\n<li>There are no overlap of EEG IDs in test set.</li>\n<li>There are no overlap of spectrogram IDs in test set.</li>\n<li>There are overlap of patient IDs in test set.</li>\n<li>EEG IDs in test csv are the same as those in test_eeg dir.</li>\n<li>Spectrogram IDs in test csv are the same as those in test_spectrogram dir.</li>\n<li>5 &lt; mean number of spectrogram/eeg per patient &lt; 6 (8.76 spectrogram/patient and 5.71 eeg/patient in train set)</li>\n<li>EEG IDs of train and test set do not overlap.</li>\n<li>Spectrogram IDs of train and test set do not overlap.</li>\n<li>Patient IDs of train and test set do not overlap.</li>\n<li>Shapes of EEGs are all [20,10000].</li>\n<li>Shapes of spectrograms are all [300, 401].</li>\n<li>2.5% &lt; mean ratio of nan data in test spectrogram &lt; 3% (1.82% in train set)</li>\n<li>There are no nan in EEG data (0.005% in train set).</li>\n</ul>\n<p>I am happy if anyone correct me if I am wrong.<br>\nI am also very happy if anyone share us other assumptions/hypothesis about test dataset.</p>\n<p>See also</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471287\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471287</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471044\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471044</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021</a></li>\n</ul>",
      "rawMarkdown": "In [this notebook](https://www.kaggle.com/code/tomooinubushi/lb-probing-notebook-for-hms), I tested some basic assumptions about test dataset.\nI demonstrated that following assumptions are all TRUE.\n\n- There are no overlap of EEG IDs in test set.\n- There are no overlap of spectrogram IDs in test set.\n- There are overlap of patient IDs in test set.\n- EEG IDs in test csv are the same as those in test_eeg dir.\n- Spectrogram IDs in test csv are the same as those in test_spectrogram dir.\n- 5 < mean number of spectrogram/eeg per patient < 6 (8.76 spectrogram/patient and 5.71 eeg/patient in train set)\n- EEG IDs of train and test set do not overlap.\n- Spectrogram IDs of train and test set do not overlap.\n- Patient IDs of train and test set do not overlap.\n- Shapes of EEGs are all [20,10000].\n- Shapes of spectrograms are all [300, 401].\n- 2.5% < mean ratio of nan data in test spectrogram < 3% (1.82% in train set)\n- There are no nan in EEG data (0.005% in train set).\n\nI am happy if anyone correct me if I am wrong.\nI am also very happy if anyone share us other assumptions/hypothesis about test dataset.\n\nSee also\n- https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471287\n- https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471044\n- https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021",
      "votes": null
    },
    {
      "id": "2626388",
      "postDate": "01/30/2024 01:26:52",
      "content": "<p><a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> Thanks for sharing, so valuable details. From this competition, i learn how to understand private dataset by probing.</p>\n<pre><code>\nhypotheses.append((test.eeg_id.unique()) == (test))\n\nhypotheses.append((test.spectrogram_id.unique()) == (test))\n\nhypotheses.append((test.patient_id.unique()) != (test))\n</code></pre>\n<blockquote>\n  <p>Can we get no of patient ID in private dataset from above analysis</p>\n  <ol>\n  <li>As we know already no of EEG's sub-sequences ~ 2640</li>\n  <li>There are no overlap of EEG IDs in test set. =&gt; ~ 2640</li>\n  <li>There are no overlap of spectrogram IDs in test set. =&gt; ~ 2640</li>\n  <li>There are overlap of patient IDs in test set. =&gt; ?</li>\n  <li>mean number of spectrogram/eeg per patient ~ [5 to 6] ~ 5.5</li>\n  </ol>\n</blockquote>\n<hr>\n<blockquote>\n  <p>Apprx No of patient ID   ~  2640/5.5 ~ <strong>480</strong> ?</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>Approx No of patent IDs in public LB   ~ 924/5.5 = <strong>168</strong> ?</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>Approx No of patent IDs in  private LB ~ 1716/5.5 = <strong>312</strong> ?</p>\n</blockquote>",
      "rawMarkdown": "tomooinubushi Thanks for sharing, so valuable details. From this competition, i learn how to understand private dataset by probing.\n\n```python\n#there are no overlap of eeg ids in test set\nhypotheses.append(len(test.eeg_id.unique()) == len(test))\n#there are no overlap of spectrogram ids in test set\nhypotheses.append(len(test.spectrogram_id.unique()) == len(test))\n#there are overlap of patient ids in test set\nhypotheses.append(len(test.patient_id.unique()) != len(test))\n```\n> Can we get no of patient ID in private dataset from above analysis\n> 1. As we know already no of EEG's sub-sequences ~ 2640\n> 2. There are no overlap of EEG IDs in test set. => ~ 2640\n> 3. There are no overlap of spectrogram IDs in test set. => ~ 2640\n> 4. There are overlap of patient IDs in test set. => ?\n> 5. mean number of spectrogram/eeg per patient ~ [5 to 6] ~ 5.5\n\n---\n> Apprx No of patient ID   ~  2640/5.5 ~ **480** ?\n\n---\n> Approx No of patent IDs in public LB   ~ 924/5.5 = **168** ?\n\n---\n>  Approx No of patent IDs in  private LB ~ 1716/5.5 = **312** ?",
      "votes": null
    },
    {
      "id": "2626443",
      "postDate": "01/30/2024 02:50:32",
      "content": "<p>Great work. Thanks for probing!</p>",
      "rawMarkdown": "Great work. Thanks for probing!",
      "votes": null
    },
    {
      "id": "2626525",
      "postDate": "01/30/2024 05:09:12",
      "content": "<p>Thank you for your comment. <br>\nAs the number of patients in test set is a bit small, I expect some shake.</p>",
      "rawMarkdown": "Thank you for your comment. \nAs the number of patients in test set is a bit small, I expect some shake.",
      "votes": null
    },
    {
      "id": "2631147",
      "postDate": "02/01/2024 16:36:38",
      "content": "<p>Very good, curious to understand what insights might emerge from this information about overlapping patient IDs.</p>",
      "rawMarkdown": "Very good, curious to understand what insights might emerge from this information about overlapping patient IDs.",
      "votes": null
    },
    {
      "id": "2631844",
      "postDate": "02/02/2024 01:48:07",
      "content": "<p>It is all up to you.<br>\nFor example, <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">one of the most important public notebook</a> used non-overlapping EEG data for training and validation, but tolerated overlap of patients (see cell 4). Related discussion is <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "It is all up to you.\nFor example, [one of the most important public notebook](https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43) used non-overlapping EEG data for training and validation, but tolerated overlap of patients (see cell 4). Related discussion is [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021).",
      "votes": null
    },
    {
      "id": "2632645",
      "postDate": "02/02/2024 14:15:25",
      "content": "<p>Very good discussion, it's indeed a topic that can bring advantages in modeling or submission. I'll do my studies, thanks for responding.</p>",
      "rawMarkdown": "Very good discussion, it's indeed a topic that can bring advantages in modeling or submission. I'll do my studies, thanks for responding.",
      "votes": null
    },
    {
      "id": "2657059",
      "postDate": "02/18/2024 09:04:09",
      "content": "<p>I wonder about higher number of spectograms vs EEGs. From Chris <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">first place discussion</a> I thought that at least in train the relationship between spectograms and EEGs is 1-N, here there would be EEGs that have multiple spectograms right?</p>\n<p>Also thanks for the probing!</p>",
      "rawMarkdown": "I wonder about higher number of spectograms vs EEGs. From Chris [first place discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010) I thought that at least in train the relationship between spectograms and EEGs is 1-N, here there would be EEGs that have multiple spectograms right?\n\nAlso thanks for the probing!",
      "votes": null
    },
    {
      "id": "2657114",
      "postDate": "02/18/2024 10:08:29",
      "content": "<p>Thank you for your comment. <br>\nThe relationship between eeg_ids and spectrogram_ids in test set is one to one. <br>\nIt is different from train set.</p>\n<blockquote>\n  <p>len(test.eeg_id.unique()) == len(test)<br>\n  len(test.spectrogram_id.unique()) == len(test)</p>\n</blockquote>",
      "rawMarkdown": "Thank you for your comment. \nThe relationship between eeg_ids and spectrogram_ids in test set is one to one. \nIt is different from train set.\n>len(test.eeg_id.unique()) == len(test)\n>len(test.spectrogram_id.unique()) == len(test)",
      "votes": null
    },
    {
      "id": "2716743",
      "postDate": "03/26/2024 07:30:12",
      "content": "<p>Great work!!!</p>",
      "rawMarkdown": "Great work!!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2626388,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "01/30/2024 01:26:52",
      "content": "<p><a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> Thanks for sharing, so valuable details. From this competition, i learn how to understand private dataset by probing.</p>\n<pre><code>\nhypotheses.append((test.eeg_id.unique()) == (test))\n\nhypotheses.append((test.spectrogram_id.unique()) == (test))\n\nhypotheses.append((test.patient_id.unique()) != (test))\n</code></pre>\n<blockquote>\n  <p>Can we get no of patient ID in private dataset from above analysis</p>\n  <ol>\n  <li>As we know already no of EEG's sub-sequences ~ 2640</li>\n  <li>There are no overlap of EEG IDs in test set. =&gt; ~ 2640</li>\n  <li>There are no overlap of spectrogram IDs in test set. =&gt; ~ 2640</li>\n  <li>There are overlap of patient IDs in test set. =&gt; ?</li>\n  <li>mean number of spectrogram/eeg per patient ~ [5 to 6] ~ 5.5</li>\n  </ol>\n</blockquote>\n<hr>\n<blockquote>\n  <p>Apprx No of patient ID   ~  2640/5.5 ~ <strong>480</strong> ?</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>Approx No of patent IDs in public LB   ~ 924/5.5 = <strong>168</strong> ?</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>Approx No of patent IDs in  private LB ~ 1716/5.5 = <strong>312</strong> ?</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 2626525,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "01/30/2024 05:09:12",
          "content": "<p>Thank you for your comment. <br>\nAs the number of patients in test set is a bit small, I expect some shake.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2626443,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/30/2024 02:50:32",
      "content": "<p>Great work. Thanks for probing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2631147,
      "author_name": "rafaelzimmermann1",
      "author_url": "",
      "post_date": "02/01/2024 16:36:38",
      "content": "<p>Very good, curious to understand what insights might emerge from this information about overlapping patient IDs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2631844,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "02/02/2024 01:48:07",
          "content": "<p>It is all up to you.<br>\nFor example, <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">one of the most important public notebook</a> used non-overlapping EEG data for training and validation, but tolerated overlap of patients (see cell 4). Related discussion is <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021\" target=\"_blank\">here</a>.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2632645,
              "author_name": "rafaelzimmermann1",
              "author_url": "",
              "post_date": "02/02/2024 14:15:25",
              "content": "<p>Very good discussion, it's indeed a topic that can bring advantages in modeling or submission. I'll do my studies, thanks for responding.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2657059,
      "author_name": "raki21",
      "author_url": "",
      "post_date": "02/18/2024 09:04:09",
      "content": "<p>I wonder about higher number of spectograms vs EEGs. From Chris <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010\" target=\"_blank\">first place discussion</a> I thought that at least in train the relationship between spectograms and EEGs is 1-N, here there would be EEGs that have multiple spectograms right?</p>\n<p>Also thanks for the probing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2657114,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "02/18/2024 10:08:29",
          "content": "<p>Thank you for your comment. <br>\nThe relationship between eeg_ids and spectrogram_ids in test set is one to one. <br>\nIt is different from train set.</p>\n<blockquote>\n  <p>len(test.eeg_id.unique()) == len(test)<br>\n  len(test.spectrogram_id.unique()) == len(test)</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2716743,
      "author_name": "psarkar1425",
      "author_url": "",
      "post_date": "03/26/2024 07:30:12",
      "content": "<p>Great work!!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2626387": "In [this notebook](https://www.kaggle.com/code/tomooinubushi/lb-probing-notebook-for-hms), I tested some basic assumptions about test dataset.\nI demonstrated that following assumptions are all TRUE.\n\n- There are no overlap of EEG IDs in test set.\n- There are no overlap of spectrogram IDs in test set.\n- There are overlap of patient IDs in test set.\n- EEG IDs in test csv are the same as those in test_eeg dir.\n- Spectrogram IDs in test csv are the same as those in test_spectrogram dir.\n- 5 < mean number of spectrogram/eeg per patient < 6 (8.76 spectrogram/patient and 5.71 eeg/patient in train set)\n- EEG IDs of train and test set do not overlap.\n- Spectrogram IDs of train and test set do not overlap.\n- Patient IDs of train and test set do not overlap.\n- Shapes of EEGs are all [20,10000].\n- Shapes of spectrograms are all [300, 401].\n- 2.5% < mean ratio of nan data in test spectrogram < 3% (1.82% in train set)\n- There are no nan in EEG data (0.005% in train set).\n\nI am happy if anyone correct me if I am wrong.\nI am also very happy if anyone share us other assumptions/hypothesis about test dataset.\n\nSee also\n- https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471287\n- https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471044\n- https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021",
    "2626388": "tomooinubushi Thanks for sharing, so valuable details. From this competition, i learn how to understand private dataset by probing.\n\n```python\n#there are no overlap of eeg ids in test set\nhypotheses.append(len(test.eeg_id.unique()) == len(test))\n#there are no overlap of spectrogram ids in test set\nhypotheses.append(len(test.spectrogram_id.unique()) == len(test))\n#there are overlap of patient ids in test set\nhypotheses.append(len(test.patient_id.unique()) != len(test))\n```\n> Can we get no of patient ID in private dataset from above analysis\n> 1. As we know already no of EEG's sub-sequences ~ 2640\n> 2. There are no overlap of EEG IDs in test set. => ~ 2640\n> 3. There are no overlap of spectrogram IDs in test set. => ~ 2640\n> 4. There are overlap of patient IDs in test set. => ?\n> 5. mean number of spectrogram/eeg per patient ~ [5 to 6] ~ 5.5\n\n---\n> Apprx No of patient ID   ~  2640/5.5 ~ **480** ?\n\n---\n> Approx No of patent IDs in public LB   ~ 924/5.5 = **168** ?\n\n---\n>  Approx No of patent IDs in  private LB ~ 1716/5.5 = **312** ?",
    "2626443": "Great work. Thanks for probing!",
    "2626525": "Thank you for your comment. \nAs the number of patients in test set is a bit small, I expect some shake.",
    "2631147": "Very good, curious to understand what insights might emerge from this information about overlapping patient IDs.",
    "2631844": "It is all up to you.\nFor example, [one of the most important public notebook](https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43) used non-overlapping EEG data for training and validation, but tolerated overlap of patients (see cell 4). Related discussion is [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021).",
    "2632645": "Very good discussion, it's indeed a topic that can bring advantages in modeling or submission. I'll do my studies, thanks for responding.",
    "2657059": "I wonder about higher number of spectograms vs EEGs. From Chris [first place discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468010) I thought that at least in train the relationship between spectograms and EEGs is 1-N, here there would be EEGs that have multiple spectograms right?\n\nAlso thanks for the probing!",
    "2657114": "Thank you for your comment. \nThe relationship between eeg_ids and spectrogram_ids in test set is one to one. \nIt is different from train set.\n>len(test.eeg_id.unique()) == len(test)\n>len(test.spectrogram_id.unique()) == len(test)",
    "2716743": "Great work!!!"
  },
  "source": "meta"
}