{
  "id": 467021,
  "title": "EDA Train.csv",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/467021",
  "author_name": "",
  "post_date": "2024-01-10T21:12:58.205273100Z",
  "votes": 100,
  "comment_count": 27,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5108b87498942a37058ab1164bf8a8de%2FScreenshot%202024-01-11%20at%202.38.42AM.png?generation=1704920935509556&amp;alt=media\"><br>\n<strong>Well balanced targets</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fbef0e17513d4ea1c81654664b7df1720%2FScreenshot%202024-01-11%20at%202.36.39AM.png?generation=1704920849053906&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5098c5d109726d95a7ff26a8fd6456a9%2FScreenshot%202024-01-11%20at%202.59.21AM.png?generation=1704922187122299&amp;alt=media\"><br>\n<strong>Significant correlation between seizure_vote, other vote and grda_vote types ( all 3 targets are frequent targets )</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F20d7d0dfbc933068f13abc37b3881443%2FScreenshot%202024-01-11%20at%202.52.19AM.png?generation=1704921765311082&amp;alt=media\"><br>\n<strong>is &lt; 1000 offset is the major information of the targets except for gpd target ?</strong></p>\n<p>with Equal probability 1/6: <strong>LB 1.09</strong></p>\n<p>With Overlapping <strong>LB: 1.1</strong></p>\n<pre><code>seizure_vote  ,\ngpd_vote  ,\nlrda_vote ,\nother_vote ,\ngrda_vote ,\nlpd_vote \n</code></pre>\n<h2>Conclusion: Test set have means closer to Train Set</h2>\n<p>with non-overlapping based on spectrogram id<br>\n<strong>LB: 1.0</strong></p>\n<pre><code>seizure_vote    \nlpd_vote        \ngpd_vote        \nlrda_vote       \ngrda_vote       \nother_vote      \n</code></pre>\n<p>with non-overlapping based on eeg id<br>\n<strong>LB: 0.97</strong></p>\n<pre><code>seizure_vote    \nlpd_vote        \ngpd_vote        \nlrda_vote       \ngrda_vote       \nother_vote      \n</code></pre>\n<h2>Conclusion: Test set have non-overlapping egg based sequences</h2>\n<p>with non-overlapping based on p id<br>\n<strong>LB: 1.28</strong></p>\n<pre><code>seizure_vote \nlpd_vote \ngpd_vote \nlrda_vote \ngrda_vote \nother_vote \n</code></pre>\n<h2>Conclusion: Test set have multiple patient_id sequences</h2>\n<h1>Overall : Test set have repeat patients with \"non overlapping eeg\".</h1>\n<h1><a href=\"https://www.kaggle.com/code/seshurajup/eda-train-csv\" target=\"_blank\">Notebook EDA train.csv</a></h1>",
  "messages": [
    {
      "id": "2596082",
      "postDate": "01/10/2024 21:12:58",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5108b87498942a37058ab1164bf8a8de%2FScreenshot%202024-01-11%20at%202.38.42AM.png?generation=1704920935509556&amp;alt=media\"><br>\n<strong>Well balanced targets</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fbef0e17513d4ea1c81654664b7df1720%2FScreenshot%202024-01-11%20at%202.36.39AM.png?generation=1704920849053906&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5098c5d109726d95a7ff26a8fd6456a9%2FScreenshot%202024-01-11%20at%202.59.21AM.png?generation=1704922187122299&amp;alt=media\"><br>\n<strong>Significant correlation between seizure_vote, other vote and grda_vote types ( all 3 targets are frequent targets )</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F20d7d0dfbc933068f13abc37b3881443%2FScreenshot%202024-01-11%20at%202.52.19AM.png?generation=1704921765311082&amp;alt=media\"><br>\n<strong>is &lt; 1000 offset is the major information of the targets except for gpd target ?</strong></p>\n<p>with Equal probability 1/6: <strong>LB 1.09</strong></p>\n<p>With Overlapping <strong>LB: 1.1</strong></p>\n<pre><code>seizure_vote  ,\ngpd_vote  ,\nlrda_vote ,\nother_vote ,\ngrda_vote ,\nlpd_vote \n</code></pre>\n<h2>Conclusion: Test set have means closer to Train Set</h2>\n<p>with non-overlapping based on spectrogram id<br>\n<strong>LB: 1.0</strong></p>\n<pre><code>seizure_vote    \nlpd_vote        \ngpd_vote        \nlrda_vote       \ngrda_vote       \nother_vote      \n</code></pre>\n<p>with non-overlapping based on eeg id<br>\n<strong>LB: 0.97</strong></p>\n<pre><code>seizure_vote    \nlpd_vote        \ngpd_vote        \nlrda_vote       \ngrda_vote       \nother_vote      \n</code></pre>\n<h2>Conclusion: Test set have non-overlapping egg based sequences</h2>\n<p>with non-overlapping based on p id<br>\n<strong>LB: 1.28</strong></p>\n<pre><code>seizure_vote \nlpd_vote \ngpd_vote \nlrda_vote \ngrda_vote \nother_vote \n</code></pre>\n<h2>Conclusion: Test set have multiple patient_id sequences</h2>\n<h1>Overall : Test set have repeat patients with \"non overlapping eeg\".</h1>\n<h1><a href=\"https://www.kaggle.com/code/seshurajup/eda-train-csv\" target=\"_blank\">Notebook EDA train.csv</a></h1>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5108b87498942a37058ab1164bf8a8de%2FScreenshot%202024-01-11%20at%202.38.42AM.png?generation=1704920935509556&alt=media)\n**Well balanced targets**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fbef0e17513d4ea1c81654664b7df1720%2FScreenshot%202024-01-11%20at%202.36.39AM.png?generation=1704920849053906&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5098c5d109726d95a7ff26a8fd6456a9%2FScreenshot%202024-01-11%20at%202.59.21AM.png?generation=1704922187122299&alt=media)\n**Significant correlation between seizure_vote, other vote and grda_vote types ( all 3 targets are frequent targets )**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F20d7d0dfbc933068f13abc37b3881443%2FScreenshot%202024-01-11%20at%202.52.19AM.png?generation=1704921765311082&alt=media)\n**is < 1000 offset is the major information of the targets except for gpd target ?**\n\n\nwith Equal probability 1/6: **LB 1.09**\n\nWith Overlapping **LB: 1.1**\n```python\nseizure_vote  0.196002,\ngpd_vote  0.156386,\nlrda_vote 0.155805,\nother_vote 0.17610,\ngrda_vote 0.17660,\nlpd_vote 0.139101\n```\n## Conclusion: Test set have means closer to Train Set\n\nwith non-overlapping based on spectrogram id\n**LB: 1.0**\n```python\nseizure_vote    0.174031\nlpd_vote        0.112700\ngpd_vote        0.090854\nlrda_vote       0.071484\ngrda_vote       0.136408\nother_vote      0.414523\n```\nwith non-overlapping based on eeg id\n**LB: 0.97**\n```python\nseizure_vote    0.152810\nlpd_vote        0.142456\ngpd_vote        0.104062\nlrda_vote       0.065407\ngrda_vote       0.114851\nother_vote      0.420414\n```\n\n## Conclusion: Test set have non-overlapping egg based sequences\nwith non-overlapping based on p id\n**LB: 1.28**\n```python\nseizure_vote 0.310718\nlpd_vote 0.046279\ngpd_vote 0.051885\nlrda_vote 0.081796\ngrda_vote 0.231471\nother_vote 0.277851\n```\n\n## Conclusion: Test set have multiple patient_id sequences\n\n# Overall : Test set have repeat patients with \"non overlapping eeg\".\n\n# [Notebook EDA train.csv](https://www.kaggle.com/code/seshurajup/eda-train-csv)",
      "votes": null
    },
    {
      "id": "2596092",
      "postDate": "01/10/2024 21:31:05",
      "content": "<p>Thanks for EDA. Have you tried submitting the train means as test predictions? I'm curious if that beats the sample submission 1/6 predictions.</p>",
      "rawMarkdown": "Thanks for EDA. Have you tried submitting the train means as test predictions? I'm curious if that beats the sample submission 1/6 predictions.",
      "votes": null
    },
    {
      "id": "2596119",
      "postDate": "01/10/2024 22:21:06",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thank you. Yes mean with overlaps and mean without overlaps</p>\n<h4><strong>with Overlaps mean: 1.1</strong></h4>\n<h3><strong>without Overlaps mean: 0.97 (beats the sample submission)</strong></h3>",
      "rawMarkdown": "cdeotte Thank you. Yes mean with overlaps and mean without overlaps\n\n#### **with Overlaps mean: 1.1**\n### **without Overlaps mean: 0.97 (beats the sample submission)**",
      "votes": null
    },
    {
      "id": "2598950",
      "postDate": "01/12/2024 15:15:01",
      "content": "<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> Have you tried submitting the means of non-overlaps spectrograms? (as shown in code below)</p>\n<pre><code> = train.groupby('spectrogram_id')[targets].sum().sum(axis=)\n = train.groupby('spectrogram_id')[targets].sum().div(total_votes_per_spec, axis=)\n = normalized_votes.mean()\n( mean_vote_ratio )\n\n    .\n        .\n        .\n       .\n       .\n      .\n</code></pre>\n<p>This is slightly different than non-overlaps eeg</p>\n<pre><code>    .\n        .\n        .\n       .\n       .\n      .\n</code></pre>\n<p>I'm curious which is better. This will give us insight about the best way to train our models.</p>",
      "rawMarkdown": "seshurajup Have you tried submitting the means of non-overlaps spectrograms? (as shown in code below)\n\n    total_votes_per_spec = train.groupby('spectrogram_id')[targets].sum().sum(axis=1)\n    normalized_votes = train.groupby('spectrogram_id')[targets].sum().div(total_votes_per_spec, axis=0)\n    mean_vote_ratio = normalized_votes.mean()\n    print( mean_vote_ratio )\n\n    seizure_vote    0.174031\n    lpd_vote        0.112700\n    gpd_vote        0.090854\n    lrda_vote       0.071484\n    grda_vote       0.136408\n    other_vote      0.414523\n\nThis is slightly different than non-overlaps eeg\n\n    seizure_vote    0.152810\n    lpd_vote        0.142456\n    gpd_vote        0.104062\n    lrda_vote       0.065407\n    grda_vote       0.114851\n    other_vote      0.420414\n\nI'm curious which is better. This will give us insight about the best way to train our models.",
      "votes": null
    },
    {
      "id": "2598973",
      "postDate": "01/12/2024 15:32:16",
      "content": "<p>What is your definition of overlap here? Overlapping 50 second entire subsample or central 10 seconds?</p>",
      "rawMarkdown": "What is your definition of overlap here? Overlapping 50 second entire subsample or central 10 seconds?",
      "votes": null
    },
    {
      "id": "2598983",
      "postDate": "01/12/2024 15:40:44",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I didn't submitted spectrograms based, only did with egg_id based. Curious how it help <strong>insight about the best way to train our models</strong> ?</p>",
      "rawMarkdown": "cdeotte I didn't submitted spectrograms based, only did with egg_id based. Curious how it help **insight about the best way to train our models** ?",
      "votes": null
    },
    {
      "id": "2598988",
      "postDate": "01/12/2024 15:41:44",
      "content": "<p>Overlap = overlapping spectrogram ids</p>\n<p>The best public notebook of 0.97 (and LGB/XGB 0.96) use \"non-overlapping eeg\". We take the average of the train targets from the 17089 unique <code>eeg_id</code> in train.csv.</p>\n<p>To submit \"non-overlapping spectrogram\" we can take the average of the train targets from the 11138 unique <code>spectrogram_id</code> in train.csv. I posted the code to do this above.</p>",
      "rawMarkdown": "Overlap = overlapping spectrogram ids\n\nThe best public notebook of 0.97 (and LGB/XGB 0.96) use \"non-overlapping eeg\". We take the average of the train targets from the 17089 unique `eeg_id` in train.csv.\n\nTo submit \"non-overlapping spectrogram\" we can take the average of the train targets from the 11138 unique `spectrogram_id` in train.csv. I posted the code to do this above.",
      "votes": null
    },
    {
      "id": "2598991",
      "postDate": "01/12/2024 15:43:21",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> for submissions I treated overlaps for eegs is 50secs</p>",
      "rawMarkdown": "gunesevitan for submissions I treated overlaps for eegs is 50secs",
      "votes": null
    },
    {
      "id": "2598994",
      "postDate": "01/12/2024 15:45:05",
      "content": "<p>Whichever has the best LB score will tell us how to set our train sample weights during training. (And inform us how to compute CV score).</p>\n<p>For example, we can give each row of <code>train.csv</code> weight 1 during training. Or we can give each unique <code>eeg_id</code> weight 1 during training. You already submitted these two cases and using unique <code>eeg_id</code> weight is better.</p>\n<p>If unique <code>spectrogram_id</code> is better, then we can give each <code>spectrogram_id</code> weight 1 during training.</p>",
      "rawMarkdown": "Whichever has the best LB score will tell us how to set our train sample weights during training. (And inform us how to compute CV score).\n\nFor example, we can give each row of `train.csv` weight 1 during training. Or we can give each unique `eeg_id` weight 1 during training. You already submitted these two cases and using unique `eeg_id` weight is better.\n\nIf unique `spectrogram_id` is better, then we can give each `spectrogram_id` weight 1 during training.",
      "votes": null
    },
    {
      "id": "2599005",
      "postDate": "01/12/2024 15:49:03",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> if I understood correctly as follow </p>\n<ol>\n<li>Average non-overlapping eeg based score -&gt; 0.97</li>\n<li>Average non-overlapped spectrogram based score -&gt; x</li>\n<li>LGB/XGB non-overlapped egg based score -&gt; 0.96</li>\n</ol>\n<p>using these scores to sample the train dataset to build cv?</p>",
      "rawMarkdown": "cdeotte if I understood correctly as follow \n1. Average non-overlapping eeg based score -> 0.97\n2. Average non-overlapped spectrogram based score -> x\n3. LGB/XGB non-overlapped egg based score -> 0.96\n\nusing these scores to sample the train dataset to build cv?",
      "votes": null
    },
    {
      "id": "2599014",
      "postDate": "01/12/2024 15:54:40",
      "content": "<p>Understanding the test data will help us both train and validate our models more accurately. </p>\n<p>The metric KL-Div is very sensitive to global means.</p>",
      "rawMarkdown": "Understanding the test data will help us both train and validate our models more accurately. \n\nThe metric KL-Div is very sensitive to global means.",
      "votes": null
    },
    {
      "id": "2599017",
      "postDate": "01/12/2024 15:55:23",
      "content": "<p>Thanks, its new learning for me <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Thanks, its new learning for me @cdeotte",
      "votes": null
    },
    {
      "id": "2599036",
      "postDate": "01/12/2024 16:08:40",
      "content": "<p><strong>LB 1.0</strong> ( updated on topic too ) <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<blockquote>\n  <p>with overlapping --&gt; 1.1<br>\n  without overlapping based spectrogram --&gt; 1.0<br>\n  <strong>without overlapping based eeg --&gt; 0.97</strong></p>\n</blockquote>",
      "rawMarkdown": "**LB 1.0** ( updated on topic too ) @cdeotte \n>with overlapping --> 1.1\nwithout overlapping based spectrogram --> 1.0\n**without overlapping based eeg --> 0.97**",
      "votes": null
    },
    {
      "id": "2599038",
      "postDate": "01/12/2024 16:11:17",
      "content": "<p>Thanks for checking <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> . I will continue to train and validate my models giving each unique <code>eeg_id</code> weight 1.</p>\n<p></p>",
      "rawMarkdown": "Thanks for checking @seshurajup . I will continue to train and validate my models giving each unique `eeg_id` weight 1.\n\n~~Note in your discussion post, you list the means of eeg_id for both eeg_id and spectrogram_id.~~",
      "votes": null
    },
    {
      "id": "2599041",
      "postDate": "01/12/2024 16:14:44",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Fixed thanks</p>",
      "rawMarkdown": "cdeotte Fixed thanks",
      "votes": null
    },
    {
      "id": "2599058",
      "postDate": "01/12/2024 16:29:37",
      "content": "<p>These are good LB probing experiments. Based on what you submitted so far, it seems that the train data has \"overlapping eeg\" whereas the test data has \"non-overlapping eeg\". This is said in the competition description, but it is always good to verify ourselves.</p>\n<blockquote>\n  <p>test.csv Metadata for the test set. As there are no overlapping samples in the test set, many columns in the train metadata don't apply.</p>\n</blockquote>\n<p>The next question i'm curious about is if the same patient appears multiple times in the test data. In the train data, the same patient does appear multiple times. For example if we compute the mean from train with \"non-overlapping patients\" we get much different means (than \"overlapping patients\"):</p>\n<pre><code> = train.groupby('patient_id')[targets].sum().sum(axis=)\n = train.groupby('patient_id')[targets].sum().div(total_votes_per_pat, axis=)\n = normalized_votes.mean()\n( mean_vote_ratio )\n\n    .\n        .\n        .\n       .\n       .\n      .\n</code></pre>\n<p>If these means produce a better LB score (than 0.97), then the test data is most likely unique patients (i.e. \"non overlapping patients\"). If these means do not produce a better LB score, then the test data is most likely repeat patients with \"non overlapping eeg\". (Of course we can also submit an \"if-else\" test to LB. Our submit code can check if a test patient repeats, then create submission.csv with 1/6 preds else create submission.csv with \"non-overlapp eeg_id\" preds. But probing LB score with means is more fun).</p>\n<p>If the test data has repeated patients, then perhaps we can use one row of test data to help us predict another row of test data.</p>",
      "rawMarkdown": "These are good LB probing experiments. Based on what you submitted so far, it seems that the train data has \"overlapping eeg\" whereas the test data has \"non-overlapping eeg\". This is said in the competition description, but it is always good to verify ourselves.\n\n>test.csv Metadata for the test set. As there are no overlapping samples in the test set, many columns in the train metadata don't apply.\n\nThe next question i'm curious about is if the same patient appears multiple times in the test data. In the train data, the same patient does appear multiple times. For example if we compute the mean from train with \"non-overlapping patients\" we get much different means (than \"overlapping patients\"):\n\n    total_votes_per_pat = train.groupby('patient_id')[targets].sum().sum(axis=1)\n    normalized_votes = train.groupby('patient_id')[targets].sum().div(total_votes_per_pat, axis=0)\n    mean_vote_ratio = normalized_votes.mean()\n    print( mean_vote_ratio )\n\n    seizure_vote    0.310718\n    lpd_vote        0.046279\n    gpd_vote        0.051885\n    lrda_vote       0.081796\n    grda_vote       0.231471\n    other_vote      0.277851\n\nIf these means produce a better LB score (than 0.97), then the test data is most likely unique patients (i.e. \"non overlapping patients\"). If these means do not produce a better LB score, then the test data is most likely repeat patients with \"non overlapping eeg\". (Of course we can also submit an \"if-else\" test to LB. Our submit code can check if a test patient repeats, then create submission.csv with 1/6 preds else create submission.csv with \"non-overlapp eeg_id\" preds. But probing LB score with means is more fun).\n\nIf the test data has repeated patients, then perhaps we can use one row of test data to help us predict another row of test data.",
      "votes": null
    },
    {
      "id": "2599250",
      "postDate": "01/12/2024 18:22:46",
      "content": "<p>Its great learning with you (@cdeotte) and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> (other competitions). Its motivate me and learn new things every time i see the knowledge share from last 4 years. Thanks for it ( Today lesson - <strong>Power of Mean</strong> )</p>\n<p><strong>LB: 1.28</strong><br>\nseizure_vote    0.310718<br>\nlpd_vote        0.046279<br>\ngpd_vote        0.051885<br>\nlrda_vote       0.081796<br>\ngrda_vote       0.231471<br>\nother_vote      0.277851</p>\n<p><strong>conclusion</strong> - repeat patients with \"non overlapping eeg\". </p>",
      "rawMarkdown": "Its great learning with you (@cdeotte) and @hengck23 (other competitions). Its motivate me and learn new things every time i see the knowledge share from last 4 years. Thanks for it ( Today lesson - **Power of Mean** )\n\n**LB: 1.28**\nseizure_vote    0.310718\nlpd_vote        0.046279\ngpd_vote        0.051885\nlrda_vote       0.081796\ngrda_vote       0.231471\nother_vote      0.277851\n\n**conclusion** - repeat patients with \"non overlapping eeg\".",
      "votes": null
    },
    {
      "id": "2599328",
      "postDate": "01/12/2024 18:59:17",
      "content": "<p>Thanks for checking this. If test has multiple patient_id, then we can use <code>patient_id count z_score</code> as a feature. From train data, it appears when patient id count z_score is small, then patient is more likely <code>seizure</code> or <code>grda</code>. And when patient count z_score is large, target is more likely <code>gpd</code> or <code>lpd</code>.</p>\n<p>We should be able to train model using only this one feature <code>patient_id count z_score</code> and beat LB 0.97. Note that we will need to standardize the count feature so that it generalizes from train to test. For example, the feature will be the <code>z_score = (count - count_mean) / count_std</code>. So during train <code>count_mean = total patients in train</code> and during infer <code>count_mean = total patients in test</code>. Likewise <code>count_std</code> is computed once during train and once during infer.</p>",
      "rawMarkdown": "Thanks for checking this. If test has multiple patient_id, then we can use `patient_id count z_score` as a feature. From train data, it appears when patient id count z_score is small, then patient is more likely `seizure` or `grda`. And when patient count z_score is large, target is more likely `gpd` or `lpd`.\n\nWe should be able to train model using only this one feature `patient_id count z_score` and beat LB 0.97. Note that we will need to standardize the count feature so that it generalizes from train to test. For example, the feature will be the `z_score = (count - count_mean) / count_std`. So during train `count_mean = total patients in train` and during infer `count_mean = total patients in test`. Likewise `count_std` is computed once during train and once during infer.",
      "votes": null
    },
    {
      "id": "2618121",
      "postDate": "01/24/2024 15:17:45",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> i have train and test the using CatBoost Classifier using feature <code>patient_id count z_score</code> and i got following results:<br>\nCV : 1.185<br>\nLB : 1.01</p>\n<p>Here is the feature calculation:</p>\n<pre><code>tmp = df.groupby().size()\neeg_id_count_per_patient_id = pd.DataFrame({:tmp.index,:tmp.values})\nmean_value = np.mean(eeg_id_count_per_patient_id[])\nstd_value = np.std(eeg_id_count_per_patient_id[])\n\neeg_id_count_per_patient_id[] = (eeg_id_count_per_patient_id[] - mean_value)/std_value\neeg_id_count_per_patient_id = eeg_id_count_per_patient_id.drop(columns=[])\n</code></pre>\n<p>So, this LB is not able to beat LB 0.97. </p>\n<p>Notebook Link <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/deepaksingh47/catboost-starter-with-patient-eeg-count-z-score#CV-Score-for-CatBoost</a></p>",
      "rawMarkdown": "cdeotte i have train and test the using CatBoost Classifier using feature `patient_id count z_score` and i got following results:\nCV : 1.185\nLB : 1.01\n\nHere is the feature calculation:\n```python\ntmp = df.groupby('patient_id').size()\neeg_id_count_per_patient_id = pd.DataFrame({'patient_id':tmp.index,'eeg_count':tmp.values})\nmean_value = np.mean(eeg_id_count_per_patient_id[\"eeg_count\"])\nstd_value = np.std(eeg_id_count_per_patient_id[\"eeg_count\"])\n\neeg_id_count_per_patient_id[\"eeg_count_z_score\"] = (eeg_id_count_per_patient_id[\"eeg_count\"] - mean_value)/std_value\neeg_id_count_per_patient_id = eeg_id_count_per_patient_id.drop(columns=[\"eeg_count\"])\n```\n\nSo, this LB is not able to beat LB 0.97. \n\nNotebook Link [https://www.kaggle.com/code/deepaksingh47/catboost-starter-with-patient-eeg-count-z-score#CV-Score-for-CatBoost](url)",
      "votes": null
    },
    {
      "id": "2618149",
      "postDate": "01/24/2024 15:31:51",
      "content": "<p>Thanks for doing the test. I also performed a test with my best model. Including patient z score improved CV score but hurt LB score.</p>",
      "rawMarkdown": "Thanks for doing the test. I also performed a test with my best model. Including patient z score improved CV score but hurt LB score.",
      "votes": null
    },
    {
      "id": "2618154",
      "postDate": "01/24/2024 15:35:19",
      "content": "<p>It is interesting that if we include z_score along with other features, then z_score becomes the most important feature. Unfortuntately it doesn't seem to work on test.</p>\n<p><a href=\"https://www.kaggle.com/deepaksingh47\" target=\"_blank\">@deepaksingh47</a> Note that your notebook has a leak. We cannot compute the z_score once. We must compute the z_score each fold. In the <code>folds for-loop</code>. We compute z_score for the train folds. And then we compute z_score for the valid folds. Then for fold 2, we delete previous z_score and compute z_score again. (In total we compute z_score on 10 dataframes for 5 KFold)</p>",
      "rawMarkdown": "It is interesting that if we include z_score along with other features, then z_score becomes the most important feature. Unfortuntately it doesn't seem to work on test.\n\n@deepaksingh47 Note that your notebook has a leak. We cannot compute the z_score once. We must compute the z_score each fold. In the `folds for-loop`. We compute z_score for the train folds. And then we compute z_score for the valid folds. Then for fold 2, we delete previous z_score and compute z_score again. (In total we compute z_score on 10 dataframes for 5 KFold)",
      "votes": null
    },
    {
      "id": "2618814",
      "postDate": "01/25/2024 02:36:59",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, thanks for pointing that out. i forget about that. </p>",
      "rawMarkdown": "cdeotte, thanks for pointing that out. i forget about that.",
      "votes": null
    },
    {
      "id": "2619937",
      "postDate": "01/25/2024 19:47:57",
      "content": "<p>Hi, Many thanks for your discussion post  and the eda nbs as they provide many insights into the data. Could you pls share how you calculated these particular values (the ones you call \"with overlapping LB: 1.1\". While I am able to understand many of the other calculations based on per eeg_id, per spectrogram_id, per row etc, I am unable to grasp how these numbers are calculated.<br>\nwith overlapping LB: 1.1</p>\n<p>seizure_vote  0.196002,<br>\ngpd_vote  0.156386,<br>\nlrda_vote 0.155805,<br>\nother_vote 0.17610,<br>\ngrda_vote 0.17660,<br>\nlpd_vote 0.139101</p>\n<p>Thanks.</p>",
      "rawMarkdown": "Hi, Many thanks for your discussion post  and the eda nbs as they provide many insights into the data. Could you pls share how you calculated these particular values (the ones you call \"with overlapping LB: 1.1\". While I am able to understand many of the other calculations based on per eeg_id, per spectrogram_id, per row etc, I am unable to grasp how these numbers are calculated.\nwith overlapping LB: 1.1\n\nseizure_vote  0.196002,\ngpd_vote  0.156386,\nlrda_vote 0.155805,\nother_vote 0.17610,\ngrda_vote 0.17660,\nlpd_vote 0.139101\n\nThanks.",
      "votes": null
    },
    {
      "id": "2619980",
      "postDate": "01/25/2024 20:17:53",
      "content": "<p><a href=\"https://www.kaggle.com/skr1125\" target=\"_blank\">@skr1125</a> from <a href=\"https://www.kaggle.com/code/seshurajup/eda-train-csv?scriptVersionId=158496297\" target=\"_blank\">Notebook v16 - EDA train.csv</a></p>\n<pre><code>  collections  Counter\ntarget_votes = Counter((train[]))\ntarget_votes = {:v  k,v  target_votes.items()}\ntotal_votes = ([v  _,v  target_votes.items()])\nmean_vote_ratio = {k:(target_votes[k]/total_votes)  k,target  target_votes.items()}\n</code></pre>",
      "rawMarkdown": "skr1125 from [Notebook v16 - EDA train.csv](https://www.kaggle.com/code/seshurajup/eda-train-csv?scriptVersionId=158496297)\n```python\n from collections import Counter\ntarget_votes = Counter(list(train['expert_consensus']))\ntarget_votes = {f\"{k.lower()}_vote\":v for k,v in target_votes.items()}\ntotal_votes = sum([v for _,v in target_votes.items()])\nmean_vote_ratio = {k:(target_votes[k]/total_votes) for k,target in target_votes.items()}\n```",
      "votes": null
    },
    {
      "id": "2633353",
      "postDate": "02/03/2024 03:09:22",
      "content": "<p>ok, understand</p>",
      "rawMarkdown": "ok, understand",
      "votes": null
    },
    {
      "id": "2641812",
      "postDate": "02/07/2024 17:19:42",
      "content": "<p>Thanks for providing this, and thanks to all the commentators. Insightful information.</p>",
      "rawMarkdown": "Thanks for providing this, and thanks to all the commentators. Insightful information.",
      "votes": null
    },
    {
      "id": "2704278",
      "postDate": "03/18/2024 17:25:00",
      "content": "<p>Could anyone be kind enough to explain the first conclusion along with what the author means by <br>\n\"with Equal probability 1/6: LB 1.09</p>\n<p>With Overlapping LB: 1.1\"</p>\n<p>I tried to get the same numbers but only in this scenario I failed. I am assuming equal probability means author put all targets value as 0.17 and submitted it. I can't figure out what he means by with overlapping. <br>\nAlso what does getting better score by equal probability mean? In the other 2 conclusions I understand that the case in which better score is achieved we can assume then test case has similar data. But for 1st conclusion I am having a hard time understanding if anyone could explain it I would be very grateful. Thanks in advance.</p>",
      "rawMarkdown": "Could anyone be kind enough to explain the first conclusion along with what the author means by \n\"with Equal probability 1/6: LB 1.09\n\nWith Overlapping LB: 1.1\"\n\nI tried to get the same numbers but only in this scenario I failed. I am assuming equal probability means author put all targets value as 0.17 and submitted it. I can't figure out what he means by with overlapping. \nAlso what does getting better score by equal probability mean? In the other 2 conclusions I understand that the case in which better score is achieved we can assume then test case has similar data. But for 1st conclusion I am having a hard time understanding if anyone could explain it I would be very grateful. Thanks in advance.",
      "votes": null
    },
    {
      "id": "2705719",
      "postDate": "03/19/2024 13:38:52",
      "content": "<p>same i also stuck with this</p>",
      "rawMarkdown": "same i also stuck with this",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2596092,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/10/2024 21:31:05",
      "content": "<p>Thanks for EDA. Have you tried submitting the train means as test predictions? I'm curious if that beats the sample submission 1/6 predictions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2596119,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "01/10/2024 22:21:06",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thank you. Yes mean with overlaps and mean without overlaps</p>\n<h4><strong>with Overlaps mean: 1.1</strong></h4>\n<h3><strong>without Overlaps mean: 0.97 (beats the sample submission)</strong></h3>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2598950,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/12/2024 15:15:01",
      "content": "<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> Have you tried submitting the means of non-overlaps spectrograms? (as shown in code below)</p>\n<pre><code> = train.groupby('spectrogram_id')[targets].sum().sum(axis=)\n = train.groupby('spectrogram_id')[targets].sum().div(total_votes_per_spec, axis=)\n = normalized_votes.mean()\n( mean_vote_ratio )\n\n    .\n        .\n        .\n       .\n       .\n      .\n</code></pre>\n<p>This is slightly different than non-overlaps eeg</p>\n<pre><code>    .\n        .\n        .\n       .\n       .\n      .\n</code></pre>\n<p>I'm curious which is better. This will give us insight about the best way to train our models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2598973,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "01/12/2024 15:32:16",
          "content": "<p>What is your definition of overlap here? Overlapping 50 second entire subsample or central 10 seconds?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2598988,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "01/12/2024 15:41:44",
              "content": "<p>Overlap = overlapping spectrogram ids</p>\n<p>The best public notebook of 0.97 (and LGB/XGB 0.96) use \"non-overlapping eeg\". We take the average of the train targets from the 17089 unique <code>eeg_id</code> in train.csv.</p>\n<p>To submit \"non-overlapping spectrogram\" we can take the average of the train targets from the 11138 unique <code>spectrogram_id</code> in train.csv. I posted the code to do this above.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2599005,
                  "author_name": "seshurajup",
                  "author_url": "",
                  "post_date": "01/12/2024 15:49:03",
                  "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> if I understood correctly as follow </p>\n<ol>\n<li>Average non-overlapping eeg based score -&gt; 0.97</li>\n<li>Average non-overlapped spectrogram based score -&gt; x</li>\n<li>LGB/XGB non-overlapped egg based score -&gt; 0.96</li>\n</ol>\n<p>using these scores to sample the train dataset to build cv?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2599014,
                      "author_name": "cdeotte",
                      "author_url": "",
                      "post_date": "01/12/2024 15:54:40",
                      "content": "<p>Understanding the test data will help us both train and validate our models more accurately. </p>\n<p>The metric KL-Div is very sensitive to global means.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            },
            {
              "id": 2598991,
              "author_name": "seshurajup",
              "author_url": "",
              "post_date": "01/12/2024 15:43:21",
              "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> for submissions I treated overlaps for eegs is 50secs</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2598983,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "01/12/2024 15:40:44",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I didn't submitted spectrograms based, only did with egg_id based. Curious how it help <strong>insight about the best way to train our models</strong> ?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2598994,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "01/12/2024 15:45:05",
              "content": "<p>Whichever has the best LB score will tell us how to set our train sample weights during training. (And inform us how to compute CV score).</p>\n<p>For example, we can give each row of <code>train.csv</code> weight 1 during training. Or we can give each unique <code>eeg_id</code> weight 1 during training. You already submitted these two cases and using unique <code>eeg_id</code> weight is better.</p>\n<p>If unique <code>spectrogram_id</code> is better, then we can give each <code>spectrogram_id</code> weight 1 during training.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2599017,
                  "author_name": "seshurajup",
                  "author_url": "",
                  "post_date": "01/12/2024 15:55:23",
                  "content": "<p>Thanks, its new learning for me <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2599036,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "01/12/2024 16:08:40",
          "content": "<p><strong>LB 1.0</strong> ( updated on topic too ) <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<blockquote>\n  <p>with overlapping --&gt; 1.1<br>\n  without overlapping based spectrogram --&gt; 1.0<br>\n  <strong>without overlapping based eeg --&gt; 0.97</strong></p>\n</blockquote>",
          "votes": null,
          "replies": [
            {
              "id": 2599038,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "01/12/2024 16:11:17",
              "content": "<p>Thanks for checking <a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> . I will continue to train and validate my models giving each unique <code>eeg_id</code> weight 1.</p>\n<p></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2599041,
                  "author_name": "seshurajup",
                  "author_url": "",
                  "post_date": "01/12/2024 16:14:44",
                  "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Fixed thanks</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2599058,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/12/2024 16:29:37",
      "content": "<p>These are good LB probing experiments. Based on what you submitted so far, it seems that the train data has \"overlapping eeg\" whereas the test data has \"non-overlapping eeg\". This is said in the competition description, but it is always good to verify ourselves.</p>\n<blockquote>\n  <p>test.csv Metadata for the test set. As there are no overlapping samples in the test set, many columns in the train metadata don't apply.</p>\n</blockquote>\n<p>The next question i'm curious about is if the same patient appears multiple times in the test data. In the train data, the same patient does appear multiple times. For example if we compute the mean from train with \"non-overlapping patients\" we get much different means (than \"overlapping patients\"):</p>\n<pre><code> = train.groupby('patient_id')[targets].sum().sum(axis=)\n = train.groupby('patient_id')[targets].sum().div(total_votes_per_pat, axis=)\n = normalized_votes.mean()\n( mean_vote_ratio )\n\n    .\n        .\n        .\n       .\n       .\n      .\n</code></pre>\n<p>If these means produce a better LB score (than 0.97), then the test data is most likely unique patients (i.e. \"non overlapping patients\"). If these means do not produce a better LB score, then the test data is most likely repeat patients with \"non overlapping eeg\". (Of course we can also submit an \"if-else\" test to LB. Our submit code can check if a test patient repeats, then create submission.csv with 1/6 preds else create submission.csv with \"non-overlapp eeg_id\" preds. But probing LB score with means is more fun).</p>\n<p>If the test data has repeated patients, then perhaps we can use one row of test data to help us predict another row of test data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2599250,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "01/12/2024 18:22:46",
          "content": "<p>Its great learning with you (@cdeotte) and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> (other competitions). Its motivate me and learn new things every time i see the knowledge share from last 4 years. Thanks for it ( Today lesson - <strong>Power of Mean</strong> )</p>\n<p><strong>LB: 1.28</strong><br>\nseizure_vote    0.310718<br>\nlpd_vote        0.046279<br>\ngpd_vote        0.051885<br>\nlrda_vote       0.081796<br>\ngrda_vote       0.231471<br>\nother_vote      0.277851</p>\n<p><strong>conclusion</strong> - repeat patients with \"non overlapping eeg\". </p>",
          "votes": null,
          "replies": [
            {
              "id": 2599328,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "01/12/2024 18:59:17",
              "content": "<p>Thanks for checking this. If test has multiple patient_id, then we can use <code>patient_id count z_score</code> as a feature. From train data, it appears when patient id count z_score is small, then patient is more likely <code>seizure</code> or <code>grda</code>. And when patient count z_score is large, target is more likely <code>gpd</code> or <code>lpd</code>.</p>\n<p>We should be able to train model using only this one feature <code>patient_id count z_score</code> and beat LB 0.97. Note that we will need to standardize the count feature so that it generalizes from train to test. For example, the feature will be the <code>z_score = (count - count_mean) / count_std</code>. So during train <code>count_mean = total patients in train</code> and during infer <code>count_mean = total patients in test</code>. Likewise <code>count_std</code> is computed once during train and once during infer.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2618121,
                  "author_name": "deepaksingh47",
                  "author_url": "",
                  "post_date": "01/24/2024 15:17:45",
                  "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> i have train and test the using CatBoost Classifier using feature <code>patient_id count z_score</code> and i got following results:<br>\nCV : 1.185<br>\nLB : 1.01</p>\n<p>Here is the feature calculation:</p>\n<pre><code>tmp = df.groupby().size()\neeg_id_count_per_patient_id = pd.DataFrame({:tmp.index,:tmp.values})\nmean_value = np.mean(eeg_id_count_per_patient_id[])\nstd_value = np.std(eeg_id_count_per_patient_id[])\n\neeg_id_count_per_patient_id[] = (eeg_id_count_per_patient_id[] - mean_value)/std_value\neeg_id_count_per_patient_id = eeg_id_count_per_patient_id.drop(columns=[])\n</code></pre>\n<p>So, this LB is not able to beat LB 0.97. </p>\n<p>Notebook Link <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/deepaksingh47/catboost-starter-with-patient-eeg-count-z-score#CV-Score-for-CatBoost</a></p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2618149,
                      "author_name": "cdeotte",
                      "author_url": "",
                      "post_date": "01/24/2024 15:31:51",
                      "content": "<p>Thanks for doing the test. I also performed a test with my best model. Including patient z score improved CV score but hurt LB score.</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2618154,
                      "author_name": "cdeotte",
                      "author_url": "",
                      "post_date": "01/24/2024 15:35:19",
                      "content": "<p>It is interesting that if we include z_score along with other features, then z_score becomes the most important feature. Unfortuntately it doesn't seem to work on test.</p>\n<p><a href=\"https://www.kaggle.com/deepaksingh47\" target=\"_blank\">@deepaksingh47</a> Note that your notebook has a leak. We cannot compute the z_score once. We must compute the z_score each fold. In the <code>folds for-loop</code>. We compute z_score for the train folds. And then we compute z_score for the valid folds. Then for fold 2, we delete previous z_score and compute z_score again. (In total we compute z_score on 10 dataframes for 5 KFold)</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2618814,
                          "author_name": "deepaksingh47",
                          "author_url": "",
                          "post_date": "01/25/2024 02:36:59",
                          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, thanks for pointing that out. i forget about that. </p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2619937,
      "author_name": "skr1125",
      "author_url": "",
      "post_date": "01/25/2024 19:47:57",
      "content": "<p>Hi, Many thanks for your discussion post  and the eda nbs as they provide many insights into the data. Could you pls share how you calculated these particular values (the ones you call \"with overlapping LB: 1.1\". While I am able to understand many of the other calculations based on per eeg_id, per spectrogram_id, per row etc, I am unable to grasp how these numbers are calculated.<br>\nwith overlapping LB: 1.1</p>\n<p>seizure_vote  0.196002,<br>\ngpd_vote  0.156386,<br>\nlrda_vote 0.155805,<br>\nother_vote 0.17610,<br>\ngrda_vote 0.17660,<br>\nlpd_vote 0.139101</p>\n<p>Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2619980,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "01/25/2024 20:17:53",
          "content": "<p><a href=\"https://www.kaggle.com/skr1125\" target=\"_blank\">@skr1125</a> from <a href=\"https://www.kaggle.com/code/seshurajup/eda-train-csv?scriptVersionId=158496297\" target=\"_blank\">Notebook v16 - EDA train.csv</a></p>\n<pre><code>  collections  Counter\ntarget_votes = Counter((train[]))\ntarget_votes = {:v  k,v  target_votes.items()}\ntotal_votes = ([v  _,v  target_votes.items()])\nmean_vote_ratio = {k:(target_votes[k]/total_votes)  k,target  target_votes.items()}\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 2633353,
              "author_name": "tiangg",
              "author_url": "",
              "post_date": "02/03/2024 03:09:22",
              "content": "<p>ok, understand</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2641812,
      "author_name": "marioandrs",
      "author_url": "",
      "post_date": "02/07/2024 17:19:42",
      "content": "<p>Thanks for providing this, and thanks to all the commentators. Insightful information.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2704278,
      "author_name": "otutsukihyuuga",
      "author_url": "",
      "post_date": "03/18/2024 17:25:00",
      "content": "<p>Could anyone be kind enough to explain the first conclusion along with what the author means by <br>\n\"with Equal probability 1/6: LB 1.09</p>\n<p>With Overlapping LB: 1.1\"</p>\n<p>I tried to get the same numbers but only in this scenario I failed. I am assuming equal probability means author put all targets value as 0.17 and submitted it. I can't figure out what he means by with overlapping. <br>\nAlso what does getting better score by equal probability mean? In the other 2 conclusions I understand that the case in which better score is achieved we can assume then test case has similar data. But for 1st conclusion I am having a hard time understanding if anyone could explain it I would be very grateful. Thanks in advance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2705719,
          "author_name": "heyviv",
          "author_url": "",
          "post_date": "03/19/2024 13:38:52",
          "content": "<p>same i also stuck with this</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2596082": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5108b87498942a37058ab1164bf8a8de%2FScreenshot%202024-01-11%20at%202.38.42AM.png?generation=1704920935509556&alt=media)\n**Well balanced targets**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fbef0e17513d4ea1c81654664b7df1720%2FScreenshot%202024-01-11%20at%202.36.39AM.png?generation=1704920849053906&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F5098c5d109726d95a7ff26a8fd6456a9%2FScreenshot%202024-01-11%20at%202.59.21AM.png?generation=1704922187122299&alt=media)\n**Significant correlation between seizure_vote, other vote and grda_vote types ( all 3 targets are frequent targets )**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F20d7d0dfbc933068f13abc37b3881443%2FScreenshot%202024-01-11%20at%202.52.19AM.png?generation=1704921765311082&alt=media)\n**is < 1000 offset is the major information of the targets except for gpd target ?**\n\n\nwith Equal probability 1/6: **LB 1.09**\n\nWith Overlapping **LB: 1.1**\n```python\nseizure_vote  0.196002,\ngpd_vote  0.156386,\nlrda_vote 0.155805,\nother_vote 0.17610,\ngrda_vote 0.17660,\nlpd_vote 0.139101\n```\n## Conclusion: Test set have means closer to Train Set\n\nwith non-overlapping based on spectrogram id\n**LB: 1.0**\n```python\nseizure_vote    0.174031\nlpd_vote        0.112700\ngpd_vote        0.090854\nlrda_vote       0.071484\ngrda_vote       0.136408\nother_vote      0.414523\n```\nwith non-overlapping based on eeg id\n**LB: 0.97**\n```python\nseizure_vote    0.152810\nlpd_vote        0.142456\ngpd_vote        0.104062\nlrda_vote       0.065407\ngrda_vote       0.114851\nother_vote      0.420414\n```\n\n## Conclusion: Test set have non-overlapping egg based sequences\nwith non-overlapping based on p id\n**LB: 1.28**\n```python\nseizure_vote 0.310718\nlpd_vote 0.046279\ngpd_vote 0.051885\nlrda_vote 0.081796\ngrda_vote 0.231471\nother_vote 0.277851\n```\n\n## Conclusion: Test set have multiple patient_id sequences\n\n# Overall : Test set have repeat patients with \"non overlapping eeg\".\n\n# [Notebook EDA train.csv](https://www.kaggle.com/code/seshurajup/eda-train-csv)",
    "2596092": "Thanks for EDA. Have you tried submitting the train means as test predictions? I'm curious if that beats the sample submission 1/6 predictions.",
    "2596119": "cdeotte Thank you. Yes mean with overlaps and mean without overlaps\n\n#### **with Overlaps mean: 1.1**\n### **without Overlaps mean: 0.97 (beats the sample submission)**",
    "2598950": "seshurajup Have you tried submitting the means of non-overlaps spectrograms? (as shown in code below)\n\n    total_votes_per_spec = train.groupby('spectrogram_id')[targets].sum().sum(axis=1)\n    normalized_votes = train.groupby('spectrogram_id')[targets].sum().div(total_votes_per_spec, axis=0)\n    mean_vote_ratio = normalized_votes.mean()\n    print( mean_vote_ratio )\n\n    seizure_vote    0.174031\n    lpd_vote        0.112700\n    gpd_vote        0.090854\n    lrda_vote       0.071484\n    grda_vote       0.136408\n    other_vote      0.414523\n\nThis is slightly different than non-overlaps eeg\n\n    seizure_vote    0.152810\n    lpd_vote        0.142456\n    gpd_vote        0.104062\n    lrda_vote       0.065407\n    grda_vote       0.114851\n    other_vote      0.420414\n\nI'm curious which is better. This will give us insight about the best way to train our models.",
    "2598973": "What is your definition of overlap here? Overlapping 50 second entire subsample or central 10 seconds?",
    "2598983": "cdeotte I didn't submitted spectrograms based, only did with egg_id based. Curious how it help **insight about the best way to train our models** ?",
    "2598988": "Overlap = overlapping spectrogram ids\n\nThe best public notebook of 0.97 (and LGB/XGB 0.96) use \"non-overlapping eeg\". We take the average of the train targets from the 17089 unique `eeg_id` in train.csv.\n\nTo submit \"non-overlapping spectrogram\" we can take the average of the train targets from the 11138 unique `spectrogram_id` in train.csv. I posted the code to do this above.",
    "2598991": "gunesevitan for submissions I treated overlaps for eegs is 50secs",
    "2598994": "Whichever has the best LB score will tell us how to set our train sample weights during training. (And inform us how to compute CV score).\n\nFor example, we can give each row of `train.csv` weight 1 during training. Or we can give each unique `eeg_id` weight 1 during training. You already submitted these two cases and using unique `eeg_id` weight is better.\n\nIf unique `spectrogram_id` is better, then we can give each `spectrogram_id` weight 1 during training.",
    "2599005": "cdeotte if I understood correctly as follow \n1. Average non-overlapping eeg based score -> 0.97\n2. Average non-overlapped spectrogram based score -> x\n3. LGB/XGB non-overlapped egg based score -> 0.96\n\nusing these scores to sample the train dataset to build cv?",
    "2599014": "Understanding the test data will help us both train and validate our models more accurately. \n\nThe metric KL-Div is very sensitive to global means.",
    "2599017": "Thanks, its new learning for me @cdeotte",
    "2599036": "**LB 1.0** ( updated on topic too ) @cdeotte \n>with overlapping --> 1.1\nwithout overlapping based spectrogram --> 1.0\n**without overlapping based eeg --> 0.97**",
    "2599038": "Thanks for checking @seshurajup . I will continue to train and validate my models giving each unique `eeg_id` weight 1.\n\n~~Note in your discussion post, you list the means of eeg_id for both eeg_id and spectrogram_id.~~",
    "2599041": "cdeotte Fixed thanks",
    "2599058": "These are good LB probing experiments. Based on what you submitted so far, it seems that the train data has \"overlapping eeg\" whereas the test data has \"non-overlapping eeg\". This is said in the competition description, but it is always good to verify ourselves.\n\n>test.csv Metadata for the test set. As there are no overlapping samples in the test set, many columns in the train metadata don't apply.\n\nThe next question i'm curious about is if the same patient appears multiple times in the test data. In the train data, the same patient does appear multiple times. For example if we compute the mean from train with \"non-overlapping patients\" we get much different means (than \"overlapping patients\"):\n\n    total_votes_per_pat = train.groupby('patient_id')[targets].sum().sum(axis=1)\n    normalized_votes = train.groupby('patient_id')[targets].sum().div(total_votes_per_pat, axis=0)\n    mean_vote_ratio = normalized_votes.mean()\n    print( mean_vote_ratio )\n\n    seizure_vote    0.310718\n    lpd_vote        0.046279\n    gpd_vote        0.051885\n    lrda_vote       0.081796\n    grda_vote       0.231471\n    other_vote      0.277851\n\nIf these means produce a better LB score (than 0.97), then the test data is most likely unique patients (i.e. \"non overlapping patients\"). If these means do not produce a better LB score, then the test data is most likely repeat patients with \"non overlapping eeg\". (Of course we can also submit an \"if-else\" test to LB. Our submit code can check if a test patient repeats, then create submission.csv with 1/6 preds else create submission.csv with \"non-overlapp eeg_id\" preds. But probing LB score with means is more fun).\n\nIf the test data has repeated patients, then perhaps we can use one row of test data to help us predict another row of test data.",
    "2599250": "Its great learning with you (@cdeotte) and @hengck23 (other competitions). Its motivate me and learn new things every time i see the knowledge share from last 4 years. Thanks for it ( Today lesson - **Power of Mean** )\n\n**LB: 1.28**\nseizure_vote    0.310718\nlpd_vote        0.046279\ngpd_vote        0.051885\nlrda_vote       0.081796\ngrda_vote       0.231471\nother_vote      0.277851\n\n**conclusion** - repeat patients with \"non overlapping eeg\".",
    "2599328": "Thanks for checking this. If test has multiple patient_id, then we can use `patient_id count z_score` as a feature. From train data, it appears when patient id count z_score is small, then patient is more likely `seizure` or `grda`. And when patient count z_score is large, target is more likely `gpd` or `lpd`.\n\nWe should be able to train model using only this one feature `patient_id count z_score` and beat LB 0.97. Note that we will need to standardize the count feature so that it generalizes from train to test. For example, the feature will be the `z_score = (count - count_mean) / count_std`. So during train `count_mean = total patients in train` and during infer `count_mean = total patients in test`. Likewise `count_std` is computed once during train and once during infer.",
    "2618121": "cdeotte i have train and test the using CatBoost Classifier using feature `patient_id count z_score` and i got following results:\nCV : 1.185\nLB : 1.01\n\nHere is the feature calculation:\n```python\ntmp = df.groupby('patient_id').size()\neeg_id_count_per_patient_id = pd.DataFrame({'patient_id':tmp.index,'eeg_count':tmp.values})\nmean_value = np.mean(eeg_id_count_per_patient_id[\"eeg_count\"])\nstd_value = np.std(eeg_id_count_per_patient_id[\"eeg_count\"])\n\neeg_id_count_per_patient_id[\"eeg_count_z_score\"] = (eeg_id_count_per_patient_id[\"eeg_count\"] - mean_value)/std_value\neeg_id_count_per_patient_id = eeg_id_count_per_patient_id.drop(columns=[\"eeg_count\"])\n```\n\nSo, this LB is not able to beat LB 0.97. \n\nNotebook Link [https://www.kaggle.com/code/deepaksingh47/catboost-starter-with-patient-eeg-count-z-score#CV-Score-for-CatBoost](url)",
    "2618149": "Thanks for doing the test. I also performed a test with my best model. Including patient z score improved CV score but hurt LB score.",
    "2618154": "It is interesting that if we include z_score along with other features, then z_score becomes the most important feature. Unfortuntately it doesn't seem to work on test.\n\n@deepaksingh47 Note that your notebook has a leak. We cannot compute the z_score once. We must compute the z_score each fold. In the `folds for-loop`. We compute z_score for the train folds. And then we compute z_score for the valid folds. Then for fold 2, we delete previous z_score and compute z_score again. (In total we compute z_score on 10 dataframes for 5 KFold)",
    "2618814": "cdeotte, thanks for pointing that out. i forget about that.",
    "2619937": "Hi, Many thanks for your discussion post  and the eda nbs as they provide many insights into the data. Could you pls share how you calculated these particular values (the ones you call \"with overlapping LB: 1.1\". While I am able to understand many of the other calculations based on per eeg_id, per spectrogram_id, per row etc, I am unable to grasp how these numbers are calculated.\nwith overlapping LB: 1.1\n\nseizure_vote  0.196002,\ngpd_vote  0.156386,\nlrda_vote 0.155805,\nother_vote 0.17610,\ngrda_vote 0.17660,\nlpd_vote 0.139101\n\nThanks.",
    "2619980": "skr1125 from [Notebook v16 - EDA train.csv](https://www.kaggle.com/code/seshurajup/eda-train-csv?scriptVersionId=158496297)\n```python\n from collections import Counter\ntarget_votes = Counter(list(train['expert_consensus']))\ntarget_votes = {f\"{k.lower()}_vote\":v for k,v in target_votes.items()}\ntotal_votes = sum([v for _,v in target_votes.items()])\nmean_vote_ratio = {k:(target_votes[k]/total_votes) for k,target in target_votes.items()}\n```",
    "2633353": "ok, understand",
    "2641812": "Thanks for providing this, and thanks to all the commentators. Insightful information.",
    "2704278": "Could anyone be kind enough to explain the first conclusion along with what the author means by \n\"with Equal probability 1/6: LB 1.09\n\nWith Overlapping LB: 1.1\"\n\nI tried to get the same numbers but only in this scenario I failed. I am assuming equal probability means author put all targets value as 0.17 and submitted it. I can't figure out what he means by with overlapping. \nAlso what does getting better score by equal probability mean? In the other 2 conclusions I understand that the case in which better score is achieved we can assume then test case has similar data. But for 1st conclusion I am having a hard time understanding if anyone could explain it I would be very grateful. Thanks in advance.",
    "2705719": "same i also stuck with this"
  },
  "source": "meta"
}