{
  "id": 471287,
  "title": "There are ~2640 unique EEG ids in the hidden test data!",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/471287",
  "author_name": "Yurnero",
  "post_date": "2024-01-27T17:14:27.650000",
  "votes": 74,
  "comment_count": 14,
  "views": 0,
  "content": "<p>There is a <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471013\" target=\"_blank\">discussion</a> about how many EEG sub-sequences there are in the hidden test data? This knowledge is very useful — it helps us understand the inference time window for stacking/blending a large amount of models / creating data for test dataset, thereby preventing us from wasting submissions due to OOT/OOM.</p>\n<p>We easily can obtain this information using a simple trick I learned from <a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a>.</p>\n<p>Just run the following code:</p>\n<pre><code> pandas  pd\n time\n\ntest = pd.read_csv()\n i  test.eeg_id.unique():\n    time.sleep()\n</code></pre>\n<p>My submission run approximately for <strong>3 hours and 40 minutes</strong> (+- 2 mins). Therefore there are</p>\n<p>$$220*60/5 = 2640 \\,\\,\\text{unique EEG ids!}$$</p>\n<p>Analogously, by using this method, you can determine the number of EEG subsequences, patients, and so on!</p>",
  "messages": [
    {
      "id": 2622736,
      "postDate": "2024-01-27T17:14:27.650Z",
      "content": "<p>There is a <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471013\" target=\"_blank\">discussion</a> about how many EEG sub-sequences there are in the hidden test data? This knowledge is very useful — it helps us understand the inference time window for stacking/blending a large amount of models / creating data for test dataset, thereby preventing us from wasting submissions due to OOT/OOM.</p>\n<p>We easily can obtain this information using a simple trick I learned from <a href=\"https://www.kaggle.com/kyakovlev\" target=\"_blank\">@kyakovlev</a>.</p>\n<p>Just run the following code:</p>\n<pre><code> pandas  pd\n time\n\ntest = pd.read_csv()\n i  test.eeg_id.unique():\n    time.sleep()\n</code></pre>\n<p>My submission run approximately for <strong>3 hours and 40 minutes</strong> (+- 2 mins). Therefore there are</p>\n<p>$$220*60/5 = 2640 \\,\\,\\text{unique EEG ids!}$$</p>\n<p>Analogously, by using this method, you can determine the number of EEG subsequences, patients, and so on!</p>",
      "rawMarkdown": "There is a [discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471013) about how many EEG sub-sequences there are in the hidden test data? This knowledge is very useful — it helps us understand the inference time window for stacking/blending a large amount of models / creating data for test dataset, thereby preventing us from wasting submissions due to OOT/OOM.\n\nWe easily can obtain this information using a simple trick I learned from @kyakovlev.\n\nJust run the following code:\n\n```python\nimport pandas as pd\nimport time\n\ntest = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/test.csv')\nfor i in test.eeg_id.unique():\n    time.sleep(5)\n```\n\nMy submission run approximately for **3 hours and 40 minutes** (+- 2 mins). Therefore there are\n\n$$220*60/5 = 2640 \\,\\,\\text{unique EEG ids!}$$\n\nAnalogously, by using this method, you can determine the number of EEG subsequences, patients, and so on!",
      "votes": 74
    },
    {
      "id": 2624804,
      "postDate": "2024-01-29T02:17:21.893Z",
      "content": "<p>Thanks for the info. This lets us determine how much shakeup there will be. The public test is only 924 eeg whereas train data has 17089 eeg. We can simulate public LBs and private LBs locally but evaluating the CV score on random subsets of 924.</p>",
      "rawMarkdown": "Thanks for the info. This lets us determine how much shakeup there will be. The public test is only 924 eeg whereas train data has 17089 eeg. We can simulate public LBs and private LBs locally but evaluating the CV score on random subsets of 924.",
      "votes": 11,
      "replies": [
        {
          "id": 2624809,
          "postDate": "2024-01-29T02:25:49.830Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in past, you mention about simulation of pub/private LB locally ( but i fail to experiment it well )</p>\n<ol>\n<li>we repeat 1000 times try to split training dataset into ~ 18 parts and ignore extra. find the min and max KL scores of each sub-set with 924 ids.</li>\n<li>are we going to select the sub-sets which correlate with 3 models LB scores ( example top 3 public models ) =&gt; this part i fail, how to do the simulations? or need to train model multiple times on these 18 folds and submit to LB?</li>\n</ol>\n<p></p>\n<p>Never mind, you already explained in detail and code snippet at <a href=\"https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/174589\" target=\"_blank\">Estimate the Shake</a></p>",
          "rawMarkdown": "@cdeotte in past, you mention about simulation of pub/private LB locally ( but i fail to experiment it well )\n1. we repeat 1000 times try to split training dataset into ~ 18 parts and ignore extra. find the min and max KL scores of each sub-set with 924 ids.\n2. are we going to select the sub-sets which correlate with 3 models LB scores ( example top 3 public models ) => this part i fail, how to do the simulations? or need to train model multiple times on these 18 folds and submit to LB?\n\n~~Could you share any reference article or code snippet for better understanding.~~\n\nNever mind, you already explained in detail and code snippet at [Estimate the Shake](https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/174589)",
          "votes": 2
        }
      ]
    },
    {
      "id": 2624791,
      "postDate": "2024-01-29T01:51:11.760Z",
      "content": "<p>Very good post, would it be possible to check if the patients in the test base are the same as those in the training base?</p>",
      "rawMarkdown": "Very good post, would it be possible to check if the patients in the test base are the same as those in the training base?",
      "votes": 1,
      "replies": [
        {
          "id": 2627881,
          "postDate": "2024-01-31T00:03:57.030Z",
          "content": "<p>I used the trick and found that there are no patient_id in the test dataset that are the same as those in the training dataset. This strongly suggests that the organizers used a grouping technique (group fold) based on patient_id during the data preparation for the competition.</p>",
          "rawMarkdown": "I used the trick and found that there are no patient_id in the test dataset that are the same as those in the training dataset. This strongly suggests that the organizers used a grouping technique (group fold) based on patient_id during the data preparation for the competition.",
          "votes": 8,
          "replies": [
            {
              "id": 2628221,
              "postDate": "2024-01-31T07:09:48.087Z",
              "content": "<p>Yea, seems like that</p>",
              "rawMarkdown": "Yea, seems like that",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2623121,
      "postDate": "2024-01-28T01:53:01.317Z",
      "content": "<p><a href=\"https://www.kaggle.com/samson8\" target=\"_blank\">@samson8</a> Thank you, it will be helpful to design cv folds</p>\n<blockquote>\n  <p>How do you track the time? with email notification or monitoring submission every 2mins ?</p>\n</blockquote>\n<ul>\n<li>So, we can know how many total unique patient_id in the hidden set. Will you share about it too.</li>\n</ul>",
      "rawMarkdown": "@samson8 Thank you, it will be helpful to design cv folds\n> How do you track the time? with email notification or monitoring submission every 2mins ?\n\n\n- So, we can know how many total unique patient_id in the hidden set. Will you share about it too.",
      "votes": 1,
      "replies": [
        {
          "id": 2623503,
          "postDate": "2024-01-28T09:12:06.193Z",
          "content": "<p>I guess I'll leave it to the community and I'll update the post if someone brings new information</p>",
          "rawMarkdown": "I guess I'll leave it to the community and I'll update the post if someone brings new information",
          "votes": 1,
          "replies": [
            {
              "id": 2623569,
              "postDate": "2024-01-28T09:50:13.470Z",
              "content": "<p><a href=\"https://www.kaggle.com/samson8\" target=\"_blank\">@samson8</a> how do you calculated time of submission?</p>",
              "rawMarkdown": "@samson8 how do you calculated time of submission?"
            },
            {
              "id": 2623581,
              "postDate": "2024-01-28T09:51:42.003Z",
              "content": "<p>I did it manually. I knew approximately it would run for more than 2 hours. And I tracked it since then</p>",
              "rawMarkdown": "I did it manually. I knew approximately it would run for more than 2 hours. And I tracked it since then",
              "votes": 1
            },
            {
              "id": 2623599,
              "postDate": "2024-01-28T09:57:56.413Z",
              "content": "<p><a href=\"https://www.kaggle.com/samson8\" target=\"_blank\">@samson8</a> Thank you for sharing and your patience </p>",
              "rawMarkdown": "@samson8 Thank you for sharing and your patience "
            }
          ]
        }
      ]
    },
    {
      "id": 2624873,
      "postDate": "2024-01-29T03:58:24.787Z",
      "content": "<p>A more exact method is to raise an exception if the count of unique EEGs is over a threshold, but this requires a bit more submissions</p>",
      "rawMarkdown": "A more exact method is to raise an exception if the count of unique EEGs is over a threshold, but this requires a bit more submissions",
      "votes": 2
    },
    {
      "id": 2623349,
      "postDate": "2024-01-28T07:15:38.593Z",
      "content": "<p>It's quite interesting, most contestants spend time improving the effectiveness of the model, and few people care about timeout, an evaluation metric that cannot be quantified in terms of scores. In my previous participation in a writing quality regression competition, someone compared the RMSE of predicted results on training data with the RMSE submitted to determine the distribution of online and offline data. I am curious, what do you think of your behavior of wasting a submission opportunity like this?</p>",
      "rawMarkdown": "It's quite interesting, most contestants spend time improving the effectiveness of the model, and few people care about timeout, an evaluation metric that cannot be quantified in terms of scores. In my previous participation in a writing quality regression competition, someone compared the RMSE of predicted results on training data with the RMSE submitted to determine the distribution of online and offline data. I am curious, what do you think of your behavior of wasting a submission opportunity like this?",
      "replies": [
        {
          "id": 2623502,
          "postDate": "2024-01-28T09:10:39.710Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2642099,
      "postDate": "2024-02-07T22:26:04.133Z",
      "content": "<p>Hi. With respect but… the private data is not private for something?</p>",
      "rawMarkdown": "Hi. With respect but... the private data is not private for something?"
    }
  ],
  "comments": [
    {
      "id": 2624804,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2024-01-29T02:17:21.893000",
      "content": "<p>Thanks for the info. This lets us determine how much shakeup there will be. The public test is only 924 eeg whereas train data has 17089 eeg. We can simulate public LBs and private LBs locally but evaluating the CV score on random subsets of 924.</p>",
      "votes": 11,
      "replies": [
        {
          "id": 2624809,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-01-29T02:25:49.830000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in past, you mention about simulation of pub/private LB locally ( but i fail to experiment it well )</p>\n<ol>\n<li>we repeat 1000 times try to split training dataset into ~ 18 parts and ignore extra. find the min and max KL scores of each sub-set with 924 ids.</li>\n<li>are we going to select the sub-sets which correlate with 3 models LB scores ( example top 3 public models ) =&gt; this part i fail, how to do the simulations? or need to train model multiple times on these 18 folds and submit to LB?</li>\n</ol>\n<p></p>\n<p>Never mind, you already explained in detail and code snippet at <a href=\"https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/174589\" target=\"_blank\">Estimate the Shake</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2624791,
      "author_name": "Rafael Zimmermann",
      "author_url": "",
      "post_date": "2024-01-29T01:51:11.760000",
      "content": "<p>Very good post, would it be possible to check if the patients in the test base are the same as those in the training base?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2627881,
          "author_name": "Rafael Zimmermann",
          "author_url": "",
          "post_date": "2024-01-31T00:03:57.030000",
          "content": "<p>I used the trick and found that there are no patient_id in the test dataset that are the same as those in the training dataset. This strongly suggests that the organizers used a grouping technique (group fold) based on patient_id during the data preparation for the competition.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2628221,
              "author_name": "Yurnero",
              "author_url": "",
              "post_date": "2024-01-31T07:09:48.087000",
              "content": "<p>Yea, seems like that</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2623121,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-01-28T01:53:01.317000",
      "content": "<p><a href=\"https://www.kaggle.com/samson8\" target=\"_blank\">@samson8</a> Thank you, it will be helpful to design cv folds</p>\n<blockquote>\n  <p>How do you track the time? with email notification or monitoring submission every 2mins ?</p>\n</blockquote>\n<ul>\n<li>So, we can know how many total unique patient_id in the hidden set. Will you share about it too.</li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 2623503,
          "author_name": "Yurnero",
          "author_url": "",
          "post_date": "2024-01-28T09:12:06.193000",
          "content": "<p>I guess I'll leave it to the community and I'll update the post if someone brings new information</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2623569,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-01-28T09:50:13.470000",
              "content": "<p><a href=\"https://www.kaggle.com/samson8\" target=\"_blank\">@samson8</a> how do you calculated time of submission?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2623581,
              "author_name": "Yurnero",
              "author_url": "",
              "post_date": "2024-01-28T09:51:42.003000",
              "content": "<p>I did it manually. I knew approximately it would run for more than 2 hours. And I tracked it since then</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2623599,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-01-28T09:57:56.413000",
              "content": "<p><a href=\"https://www.kaggle.com/samson8\" target=\"_blank\">@samson8</a> Thank you for sharing and your patience </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2624873,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2024-01-29T03:58:24.787000",
      "content": "<p>A more exact method is to raise an exception if the count of unique EEGs is over a threshold, but this requires a bit more submissions</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2623349,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "2024-01-28T07:15:38.593000",
      "content": "<p>It's quite interesting, most contestants spend time improving the effectiveness of the model, and few people care about timeout, an evaluation metric that cannot be quantified in terms of scores. In my previous participation in a writing quality regression competition, someone compared the RMSE of predicted results on training data with the RMSE submitted to determine the distribution of online and offline data. I am curious, what do you think of your behavior of wasting a submission opportunity like this?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2623502,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-28T09:10:39.710000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2642099,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2024-02-07T22:26:04.133000",
      "content": "<p>Hi. With respect but… the private data is not private for something?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2622736": "There is a [discussion](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471013) about how many EEG sub-sequences there are in the hidden test data? This knowledge is very useful — it helps us understand the inference time window for stacking/blending a large amount of models / creating data for test dataset, thereby preventing us from wasting submissions due to OOT/OOM.\n\nWe easily can obtain this information using a simple trick I learned from @kyakovlev.\n\nJust run the following code:\n\n```python\nimport pandas as pd\nimport time\n\ntest = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/test.csv')\nfor i in test.eeg_id.unique():\n    time.sleep(5)\n```\n\nMy submission run approximately for **3 hours and 40 minutes** (+- 2 mins). Therefore there are\n\n$$220*60/5 = 2640 \\,\\,\\text{unique EEG ids!}$$\n\nAnalogously, by using this method, you can determine the number of EEG subsequences, patients, and so on!",
    "2624804": "Thanks for the info. This lets us determine how much shakeup there will be. The public test is only 924 eeg whereas train data has 17089 eeg. We can simulate public LBs and private LBs locally but evaluating the CV score on random subsets of 924.",
    "2624791": "Very good post, would it be possible to check if the patients in the test base are the same as those in the training base?",
    "2623121": "@samson8 Thank you, it will be helpful to design cv folds\n> How do you track the time? with email notification or monitoring submission every 2mins ?\n\n\n- So, we can know how many total unique patient_id in the hidden set. Will you share about it too.",
    "2624873": "A more exact method is to raise an exception if the count of unique EEGs is over a threshold, but this requires a bit more submissions",
    "2623349": "It's quite interesting, most contestants spend time improving the effectiveness of the model, and few people care about timeout, an evaluation metric that cannot be quantified in terms of scores. In my previous participation in a writing quality regression competition, someone compared the RMSE of predicted results on training data with the RMSE submitted to determine the distribution of online and offline data. I am curious, what do you think of your behavior of wasting a submission opportunity like this?",
    "2642099": "Hi. With respect but... the private data is not private for something?"
  }
}