{
  "id": 477461,
  "title": "Hard samples are more important - [LB 0.37]",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/477461",
  "author_name": "Koolo",
  "post_date": "2024-02-16T07:38:41.517000",
  "votes": 114,
  "comment_count": 33,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/seanbearden\" target=\"_blank\">@Sean R.B. Bearden</a> provides a great idea in his insightful <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135\" target=\"_blank\">notebook</a>. He proposes a two-stage training method to address the KL-divergence issue caused by samples with 1-7 votes. </p>\n<ul>\n<li><strong>Motivation</strong>: samples with peaked distributions lead models to predict with exaggerated confidence, resulting in worse performance.</li>\n<li><strong>Method</strong>: two-stage training method (the first stage only uses samples with 1-7 votes, then samples with 10-28 votes are used to fine-tune models).</li>\n</ul>\n<p>However, there is no clear correlation between the number of votes and the distribution of votes.  As evidence, the scatter diagram of total votes - KL Loss (KL Loss between votes and uniform distribution) is as follows.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15433644%2F19065e32cf0f6cabdec631924707a537%2FUntitled.png?generation=1708067270830919&amp;alt=media\" alt=\"scatter diagram of total votes - KL Loss\"></p>\n<p>The calculation of KL Loss is as follows:</p>\n<pre><code>y_data = data[label_cols].values\ny_data = y_data / y_data.(axis=, keepdims=)\ndata[label_cols] = y_data\nlabels = data[[, , , , , ]].values + \n\n\ndata[] = torch.nn.functional.kl_div(\n    torch.log(torch.tensor(labels)),\n    torch.tensor([ / ] * ),\n    reduction=\n).(dim=).numpy()\n(, data[].(), , data[].())\n</code></pre>\n<p>There is an important observation: samples with peaked distributions (KL Loss &gt; 7) present in both groups</p>\n<p>Therefore, according to the motivation and the observation, I suggest dividing train data according to KL Loss. My experimental configs are as follows:</p>\n<ul>\n<li>EffNetB2 (based on <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>'s <a href=\"https://www.kaggle.com/code/alejopaullier/hms-efficientnetb0-pytorch-train/notebook\" target=\"_blank\">notebook</a>)</li>\n<li>First stage: training with all data</li>\n<li>Second stage: training with samples (KL Loss &lt; 5.5).</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>CV (first stage)</th>\n<th>CV (second stage)</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>single-stage</td>\n<td>0.60</td>\n<td>-</td>\n<td>0.41</td>\n</tr>\n<tr>\n<td>two-stage (KL Loss)</td>\n<td>0.60</td>\n<td>0.33</td>\n<td><strong>0.37</strong></td>\n</tr>\n<tr>\n<td>two-stage (total votes)</td>\n<td>0.60</td>\n<td>0.38</td>\n<td>0.41</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p>Please note that CVs of different stages cannot be compared. Because the training and validation of the 2nd stage are based on filtered data. This is a fast implementation. For a more detailed discussion, please refer to <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472\" target=\"_blank\">my comment</a>. I will fix this issue and update the results.</p>\n<hr>\n<p>I fix the issue and the training framework is as follows:</p>\n<pre><code> ():\n    ....\n    total_df = read_train_metadata(data_root)  \n    split_train_validation(total_df, **self.data_conf[])  \n\n    \n    train_df = total_df[total_df[] != self.data_conf[]].reset_index(drop=)\n    val_df = total_df[total_df[] == self.data_conf[]].reset_index(drop=)\n\n     self.optim_conf[][] == :\n        .... \n    ....\n</code></pre>\n<p>As long as the tagging of training and validation is the same (this can be achieved by fixing random seeds), these two stages use the same validation data. The result are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>CV (first stage)</th>\n<th>CV (second stage)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>two-stage (KL Loss)</td>\n<td>0.59836</td>\n<td>0.58767</td>\n</tr>\n</tbody>\n</table>\n<p>CV has been improved.</p>\n<p>More detailed result:</p>\n<table>\n<thead>\n<tr>\n<th>stage</th>\n<th>fold-1</th>\n<th>fold-2</th>\n<th>fold-3</th>\n<th>fold-4</th>\n<th>fold-5</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1st</td>\n<td>0.63056</td>\n<td>0.55998</td>\n<td>0.53577</td>\n<td><strong>0.58305</strong></td>\n<td>0.68245</td>\n</tr>\n<tr>\n<td>2nd</td>\n<td><strong>0.60282</strong></td>\n<td><strong>0.55192</strong></td>\n<td><strong>0.53050</strong></td>\n<td>0.60580</td>\n<td><strong>0.64730</strong></td>\n</tr>\n</tbody>\n</table>\n<p><strong>This result is surprising</strong>. Although there is an improvement in CV (0.60-&gt;0.59), the magnitude of the improvement is much smaller than that of LB (0.41-&gt;0.37). I'm not sure if there's any special reason.</p>",
  "messages": [
    {
      "id": 2654532,
      "postDate": "2024-02-16T07:38:41.517Z",
      "content": "<p><a href=\"https://www.kaggle.com/seanbearden\" target=\"_blank\">@Sean R.B. Bearden</a> provides a great idea in his insightful <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135\" target=\"_blank\">notebook</a>. He proposes a two-stage training method to address the KL-divergence issue caused by samples with 1-7 votes. </p>\n<ul>\n<li><strong>Motivation</strong>: samples with peaked distributions lead models to predict with exaggerated confidence, resulting in worse performance.</li>\n<li><strong>Method</strong>: two-stage training method (the first stage only uses samples with 1-7 votes, then samples with 10-28 votes are used to fine-tune models).</li>\n</ul>\n<p>However, there is no clear correlation between the number of votes and the distribution of votes.  As evidence, the scatter diagram of total votes - KL Loss (KL Loss between votes and uniform distribution) is as follows.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15433644%2F19065e32cf0f6cabdec631924707a537%2FUntitled.png?generation=1708067270830919&amp;alt=media\" alt=\"scatter diagram of total votes - KL Loss\"></p>\n<p>The calculation of KL Loss is as follows:</p>\n<pre><code>y_data = data[label_cols].values\ny_data = y_data / y_data.(axis=, keepdims=)\ndata[label_cols] = y_data\nlabels = data[[, , , , , ]].values + \n\n\ndata[] = torch.nn.functional.kl_div(\n    torch.log(torch.tensor(labels)),\n    torch.tensor([ / ] * ),\n    reduction=\n).(dim=).numpy()\n(, data[].(), , data[].())\n</code></pre>\n<p>There is an important observation: samples with peaked distributions (KL Loss &gt; 7) present in both groups</p>\n<p>Therefore, according to the motivation and the observation, I suggest dividing train data according to KL Loss. My experimental configs are as follows:</p>\n<ul>\n<li>EffNetB2 (based on <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>'s <a href=\"https://www.kaggle.com/code/alejopaullier/hms-efficientnetb0-pytorch-train/notebook\" target=\"_blank\">notebook</a>)</li>\n<li>First stage: training with all data</li>\n<li>Second stage: training with samples (KL Loss &lt; 5.5).</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>CV (first stage)</th>\n<th>CV (second stage)</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>single-stage</td>\n<td>0.60</td>\n<td>-</td>\n<td>0.41</td>\n</tr>\n<tr>\n<td>two-stage (KL Loss)</td>\n<td>0.60</td>\n<td>0.33</td>\n<td><strong>0.37</strong></td>\n</tr>\n<tr>\n<td>two-stage (total votes)</td>\n<td>0.60</td>\n<td>0.38</td>\n<td>0.41</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p>Please note that CVs of different stages cannot be compared. Because the training and validation of the 2nd stage are based on filtered data. This is a fast implementation. For a more detailed discussion, please refer to <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472\" target=\"_blank\">my comment</a>. I will fix this issue and update the results.</p>\n<hr>\n<p>I fix the issue and the training framework is as follows:</p>\n<pre><code> ():\n    ....\n    total_df = read_train_metadata(data_root)  \n    split_train_validation(total_df, **self.data_conf[])  \n\n    \n    train_df = total_df[total_df[] != self.data_conf[]].reset_index(drop=)\n    val_df = total_df[total_df[] == self.data_conf[]].reset_index(drop=)\n\n     self.optim_conf[][] == :\n        .... \n    ....\n</code></pre>\n<p>As long as the tagging of training and validation is the same (this can be achieved by fixing random seeds), these two stages use the same validation data. The result are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>CV (first stage)</th>\n<th>CV (second stage)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>two-stage (KL Loss)</td>\n<td>0.59836</td>\n<td>0.58767</td>\n</tr>\n</tbody>\n</table>\n<p>CV has been improved.</p>\n<p>More detailed result:</p>\n<table>\n<thead>\n<tr>\n<th>stage</th>\n<th>fold-1</th>\n<th>fold-2</th>\n<th>fold-3</th>\n<th>fold-4</th>\n<th>fold-5</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1st</td>\n<td>0.63056</td>\n<td>0.55998</td>\n<td>0.53577</td>\n<td><strong>0.58305</strong></td>\n<td>0.68245</td>\n</tr>\n<tr>\n<td>2nd</td>\n<td><strong>0.60282</strong></td>\n<td><strong>0.55192</strong></td>\n<td><strong>0.53050</strong></td>\n<td>0.60580</td>\n<td><strong>0.64730</strong></td>\n</tr>\n</tbody>\n</table>\n<p><strong>This result is surprising</strong>. Although there is an improvement in CV (0.60-&gt;0.59), the magnitude of the improvement is much smaller than that of LB (0.41-&gt;0.37). I'm not sure if there's any special reason.</p>",
      "rawMarkdown": "[@Sean R.B. Bearden](https://www.kaggle.com/seanbearden) provides a great idea in his insightful [notebook](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135). He proposes a two-stage training method to address the KL-divergence issue caused by samples with 1-7 votes. \n\n- **Motivation**: samples with peaked distributions lead models to predict with exaggerated confidence, resulting in worse performance.\n- **Method**: two-stage training method (the first stage only uses samples with 1-7 votes, then samples with 10-28 votes are used to fine-tune models).\n\nHowever, there is no clear correlation between the number of votes and the distribution of votes.  As evidence, the scatter diagram of total votes - KL Loss (KL Loss between votes and uniform distribution) is as follows.\n\n![scatter diagram of total votes - KL Loss](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15433644%2F19065e32cf0f6cabdec631924707a537%2FUntitled.png?generation=1708067270830919&alt=media)\n\nThe calculation of KL Loss is as follows:\n```python\ny_data = data[label_cols].values\ny_data = y_data / y_data.sum(axis=1, keepdims=True)\ndata[label_cols] = y_data\nlabels = data[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].values + 1e-5\n\n# compute kl-loss with uniform distribution by pytorch\ndata['kl'] = torch.nn.functional.kl_div(\n    torch.log(torch.tensor(labels)),\n    torch.tensor([1 / 6] * 6),\n    reduction='none'\n).sum(dim=1).numpy()\nprint('min:', data['kl'].min(), 'max:', data['kl'].max())\n```\n\nThere is an important observation: samples with peaked distributions (KL Loss > 7) present in both groups\n\nTherefore, according to the motivation and the observation, I suggest dividing train data according to KL Loss. My experimental configs are as follows:\n- EffNetB2 (based on [@alejopaullier](https://www.kaggle.com/alejopaullier)'s [notebook](https://www.kaggle.com/code/alejopaullier/hms-efficientnetb0-pytorch-train/notebook))\n- First stage: training with all data\n- Second stage: training with samples (KL Loss < 5.5).\n\n| Method | CV (first stage) | CV (second stage) | LB |\n| --- | --- | --- | --- |\n| single-stage |  0.60 | - | 0.41 |\n|two-stage (KL Loss)|0.60|0.33|**0.37**|\n|two-stage (total votes)|0.60|0.38|0.41|\n\n---\nPlease note that CVs of different stages cannot be compared. Because the training and validation of the 2nd stage are based on filtered data. This is a fast implementation. For a more detailed discussion, please refer to [my comment](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472). I will fix this issue and update the results.\n\n---\nI fix the issue and the training framework is as follows:\n```python\ndef training():\n    ....\n    total_df = read_train_metadata(data_root)  # load metadata\n    split_train_validation(total_df, **self.data_conf['fold'])  # tag training and validation\n\n    # create train_df & val_df\n    train_df = total_df[total_df['fold'] != self.data_conf['val_fold_id']].reset_index(drop=True)\n    val_df = total_df[total_df['fold'] == self.data_conf['val_fold_id']].reset_index(drop=True)\n\n    if self.optim_conf['train']['stage'] == 2:\n        .... # filter only on train_df\n    ....\n```\n\nAs long as the tagging of training and validation is the same (this can be achieved by fixing random seeds), these two stages use the same validation data. The result are as follows:\n\n| Method | CV (first stage) | CV (second stage) |\n| --- | --- | --- |\n|two-stage (KL Loss)|0.59836|0.58767|\n\nCV has been improved.\n\nMore detailed result:\n\n| stage | fold-1 | fold-2 | fold-3 | fold-4 | fold-5 |\n| --- | --- | --- | --- | --- | --- |\n|1st|0.63056|0.55998|0.53577|**0.58305**|0.68245|\n|2nd|**0.60282**|**0.55192**|**0.53050**|0.60580|**0.64730**|\n\n\n**This result is surprising**. Although there is an improvement in CV (0.60->0.59), the magnitude of the improvement is much smaller than that of LB (0.41->0.37). I'm not sure if there's any special reason.",
      "votes": 112
    },
    {
      "id": 2654899,
      "postDate": "2024-02-16T14:41:04.903Z",
      "content": "<p>did you just train and validated with filtered record in stage 2 ?  I would have thought that you would filter out the &lt;5.5 KLLoss records for train set and continue with the same validation set as stage1 ? </p>",
      "rawMarkdown": "did you just train and validated with filtered record in stage 2 ?  I would have thought that you would filter out the <5.5 KLLoss records for train set and continue with the same validation set as stage1 ? ",
      "votes": 5,
      "replies": [
        {
          "id": 2655477,
          "postDate": "2024-02-16T22:51:04.567Z",
          "content": "<p>Training and validation are based on filtered data in the 2nd stage. I discuss my training framework in this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472\" target=\"_blank\">comment</a>. And you are right, using the same validation set is a better method, I will fix it and update CVs.</p>",
          "rawMarkdown": "Training and validation are based on filtered data in the 2nd stage. I discuss my training framework in this [comment](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472). And you are right, using the same validation set is a better method, I will fix it and update CVs.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2654920,
      "postDate": "2024-02-16T15:04:58.210Z",
      "content": "<p>Do you make sure you are using same split in stage 1 and stage 2? Looks to me that choosing samples with KL&lt;5.5 is causing mild leakage to your strategy, as CV is too optimal.</p>",
      "rawMarkdown": "Do you make sure you are using same split in stage 1 and stage 2? Looks to me that choosing samples with KL<5.5 is causing mild leakage to your strategy, as CV is too optimal.",
      "votes": 6,
      "replies": [
        {
          "id": 2654949,
          "postDate": "2024-02-16T15:37:23.423Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2654953,
          "postDate": "2024-02-16T15:37:57.683Z",
          "content": "<p>stage 1: all data<br>\nstage 2: all data[kl&lt;5.5]<br>\nWouldn't this strategy create leaks?</p>",
          "rawMarkdown": "stage 1: all data\nstage 2: all data[kl<5.5]\nWouldn't this strategy create leaks?",
          "replies": [
            {
              "id": 2654973,
              "postDate": "2024-02-16T15:53:09.430Z",
              "content": "<p>But how do you create splits?<br>\nlet's say you create splits in stage-1  and then in stage-2, you need to make sure both have the same sample that is (all samples of fold0 stage-1 must remain in all samples of fold 0 in stage-2, same for all other folds).</p>",
              "rawMarkdown": "But how do you create splits?\nlet's say you create splits in stage-1  and then in stage-2, you need to make sure both have the same sample that is (all samples of fold0 stage-1 must remain in all samples of fold 0 in stage-2, same for all other folds).",
              "votes": 2
            },
            {
              "id": 2654991,
              "postDate": "2024-02-16T16:11:04.910Z",
              "content": "<p>You are right, the validation set with the same distribution should be used in the second stage of training</p>",
              "rawMarkdown": "You are right, the validation set with the same distribution should be used in the second stage of training"
            }
          ]
        },
        {
          "id": 2655480,
          "postDate": "2024-02-16T22:56:32.917Z",
          "content": "<p>I discuss my training framework in this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472\" target=\"_blank\">comment</a>. Overall, there is no leakage. However, training and validation are based on filtered data in stage 2, it is better to use the same validation data in both stages. I will fix it.</p>",
          "rawMarkdown": "I discuss my training framework in this [comment](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472). Overall, there is no leakage. However, training and validation are based on filtered data in stage 2, it is better to use the same validation data in both stages. I will fix it."
        }
      ]
    },
    {
      "id": 2655613,
      "postDate": "2024-02-17T04:53:28.110Z",
      "content": "<p>My experience is that the Public LB score for one-stage training is 0.43, while the Public LB score for two-stage training is 0.36, which is a strange result.</p>",
      "rawMarkdown": "My experience is that the Public LB score for one-stage training is 0.43, while the Public LB score for two-stage training is 0.36, which is a strange result.",
      "votes": 5,
      "replies": [
        {
          "id": 2655625,
          "postDate": "2024-02-17T05:10:10.710Z",
          "content": "<p>Is it possible that it is caused by the evaluation metric (KL divergence), which is sensitive to extreme predictions? Experiments seem to imply that smooth prediction results have more obvious advantages.</p>",
          "rawMarkdown": "Is it possible that it is caused by the evaluation metric (KL divergence), which is sensitive to extreme predictions? Experiments seem to imply that smooth prediction results have more obvious advantages.",
          "votes": 3,
          "replies": [
            {
              "id": 2655633,
              "postDate": "2024-02-17T05:21:59.050Z",
              "content": "<p>Maybe hard samples like your title describes are more important in this competition metrics. However, the distribution of Private and Public may lead to large differences in the final results.</p>",
              "rawMarkdown": "Maybe hard samples like your title describes are more important in this competition metrics. However, the distribution of Private and Public may lead to large differences in the final results.",
              "votes": 3
            },
            {
              "id": 2655654,
              "postDate": "2024-02-17T05:45:47.047Z",
              "content": "<p>This is indeed a crucial issue. <br>\nHave you tried using the same validation data in two stages, or is it similar to my implementation method (validation of the 2nd stage is filtered data)? I think using the same validation data can better evaluate the importance of hard samples.</p>",
              "rawMarkdown": "This is indeed a crucial issue. \nHave you tried using the same validation data in two stages, or is it similar to my implementation method (validation of the 2nd stage is filtered data)? I think using the same validation data can better evaluate the importance of hard samples."
            },
            {
              "id": 2655709,
              "postDate": "2024-02-17T06:52:02.947Z",
              "content": "<p>I'm now validating with validation data from the respective stages.</p>",
              "rawMarkdown": "I'm now validating with validation data from the respective stages."
            },
            {
              "id": 2655851,
              "postDate": "2024-02-17T08:44:12.513Z",
              "content": "<p>I update the result with the same validation data. The result is indeed quite surprising. The improvement of CV is much smaller than LB.</p>",
              "rawMarkdown": "I update the result with the same validation data. The result is indeed quite surprising. The improvement of CV is much smaller than LB.",
              "votes": 1
            },
            {
              "id": 2655895,
              "postDate": "2024-02-17T09:08:22.637Z",
              "content": "<p>I tried the same configuration as you, stage1: train on all data, stage2: train and verify on data with kl&lt;5.5, this did not improve my single-mode lb</p>",
              "rawMarkdown": "I tried the same configuration as you, stage1: train on all data, stage2: train and verify on data with kl<5.5, this did not improve my single-mode lb"
            },
            {
              "id": 2656689,
              "postDate": "2024-02-18T00:14:31.390Z",
              "content": "<p><a href=\"https://www.kaggle.com/gentlezdh\" target=\"_blank\">@gentlezdh</a> have you tried to lower the learning rate?<br>\nI tried the Two stages on - kaggle specs only - model, but I've reset the learning rate for the stages, CV and LB got worst.<br>\nMaybe the learning rate for the second stage should be much lower.</p>",
              "rawMarkdown": "@gentlezdh have you tried to lower the learning rate?\nI tried the Two stages on - kaggle specs only - model, but I've reset the learning rate for the stages, CV and LB got worst.\nMaybe the learning rate for the second stage should be much lower."
            },
            {
              "id": 2656789,
              "postDate": "2024-02-18T02:38:27.037Z",
              "content": "<p>I conducted more experiments based on different backbones (effnetB0/1/2, resnet18/34) and training configs. Two-stage training seems to have limitations. The efficiency is limited for baselines with LB&lt;=0.39 (same or worse). But for baselines with LB&gt;0.41, there is an improvement in most cases.</p>",
              "rawMarkdown": "I conducted more experiments based on different backbones (effnetB0/1/2, resnet18/34) and training configs. Two-stage training seems to have limitations. The efficiency is limited for baselines with LB<=0.39 (same or worse). But for baselines with LB>0.41, there is an improvement in most cases.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2670651,
      "postDate": "2024-02-27T03:39:06.027Z",
      "content": "<p>Nice work!!<br>\nsingle-stage     CV (first stage) [0.485]                 LB [0.34]<br>\ntwo-stage       CV (first stage) [0.485]      CV (second stage) [0.325]    LB [0.32]</p>",
      "rawMarkdown": "Nice work!!\nsingle-stage \tCV (first stage) [0.485]                 LB [0.34]\ntwo-stage       CV (first stage) [0.485]      CV (second stage) [0.325]\tLB [0.32]\n",
      "votes": 4,
      "replies": [
        {
          "id": 2670660,
          "postDate": "2024-02-27T03:57:22.567Z",
          "content": "<p>Its result is Single Models? or ensemble?</p>",
          "rawMarkdown": "Its result is Single Models? or ensemble?",
          "replies": [
            {
              "id": 2670727,
              "postDate": "2024-02-27T05:01:43.800Z",
              "content": "<p>ensemble      spe + raweeg</p>",
              "rawMarkdown": "ensemble      spe + raweeg",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2657511,
      "postDate": "2024-02-18T14:52:01.500Z",
      "content": "<p>I have to say that I am impressed, it worked!!!<br>\nI managed to get the - kaggle only - model score from [CV 0.62 – LB 0.47] to [CV 0.6365 – LB 0.43], CV got worst, but LB got a lot better. I am planning on retraining the other model types then rensemble.<br>\nI released the new version of the code with LB 0.36 <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "I have to say that I am impressed, it worked!!!\nI managed to get the - kaggle only - model score from [CV 0.62 – LB 0.47] to [CV 0.6365 – LB 0.43], CV got worst, but LB got a lot better. I am planning on retraining the other model types then rensemble.\nI released the new version of the code with LB 0.36 [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666).",
      "votes": 4,
      "replies": [
        {
          "id": 2657556,
          "postDate": "2024-02-18T15:29:05.867Z",
          "content": "<p><a href=\"https://www.kaggle.com/nartaa\" target=\"_blank\">@nartaa</a> Good job! This really helps me a lot.</p>",
          "rawMarkdown": "@nartaa Good job! This really helps me a lot.",
          "votes": 1,
          "replies": [
            {
              "id": 2657893,
              "postDate": "2024-02-18T20:16:01.680Z",
              "content": "<p><a href=\"https://www.kaggle.com/gentlezdh\" target=\"_blank\">@gentlezdh</a> Thank you! You can check out the code from <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666\" target=\"_blank\">here</a></p>",
              "rawMarkdown": "@gentlezdh Thank you! You can check out the code from [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666)",
              "votes": 1
            },
            {
              "id": 2658135,
              "postDate": "2024-02-19T01:59:14.270Z",
              "content": "<p>I'm glad this method works for you! The code shared by you is insightful. I think this method is worth exploring further.</p>",
              "rawMarkdown": "I'm glad this method works for you! The code shared by you is insightful. I think this method is worth exploring further.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2655472,
      "postDate": "2024-02-16T22:42:58.260Z",
      "content": "<p>Hi. The training loop is as follows (I use Hydra to manage hyper-parameters): </p>\n<pre><code> fold_id  (n_fold):\n    current_cfg = cfg.copy()\n    current_cfg[][] = fold_id\n\n    \n    current_cfg[][][] = \n    best_model_score_1, best_model_path_1 = training(current_cfg)\n\n    \n    current_cfg[][][] = \n    best_model_score_2, _ = training(current_cfg)\n</code></pre>\n<p>All seeds have also been fixed. Filtering is performed after tag training and validation. </p>\n<pre><code> ():\n    ....\n    total_df = read_train_metadata(data_root)  \n    split_train_validation(total_df, **self.data_conf[])  \n     self.optim_conf[][] == :\n        .... \n    ....\n</code></pre>\n<p>So I am certain that there is no leakage.</p>\n<p>The reason for the small CV in the second stage is that it was only trained and tested on the filtered data (as with motivation, samples with peaked distributions are more sensitive in evaluation). Therefore, there are some minor flaws: the data used for validation in the second stage is a <strong>subset</strong> (KL Loss &lt; 5.5) of the data used in the first stage.</p>\n<p>For my training framework, this is a fast implementation because it only requires filtering after tagging (a few lines of code). CVs of different stages don't have comparability (at the same stage, it is possible to).</p>\n<p>Thank you for pointing out this issue. I will fix it and update my experimental results.</p>",
      "rawMarkdown": "Hi. The training loop is as follows (I use Hydra to manage hyper-parameters): \n```python\nfor fold_id in range(n_fold):\n    current_cfg = cfg.copy()\n    current_cfg['data_conf']['val_fold_id'] = fold_id\n\n    # stage I\n    current_cfg['optim_conf']['train']['stage'] = 1\n    best_model_score_1, best_model_path_1 = training(current_cfg)\n\n    # stage II\n    current_cfg['optim_conf']['train']['stage'] = 2\n    best_model_score_2, _ = training(current_cfg)\n```\nAll seeds have also been fixed. Filtering is performed after tag training and validation. \n```python\ndef training():\n    ....\n    total_df = read_train_metadata(data_root)  # load metadata\n    split_train_validation(total_df, **self.data_conf['fold'])  # tag training and validation\n    if self.optim_conf['train']['stage'] == 2:\n        .... # filter\n    ....\n```\nSo I am certain that there is no leakage.\n\nThe reason for the small CV in the second stage is that it was only trained and tested on the filtered data (as with motivation, samples with peaked distributions are more sensitive in evaluation). Therefore, there are some minor flaws: the data used for validation in the second stage is a **subset** (KL Loss < 5.5) of the data used in the first stage.\n\nFor my training framework, this is a fast implementation because it only requires filtering after tagging (a few lines of code). CVs of different stages don't have comparability (at the same stage, it is possible to).\n\nThank you for pointing out this issue. I will fix it and update my experimental results.",
      "votes": 1
    },
    {
      "id": 2703636,
      "postDate": "2024-03-18T10:14:18.380Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/zijiangyang1116\" target=\"_blank\">@zijiangyang1116</a><br>\nShouldn't the title be <strong>Soft samples are more important</strong> , since we are training the second stage on soft samples, I m assuming by a hard sample you mean - the ones closer to a degenerate distribution.<br>\nAnd the soft samples are the ones closer to a uniform distribution.</p>",
      "rawMarkdown": "Hello @zijiangyang1116\nShouldn't the title be **Soft samples are more important** , since we are training the second stage on soft samples, I m assuming by a hard sample you mean - the ones closer to a degenerate distribution.\nAnd the soft samples are the ones closer to a uniform distribution.",
      "replies": [
        {
          "id": 2704753,
          "postDate": "2024-03-18T22:42:36.637Z",
          "content": "<p>I think he meant hard sample in the sense that the experts find it hard to classify. I.e. they are unsure which option it is and the distribution is closer to uniform</p>",
          "rawMarkdown": "I think he meant hard sample in the sense that the experts find it hard to classify. I.e. they are unsure which option it is and the distribution is closer to uniform"
        }
      ]
    },
    {
      "id": 2665331,
      "postDate": "2024-02-23T15:42:47.610Z",
      "content": "<p>Thanks for sharing. I have two questions.</p>\n<ol>\n<li>How did you find threshold=5.5 ? By a lot of experiment?</li>\n<li>About two-stage (total votes) in result table, how did you train model ? Did you train in the same way as Sean R.B. Bearden ? or Did you train with all data and then filter like total_evaluators &gt; threshold ?</li>\n</ol>",
      "rawMarkdown": "Thanks for sharing. I have two questions.\n1. How did you find threshold=5.5 ? By a lot of experiment?\n2. About two-stage (total votes) in result table, how did you train model ? Did you train in the same way as Sean R.B. Bearden ? or Did you train with all data and then filter like total_evaluators > threshold ?",
      "replies": [
        {
          "id": 2666521,
          "postDate": "2024-02-24T12:39:54.263Z",
          "content": "<ol>\n<li><p>At first, it was just intuition. Because threshold=5.5 can retain enough samples for training while excluding extreme samples. I also conducted many experiments later, and threshold=5.5 is an optimal choice in most cases.</p></li>\n<li><p>I train with all data in the 1st stage and then with filtered data (the 2nd stage is the same as Sean R.B. Bearden's method).</p></li>\n</ol>",
          "rawMarkdown": "1. At first, it was just intuition. Because threshold=5.5 can retain enough samples for training while excluding extreme samples. I also conducted many experiments later, and threshold=5.5 is an optimal choice in most cases.\n\n2. I train with all data in the 1st stage and then with filtered data (the 2nd stage is the same as Sean R.B. Bearden's method).",
          "replies": [
            {
              "id": 2666583,
              "postDate": "2024-02-24T13:18:52.367Z",
              "content": "<p>Thanks. Filtering train_df by kl_div worked for me</p>",
              "rawMarkdown": "Thanks. Filtering train_df by kl_div worked for me",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2655261,
      "postDate": "2024-02-16T19:05:05.063Z",
      "content": "<p>did you just train and validated with filtered record in stage 2!! plz sure once again</p>",
      "rawMarkdown": "did you just train and validated with filtered record in stage 2!! plz sure once again",
      "replies": [
        {
          "id": 2655473,
          "postDate": "2024-02-16T22:47:43.877Z",
          "content": "<p>Yes, training and validation are based on filtered data. I discuss my training framework in this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472\" target=\"_blank\">comment</a>. You can refer to it.</p>",
          "rawMarkdown": "Yes, training and validation are based on filtered data. I discuss my training framework in this [comment](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472). You can refer to it."
        }
      ]
    },
    {
      "id": 2663269,
      "postDate": "2024-02-22T12:01:37.273Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2654899,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2024-02-16T14:41:04.903000",
      "content": "<p>did you just train and validated with filtered record in stage 2 ?  I would have thought that you would filter out the &lt;5.5 KLLoss records for train set and continue with the same validation set as stage1 ? </p>",
      "votes": 5,
      "replies": [
        {
          "id": 2655477,
          "author_name": "Koolo",
          "author_url": "",
          "post_date": "2024-02-16T22:51:04.567000",
          "content": "<p>Training and validation are based on filtered data in the 2nd stage. I discuss my training framework in this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472\" target=\"_blank\">comment</a>. And you are right, using the same validation set is a better method, I will fix it and update CVs.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2654920,
      "author_name": "Priyanshu Chaudhary",
      "author_url": "",
      "post_date": "2024-02-16T15:04:58.210000",
      "content": "<p>Do you make sure you are using same split in stage 1 and stage 2? Looks to me that choosing samples with KL&lt;5.5 is causing mild leakage to your strategy, as CV is too optimal.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2654949,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-16T15:37:23.423000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2654953,
          "author_name": "Donghui Zhang",
          "author_url": "",
          "post_date": "2024-02-16T15:37:57.683000",
          "content": "<p>stage 1: all data<br>\nstage 2: all data[kl&lt;5.5]<br>\nWouldn't this strategy create leaks?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2654973,
              "author_name": "Priyanshu Chaudhary",
              "author_url": "",
              "post_date": "2024-02-16T15:53:09.430000",
              "content": "<p>But how do you create splits?<br>\nlet's say you create splits in stage-1  and then in stage-2, you need to make sure both have the same sample that is (all samples of fold0 stage-1 must remain in all samples of fold 0 in stage-2, same for all other folds).</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2654991,
              "author_name": "Donghui Zhang",
              "author_url": "",
              "post_date": "2024-02-16T16:11:04.910000",
              "content": "<p>You are right, the validation set with the same distribution should be used in the second stage of training</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2655480,
          "author_name": "Koolo",
          "author_url": "",
          "post_date": "2024-02-16T22:56:32.917000",
          "content": "<p>I discuss my training framework in this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472\" target=\"_blank\">comment</a>. Overall, there is no leakage. However, training and validation are based on filtered data in stage 2, it is better to use the same validation data in both stages. I will fix it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2655613,
      "author_name": "ynhuhu",
      "author_url": "",
      "post_date": "2024-02-17T04:53:28.110000",
      "content": "<p>My experience is that the Public LB score for one-stage training is 0.43, while the Public LB score for two-stage training is 0.36, which is a strange result.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2655625,
          "author_name": "Koolo",
          "author_url": "",
          "post_date": "2024-02-17T05:10:10.710000",
          "content": "<p>Is it possible that it is caused by the evaluation metric (KL divergence), which is sensitive to extreme predictions? Experiments seem to imply that smooth prediction results have more obvious advantages.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2655633,
              "author_name": "ynhuhu",
              "author_url": "",
              "post_date": "2024-02-17T05:21:59.050000",
              "content": "<p>Maybe hard samples like your title describes are more important in this competition metrics. However, the distribution of Private and Public may lead to large differences in the final results.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2655654,
              "author_name": "Koolo",
              "author_url": "",
              "post_date": "2024-02-17T05:45:47.047000",
              "content": "<p>This is indeed a crucial issue. <br>\nHave you tried using the same validation data in two stages, or is it similar to my implementation method (validation of the 2nd stage is filtered data)? I think using the same validation data can better evaluate the importance of hard samples.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2655709,
              "author_name": "ynhuhu",
              "author_url": "",
              "post_date": "2024-02-17T06:52:02.947000",
              "content": "<p>I'm now validating with validation data from the respective stages.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2655851,
              "author_name": "Koolo",
              "author_url": "",
              "post_date": "2024-02-17T08:44:12.513000",
              "content": "<p>I update the result with the same validation data. The result is indeed quite surprising. The improvement of CV is much smaller than LB.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2655895,
              "author_name": "Donghui Zhang",
              "author_url": "",
              "post_date": "2024-02-17T09:08:22.637000",
              "content": "<p>I tried the same configuration as you, stage1: train on all data, stage2: train and verify on data with kl&lt;5.5, this did not improve my single-mode lb</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2656689,
              "author_name": "Danial Zakaria",
              "author_url": "",
              "post_date": "2024-02-18T00:14:31.390000",
              "content": "<p><a href=\"https://www.kaggle.com/gentlezdh\" target=\"_blank\">@gentlezdh</a> have you tried to lower the learning rate?<br>\nI tried the Two stages on - kaggle specs only - model, but I've reset the learning rate for the stages, CV and LB got worst.<br>\nMaybe the learning rate for the second stage should be much lower.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2656789,
              "author_name": "Koolo",
              "author_url": "",
              "post_date": "2024-02-18T02:38:27.037000",
              "content": "<p>I conducted more experiments based on different backbones (effnetB0/1/2, resnet18/34) and training configs. Two-stage training seems to have limitations. The efficiency is limited for baselines with LB&lt;=0.39 (same or worse). But for baselines with LB&gt;0.41, there is an improvement in most cases.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2670651,
      "author_name": "Yuri Sun",
      "author_url": "",
      "post_date": "2024-02-27T03:39:06.027000",
      "content": "<p>Nice work!!<br>\nsingle-stage     CV (first stage) [0.485]                 LB [0.34]<br>\ntwo-stage       CV (first stage) [0.485]      CV (second stage) [0.325]    LB [0.32]</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2670660,
          "author_name": "Haru",
          "author_url": "",
          "post_date": "2024-02-27T03:57:22.567000",
          "content": "<p>Its result is Single Models? or ensemble?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2670727,
              "author_name": "Yuri Sun",
              "author_url": "",
              "post_date": "2024-02-27T05:01:43.800000",
              "content": "<p>ensemble      spe + raweeg</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2657511,
      "author_name": "Danial Zakaria",
      "author_url": "",
      "post_date": "2024-02-18T14:52:01.500000",
      "content": "<p>I have to say that I am impressed, it worked!!!<br>\nI managed to get the - kaggle only - model score from [CV 0.62 – LB 0.47] to [CV 0.6365 – LB 0.43], CV got worst, but LB got a lot better. I am planning on retraining the other model types then rensemble.<br>\nI released the new version of the code with LB 0.36 <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666\" target=\"_blank\">here</a>.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2657556,
          "author_name": "Donghui Zhang",
          "author_url": "",
          "post_date": "2024-02-18T15:29:05.867000",
          "content": "<p><a href=\"https://www.kaggle.com/nartaa\" target=\"_blank\">@nartaa</a> Good job! This really helps me a lot.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2657893,
              "author_name": "Danial Zakaria",
              "author_url": "",
              "post_date": "2024-02-18T20:16:01.680000",
              "content": "<p><a href=\"https://www.kaggle.com/gentlezdh\" target=\"_blank\">@gentlezdh</a> Thank you! You can check out the code from <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666\" target=\"_blank\">here</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2658135,
              "author_name": "Koolo",
              "author_url": "",
              "post_date": "2024-02-19T01:59:14.270000",
              "content": "<p>I'm glad this method works for you! The code shared by you is insightful. I think this method is worth exploring further.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2655472,
      "author_name": "Koolo",
      "author_url": "",
      "post_date": "2024-02-16T22:42:58.260000",
      "content": "<p>Hi. The training loop is as follows (I use Hydra to manage hyper-parameters): </p>\n<pre><code> fold_id  (n_fold):\n    current_cfg = cfg.copy()\n    current_cfg[][] = fold_id\n\n    \n    current_cfg[][][] = \n    best_model_score_1, best_model_path_1 = training(current_cfg)\n\n    \n    current_cfg[][][] = \n    best_model_score_2, _ = training(current_cfg)\n</code></pre>\n<p>All seeds have also been fixed. Filtering is performed after tag training and validation. </p>\n<pre><code> ():\n    ....\n    total_df = read_train_metadata(data_root)  \n    split_train_validation(total_df, **self.data_conf[])  \n     self.optim_conf[][] == :\n        .... \n    ....\n</code></pre>\n<p>So I am certain that there is no leakage.</p>\n<p>The reason for the small CV in the second stage is that it was only trained and tested on the filtered data (as with motivation, samples with peaked distributions are more sensitive in evaluation). Therefore, there are some minor flaws: the data used for validation in the second stage is a <strong>subset</strong> (KL Loss &lt; 5.5) of the data used in the first stage.</p>\n<p>For my training framework, this is a fast implementation because it only requires filtering after tagging (a few lines of code). CVs of different stages don't have comparability (at the same stage, it is possible to).</p>\n<p>Thank you for pointing out this issue. I will fix it and update my experimental results.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2703636,
      "author_name": "Danial Zakaria",
      "author_url": "",
      "post_date": "2024-03-18T10:14:18.380000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/zijiangyang1116\" target=\"_blank\">@zijiangyang1116</a><br>\nShouldn't the title be <strong>Soft samples are more important</strong> , since we are training the second stage on soft samples, I m assuming by a hard sample you mean - the ones closer to a degenerate distribution.<br>\nAnd the soft samples are the ones closer to a uniform distribution.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2704753,
          "author_name": "Danilo Jr Dela Cruz",
          "author_url": "",
          "post_date": "2024-03-18T22:42:36.637000",
          "content": "<p>I think he meant hard sample in the sense that the experts find it hard to classify. I.e. they are unsure which option it is and the distribution is closer to uniform</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2665331,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2024-02-23T15:42:47.610000",
      "content": "<p>Thanks for sharing. I have two questions.</p>\n<ol>\n<li>How did you find threshold=5.5 ? By a lot of experiment?</li>\n<li>About two-stage (total votes) in result table, how did you train model ? Did you train in the same way as Sean R.B. Bearden ? or Did you train with all data and then filter like total_evaluators &gt; threshold ?</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 2666521,
          "author_name": "Koolo",
          "author_url": "",
          "post_date": "2024-02-24T12:39:54.263000",
          "content": "<ol>\n<li><p>At first, it was just intuition. Because threshold=5.5 can retain enough samples for training while excluding extreme samples. I also conducted many experiments later, and threshold=5.5 is an optimal choice in most cases.</p></li>\n<li><p>I train with all data in the 1st stage and then with filtered data (the 2nd stage is the same as Sean R.B. Bearden's method).</p></li>\n</ol>",
          "votes": 0,
          "replies": [
            {
              "id": 2666583,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-02-24T13:18:52.367000",
              "content": "<p>Thanks. Filtering train_df by kl_div worked for me</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2655261,
      "author_name": "Tanishq dublish",
      "author_url": "",
      "post_date": "2024-02-16T19:05:05.063000",
      "content": "<p>did you just train and validated with filtered record in stage 2!! plz sure once again</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2655473,
          "author_name": "Koolo",
          "author_url": "",
          "post_date": "2024-02-16T22:47:43.877000",
          "content": "<p>Yes, training and validation are based on filtered data. I discuss my training framework in this <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472\" target=\"_blank\">comment</a>. You can refer to it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2663269,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-22T12:01:37.273000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2654532": "[@Sean R.B. Bearden](https://www.kaggle.com/seanbearden) provides a great idea in his insightful [notebook](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135). He proposes a two-stage training method to address the KL-divergence issue caused by samples with 1-7 votes. \n\n- **Motivation**: samples with peaked distributions lead models to predict with exaggerated confidence, resulting in worse performance.\n- **Method**: two-stage training method (the first stage only uses samples with 1-7 votes, then samples with 10-28 votes are used to fine-tune models).\n\nHowever, there is no clear correlation between the number of votes and the distribution of votes.  As evidence, the scatter diagram of total votes - KL Loss (KL Loss between votes and uniform distribution) is as follows.\n\n![scatter diagram of total votes - KL Loss](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15433644%2F19065e32cf0f6cabdec631924707a537%2FUntitled.png?generation=1708067270830919&alt=media)\n\nThe calculation of KL Loss is as follows:\n```python\ny_data = data[label_cols].values\ny_data = y_data / y_data.sum(axis=1, keepdims=True)\ndata[label_cols] = y_data\nlabels = data[['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']].values + 1e-5\n\n# compute kl-loss with uniform distribution by pytorch\ndata['kl'] = torch.nn.functional.kl_div(\n    torch.log(torch.tensor(labels)),\n    torch.tensor([1 / 6] * 6),\n    reduction='none'\n).sum(dim=1).numpy()\nprint('min:', data['kl'].min(), 'max:', data['kl'].max())\n```\n\nThere is an important observation: samples with peaked distributions (KL Loss > 7) present in both groups\n\nTherefore, according to the motivation and the observation, I suggest dividing train data according to KL Loss. My experimental configs are as follows:\n- EffNetB2 (based on [@alejopaullier](https://www.kaggle.com/alejopaullier)'s [notebook](https://www.kaggle.com/code/alejopaullier/hms-efficientnetb0-pytorch-train/notebook))\n- First stage: training with all data\n- Second stage: training with samples (KL Loss < 5.5).\n\n| Method | CV (first stage) | CV (second stage) | LB |\n| --- | --- | --- | --- |\n| single-stage |  0.60 | - | 0.41 |\n|two-stage (KL Loss)|0.60|0.33|**0.37**|\n|two-stage (total votes)|0.60|0.38|0.41|\n\n---\nPlease note that CVs of different stages cannot be compared. Because the training and validation of the 2nd stage are based on filtered data. This is a fast implementation. For a more detailed discussion, please refer to [my comment](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461#2655472). I will fix this issue and update the results.\n\n---\nI fix the issue and the training framework is as follows:\n```python\ndef training():\n    ....\n    total_df = read_train_metadata(data_root)  # load metadata\n    split_train_validation(total_df, **self.data_conf['fold'])  # tag training and validation\n\n    # create train_df & val_df\n    train_df = total_df[total_df['fold'] != self.data_conf['val_fold_id']].reset_index(drop=True)\n    val_df = total_df[total_df['fold'] == self.data_conf['val_fold_id']].reset_index(drop=True)\n\n    if self.optim_conf['train']['stage'] == 2:\n        .... # filter only on train_df\n    ....\n```\n\nAs long as the tagging of training and validation is the same (this can be achieved by fixing random seeds), these two stages use the same validation data. The result are as follows:\n\n| Method | CV (first stage) | CV (second stage) |\n| --- | --- | --- |\n|two-stage (KL Loss)|0.59836|0.58767|\n\nCV has been improved.\n\nMore detailed result:\n\n| stage | fold-1 | fold-2 | fold-3 | fold-4 | fold-5 |\n| --- | --- | --- | --- | --- | --- |\n|1st|0.63056|0.55998|0.53577|**0.58305**|0.68245|\n|2nd|**0.60282**|**0.55192**|**0.53050**|0.60580|**0.64730**|\n\n\n**This result is surprising**. Although there is an improvement in CV (0.60->0.59), the magnitude of the improvement is much smaller than that of LB (0.41->0.37). I'm not sure if there's any special reason.",
    "2654899": "did you just train and validated with filtered record in stage 2 ?  I would have thought that you would filter out the <5.5 KLLoss records for train set and continue with the same validation set as stage1 ? ",
    "2654920": "Do you make sure you are using same split in stage 1 and stage 2? Looks to me that choosing samples with KL<5.5 is causing mild leakage to your strategy, as CV is too optimal.",
    "2655613": "My experience is that the Public LB score for one-stage training is 0.43, while the Public LB score for two-stage training is 0.36, which is a strange result.",
    "2670651": "Nice work!!\nsingle-stage \tCV (first stage) [0.485]                 LB [0.34]\ntwo-stage       CV (first stage) [0.485]      CV (second stage) [0.325]\tLB [0.32]\n",
    "2657511": "I have to say that I am impressed, it worked!!!\nI managed to get the - kaggle only - model score from [CV 0.62 – LB 0.47] to [CV 0.6365 – LB 0.43], CV got worst, but LB got a lot better. I am planning on retraining the other model types then rensemble.\nI released the new version of the code with LB 0.36 [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469666).",
    "2655472": "Hi. The training loop is as follows (I use Hydra to manage hyper-parameters): \n```python\nfor fold_id in range(n_fold):\n    current_cfg = cfg.copy()\n    current_cfg['data_conf']['val_fold_id'] = fold_id\n\n    # stage I\n    current_cfg['optim_conf']['train']['stage'] = 1\n    best_model_score_1, best_model_path_1 = training(current_cfg)\n\n    # stage II\n    current_cfg['optim_conf']['train']['stage'] = 2\n    best_model_score_2, _ = training(current_cfg)\n```\nAll seeds have also been fixed. Filtering is performed after tag training and validation. \n```python\ndef training():\n    ....\n    total_df = read_train_metadata(data_root)  # load metadata\n    split_train_validation(total_df, **self.data_conf['fold'])  # tag training and validation\n    if self.optim_conf['train']['stage'] == 2:\n        .... # filter\n    ....\n```\nSo I am certain that there is no leakage.\n\nThe reason for the small CV in the second stage is that it was only trained and tested on the filtered data (as with motivation, samples with peaked distributions are more sensitive in evaluation). Therefore, there are some minor flaws: the data used for validation in the second stage is a **subset** (KL Loss < 5.5) of the data used in the first stage.\n\nFor my training framework, this is a fast implementation because it only requires filtering after tagging (a few lines of code). CVs of different stages don't have comparability (at the same stage, it is possible to).\n\nThank you for pointing out this issue. I will fix it and update my experimental results.",
    "2703636": "Hello @zijiangyang1116\nShouldn't the title be **Soft samples are more important** , since we are training the second stage on soft samples, I m assuming by a hard sample you mean - the ones closer to a degenerate distribution.\nAnd the soft samples are the ones closer to a uniform distribution.",
    "2665331": "Thanks for sharing. I have two questions.\n1. How did you find threshold=5.5 ? By a lot of experiment?\n2. About two-stage (total votes) in result table, how did you train model ? Did you train in the same way as Sean R.B. Bearden ? or Did you train with all data and then filter like total_evaluators > threshold ?",
    "2655261": "did you just train and validated with filtered record in stage 2!! plz sure once again",
    "2663269": ""
  }
}