{
  "id": 266501,
  "title": "208th Solution - Common Approach, But Hope for Someone's Help",
  "url": "/competitions/seti-breakthrough-listen/discussion/266501",
  "author_name": "Bilzard",
  "post_date": "2021-08-19T10:48:03.695000",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Thanks to the all hosts &amp; participants.<br>\nEspecially for the <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> 's baseline[1]. I reused the notebook throughout my experiments.</p>\n<p>Also, Google Colabo Pro &amp; Kaggle Notebook.</p>\n<h2>Final Model</h2>\n<p>Stacking of 5 models (Table 1).</p>\n<p><b>Table 1: Models and Score of Final Submission</b></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Private LB</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B2, 512x512, 5-fold</td>\n<td>0.763</td>\n<td>0.764</td>\n</tr>\n<tr>\n<td>EfficientNet-B0, 512x512, 5-fold</td>\n<td>0.757</td>\n<td>0.759</td>\n</tr>\n<tr>\n<td>EfficientNet-B0, 512x512, 3-fold, controlling sampling rate, logical-or on mix-upped target</td>\n<td>0.730</td>\n<td>0.733</td>\n</tr>\n<tr>\n<td>ResNet18d, 512x512, 5-fold, controlling sampling rate</td>\n<td>0.745</td>\n<td>0.744</td>\n</tr>\n<tr>\n<td>ResNet18d, 320x320, 5-fold</td>\n<td>0.719</td>\n<td>0.720</td>\n</tr>\n<tr>\n<td>Stacking 5 models (1-hidden-layer MLP)</td>\n<td>0.764</td>\n<td>0.765</td>\n</tr>\n</tbody>\n</table>\n<p>In the early trials, I was wondering about class imbalance problem. Since AUC-ROC equally evaluates the minority and majority class, I thought simple BCE loss over-penalize the positive set, and make it hard to detect them (which I finally found that's NOT TRUE).<br>\nSo I balanced class by controlling sampling rate of training set (i.e. oversampling positive set or/and downsampling negative set). I also added the train/test's B-observations (1, 3, 5 channels) as the negative sets to help models learning the background information. A typical settings are shown in table2.</p>\n<p>The trial boosted the training speed, but at the same time it suppress the score: feeding the whole dataset improves score about +0.03. I think it is because \"knowing background\" is  an important information for the competition.</p>\n<p><b>Table 2: Number of Samples for Train Dataset</b></p>\n<table>\n<thead>\n<tr>\n<th>target</th>\n<th>train-pos-A</th>\n<th>train-pos-B</th>\n<th>train-neg-A</th>\n<th>total</th>\n<th>ratio</th>\n<th>after mix-up</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0</td>\n<td>4,000</td>\n<td>4,000</td>\n<td>8,000</td>\n<td>0.67</td>\n<td>0.45</td>\n</tr>\n<tr>\n<td>1</td>\n<td>4,000</td>\n<td>0</td>\n<td>0</td>\n<td>4,000</td>\n<td>0.33</td>\n<td>0.55</td>\n</tr>\n<tr>\n<td>total</td>\n<td>4,000</td>\n<td>4,000</td>\n<td>4,000</td>\n<td>12,000</td>\n<td>1.0</td>\n<td>1.0</td>\n</tr>\n</tbody>\n</table>\n<h2>Other Settings (of the Best Score Model)</h2>\n<ul>\n<li>Augmentations: <ul>\n<li>Resize 512x512</li>\n<li>HorizontalFlip</li>\n<li>VerticalFlip</li>\n<li>ShiftScaleRotate (rotate_limit: 0)</li>\n<li>MotionBlur</li>\n<li>Sharpen</li>\n<li>MixUp</li></ul></li>\n<li>Loss<ul>\n<li>BCE loss</li></ul></li>\n<li>Epoch: 20</li>\n<li>Inference<ul>\n<li>5-fold average</li>\n<li>TTA x4: original, vertical flip, horizontal flip, both</li></ul></li>\n</ul>\n<h2>Not Worked</h2>\n<ul>\n<li>Weighed BCE Loss (Class-Balanced Loss) [2]</li>\n<li>learn image scaling[3]</li>\n<li>logical-OR of mix-upped target</li>\n<li>Image preprocessing (column-wise diff, log1p transform, and min/max-clipping)</li>\n</ul>\n<h2>Reference</h2>\n<p>[1] <a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-resnet18d-baseline\" target=\"_blank\">https://www.kaggle.com/ttahara/rerun-seti-e-t-resnet18d-baseline</a><br>\n[2] <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/265973\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/265973</a><br>\n[3] <a href=\"https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing\" target=\"_blank\">https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing</a></p>",
  "messages": [
    {
      "id": 1481156,
      "postDate": "2021-08-19T10:48:03.697Z",
      "content": "<p>Thanks to the all hosts &amp; participants.<br>\nEspecially for the <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> 's baseline[1]. I reused the notebook throughout my experiments.</p>\n<p>Also, Google Colabo Pro &amp; Kaggle Notebook.</p>\n<h2>Final Model</h2>\n<p>Stacking of 5 models (Table 1).</p>\n<p><b>Table 1: Models and Score of Final Submission</b></p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Private LB</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EfficientNet-B2, 512x512, 5-fold</td>\n<td>0.763</td>\n<td>0.764</td>\n</tr>\n<tr>\n<td>EfficientNet-B0, 512x512, 5-fold</td>\n<td>0.757</td>\n<td>0.759</td>\n</tr>\n<tr>\n<td>EfficientNet-B0, 512x512, 3-fold, controlling sampling rate, logical-or on mix-upped target</td>\n<td>0.730</td>\n<td>0.733</td>\n</tr>\n<tr>\n<td>ResNet18d, 512x512, 5-fold, controlling sampling rate</td>\n<td>0.745</td>\n<td>0.744</td>\n</tr>\n<tr>\n<td>ResNet18d, 320x320, 5-fold</td>\n<td>0.719</td>\n<td>0.720</td>\n</tr>\n<tr>\n<td>Stacking 5 models (1-hidden-layer MLP)</td>\n<td>0.764</td>\n<td>0.765</td>\n</tr>\n</tbody>\n</table>\n<p>In the early trials, I was wondering about class imbalance problem. Since AUC-ROC equally evaluates the minority and majority class, I thought simple BCE loss over-penalize the positive set, and make it hard to detect them (which I finally found that's NOT TRUE).<br>\nSo I balanced class by controlling sampling rate of training set (i.e. oversampling positive set or/and downsampling negative set). I also added the train/test's B-observations (1, 3, 5 channels) as the negative sets to help models learning the background information. A typical settings are shown in table2.</p>\n<p>The trial boosted the training speed, but at the same time it suppress the score: feeding the whole dataset improves score about +0.03. I think it is because \"knowing background\" is  an important information for the competition.</p>\n<p><b>Table 2: Number of Samples for Train Dataset</b></p>\n<table>\n<thead>\n<tr>\n<th>target</th>\n<th>train-pos-A</th>\n<th>train-pos-B</th>\n<th>train-neg-A</th>\n<th>total</th>\n<th>ratio</th>\n<th>after mix-up</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0</td>\n<td>4,000</td>\n<td>4,000</td>\n<td>8,000</td>\n<td>0.67</td>\n<td>0.45</td>\n</tr>\n<tr>\n<td>1</td>\n<td>4,000</td>\n<td>0</td>\n<td>0</td>\n<td>4,000</td>\n<td>0.33</td>\n<td>0.55</td>\n</tr>\n<tr>\n<td>total</td>\n<td>4,000</td>\n<td>4,000</td>\n<td>4,000</td>\n<td>12,000</td>\n<td>1.0</td>\n<td>1.0</td>\n</tr>\n</tbody>\n</table>\n<h2>Other Settings (of the Best Score Model)</h2>\n<ul>\n<li>Augmentations: <ul>\n<li>Resize 512x512</li>\n<li>HorizontalFlip</li>\n<li>VerticalFlip</li>\n<li>ShiftScaleRotate (rotate_limit: 0)</li>\n<li>MotionBlur</li>\n<li>Sharpen</li>\n<li>MixUp</li></ul></li>\n<li>Loss<ul>\n<li>BCE loss</li></ul></li>\n<li>Epoch: 20</li>\n<li>Inference<ul>\n<li>5-fold average</li>\n<li>TTA x4: original, vertical flip, horizontal flip, both</li></ul></li>\n</ul>\n<h2>Not Worked</h2>\n<ul>\n<li>Weighed BCE Loss (Class-Balanced Loss) [2]</li>\n<li>learn image scaling[3]</li>\n<li>logical-OR of mix-upped target</li>\n<li>Image preprocessing (column-wise diff, log1p transform, and min/max-clipping)</li>\n</ul>\n<h2>Reference</h2>\n<p>[1] <a href=\"https://www.kaggle.com/ttahara/rerun-seti-e-t-resnet18d-baseline\" target=\"_blank\">https://www.kaggle.com/ttahara/rerun-seti-e-t-resnet18d-baseline</a><br>\n[2] <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/265973\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/265973</a><br>\n[3] <a href=\"https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing\" target=\"_blank\">https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing</a></p>",
      "rawMarkdown": "Thanks to the all hosts & participants.\nEspecially for the @ttahara 's baseline[1]. I reused the notebook throughout my experiments.\n\nAlso, Google Colabo Pro & Kaggle Notebook.\n\n## Final Model\n\nStacking of 5 models (Table 1).\n\n<center><b>Table 1: Models and Score of Final Submission</b></center>\nModel|Private LB|Public LB\n---|---|---\nEfficientNet-B2, 512x512, 5-fold | 0.763 | 0.764\nEfficientNet-B0, 512x512, 5-fold | 0.757 | 0.759\nEfficientNet-B0, 512x512, 3-fold, controlling sampling rate, logical-or on mix-upped target | 0.730 | 0.733\nResNet18d, 512x512, 5-fold, controlling sampling rate | 0.745 | 0.744\nResNet18d, 320x320, 5-fold | 0.719 | 0.720\nStacking 5 models (1-hidden-layer MLP) | 0.764 | 0.765\n\nIn the early trials, I was wondering about class imbalance problem. Since AUC-ROC equally evaluates the minority and majority class, I thought simple BCE loss over-penalize the positive set, and make it hard to detect them (which I finally found that's NOT TRUE).\nSo I balanced class by controlling sampling rate of training set (i.e. oversampling positive set or/and downsampling negative set). I also added the train/test's B-observations (1, 3, 5 channels) as the negative sets to help models learning the background information. A typical settings are shown in table2.\n\nThe trial boosted the training speed, but at the same time it suppress the score: feeding the whole dataset improves score about +0.03. I think it is because \"knowing background\" is  an important information for the competition.\n\n<center><b>Table 2: Number of Samples for Train Dataset</b></center>\ntarget|train-pos-A|train-pos-B|train-neg-A|total|ratio|after mix-up\n---|---:|---:|---:|---:|---:|----:\n0|0|4,000|4,000|8,000|0.67|0.45\n1|4,000|0|0|4,000|0.33|0.55\ntotal|4,000|4,000|4,000|12,000|1.0|1.0\n\n## Other Settings (of the Best Score Model)\n\n- Augmentations: \n  - Resize 512x512\n  - HorizontalFlip\n  - VerticalFlip\n  - ShiftScaleRotate (rotate_limit: 0)\n  - MotionBlur\n  - Sharpen\n  - MixUp\n- Loss\n  - BCE loss\n- Epoch: 20\n- Inference\n  - 5-fold average\n  - TTA x4: original, vertical flip, horizontal flip, both\n\n## Not Worked\n\n- Weighed BCE Loss (Class-Balanced Loss) [2]\n- learn image scaling[3]\n- logical-OR of mix-upped target\n- Image preprocessing (column-wise diff, log1p transform, and min/max-clipping)\n\n## Reference\n\n[1] https://www.kaggle.com/ttahara/rerun-seti-e-t-resnet18d-baseline\n[2] https://www.kaggle.com/c/seti-breakthrough-listen/discussion/265973\n[3] https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing",
      "votes": 4
    },
    {
      "id": 1481235,
      "postDate": "2021-08-19T11:34:05.370Z",
      "content": "<p>As for top ranker's solutions, it seems larger model performs better.<br>\nI am curious about anyone who achieves high score with relatively small models.<br>\nIf anyone who can achieve the high score with small model, please share the approach.</p>",
      "rawMarkdown": "As for top ranker's solutions, it seems larger model performs better.\nI am curious about anyone who achieves high score with relatively small models.\nIf anyone who can achieve the high score with small model, please share the approach.",
      "replies": [
        {
          "id": 1481795,
          "postDate": "2021-08-19T17:05:59.033Z",
          "content": "<p>nf-regnet b1 and efficientnet b2 are about same capacity</p>",
          "rawMarkdown": "nf-regnet b1 and efficientnet b2 are about same capacity",
          "votes": 1
        },
        {
          "id": 1483128,
          "postDate": "2021-08-20T13:07:39.687Z",
          "content": "<p><a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> Thank you for sharing.</p>\n<p>Do you think the score difference is due to the model choice or other trials (or both)?</p>\n<p>According to your solution, you tried other trials like larger image size, logical-OR on mix-upped target, freezing backbone etc. Which trial do you think is the most effective to the score?</p>",
          "rawMarkdown": "@bakeryproducts Thank you for sharing.\n\nDo you think the score difference is due to the model choice or other trials (or both)?\n\nAccording to your solution, you tried other trials like larger image size, logical-OR on mix-upped target, freezing backbone etc. Which trial do you think is the most effective to the score?"
        },
        {
          "id": 1483156,
          "postDate": "2021-08-20T13:29:04.110Z",
          "content": "<p>Hi bilzard. Nice job, you did well.</p>\n<p>I am similarly confused. I used big models such as EfficientNetB6, B4 and B3. And I used large image sizes like 384, 512, and 768. None-the-less, i barely got Bronze medal. My public LB only reached LB 779.</p>\n<p>I think mixup, augmentation, regularization are the key factors to prevent models from overfitting train data and generalize to test data. I did not use much of these, so I'm thinking this was my shortcoming. I'm currently running experiments now to see if this is indeed the missing factor.</p>",
          "rawMarkdown": "Hi bilzard. Nice job, you did well.\n\nI am similarly confused. I used big models such as EfficientNetB6, B4 and B3. And I used large image sizes like 384, 512, and 768. None-the-less, i barely got Bronze medal. My public LB only reached LB 779.\n\nI think mixup, augmentation, regularization are the key factors to prevent models from overfitting train data and generalize to test data. I did not use much of these, so I'm thinking this was my shortcoming. I'm currently running experiments now to see if this is indeed the missing factor.",
          "votes": 1
        },
        {
          "id": 1483164,
          "postDate": "2021-08-20T13:32:44.373Z",
          "content": "<p>Image size, epochs,  are crucial, tricks like MSDA( mixup, etc), freezing bb, are boosting good model even further.  20 epochs is nothing, my submit models never early exited until epoch ~300.<br>\n(for now) You will get much more from doing right things and monitoring everything, then from tricks. And you gotta be careful with augmentations, they are good, but sometimes break logic of sample. For example you cant vertical flip image from car dashcam.</p>\n<p>Update. to be fair number of epochs to converge depends on a lot of things and you probably should just monitor your validation.</p>",
          "rawMarkdown": "Image size, epochs,  are crucial, tricks like MSDA( mixup, etc), freezing bb, are boosting good model even further.  20 epochs is nothing, my submit models never early exited until epoch ~300.\n(for now) You will get much more from doing right things and monitoring everything, then from tricks. And you gotta be careful with augmentations, they are good, but sometimes break logic of sample. For example you cant vertical flip image from car dashcam.\n\nUpdate. to be fair number of epochs to converge depends on a lot of things and you probably should just monitor your validation.",
          "votes": 1
        },
        {
          "id": 1483180,
          "postDate": "2021-08-20T13:44:58.577Z",
          "content": "<p>Our models converged after 20 epochs (also without cleaning and early 0.800 score), but we used all data (also old).</p>",
          "rawMarkdown": "Our models converged after 20 epochs (also without cleaning and early 0.800 score), but we used all data (also old).",
          "votes": 2
        },
        {
          "id": 1486594,
          "postDate": "2021-08-23T05:09:26.153Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thank you for sharing your model and score. Larger model, larger image size was the key to the score LB 0.779. In my situation, I don't have my own workstation, I have to train within the limit of cloud service like Google Colab and Kaggle notebook. So I didn't dare to train larger model than EfficientNet-B2. I think sufficient computing resource is de facto standard for today's kaggle competition.</p>\n<p><a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> I didn't come up with training over hundreds of epochs. It's like training the model from scratch. Thank you for the advice.</p>\n<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Variation of data was also crucial to this competition. </p>\n<p>Personally, one more important thing is EDA. I did some EDA and error analysis, but I couldn't find out useful information from them, or, if I find something (S-signal, CV-LB difference), I didn't dig further or doing something by utilizing the fact. </p>",
          "rawMarkdown": "@cdeotte Thank you for sharing your model and score. Larger model, larger image size was the key to the score LB 0.779. In my situation, I don't have my own workstation, I have to train within the limit of cloud service like Google Colab and Kaggle notebook. So I didn't dare to train larger model than EfficientNet-B2. I think sufficient computing resource is de facto standard for today's kaggle competition.\n\n@bakeryproducts I didn't come up with training over hundreds of epochs. It's like training the model from scratch. Thank you for the advice.\n\n@philippsinger Variation of data was also crucial to this competition. \n\nPersonally, one more important thing is EDA. I did some EDA and error analysis, but I couldn't find out useful information from them, or, if I find something (S-signal, CV-LB difference), I didn't dig further or doing something by utilizing the fact. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1481235,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2021-08-19T11:34:05.370000",
      "content": "<p>As for top ranker's solutions, it seems larger model performs better.<br>\nI am curious about anyone who achieves high score with relatively small models.<br>\nIf anyone who can achieve the high score with small model, please share the approach.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1481795,
          "author_name": "Gleb",
          "author_url": "",
          "post_date": "2021-08-19T17:05:59.033000",
          "content": "<p>nf-regnet b1 and efficientnet b2 are about same capacity</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1483128,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2021-08-20T13:07:39.687000",
          "content": "<p><a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> Thank you for sharing.</p>\n<p>Do you think the score difference is due to the model choice or other trials (or both)?</p>\n<p>According to your solution, you tried other trials like larger image size, logical-OR on mix-upped target, freezing backbone etc. Which trial do you think is the most effective to the score?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1483156,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T13:29:04.110000",
          "content": "<p>Hi bilzard. Nice job, you did well.</p>\n<p>I am similarly confused. I used big models such as EfficientNetB6, B4 and B3. And I used large image sizes like 384, 512, and 768. None-the-less, i barely got Bronze medal. My public LB only reached LB 779.</p>\n<p>I think mixup, augmentation, regularization are the key factors to prevent models from overfitting train data and generalize to test data. I did not use much of these, so I'm thinking this was my shortcoming. I'm currently running experiments now to see if this is indeed the missing factor.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1483164,
          "author_name": "Gleb",
          "author_url": "",
          "post_date": "2021-08-20T13:32:44.373000",
          "content": "<p>Image size, epochs,  are crucial, tricks like MSDA( mixup, etc), freezing bb, are boosting good model even further.  20 epochs is nothing, my submit models never early exited until epoch ~300.<br>\n(for now) You will get much more from doing right things and monitoring everything, then from tricks. And you gotta be careful with augmentations, they are good, but sometimes break logic of sample. For example you cant vertical flip image from car dashcam.</p>\n<p>Update. to be fair number of epochs to converge depends on a lot of things and you probably should just monitor your validation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1483180,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-20T13:44:58.577000",
          "content": "<p>Our models converged after 20 epochs (also without cleaning and early 0.800 score), but we used all data (also old).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1486594,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2021-08-23T05:09:26.153000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thank you for sharing your model and score. Larger model, larger image size was the key to the score LB 0.779. In my situation, I don't have my own workstation, I have to train within the limit of cloud service like Google Colab and Kaggle notebook. So I didn't dare to train larger model than EfficientNet-B2. I think sufficient computing resource is de facto standard for today's kaggle competition.</p>\n<p><a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> I didn't come up with training over hundreds of epochs. It's like training the model from scratch. Thank you for the advice.</p>\n<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Variation of data was also crucial to this competition. </p>\n<p>Personally, one more important thing is EDA. I did some EDA and error analysis, but I couldn't find out useful information from them, or, if I find something (S-signal, CV-LB difference), I didn't dig further or doing something by utilizing the fact. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1481156": "Thanks to the all hosts & participants.\nEspecially for the @ttahara 's baseline[1]. I reused the notebook throughout my experiments.\n\nAlso, Google Colabo Pro & Kaggle Notebook.\n\n## Final Model\n\nStacking of 5 models (Table 1).\n\n<center><b>Table 1: Models and Score of Final Submission</b></center>\nModel|Private LB|Public LB\n---|---|---\nEfficientNet-B2, 512x512, 5-fold | 0.763 | 0.764\nEfficientNet-B0, 512x512, 5-fold | 0.757 | 0.759\nEfficientNet-B0, 512x512, 3-fold, controlling sampling rate, logical-or on mix-upped target | 0.730 | 0.733\nResNet18d, 512x512, 5-fold, controlling sampling rate | 0.745 | 0.744\nResNet18d, 320x320, 5-fold | 0.719 | 0.720\nStacking 5 models (1-hidden-layer MLP) | 0.764 | 0.765\n\nIn the early trials, I was wondering about class imbalance problem. Since AUC-ROC equally evaluates the minority and majority class, I thought simple BCE loss over-penalize the positive set, and make it hard to detect them (which I finally found that's NOT TRUE).\nSo I balanced class by controlling sampling rate of training set (i.e. oversampling positive set or/and downsampling negative set). I also added the train/test's B-observations (1, 3, 5 channels) as the negative sets to help models learning the background information. A typical settings are shown in table2.\n\nThe trial boosted the training speed, but at the same time it suppress the score: feeding the whole dataset improves score about +0.03. I think it is because \"knowing background\" is  an important information for the competition.\n\n<center><b>Table 2: Number of Samples for Train Dataset</b></center>\ntarget|train-pos-A|train-pos-B|train-neg-A|total|ratio|after mix-up\n---|---:|---:|---:|---:|---:|----:\n0|0|4,000|4,000|8,000|0.67|0.45\n1|4,000|0|0|4,000|0.33|0.55\ntotal|4,000|4,000|4,000|12,000|1.0|1.0\n\n## Other Settings (of the Best Score Model)\n\n- Augmentations: \n  - Resize 512x512\n  - HorizontalFlip\n  - VerticalFlip\n  - ShiftScaleRotate (rotate_limit: 0)\n  - MotionBlur\n  - Sharpen\n  - MixUp\n- Loss\n  - BCE loss\n- Epoch: 20\n- Inference\n  - 5-fold average\n  - TTA x4: original, vertical flip, horizontal flip, both\n\n## Not Worked\n\n- Weighed BCE Loss (Class-Balanced Loss) [2]\n- learn image scaling[3]\n- logical-OR of mix-upped target\n- Image preprocessing (column-wise diff, log1p transform, and min/max-clipping)\n\n## Reference\n\n[1] https://www.kaggle.com/ttahara/rerun-seti-e-t-resnet18d-baseline\n[2] https://www.kaggle.com/c/seti-breakthrough-listen/discussion/265973\n[3] https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing",
    "1481235": "As for top ranker's solutions, it seems larger model performs better.\nI am curious about anyone who achieves high score with relatively small models.\nIf anyone who can achieve the high score with small model, please share the approach."
  }
}