{
  "id": 242644,
  "title": "small models collection",
  "url": "/competitions/seti-breakthrough-listen/discussion/242644",
  "author_name": "Tawara",
  "post_date": "2021-05-30T02:50:22.627000",
  "votes": 60,
  "comment_count": 30,
  "views": 0,
  "content": "<p>In this competition, it seems to me that small models are sufficient to get decent score.  <br>\nThis is a welcome situation because it allows us to increase the number of trials.</p>\n<p>Here I'd like to share some relatively small models' results. </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>OOF Score</th>\n<th>order in LB(all models get 0.97x)</th>\n<th>training time per epoch (sec)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>resnet18d</td>\n<td>0.9868</td>\n<td>7</td>\n<td>186</td>\n</tr>\n<tr>\n<td>resnet34d</td>\n<td>0.9869</td>\n<td>5</td>\n<td>260</td>\n</tr>\n<tr>\n<td>resnet50d</td>\n<td>0.9862</td>\n<td>6</td>\n<td>444</td>\n</tr>\n<tr>\n<td>resnest14d</td>\n<td>0.9864</td>\n<td>4</td>\n<td>393</td>\n</tr>\n<tr>\n<td>resnest26d</td>\n<td>0.9866</td>\n<td><strong>1</strong></td>\n<td>516</td>\n</tr>\n<tr>\n<td>efficientnetv2_s</td>\n<td><strong>0.9884</strong></td>\n<td>2</td>\n<td>533</td>\n</tr>\n<tr>\n<td>mobilenetv3</td>\n<td>0.9860</td>\n<td>3</td>\n<td>217</td>\n</tr>\n</tbody>\n</table>\n<h5>Settings:</h5>\n<ul>\n<li>Input size: <strong>1x512x512</strong></li>\n<li>Initial Weights: ImageNet</li>\n<li>Train-Val Split: StratifiedKFold(K=5)</li>\n<li>Augmentation: HFlip, VFlip, ShiftScaleRotate, RandomResizedCrop</li>\n<li>Loss: Binary Cross Entropy</li>\n<li>Optimizer: AdamW</li>\n<li>Scheduler: OneCycleLR</li>\n<li>Mixed Precision: <strong>Enabled</strong></li>\n</ul>\n<h5>Hardware</h5>\n<ul>\n<li>GPU: TitanRTX</li>\n<li>CPU: Intel(R) Core(TM) i7-7820X @ 3.60GHz (8 core)</li>\n</ul>\n<p>I'll add more models later.</p>",
  "messages": [
    {
      "id": 1328187,
      "postDate": "2021-05-30T02:50:22.627Z",
      "content": "<p>In this competition, it seems to me that small models are sufficient to get decent score.  <br>\nThis is a welcome situation because it allows us to increase the number of trials.</p>\n<p>Here I'd like to share some relatively small models' results. </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>OOF Score</th>\n<th>order in LB(all models get 0.97x)</th>\n<th>training time per epoch (sec)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>resnet18d</td>\n<td>0.9868</td>\n<td>7</td>\n<td>186</td>\n</tr>\n<tr>\n<td>resnet34d</td>\n<td>0.9869</td>\n<td>5</td>\n<td>260</td>\n</tr>\n<tr>\n<td>resnet50d</td>\n<td>0.9862</td>\n<td>6</td>\n<td>444</td>\n</tr>\n<tr>\n<td>resnest14d</td>\n<td>0.9864</td>\n<td>4</td>\n<td>393</td>\n</tr>\n<tr>\n<td>resnest26d</td>\n<td>0.9866</td>\n<td><strong>1</strong></td>\n<td>516</td>\n</tr>\n<tr>\n<td>efficientnetv2_s</td>\n<td><strong>0.9884</strong></td>\n<td>2</td>\n<td>533</td>\n</tr>\n<tr>\n<td>mobilenetv3</td>\n<td>0.9860</td>\n<td>3</td>\n<td>217</td>\n</tr>\n</tbody>\n</table>\n<h5>Settings:</h5>\n<ul>\n<li>Input size: <strong>1x512x512</strong></li>\n<li>Initial Weights: ImageNet</li>\n<li>Train-Val Split: StratifiedKFold(K=5)</li>\n<li>Augmentation: HFlip, VFlip, ShiftScaleRotate, RandomResizedCrop</li>\n<li>Loss: Binary Cross Entropy</li>\n<li>Optimizer: AdamW</li>\n<li>Scheduler: OneCycleLR</li>\n<li>Mixed Precision: <strong>Enabled</strong></li>\n</ul>\n<h5>Hardware</h5>\n<ul>\n<li>GPU: TitanRTX</li>\n<li>CPU: Intel(R) Core(TM) i7-7820X @ 3.60GHz (8 core)</li>\n</ul>\n<p>I'll add more models later.</p>",
      "rawMarkdown": "In this competition, it seems to me that small models are sufficient to get decent score.  \nThis is a welcome situation because it allows us to increase the number of trials.\n\nHere I'd like to share some relatively small models' results. \n\n| Model | OOF Score | order in LB(all models get 0.97x) | training time per epoch (sec) |\n|:---:|:----:|:----:|:----:|\n| resnet18d | 0.9868 | 7 | 186 |\n| resnet34d | 0.9869 | 5 | 260 |\n| resnet50d | 0.9862 | 6 | 444 |\n| resnest14d | 0.9864 | 4 | 393 |\n| resnest26d | 0.9866 | **1** | 516 |\n| efficientnetv2_s | **0.9884** | 2 | 533 |\n| mobilenetv3 | 0.9860 | 3 | 217 |\n\n##### Settings:\n* Input size: **1x512x512**\n* Initial Weights: ImageNet\n* Train-Val Split: StratifiedKFold(K=5)\n* Augmentation: HFlip, VFlip, ShiftScaleRotate, RandomResizedCrop\n* Loss: Binary Cross Entropy\n* Optimizer: AdamW\n* Scheduler: OneCycleLR\n* Mixed Precision: **Enabled**\n\n##### Hardware\n* GPU: TitanRTX\n* CPU: Intel(R) Core(TM) i7-7820X @ 3.60GHz (8 core)\n\n\nI'll add more models later.",
      "votes": 60
    },
    {
      "id": 1329498,
      "postDate": "2021-05-31T07:29:54.280Z",
      "content": "<p>Currently, my best LB score is provided by 'efficientnet-b0'. In other words, it is possible to get the equivalent of a 3rd place score with 'efficientnet-b0'. For your reference…</p>",
      "rawMarkdown": "Currently, my best LB score is provided by 'efficientnet-b0'. In other words, it is possible to get the equivalent of a 3rd place score with 'efficientnet-b0'. For your reference...",
      "votes": 19,
      "replies": [
        {
          "id": 1329680,
          "postDate": "2021-05-31T09:50:26.600Z",
          "content": "<p>Think you for information !<br>\nThat is good news for me because I believe this competition doesn't requre big models. </p>\n<p>BTW, is there any trick on your 3rd place score such as pre/postprocessing, augmentation, special architecture ? </p>",
          "rawMarkdown": "Think you for information !\nThat is good news for me because I believe this competition doesn't requre big models. \n\nBTW, is there any trick on your 3rd place score such as pre/postprocessing, augmentation, special architecture ? ",
          "votes": 2
        },
        {
          "id": 1329764,
          "postDate": "2021-05-31T11:02:16.820Z",
          "content": "<p>Thank you for sharing, would you mind share your CV score?</p>",
          "rawMarkdown": "Thank you for sharing, would you mind share your CV score?",
          "votes": 1
        },
        {
          "id": 1329996,
          "postDate": "2021-05-31T14:16:52.267Z",
          "content": "<p>I can't tell you about the detailed tricks yet, but most of them are almost the same setup as Tawara's at the beginning of this topic. 512x512x1 is the size I use too! Also, my cv score is 0.9924 (5 fold avg).</p>",
          "rawMarkdown": "I can't tell you about the detailed tricks yet, but most of them are almost the same setup as Tawara's at the beginning of this topic. 512x512x1 is the size I use too! Also, my cv score is 0.9924 (5 fold avg).",
          "votes": 12
        },
        {
          "id": 1330039,
          "postDate": "2021-05-31T14:52:53.783Z",
          "content": "<p>Thanks! <br>\nI'm looking forward to finding tricks😃</p>",
          "rawMarkdown": "Thanks! \nI'm looking forward to finding tricks😃"
        },
        {
          "id": 1336446,
          "postDate": "2021-06-04T23:08:49.340Z",
          "content": "<p>Mathematically speaking, is there any difference if I use BCE loss (class = 1) VS Cross entropy loss (class = 2)?</p>",
          "rawMarkdown": "Mathematically speaking, is there any difference if I use BCE loss (class = 1) VS Cross entropy loss (class = 2)?",
          "votes": 1
        },
        {
          "id": 1336669,
          "postDate": "2021-06-05T06:05:37.127Z",
          "content": "<p><a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> how many epochs you are training?</p>",
          "rawMarkdown": "@hirune924 how many epochs you are training?"
        },
        {
          "id": 1337465,
          "postDate": "2021-06-05T16:12:19.087Z",
          "content": "<p><a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> I am also trying to figure out some tricks. Can you please share how much time it takes you for your complete training?<br>\nI just want to compare it to mine. </p>",
          "rawMarkdown": "@hirune924 I am also trying to figure out some tricks. Can you please share how much time it takes you for your complete training?\nI just want to compare it to mine. ",
          "votes": -1
        },
        {
          "id": 1338004,
          "postDate": "2021-06-06T05:15:29.673Z",
          "content": "<p>20-40epoch per fold. it takes about one day for 5 fold. </p>",
          "rawMarkdown": "20-40epoch per fold. it takes about one day for 5 fold. "
        }
      ]
    },
    {
      "id": 1331460,
      "postDate": "2021-06-01T13:41:33.447Z",
      "content": "<p>Input size: 1x512x512, how you convert 6x273x256 to that size?</p>",
      "rawMarkdown": "Input size: 1x512x512, how you convert 6x273x256 to that size?",
      "votes": 3,
      "replies": [
        {
          "id": 1331550,
          "postDate": "2021-06-01T14:33:02.370Z",
          "content": "<p>I resize cadence snippets after stacking its channels into 1 channel along time axis, and transposing it.</p>",
          "rawMarkdown": "I resize cadence snippets after stacking its channels into 1 channel along time axis, and transposing it.",
          "votes": 1
        },
        {
          "id": 1332203,
          "postDate": "2021-06-02T01:06:45.170Z",
          "content": "<p>Do you mean you transpose 1638x256 to 256x1638 then resize?</p>",
          "rawMarkdown": "Do you mean you transpose 1638x256 to 256x1638 then resize?"
        },
        {
          "id": 1332212,
          "postDate": "2021-06-02T01:26:43.680Z",
          "content": "<p>Yes, I train models with (height, width) = (Frequency, TIme). But I think it's a meaningless operation.</p>\n<p>For example, <a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a> got good score with (height, width) = (TIme, Frequency):<br>\n<a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/241195#1325779\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/241195#1325779</a></p>",
          "rawMarkdown": "Yes, I train models with (height, width) = (Frequency, TIme). But I think it's a meaningless operation.\n\nFor example, @yukia18 got good score with (height, width) = (TIme, Frequency):\nhttps://www.kaggle.com/c/seti-breakthrough-listen/discussion/241195#1325779",
          "votes": 2
        }
      ]
    },
    {
      "id": 1334665,
      "postDate": "2021-06-03T17:08:48.433Z",
      "content": "<p>You can also use RegNetX/RegNetY models from <a href=\"https://github.com/facebookresearch/pycls\" target=\"_blank\">https://github.com/facebookresearch/pycls</a>. 200MF/400MF/800MF are quite small.</p>",
      "rawMarkdown": "You can also use RegNetX/RegNetY models from https://github.com/facebookresearch/pycls. 200MF/400MF/800MF are quite small.",
      "votes": 1
    },
    {
      "id": 1331581,
      "postDate": "2021-06-01T15:03:05.517Z",
      "content": "<p>Personally, I wonder if the ViT models works for this dataset.</p>\n<p>FYI:<br>\nI tried <code>ResMLP</code> early in the competition, but it didn't perform well. (efficientnet_b0 got 0.97 with same settings.)<br>\nAfter that, I put off ViT and focused on other experiments.</p>",
      "rawMarkdown": "Personally, I wonder if the ViT models works for this dataset.\n\nFYI:\nI tried `ResMLP` early in the competition, but it didn't perform well. (efficientnet_b0 got 0.97 with same settings.)\nAfter that, I put off ViT and focused on other experiments.\n",
      "votes": 1
    },
    {
      "id": 1328401,
      "postDate": "2021-05-30T08:14:49.783Z",
      "content": "<p>Thank you for sharing info, just curious did you train these models with channel-wise samples or spatial?</p>",
      "rawMarkdown": "Thank you for sharing info, just curious did you train these models with channel-wise samples or spatial?",
      "votes": 1,
      "replies": [
        {
          "id": 1328570,
          "postDate": "2021-05-30T11:13:25.903Z",
          "content": "<p>spatial wise</p>",
          "rawMarkdown": "spatial wise",
          "votes": 1
        },
        {
          "id": 1329288,
          "postDate": "2021-05-31T03:18:44.907Z",
          "content": "<p>I train models with spatial samples (cf: <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/238611\" target=\"_blank\">Channel-wise Vs Spatial</a>)</p>",
          "rawMarkdown": "I train models with spatial samples (cf: [Channel-wise Vs Spatial](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/238611))"
        }
      ]
    },
    {
      "id": 1328221,
      "postDate": "2021-05-30T03:51:48.510Z",
      "content": "<p>For me smaller models just don't work. But very big models like EfficientNet b5 also don't work. The sweet spot for me is  EfficientNet b4</p>",
      "rawMarkdown": "For me smaller models just don't work. But very big models like EfficientNet b5 also don't work. The sweet spot for me is  EfficientNet b4",
      "votes": 2,
      "replies": [
        {
          "id": 1328230,
          "postDate": "2021-05-30T04:21:54.937Z",
          "content": "<p>If the model is too small or too large, it will not benefit the competition. Most of the models in this competition should be medium model, so I agree with you</p>",
          "rawMarkdown": "If the model is too small or too large, it will not benefit the competition. Most of the models in this competition should be medium model, so I agree with you",
          "votes": 3
        },
        {
          "id": 1328308,
          "postDate": "2021-05-30T06:28:57.017Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> <a href=\"https://www.kaggle.com/zhangeng\" target=\"_blank\">@zhangeng</a> </p>\n<p>I think so, too.  Bigger models require more resources but score improvement is very small.</p>",
          "rawMarkdown": "Thank you @mithilsalunkhe @zhangeng \n\nI think so, too.  Bigger models require more resources but score improvement is very small.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1331508,
      "postDate": "2021-06-01T14:08:04.707Z",
      "content": "<p>what batch size are you using my notebook is taking 1500 sec per epoch for 512x512 with batch size 16</p>",
      "rawMarkdown": "what batch size are you using my notebook is taking 1500 sec per epoch for 512x512 with batch size 16",
      "replies": [
        {
          "id": 1331555,
          "postDate": "2021-06-01T14:36:19.503Z",
          "content": "<p>32 for efficientnetv2_s and 64 for others</p>",
          "rawMarkdown": "32 for efficientnetv2_s and 64 for others",
          "votes": -1
        },
        {
          "id": 1331799,
          "postDate": "2021-06-01T18:18:14.403Z",
          "content": "<p>with a batch size of 64 aren't you running into OOM issues?</p>",
          "rawMarkdown": "with a batch size of 64 aren't you running into OOM issues?"
        },
        {
          "id": 1332202,
          "postDate": "2021-06-02T01:05:55.120Z",
          "content": "<p>No.</p>\n<p>I think you can train at least resnet18d, resnet34d and mobilenetv3 with batch_size=64 on Kaggle Notebook if you use MIxed Precision Training.</p>",
          "rawMarkdown": "No.\n\nI think you can train at least resnet18d, resnet34d and mobilenetv3 with batch_size=64 on Kaggle Notebook if you use MIxed Precision Training.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1328206,
      "postDate": "2021-05-30T03:33:38.803Z",
      "content": "<p>train with same epochs?<br>\n or same training time? (example: resnet34d with 20 epoch and efficientnetv2_s with 10 epoch)</p>\n<p>However, this result must be the most important experimental results I had seen</p>",
      "rawMarkdown": "train with same epochs?\n or same training time? (example: resnet34d with 20 epoch and efficientnetv2_s with 10 epoch)\n\nHowever, this result must be the most important experimental results I had seen",
      "replies": [
        {
          "id": 1328214,
          "postDate": "2021-05-30T03:45:45.647Z",
          "content": "<p>Thanks.</p>\n<p>I used same epochs (16 eopch).</p>",
          "rawMarkdown": "Thanks.\n\nI used same epochs (16 eopch).",
          "votes": 1
        },
        {
          "id": 1328239,
          "postDate": "2021-05-30T04:28:03.443Z",
          "content": "<p>Adjust num_ Epochs, also need to adjust LR. Otherwise, the data may not reach the minimum or maximum value. Of course, it depends on the specific changes of LR. For example, you use the following LR algorithms: ['reduce lronplateau ',' cosine networking LR ',' cosine networking warm restarts']</p>",
          "rawMarkdown": "Adjust num_ Epochs, also need to adjust LR. Otherwise, the data may not reach the minimum or maximum value. Of course, it depends on the specific changes of LR. For example, you use the following LR algorithms: ['reduce lronplateau ',' cosine networking LR ',' cosine networking warm restarts']"
        }
      ]
    },
    {
      "id": 1332164,
      "postDate": "2021-06-02T00:27:11.960Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1328298,
      "postDate": "2021-05-30T06:08:56.527Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1329498,
      "author_name": "hirune924",
      "author_url": "",
      "post_date": "2021-05-31T07:29:54.280000",
      "content": "<p>Currently, my best LB score is provided by 'efficientnet-b0'. In other words, it is possible to get the equivalent of a 3rd place score with 'efficientnet-b0'. For your reference…</p>",
      "votes": 19,
      "replies": [
        {
          "id": 1329680,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-31T09:50:26.600000",
          "content": "<p>Think you for information !<br>\nThat is good news for me because I believe this competition doesn't requre big models. </p>\n<p>BTW, is there any trick on your 3rd place score such as pre/postprocessing, augmentation, special architecture ? </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1329764,
          "author_name": "He",
          "author_url": "",
          "post_date": "2021-05-31T11:02:16.820000",
          "content": "<p>Thank you for sharing, would you mind share your CV score?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1329996,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2021-05-31T14:16:52.267000",
          "content": "<p>I can't tell you about the detailed tricks yet, but most of them are almost the same setup as Tawara's at the beginning of this topic. 512x512x1 is the size I use too! Also, my cv score is 0.9924 (5 fold avg).</p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 1330039,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-31T14:52:53.783000",
          "content": "<p>Thanks! <br>\nI'm looking forward to finding tricks😃</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1336446,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-06-04T23:08:49.340000",
          "content": "<p>Mathematically speaking, is there any difference if I use BCE loss (class = 1) VS Cross entropy loss (class = 2)?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1336669,
          "author_name": "Aman Deep Gupta",
          "author_url": "",
          "post_date": "2021-06-05T06:05:37.127000",
          "content": "<p><a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> how many epochs you are training?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1337465,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-06-05T16:12:19.087000",
          "content": "<p><a href=\"https://www.kaggle.com/hirune924\" target=\"_blank\">@hirune924</a> I am also trying to figure out some tricks. Can you please share how much time it takes you for your complete training?<br>\nI just want to compare it to mine. </p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1338004,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2021-06-06T05:15:29.673000",
          "content": "<p>20-40epoch per fold. it takes about one day for 5 fold. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1331460,
      "author_name": "Lolik",
      "author_url": "",
      "post_date": "2021-06-01T13:41:33.447000",
      "content": "<p>Input size: 1x512x512, how you convert 6x273x256 to that size?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1331550,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-06-01T14:33:02.370000",
          "content": "<p>I resize cadence snippets after stacking its channels into 1 channel along time axis, and transposing it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1332203,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-06-02T01:06:45.170000",
          "content": "<p>Do you mean you transpose 1638x256 to 256x1638 then resize?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1332212,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-06-02T01:26:43.680000",
          "content": "<p>Yes, I train models with (height, width) = (Frequency, TIme). But I think it's a meaningless operation.</p>\n<p>For example, <a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a> got good score with (height, width) = (TIme, Frequency):<br>\n<a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/241195#1325779\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/241195#1325779</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1334665,
      "author_name": "Araik Tamazian",
      "author_url": "",
      "post_date": "2021-06-03T17:08:48.433000",
      "content": "<p>You can also use RegNetX/RegNetY models from <a href=\"https://github.com/facebookresearch/pycls\" target=\"_blank\">https://github.com/facebookresearch/pycls</a>. 200MF/400MF/800MF are quite small.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1331581,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2021-06-01T15:03:05.517000",
      "content": "<p>Personally, I wonder if the ViT models works for this dataset.</p>\n<p>FYI:<br>\nI tried <code>ResMLP</code> early in the competition, but it didn't perform well. (efficientnet_b0 got 0.97 with same settings.)<br>\nAfter that, I put off ViT and focused on other experiments.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1328401,
      "author_name": "Varun Dutt",
      "author_url": "",
      "post_date": "2021-05-30T08:14:49.783000",
      "content": "<p>Thank you for sharing info, just curious did you train these models with channel-wise samples or spatial?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1328570,
          "author_name": "Aman Deep Gupta",
          "author_url": "",
          "post_date": "2021-05-30T11:13:25.903000",
          "content": "<p>spatial wise</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1329288,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-31T03:18:44.907000",
          "content": "<p>I train models with spatial samples (cf: <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/238611\" target=\"_blank\">Channel-wise Vs Spatial</a>)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1328221,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-05-30T03:51:48.510000",
      "content": "<p>For me smaller models just don't work. But very big models like EfficientNet b5 also don't work. The sweet spot for me is  EfficientNet b4</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1328230,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "2021-05-30T04:21:54.937000",
          "content": "<p>If the model is too small or too large, it will not benefit the competition. Most of the models in this competition should be medium model, so I agree with you</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1328308,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-30T06:28:57.017000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> <a href=\"https://www.kaggle.com/zhangeng\" target=\"_blank\">@zhangeng</a> </p>\n<p>I think so, too.  Bigger models require more resources but score improvement is very small.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1331508,
      "author_name": "Varun Dutt",
      "author_url": "",
      "post_date": "2021-06-01T14:08:04.707000",
      "content": "<p>what batch size are you using my notebook is taking 1500 sec per epoch for 512x512 with batch size 16</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1331555,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-06-01T14:36:19.503000",
          "content": "<p>32 for efficientnetv2_s and 64 for others</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1331799,
          "author_name": "Varun Dutt",
          "author_url": "",
          "post_date": "2021-06-01T18:18:14.403000",
          "content": "<p>with a batch size of 64 aren't you running into OOM issues?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1332202,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-06-02T01:05:55.120000",
          "content": "<p>No.</p>\n<p>I think you can train at least resnet18d, resnet34d and mobilenetv3 with batch_size=64 on Kaggle Notebook if you use MIxed Precision Training.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1328206,
      "author_name": "assign",
      "author_url": "",
      "post_date": "2021-05-30T03:33:38.803000",
      "content": "<p>train with same epochs?<br>\n or same training time? (example: resnet34d with 20 epoch and efficientnetv2_s with 10 epoch)</p>\n<p>However, this result must be the most important experimental results I had seen</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1328214,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-30T03:45:45.647000",
          "content": "<p>Thanks.</p>\n<p>I used same epochs (16 eopch).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1328239,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "2021-05-30T04:28:03.443000",
          "content": "<p>Adjust num_ Epochs, also need to adjust LR. Otherwise, the data may not reach the minimum or maximum value. Of course, it depends on the specific changes of LR. For example, you use the following LR algorithms: ['reduce lronplateau ',' cosine networking LR ',' cosine networking warm restarts']</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1332164,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-02T00:27:11.960000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1328298,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-30T06:08:56.527000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1328187": "In this competition, it seems to me that small models are sufficient to get decent score.  \nThis is a welcome situation because it allows us to increase the number of trials.\n\nHere I'd like to share some relatively small models' results. \n\n| Model | OOF Score | order in LB(all models get 0.97x) | training time per epoch (sec) |\n|:---:|:----:|:----:|:----:|\n| resnet18d | 0.9868 | 7 | 186 |\n| resnet34d | 0.9869 | 5 | 260 |\n| resnet50d | 0.9862 | 6 | 444 |\n| resnest14d | 0.9864 | 4 | 393 |\n| resnest26d | 0.9866 | **1** | 516 |\n| efficientnetv2_s | **0.9884** | 2 | 533 |\n| mobilenetv3 | 0.9860 | 3 | 217 |\n\n##### Settings:\n* Input size: **1x512x512**\n* Initial Weights: ImageNet\n* Train-Val Split: StratifiedKFold(K=5)\n* Augmentation: HFlip, VFlip, ShiftScaleRotate, RandomResizedCrop\n* Loss: Binary Cross Entropy\n* Optimizer: AdamW\n* Scheduler: OneCycleLR\n* Mixed Precision: **Enabled**\n\n##### Hardware\n* GPU: TitanRTX\n* CPU: Intel(R) Core(TM) i7-7820X @ 3.60GHz (8 core)\n\n\nI'll add more models later.",
    "1329498": "Currently, my best LB score is provided by 'efficientnet-b0'. In other words, it is possible to get the equivalent of a 3rd place score with 'efficientnet-b0'. For your reference...",
    "1331460": "Input size: 1x512x512, how you convert 6x273x256 to that size?",
    "1334665": "You can also use RegNetX/RegNetY models from https://github.com/facebookresearch/pycls. 200MF/400MF/800MF are quite small.",
    "1331581": "Personally, I wonder if the ViT models works for this dataset.\n\nFYI:\nI tried `ResMLP` early in the competition, but it didn't perform well. (efficientnet_b0 got 0.97 with same settings.)\nAfter that, I put off ViT and focused on other experiments.\n",
    "1328401": "Thank you for sharing info, just curious did you train these models with channel-wise samples or spatial?",
    "1328221": "For me smaller models just don't work. But very big models like EfficientNet b5 also don't work. The sweet spot for me is  EfficientNet b4",
    "1331508": "what batch size are you using my notebook is taking 1500 sec per epoch for 512x512 with batch size 16",
    "1328206": "train with same epochs?\n or same training time? (example: resnet34d with 20 epoch and efficientnetv2_s with 10 epoch)\n\nHowever, this result must be the most important experimental results I had seen",
    "1332164": "",
    "1328298": ""
  }
}