{
  "id": 245608,
  "title": "0.98 LB Starter Pack",
  "url": "/competitions/seti-breakthrough-listen/discussion/245608",
  "author_name": "Salman",
  "post_date": "2021-06-11T14:18:53.893000",
  "votes": 39,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Thanks to <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for sharing such amazing notebook.<br>\nJust follow the following steps:</p>\n<ol>\n<li>Copy <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> 's recent <a href=\"https://www.kaggle.com/ttahara/seti-e-t-resnet18d-baseline\" target=\"_blank\">notebook</a></li>\n<li>Change number of epochs to 70</li>\n<li>Change 5 Folds to 4 Folds</li>\n<li>Change Mixup value to 0.5</li>\n<li>Change backbone to efficientnet_b0</li>\n</ol>\n<p>Log Files : <a href=\"https://drive.google.com/drive/folders/1Nnfi3sIyJpTOB0RJ9lsQo6KjfAILdRaj?usp=sharing\" target=\"_blank\">Logs Drive Link</a></p>\n<p>Fold 0: <br>\n<img src=\"https://drive.google.com/file/d/1U2YOtHR5Sqjz7roUPUuJyzSM27WdWxp7/view?usp=sharing\" alt=\"Fold 0\"></p>\n<p>epoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time<br>\n61          35807       0.000247583  0.111056    0.0415692   0.988942    16449.8       <br>\n62          36394       0.000196314  0.107355    0.0460939   0.988797    16719.2       <br>\n63          36981       0.000150773  0.108023    0.0410305   0.989163    16988         <br>\n64          37568       0.000111073  0.110912    0.0457478   0.989234    17257.5       <br>\n65          38155       7.73117e-05  0.108495    0.043509    0.989056    17526.2       <br>\n66          38742       4.95737e-05  0.107411    0.0436515   0.98932     17794.6       <br>\n67          39329       2.79279e-05  0.106854    0.0466319   0.989144    18063.8       <br>\n68          39916       1.2428e-05  0.108483    0.0406972   0.989437    18332.7       <br>\n69          40503       3.11269e-06  0.105295    0.0466031   0.989057    18601.5       <br>\n70          41090       5e-09       0.109043    0.0511871   0.989061    18871 </p>\n<p>Fold 1:<br>\n<img src=\"https://drive.google.com/file/d/1Pndy5uRmztvBG5-cBPCK2p7R98J9_dIk/view?usp=sharing\" alt=\"Fold 1\"><br>\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time<br>\n61          35807       0.000247583  0.109845    0.0439557   0.991264    16410.7       <br>\n62          36394       0.000196314  0.108121    0.0505708   0.991334    16679         <br>\n63          36981       0.000150773  0.110857    0.0397121   0.99074     16947.8       <br>\n64          37568       0.000111073  0.107679    0.0514122   0.99078     17216.6       <br>\n65          38155       7.73117e-05  0.107827    0.0483236   0.991033    17485.4       <br>\n66          38742       4.95737e-05  0.111485    0.0459637   0.990862    17756.1       <br>\n67          39329       2.79279e-05  0.106903    0.049502    0.991183    18025.4       <br>\n68          39916       1.2428e-05  0.110058    0.0412526   0.990907    18295         <br>\n69          40503       3.11269e-06  0.104129    0.052139    0.991297    18563.6       <br>\n70          41090       5e-09       0.111615    0.055737    0.99116     18833.4 </p>\n<p>Fold 2:<br>\n<img src=\"https://drive.google.com/file/d/1VrysfwmMEf8nILwvwf6-DUiIKG7gtKEL/view?usp=sharing\" alt=\"Fold 2\"><br>\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time<br>\n61          35807       0.000247583  0.108603    0.0462216   0.989612    16454.6       <br>\n62          36394       0.000196314  0.106998    0.048744    0.989524    16725.3       <br>\n63          36981       0.000150773  0.108849    0.043051    0.989977    16995.5       <br>\n64          37568       0.000111073  0.10695     0.0450218   0.990057    17265.9       <br>\n65          38155       7.73117e-05  0.105442    0.0438167   0.989921    17536.3       <br>\n66          38742       4.95737e-05  0.10841     0.0462688   0.990048    17807.2       <br>\n67          39329       2.79279e-05  0.104094    0.0478542   0.99        18077.9       <br>\n68          39916       1.2428e-05  0.109336    0.0421842   0.989957    18350.4       <br>\n69          40503       3.11269e-06  0.100248    0.0480441   0.990076    18623         <br>\n70          41090       5e-09       0.108553    0.0534302   0.989969    18894.5 </p>\n<p>Fold 3:<br>\n<img src=\"&lt;a href=\">\" alt=\"Fold 3\" /&gt;<br>\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time<br>\n46          27002       0.00158665  0.122813    0.0497365   0.990732    12479.8       <br>\n47          27589       0.00147179  0.120656    0.0501983   0.989932    12751         <br>\n48          28176       0.00135948  0.11971     0.0463072   0.98993     13022.5       <br>\n49          28763       0.00125     0.117015    0.054712    0.990123    13293.9       <br>\n50          29350       0.00114364  0.118708    0.0424985   0.990918    13565.2       <br>\n51          29937       0.00104064  0.114366    0.0409169   0.991009    13836.5       <br>\n52          30524       0.00094128  0.11631     0.0440744   0.99088     14106.5       <br>\n53          31111       0.00084579  0.114738    0.0431455   0.990284    14377         <br>\n54          31698       0.000754412  0.11416     0.0457304   0.990738    14647.7       <br>\n55          32285       0.000667375  0.110311    0.0431134   0.990832    14918.4       <br>\n56          32872       0.000584893  0.118604    0.0613564   0.989212    15189.6  </p>\n<p>Enjoy Kaggling. </p>",
  "messages": [
    {
      "id": 1345371,
      "postDate": "2021-06-11T14:18:53.893Z",
      "content": "<p>Thanks to <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for sharing such amazing notebook.<br>\nJust follow the following steps:</p>\n<ol>\n<li>Copy <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> 's recent <a href=\"https://www.kaggle.com/ttahara/seti-e-t-resnet18d-baseline\" target=\"_blank\">notebook</a></li>\n<li>Change number of epochs to 70</li>\n<li>Change 5 Folds to 4 Folds</li>\n<li>Change Mixup value to 0.5</li>\n<li>Change backbone to efficientnet_b0</li>\n</ol>\n<p>Log Files : <a href=\"https://drive.google.com/drive/folders/1Nnfi3sIyJpTOB0RJ9lsQo6KjfAILdRaj?usp=sharing\" target=\"_blank\">Logs Drive Link</a></p>\n<p>Fold 0: <br>\n<img src=\"https://drive.google.com/file/d/1U2YOtHR5Sqjz7roUPUuJyzSM27WdWxp7/view?usp=sharing\" alt=\"Fold 0\"></p>\n<p>epoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time<br>\n61          35807       0.000247583  0.111056    0.0415692   0.988942    16449.8       <br>\n62          36394       0.000196314  0.107355    0.0460939   0.988797    16719.2       <br>\n63          36981       0.000150773  0.108023    0.0410305   0.989163    16988         <br>\n64          37568       0.000111073  0.110912    0.0457478   0.989234    17257.5       <br>\n65          38155       7.73117e-05  0.108495    0.043509    0.989056    17526.2       <br>\n66          38742       4.95737e-05  0.107411    0.0436515   0.98932     17794.6       <br>\n67          39329       2.79279e-05  0.106854    0.0466319   0.989144    18063.8       <br>\n68          39916       1.2428e-05  0.108483    0.0406972   0.989437    18332.7       <br>\n69          40503       3.11269e-06  0.105295    0.0466031   0.989057    18601.5       <br>\n70          41090       5e-09       0.109043    0.0511871   0.989061    18871 </p>\n<p>Fold 1:<br>\n<img src=\"https://drive.google.com/file/d/1Pndy5uRmztvBG5-cBPCK2p7R98J9_dIk/view?usp=sharing\" alt=\"Fold 1\"><br>\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time<br>\n61          35807       0.000247583  0.109845    0.0439557   0.991264    16410.7       <br>\n62          36394       0.000196314  0.108121    0.0505708   0.991334    16679         <br>\n63          36981       0.000150773  0.110857    0.0397121   0.99074     16947.8       <br>\n64          37568       0.000111073  0.107679    0.0514122   0.99078     17216.6       <br>\n65          38155       7.73117e-05  0.107827    0.0483236   0.991033    17485.4       <br>\n66          38742       4.95737e-05  0.111485    0.0459637   0.990862    17756.1       <br>\n67          39329       2.79279e-05  0.106903    0.049502    0.991183    18025.4       <br>\n68          39916       1.2428e-05  0.110058    0.0412526   0.990907    18295         <br>\n69          40503       3.11269e-06  0.104129    0.052139    0.991297    18563.6       <br>\n70          41090       5e-09       0.111615    0.055737    0.99116     18833.4 </p>\n<p>Fold 2:<br>\n<img src=\"https://drive.google.com/file/d/1VrysfwmMEf8nILwvwf6-DUiIKG7gtKEL/view?usp=sharing\" alt=\"Fold 2\"><br>\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time<br>\n61          35807       0.000247583  0.108603    0.0462216   0.989612    16454.6       <br>\n62          36394       0.000196314  0.106998    0.048744    0.989524    16725.3       <br>\n63          36981       0.000150773  0.108849    0.043051    0.989977    16995.5       <br>\n64          37568       0.000111073  0.10695     0.0450218   0.990057    17265.9       <br>\n65          38155       7.73117e-05  0.105442    0.0438167   0.989921    17536.3       <br>\n66          38742       4.95737e-05  0.10841     0.0462688   0.990048    17807.2       <br>\n67          39329       2.79279e-05  0.104094    0.0478542   0.99        18077.9       <br>\n68          39916       1.2428e-05  0.109336    0.0421842   0.989957    18350.4       <br>\n69          40503       3.11269e-06  0.100248    0.0480441   0.990076    18623         <br>\n70          41090       5e-09       0.108553    0.0534302   0.989969    18894.5 </p>\n<p>Fold 3:<br>\n<img src=\"&lt;a href=\">\" alt=\"Fold 3\" /&gt;<br>\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time<br>\n46          27002       0.00158665  0.122813    0.0497365   0.990732    12479.8       <br>\n47          27589       0.00147179  0.120656    0.0501983   0.989932    12751         <br>\n48          28176       0.00135948  0.11971     0.0463072   0.98993     13022.5       <br>\n49          28763       0.00125     0.117015    0.054712    0.990123    13293.9       <br>\n50          29350       0.00114364  0.118708    0.0424985   0.990918    13565.2       <br>\n51          29937       0.00104064  0.114366    0.0409169   0.991009    13836.5       <br>\n52          30524       0.00094128  0.11631     0.0440744   0.99088     14106.5       <br>\n53          31111       0.00084579  0.114738    0.0431455   0.990284    14377         <br>\n54          31698       0.000754412  0.11416     0.0457304   0.990738    14647.7       <br>\n55          32285       0.000667375  0.110311    0.0431134   0.990832    14918.4       <br>\n56          32872       0.000584893  0.118604    0.0613564   0.989212    15189.6  </p>\n<p>Enjoy Kaggling. </p>",
      "rawMarkdown": "Thanks to @ttahara for sharing such amazing notebook.\nJust follow the following steps:\n1. Copy @ttahara 's recent [notebook](https://www.kaggle.com/ttahara/seti-e-t-resnet18d-baseline)\n2. Change number of epochs to 70\n3. Change 5 Folds to 4 Folds\n4. Change Mixup value to 0.5\n5. Change backbone to efficientnet_b0\n\nLog Files : [Logs Drive Link](https://drive.google.com/drive/folders/1Nnfi3sIyJpTOB0RJ9lsQo6KjfAILdRaj?usp=sharing)\n\nFold 0: \n![Fold 0](https://drive.google.com/file/d/1U2YOtHR5Sqjz7roUPUuJyzSM27WdWxp7/view?usp=sharing)\n\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time\n61          35807       0.000247583  0.111056    0.0415692   0.988942    16449.8       \n62          36394       0.000196314  0.107355    0.0460939   0.988797    16719.2       \n63          36981       0.000150773  0.108023    0.0410305   0.989163    16988         \n64          37568       0.000111073  0.110912    0.0457478   0.989234    17257.5       \n65          38155       7.73117e-05  0.108495    0.043509    0.989056    17526.2       \n66          38742       4.95737e-05  0.107411    0.0436515   0.98932     17794.6       \n67          39329       2.79279e-05  0.106854    0.0466319   0.989144    18063.8       \n68          39916       1.2428e-05  0.108483    0.0406972   0.989437    18332.7       \n69          40503       3.11269e-06  0.105295    0.0466031   0.989057    18601.5       \n70          41090       5e-09       0.109043    0.0511871   0.989061    18871 \n\nFold 1:\n![Fold 1](https://drive.google.com/file/d/1Pndy5uRmztvBG5-cBPCK2p7R98J9_dIk/view?usp=sharing)\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time\n61          35807       0.000247583  0.109845    0.0439557   0.991264    16410.7       \n62          36394       0.000196314  0.108121    0.0505708   0.991334    16679         \n63          36981       0.000150773  0.110857    0.0397121   0.99074     16947.8       \n64          37568       0.000111073  0.107679    0.0514122   0.99078     17216.6       \n65          38155       7.73117e-05  0.107827    0.0483236   0.991033    17485.4       \n66          38742       4.95737e-05  0.111485    0.0459637   0.990862    17756.1       \n67          39329       2.79279e-05  0.106903    0.049502    0.991183    18025.4       \n68          39916       1.2428e-05  0.110058    0.0412526   0.990907    18295         \n69          40503       3.11269e-06  0.104129    0.052139    0.991297    18563.6       \n70          41090       5e-09       0.111615    0.055737    0.99116     18833.4 \n\nFold 2:\n![Fold 2](https://drive.google.com/file/d/1VrysfwmMEf8nILwvwf6-DUiIKG7gtKEL/view?usp=sharing)\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time\n61          35807       0.000247583  0.108603    0.0462216   0.989612    16454.6       \n62          36394       0.000196314  0.106998    0.048744    0.989524    16725.3       \n63          36981       0.000150773  0.108849    0.043051    0.989977    16995.5       \n64          37568       0.000111073  0.10695     0.0450218   0.990057    17265.9       \n65          38155       7.73117e-05  0.105442    0.0438167   0.989921    17536.3       \n66          38742       4.95737e-05  0.10841     0.0462688   0.990048    17807.2       \n67          39329       2.79279e-05  0.104094    0.0478542   0.99        18077.9       \n68          39916       1.2428e-05  0.109336    0.0421842   0.989957    18350.4       \n69          40503       3.11269e-06  0.100248    0.0480441   0.990076    18623         \n70          41090       5e-09       0.108553    0.0534302   0.989969    18894.5 \n\nFold 3:\n![Fold 3]([](url))\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time\n46          27002       0.00158665  0.122813    0.0497365   0.990732    12479.8       \n47          27589       0.00147179  0.120656    0.0501983   0.989932    12751         \n48          28176       0.00135948  0.11971     0.0463072   0.98993     13022.5       \n49          28763       0.00125     0.117015    0.054712    0.990123    13293.9       \n50          29350       0.00114364  0.118708    0.0424985   0.990918    13565.2       \n51          29937       0.00104064  0.114366    0.0409169   0.991009    13836.5       \n52          30524       0.00094128  0.11631     0.0440744   0.99088     14106.5       \n53          31111       0.00084579  0.114738    0.0431455   0.990284    14377         \n54          31698       0.000754412  0.11416     0.0457304   0.990738    14647.7       \n55          32285       0.000667375  0.110311    0.0431134   0.990832    14918.4       \n56          32872       0.000584893  0.118604    0.0613564   0.989212    15189.6  \n\n\nEnjoy Kaggling. ",
      "votes": 38
    },
    {
      "id": 1347840,
      "postDate": "2021-06-13T14:03:11.423Z",
      "content": "<p>Thanks for Sharing!</p>\n<p>Tried this with some variations (Only have enough GPU hours for 1 fold for everything below) :</p>\n<p><strong>Variation 1</strong></p>\n<ul>\n<li><p>Backbone : resnet18</p></li>\n<li><p>Epochs : 80</p></li>\n<li><p>Folds : 4</p></li>\n<li><p>Mixup Alpha : 0.5</p></li>\n<li><p>Best CV : 0.983 (~50-60th Epoch)</p></li>\n</ul>\n<p><strong>Variation 2</strong></p>\n<ul>\n<li><p>Backbone : efficientnetb0</p></li>\n<li><p>Epochs : 70</p></li>\n<li><p>Folds : 4</p></li>\n<li><p>Mixup Alpha : 0.5</p></li>\n<li><p>Best CV : 0.985 (59th Epoch)</p></li>\n</ul>\n<p><strong>Variation 3</strong></p>\n<ul>\n<li><p>Backbone  : efficientnetb0</p></li>\n<li><p>Epochs : 70</p></li>\n<li><p>Folds : 5</p></li>\n<li><p>Mixup Alpha : 1</p></li>\n<li><p>Best CV : 0.987 (55th Epoch)</p></li>\n</ul>\n<p>In Variation 2, I trained a similar setup as <a href=\"https://www.kaggle.com/micheomaano\" target=\"_blank\">@micheomaano</a> but using Kaggle GPU (Tesla P100), quite surprised that the results differ so much. Just curious if different GPU can affect results of model trng? </p>\n<p>NOTE: There is also a tendency for my models to converge at 50 - 60 Epochs as well so i guess reducing the number of Epochs may help with training time 😉</p>",
      "rawMarkdown": "Thanks for Sharing!\n\nTried this with some variations (Only have enough GPU hours for 1 fold for everything below) :\n\n**Variation 1**\n- Backbone : resnet18\n- Epochs : 80\n- Folds : 4\n- Mixup Alpha : 0.5\n\n- Best CV : 0.983 (~50-60th Epoch)\n\n**Variation 2**\n- Backbone : efficientnetb0\n- Epochs : 70\n- Folds : 4\n- Mixup Alpha : 0.5\n\n- Best CV : 0.985 (59th Epoch)\n\n**Variation 3**\n- Backbone  : efficientnetb0\n- Epochs : 70\n- Folds : 5\n- Mixup Alpha : 1\n\n- Best CV : 0.987 (55th Epoch)\n\n\nIn Variation 2, I trained a similar setup as @micheomaano but using Kaggle GPU (Tesla P100), quite surprised that the results differ so much. Just curious if different GPU can affect results of model trng? \n\nNOTE: There is also a tendency for my models to converge at 50 - 60 Epochs as well so i guess reducing the number of Epochs may help with training time 😉",
      "votes": 5,
      "replies": [
        {
          "id": 1348151,
          "postDate": "2021-06-13T19:11:07.657Z",
          "content": "<p>👍          </p>",
          "rawMarkdown": "👍          "
        },
        {
          "id": 1349619,
          "postDate": "2021-06-15T01:14:58.203Z",
          "content": "<p>Different gpu results may vary. But by how much I’m unsure </p>",
          "rawMarkdown": "Different gpu results may vary. But by how much I’m unsure ",
          "votes": 1
        },
        {
          "id": 1349841,
          "postDate": "2021-06-15T05:59:02.443Z",
          "content": "<p>Exactly. When I use multiple GPUs for training. My network doesn't even converge. :( </p>",
          "rawMarkdown": "Exactly. When I use multiple GPUs for training. My network doesn't even converge. :( ",
          "votes": 1
        },
        {
          "id": 1349877,
          "postDate": "2021-06-15T06:41:35.987Z",
          "content": "<p>you try batchsize under 20 per replica?<br>\nbatch normalization is unstable on distribution learning</p>",
          "rawMarkdown": "you try batchsize under 20 per replica?\nbatch normalization is unstable on distribution learning"
        },
        {
          "id": 1349893,
          "postDate": "2021-06-15T07:06:52.483Z",
          "content": "<p>My batch size is 64 with Mixed Precision.</p>",
          "rawMarkdown": "My batch size is 64 with Mixed Precision."
        }
      ]
    },
    {
      "id": 1350836,
      "postDate": "2021-06-15T20:07:19.190Z",
      "content": "<p>Cool for starter pack! Did you use only even channels of sample?</p>",
      "rawMarkdown": "Cool for starter pack! Did you use only even channels of sample?",
      "replies": [
        {
          "id": 1350863,
          "postDate": "2021-06-15T21:23:34.777Z",
          "content": "<p>Yeah My Bad.<br>\nI guess I did something like this.<br>\n**<br>\nimg_on = np.vstack(img[[0, 2, 4]])<br>\nimg_off = np.vstack(img[[1, 3, 5]])<br>\nimg = np.stack([img_on, img_off]).transpose(2, 1, 0)**</p>",
          "rawMarkdown": "Yeah My Bad.\nI guess I did something like this.\n**\nimg_on = np.vstack(img[[0, 2, 4]])\nimg_off = np.vstack(img[[1, 3, 5]])\nimg = np.stack([img_on, img_off]).transpose(2, 1, 0)**",
          "votes": 2
        },
        {
          "id": 1353290,
          "postDate": "2021-06-17T03:13:24.687Z",
          "content": "<p>Hi, DId you stack even channels and odd channels together such that the model should have input channel number of 2?</p>",
          "rawMarkdown": "Hi, DId you stack even channels and odd channels together such that the model should have input channel number of 2?"
        },
        {
          "id": 1353449,
          "postDate": "2021-06-17T05:40:30.377Z",
          "content": "<p>Yes              </p>",
          "rawMarkdown": "Yes              "
        }
      ]
    },
    {
      "id": 1346084,
      "postDate": "2021-06-12T05:32:25.707Z",
      "content": "<p>all folds from same initialized weights or random weights? </p>",
      "rawMarkdown": "all folds from same initialized weights or random weights? ",
      "replies": [
        {
          "id": 1346089,
          "postDate": "2021-06-12T05:36:27.910Z",
          "content": "<p>Exactly same as in the notebook</p>",
          "rawMarkdown": "Exactly same as in the notebook"
        },
        {
          "id": 1346129,
          "postDate": "2021-06-12T06:06:23.523Z",
          "content": "<p>thank you!<br>\n(+ it is random initialized)</p>",
          "rawMarkdown": "thank you!\n(+ it is random initialized)"
        }
      ]
    },
    {
      "id": 1345481,
      "postDate": "2021-06-11T16:01:00.770Z",
      "content": "<p>Thanks for sharing!</p>\n<p>1) I guess your image size (for these logs) is 512 x 512 (?) Since it takes about 5h per fold, ie one full model/day. </p>\n<p>2) Maybe running for more epochs makes it more stable to LB (due to mixup etc) but I can get similar valid scores for less epochs  </p>\n<p>Edit: although (2) is not directly comparable since we have diff splits probably </p>",
      "rawMarkdown": "Thanks for sharing!\n\n1) I guess your image size (for these logs) is 512 x 512 (?) Since it takes about 5h per fold, ie one full model/day. \n\n2) Maybe running for more epochs makes it more stable to LB (due to mixup etc) but I can get similar valid scores for less epochs  \n\nEdit: although (2) is not directly comparable since we have diff splits probably ",
      "replies": [
        {
          "id": 1345492,
          "postDate": "2021-06-11T16:08:19.260Z",
          "content": "<p>Yes it is 512*512.<br>\nAbout 2 I am using same seed split as in <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> notebook that I mentioned so its same if you want to reproduce results. </p>",
          "rawMarkdown": "Yes it is 512*512.\nAbout 2 I am using same seed split as in @ttahara notebook that I mentioned so its same if you want to reproduce results. ",
          "votes": 1
        },
        {
          "id": 1345985,
          "postDate": "2021-06-12T03:04:47.070Z",
          "content": "<p>Hello! Is 70 epoch necessary? As I am afraid May time out </p>",
          "rawMarkdown": "Hello! Is 70 epoch necessary? As I am afraid May time out ",
          "votes": 2
        },
        {
          "id": 1346022,
          "postDate": "2021-06-12T03:53:33.890Z",
          "content": "<p>For me it starts converging after 50 epochs.</p>",
          "rawMarkdown": "For me it starts converging after 50 epochs."
        },
        {
          "id": 1346044,
          "postDate": "2021-06-12T04:33:55.613Z",
          "content": "<p>May I know your GPU setup? I’m wondering if it’s feasible on colab pro to train 70 epochs</p>",
          "rawMarkdown": "May I know your GPU setup? I’m wondering if it’s feasible on colab pro to train 70 epochs"
        },
        {
          "id": 1346047,
          "postDate": "2021-06-12T04:44:29.850Z",
          "content": "<p>I am using RTX 3090 for now.</p>",
          "rawMarkdown": "I am using RTX 3090 for now.",
          "votes": 1
        },
        {
          "id": 1346082,
          "postDate": "2021-06-12T05:30:24.473Z",
          "content": "<p>Got it. Thanks!</p>",
          "rawMarkdown": "Got it. Thanks!"
        },
        {
          "id": 1346136,
          "postDate": "2021-06-12T06:11:18.283Z",
          "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> using tpu might be a one of choice. But, large batch for speed up seems to make lower score </p>\n<p>256 x 256 tpu with bfloat16, gpu with float32<br>\nkaggle with gpu P100: 230s per epochs ( 64 batch size), <br>\ncolab with gpu V100: 240s per epochs ( 64 batch size), (only try 1 epochs, val_loss = 0.3x val_auc = 0.95x)<br>\ncolab with tpuv2 x8 : 48s per epochs ( 16x8 batch size), val_loss = 3.xxx val_auc = 0.5<br>\ncolab with tpuv2 x8 : 38s per epochs ( 32x8 batch size), val_loss = 0.22xx val_auc = 0.99xx<br>\ncolab with tpuv2 x8 : 21s per epochs ( 64x8 batch size), val_loss = 0.21xx val_auc = 0.99xx<br>\ncolab with tpuv2 x8 : 14s per epochs ( 128x8 batch size), val_loss = 0.22xx val_auc = 0.992<br>\nkaggle with tpuv3 x8 : 12s per epochs ( 128x8 batch size) (only try 2 epochs, val_loss = 0.7x val_auc = 0.6x)<br>\nkaggle with tpuv3 x8 : 13s per epochs (256x8 batch size) (only try 2 epochs, val_loss = 0.8x val_auc = 0.4x)</p>\n<p>data load &amp; augmentation time = 5.8 ~ 6.5s (on tpuv2)<br>\ndata load &amp; augmentation time = 65 ~ 70s (on P100)</p>\n<p>gpu single model epoch 20 : LB .97 (maybe .970~973?) (no validation)<br>\ntpu 16 folding 128x8 epoch 100 : LB .96(maybe .961~963?)     (16 mean fold , best validation loss)</p>\n<p>*EDIT:  P100: 1300 -&gt; 230s <br>\n*EDIT: V100: 1400 -&gt; 240s (both are typo)</p>",
          "rawMarkdown": "@reighns using tpu might be a one of choice. But, large batch for speed up seems to make lower score \n\n256 x 256 tpu with bfloat16, gpu with float32\nkaggle with gpu P100: 230s per epochs ( 64 batch size), \ncolab with gpu V100: 240s per epochs ( 64 batch size), (only try 1 epochs, val_loss = 0.3x val_auc = 0.95x)\ncolab with tpuv2 x8 : 48s per epochs ( 16x8 batch size), val_loss = 3.xxx val_auc = 0.5\ncolab with tpuv2 x8 : 38s per epochs ( 32x8 batch size), val_loss = 0.22xx val_auc = 0.99xx\ncolab with tpuv2 x8 : 21s per epochs ( 64x8 batch size), val_loss = 0.21xx val_auc = 0.99xx\ncolab with tpuv2 x8 : 14s per epochs ( 128x8 batch size), val_loss = 0.22xx val_auc = 0.992\nkaggle with tpuv3 x8 : 12s per epochs ( 128x8 batch size) (only try 2 epochs, val_loss = 0.7x val_auc = 0.6x)\nkaggle with tpuv3 x8 : 13s per epochs (256x8 batch size) (only try 2 epochs, val_loss = 0.8x val_auc = 0.4x)\n\ndata load & augmentation time = 5.8 ~ 6.5s (on tpuv2)\ndata load & augmentation time = 65 ~ 70s (on P100)\n\ngpu single model epoch 20 : LB .97 (maybe .970~973?) (no validation)\ntpu 16 folding 128x8 epoch 100 : LB .96(maybe .961~963?)     (16 mean fold , best validation loss)\n\n*EDIT:  P100: 1300 -> 230s \n*EDIT: V100: 1400 -> 240s (both are typo)\n",
          "votes": 4
        },
        {
          "id": 1349870,
          "postDate": "2021-06-15T06:36:40.880Z",
          "content": "<p><img src=\"https://i.imgflip.com/5db71x.jpg\" alt=\"\"></p>",
          "rawMarkdown": "![](https://i.imgflip.com/5db71x.jpg)",
          "votes": 7
        },
        {
          "id": 1350544,
          "postDate": "2021-06-15T14:25:07.817Z",
          "content": "<p>😂😂😂😂😂😂😂😂😂😂</p>",
          "rawMarkdown": "😂😂😂😂😂😂😂😂😂😂"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1347840,
      "author_name": "maze508",
      "author_url": "",
      "post_date": "2021-06-13T14:03:11.423000",
      "content": "<p>Thanks for Sharing!</p>\n<p>Tried this with some variations (Only have enough GPU hours for 1 fold for everything below) :</p>\n<p><strong>Variation 1</strong></p>\n<ul>\n<li><p>Backbone : resnet18</p></li>\n<li><p>Epochs : 80</p></li>\n<li><p>Folds : 4</p></li>\n<li><p>Mixup Alpha : 0.5</p></li>\n<li><p>Best CV : 0.983 (~50-60th Epoch)</p></li>\n</ul>\n<p><strong>Variation 2</strong></p>\n<ul>\n<li><p>Backbone : efficientnetb0</p></li>\n<li><p>Epochs : 70</p></li>\n<li><p>Folds : 4</p></li>\n<li><p>Mixup Alpha : 0.5</p></li>\n<li><p>Best CV : 0.985 (59th Epoch)</p></li>\n</ul>\n<p><strong>Variation 3</strong></p>\n<ul>\n<li><p>Backbone  : efficientnetb0</p></li>\n<li><p>Epochs : 70</p></li>\n<li><p>Folds : 5</p></li>\n<li><p>Mixup Alpha : 1</p></li>\n<li><p>Best CV : 0.987 (55th Epoch)</p></li>\n</ul>\n<p>In Variation 2, I trained a similar setup as <a href=\"https://www.kaggle.com/micheomaano\" target=\"_blank\">@micheomaano</a> but using Kaggle GPU (Tesla P100), quite surprised that the results differ so much. Just curious if different GPU can affect results of model trng? </p>\n<p>NOTE: There is also a tendency for my models to converge at 50 - 60 Epochs as well so i guess reducing the number of Epochs may help with training time 😉</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1348151,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-06-13T19:11:07.657000",
          "content": "<p>👍          </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1349619,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-06-15T01:14:58.203000",
          "content": "<p>Different gpu results may vary. But by how much I’m unsure </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1349841,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-06-15T05:59:02.443000",
          "content": "<p>Exactly. When I use multiple GPUs for training. My network doesn't even converge. :( </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1349877,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-06-15T06:41:35.987000",
          "content": "<p>you try batchsize under 20 per replica?<br>\nbatch normalization is unstable on distribution learning</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1349893,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-06-15T07:06:52.483000",
          "content": "<p>My batch size is 64 with Mixed Precision.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1350836,
      "author_name": "Sergey Bryansky",
      "author_url": "",
      "post_date": "2021-06-15T20:07:19.190000",
      "content": "<p>Cool for starter pack! Did you use only even channels of sample?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1350863,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-06-15T21:23:34.777000",
          "content": "<p>Yeah My Bad.<br>\nI guess I did something like this.<br>\n**<br>\nimg_on = np.vstack(img[[0, 2, 4]])<br>\nimg_off = np.vstack(img[[1, 3, 5]])<br>\nimg = np.stack([img_on, img_off]).transpose(2, 1, 0)**</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1353290,
          "author_name": "Coin",
          "author_url": "",
          "post_date": "2021-06-17T03:13:24.687000",
          "content": "<p>Hi, DId you stack even channels and odd channels together such that the model should have input channel number of 2?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1353449,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-06-17T05:40:30.377000",
          "content": "<p>Yes              </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1346084,
      "author_name": "assign",
      "author_url": "",
      "post_date": "2021-06-12T05:32:25.707000",
      "content": "<p>all folds from same initialized weights or random weights? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1346089,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-06-12T05:36:27.910000",
          "content": "<p>Exactly same as in the notebook</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1346129,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-06-12T06:06:23.523000",
          "content": "<p>thank you!<br>\n(+ it is random initialized)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1345481,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2021-06-11T16:01:00.770000",
      "content": "<p>Thanks for sharing!</p>\n<p>1) I guess your image size (for these logs) is 512 x 512 (?) Since it takes about 5h per fold, ie one full model/day. </p>\n<p>2) Maybe running for more epochs makes it more stable to LB (due to mixup etc) but I can get similar valid scores for less epochs  </p>\n<p>Edit: although (2) is not directly comparable since we have diff splits probably </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1345492,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-06-11T16:08:19.260000",
          "content": "<p>Yes it is 512*512.<br>\nAbout 2 I am using same seed split as in <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> notebook that I mentioned so its same if you want to reproduce results. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1345985,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-06-12T03:04:47.070000",
          "content": "<p>Hello! Is 70 epoch necessary? As I am afraid May time out </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1346022,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-06-12T03:53:33.890000",
          "content": "<p>For me it starts converging after 50 epochs.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1346044,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-06-12T04:33:55.613000",
          "content": "<p>May I know your GPU setup? I’m wondering if it’s feasible on colab pro to train 70 epochs</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1346047,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-06-12T04:44:29.850000",
          "content": "<p>I am using RTX 3090 for now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1346082,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-06-12T05:30:24.473000",
          "content": "<p>Got it. Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1346136,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-06-12T06:11:18.283000",
          "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> using tpu might be a one of choice. But, large batch for speed up seems to make lower score </p>\n<p>256 x 256 tpu with bfloat16, gpu with float32<br>\nkaggle with gpu P100: 230s per epochs ( 64 batch size), <br>\ncolab with gpu V100: 240s per epochs ( 64 batch size), (only try 1 epochs, val_loss = 0.3x val_auc = 0.95x)<br>\ncolab with tpuv2 x8 : 48s per epochs ( 16x8 batch size), val_loss = 3.xxx val_auc = 0.5<br>\ncolab with tpuv2 x8 : 38s per epochs ( 32x8 batch size), val_loss = 0.22xx val_auc = 0.99xx<br>\ncolab with tpuv2 x8 : 21s per epochs ( 64x8 batch size), val_loss = 0.21xx val_auc = 0.99xx<br>\ncolab with tpuv2 x8 : 14s per epochs ( 128x8 batch size), val_loss = 0.22xx val_auc = 0.992<br>\nkaggle with tpuv3 x8 : 12s per epochs ( 128x8 batch size) (only try 2 epochs, val_loss = 0.7x val_auc = 0.6x)<br>\nkaggle with tpuv3 x8 : 13s per epochs (256x8 batch size) (only try 2 epochs, val_loss = 0.8x val_auc = 0.4x)</p>\n<p>data load &amp; augmentation time = 5.8 ~ 6.5s (on tpuv2)<br>\ndata load &amp; augmentation time = 65 ~ 70s (on P100)</p>\n<p>gpu single model epoch 20 : LB .97 (maybe .970~973?) (no validation)<br>\ntpu 16 folding 128x8 epoch 100 : LB .96(maybe .961~963?)     (16 mean fold , best validation loss)</p>\n<p>*EDIT:  P100: 1300 -&gt; 230s <br>\n*EDIT: V100: 1400 -&gt; 240s (both are typo)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1349870,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2021-06-15T06:36:40.880000",
          "content": "<p><img src=\"https://i.imgflip.com/5db71x.jpg\" alt=\"\"></p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1350544,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-06-15T14:25:07.817000",
          "content": "<p>😂😂😂😂😂😂😂😂😂😂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1345371": "Thanks to @ttahara for sharing such amazing notebook.\nJust follow the following steps:\n1. Copy @ttahara 's recent [notebook](https://www.kaggle.com/ttahara/seti-e-t-resnet18d-baseline)\n2. Change number of epochs to 70\n3. Change 5 Folds to 4 Folds\n4. Change Mixup value to 0.5\n5. Change backbone to efficientnet_b0\n\nLog Files : [Logs Drive Link](https://drive.google.com/drive/folders/1Nnfi3sIyJpTOB0RJ9lsQo6KjfAILdRaj?usp=sharing)\n\nFold 0: \n![Fold 0](https://drive.google.com/file/d/1U2YOtHR5Sqjz7roUPUuJyzSM27WdWxp7/view?usp=sharing)\n\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time\n61          35807       0.000247583  0.111056    0.0415692   0.988942    16449.8       \n62          36394       0.000196314  0.107355    0.0460939   0.988797    16719.2       \n63          36981       0.000150773  0.108023    0.0410305   0.989163    16988         \n64          37568       0.000111073  0.110912    0.0457478   0.989234    17257.5       \n65          38155       7.73117e-05  0.108495    0.043509    0.989056    17526.2       \n66          38742       4.95737e-05  0.107411    0.0436515   0.98932     17794.6       \n67          39329       2.79279e-05  0.106854    0.0466319   0.989144    18063.8       \n68          39916       1.2428e-05  0.108483    0.0406972   0.989437    18332.7       \n69          40503       3.11269e-06  0.105295    0.0466031   0.989057    18601.5       \n70          41090       5e-09       0.109043    0.0511871   0.989061    18871 \n\nFold 1:\n![Fold 1](https://drive.google.com/file/d/1Pndy5uRmztvBG5-cBPCK2p7R98J9_dIk/view?usp=sharing)\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time\n61          35807       0.000247583  0.109845    0.0439557   0.991264    16410.7       \n62          36394       0.000196314  0.108121    0.0505708   0.991334    16679         \n63          36981       0.000150773  0.110857    0.0397121   0.99074     16947.8       \n64          37568       0.000111073  0.107679    0.0514122   0.99078     17216.6       \n65          38155       7.73117e-05  0.107827    0.0483236   0.991033    17485.4       \n66          38742       4.95737e-05  0.111485    0.0459637   0.990862    17756.1       \n67          39329       2.79279e-05  0.106903    0.049502    0.991183    18025.4       \n68          39916       1.2428e-05  0.110058    0.0412526   0.990907    18295         \n69          40503       3.11269e-06  0.104129    0.052139    0.991297    18563.6       \n70          41090       5e-09       0.111615    0.055737    0.99116     18833.4 \n\nFold 2:\n![Fold 2](https://drive.google.com/file/d/1VrysfwmMEf8nILwvwf6-DUiIKG7gtKEL/view?usp=sharing)\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time\n61          35807       0.000247583  0.108603    0.0462216   0.989612    16454.6       \n62          36394       0.000196314  0.106998    0.048744    0.989524    16725.3       \n63          36981       0.000150773  0.108849    0.043051    0.989977    16995.5       \n64          37568       0.000111073  0.10695     0.0450218   0.990057    17265.9       \n65          38155       7.73117e-05  0.105442    0.0438167   0.989921    17536.3       \n66          38742       4.95737e-05  0.10841     0.0462688   0.990048    17807.2       \n67          39329       2.79279e-05  0.104094    0.0478542   0.99        18077.9       \n68          39916       1.2428e-05  0.109336    0.0421842   0.989957    18350.4       \n69          40503       3.11269e-06  0.100248    0.0480441   0.990076    18623         \n70          41090       5e-09       0.108553    0.0534302   0.989969    18894.5 \n\nFold 3:\n![Fold 3]([](url))\nepoch       iteration   lr          train/loss  val/loss    val/metric  elapsed_time\n46          27002       0.00158665  0.122813    0.0497365   0.990732    12479.8       \n47          27589       0.00147179  0.120656    0.0501983   0.989932    12751         \n48          28176       0.00135948  0.11971     0.0463072   0.98993     13022.5       \n49          28763       0.00125     0.117015    0.054712    0.990123    13293.9       \n50          29350       0.00114364  0.118708    0.0424985   0.990918    13565.2       \n51          29937       0.00104064  0.114366    0.0409169   0.991009    13836.5       \n52          30524       0.00094128  0.11631     0.0440744   0.99088     14106.5       \n53          31111       0.00084579  0.114738    0.0431455   0.990284    14377         \n54          31698       0.000754412  0.11416     0.0457304   0.990738    14647.7       \n55          32285       0.000667375  0.110311    0.0431134   0.990832    14918.4       \n56          32872       0.000584893  0.118604    0.0613564   0.989212    15189.6  \n\n\nEnjoy Kaggling. ",
    "1347840": "Thanks for Sharing!\n\nTried this with some variations (Only have enough GPU hours for 1 fold for everything below) :\n\n**Variation 1**\n- Backbone : resnet18\n- Epochs : 80\n- Folds : 4\n- Mixup Alpha : 0.5\n\n- Best CV : 0.983 (~50-60th Epoch)\n\n**Variation 2**\n- Backbone : efficientnetb0\n- Epochs : 70\n- Folds : 4\n- Mixup Alpha : 0.5\n\n- Best CV : 0.985 (59th Epoch)\n\n**Variation 3**\n- Backbone  : efficientnetb0\n- Epochs : 70\n- Folds : 5\n- Mixup Alpha : 1\n\n- Best CV : 0.987 (55th Epoch)\n\n\nIn Variation 2, I trained a similar setup as @micheomaano but using Kaggle GPU (Tesla P100), quite surprised that the results differ so much. Just curious if different GPU can affect results of model trng? \n\nNOTE: There is also a tendency for my models to converge at 50 - 60 Epochs as well so i guess reducing the number of Epochs may help with training time 😉",
    "1350836": "Cool for starter pack! Did you use only even channels of sample?",
    "1346084": "all folds from same initialized weights or random weights? ",
    "1345481": "Thanks for sharing!\n\n1) I guess your image size (for these logs) is 512 x 512 (?) Since it takes about 5h per fold, ie one full model/day. \n\n2) Maybe running for more epochs makes it more stable to LB (due to mixup etc) but I can get similar valid scores for less epochs  \n\nEdit: although (2) is not directly comparable since we have diff splits probably "
  }
}