{
  "id": 252754,
  "title": "My experiments with EfficientNet",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/252754",
  "author_name": "Old Monk",
  "post_date": "2021-07-13T16:15:12.765000",
  "votes": 39,
  "comment_count": 35,
  "views": 0,
  "content": "<p>Hello Kagglers,</p>\n<p>I have tried experimenting with the entire suite of EfficientNet from B0 to B7. We are getting improved LB score for this competition from the higher end of EfficientNet that is B7 compared to B0. For B0 the LB score was in the range of 0.830 while for B7 it was in the range of 0.860. There is also a newer version of EfficientNet B8 which has additional option of adversarial training. Hopefully the LB score and CV score improves with this. The only negative part is a lot of GPU hours went away due to this experimentation, so wanted to share findings in larger forum, in case someone is experimenting on similar lines. Eager to know your thoughts as well.</p>\n<p>Thanks and Regards,<br>\nOld Monk</p>",
  "messages": [
    {
      "id": 1386690,
      "postDate": "2021-07-13T16:15:12.767Z",
      "content": "<p>Hello Kagglers,</p>\n<p>I have tried experimenting with the entire suite of EfficientNet from B0 to B7. We are getting improved LB score for this competition from the higher end of EfficientNet that is B7 compared to B0. For B0 the LB score was in the range of 0.830 while for B7 it was in the range of 0.860. There is also a newer version of EfficientNet B8 which has additional option of adversarial training. Hopefully the LB score and CV score improves with this. The only negative part is a lot of GPU hours went away due to this experimentation, so wanted to share findings in larger forum, in case someone is experimenting on similar lines. Eager to know your thoughts as well.</p>\n<p>Thanks and Regards,<br>\nOld Monk</p>",
      "rawMarkdown": "Hello Kagglers,\n\nI have tried experimenting with the entire suite of EfficientNet from B0 to B7. We are getting improved LB score for this competition from the higher end of EfficientNet that is B7 compared to B0. For B0 the LB score was in the range of 0.830 while for B7 it was in the range of 0.860. There is also a newer version of EfficientNet B8 which has additional option of adversarial training. Hopefully the LB score and CV score improves with this. The only negative part is a lot of GPU hours went away due to this experimentation, so wanted to share findings in larger forum, in case someone is experimenting on similar lines. Eager to know your thoughts as well.\n\nThanks and Regards,\nOld Monk",
      "votes": 38
    },
    {
      "id": 1395264,
      "postDate": "2021-07-21T04:59:51.923Z",
      "content": "<p>I'm not quite sure why do you use so large models… My current LB, 0.875, is just a simple single ResNeXt50 model. I also tried ResNeXt101 that gives only a tiny, ~0.0005 CV boost. Likely for super large models the CV boost could be just ~0.001 with several times longer training…</p>",
      "rawMarkdown": "I'm not quite sure why do you use so large models... My current LB, 0.875, is just a simple single ResNeXt50 model. I also tried ResNeXt101 that gives only a tiny, ~0.0005 CV boost. Likely for super large models the CV boost could be just ~0.001 with several times longer training...",
      "votes": 15,
      "replies": [
        {
          "id": 1395355,
          "postDate": "2021-07-21T06:49:52.780Z",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> , will try out the resnet and resnext series now. </p>",
          "rawMarkdown": "Thanks for sharing @iafoss , will try out the resnet and resnext series now. ",
          "votes": 1
        },
        {
          "id": 1395682,
          "postDate": "2021-07-21T13:00:33.993Z",
          "content": "<p>I think the more important thing here may be the proper treatment of the data. It might be a way how some ppl got 0.879.</p>",
          "rawMarkdown": "I think the more important thing here may be the proper treatment of the data. It might be a way how some ppl got 0.879.",
          "votes": 5
        },
        {
          "id": 1397248,
          "postDate": "2021-07-23T01:15:42.873Z",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, maybe i'm doing something wrong but my B7 model performs worse than my ResNeXt50. Thoughts ?</p>",
          "rawMarkdown": "@iafoss, maybe i'm doing something wrong but my B7 model performs worse than my ResNeXt50. Thoughts ?",
          "votes": 1
        },
        {
          "id": 1397278,
          "postDate": "2021-07-23T02:51:41.977Z",
          "content": "<p>The thing I usually say in this case: \"EfficientNet is build based on architecture search, and the architecture itself may be overfitted to the specific task: working with ImageNet. Other model architecture may be more flexible in using for something very different from ImageNet.\" But in your case poor performance of B7 may also be resulted by some issues with the training setup, not only because of the model itself.</p>\n<p>To be honest, EfficientNet worked well for me only once, and usually I prefer more traditional models (built without automatic architecture search), like ResNeXt, SWIN, ResNeSt. But given that so many people use EfficientNet everywhere, I may be just a person who doesn't know how to train it properly( or know how to train properly other networks))</p>",
          "rawMarkdown": "The thing I usually say in this case: \"EfficientNet is build based on architecture search, and the architecture itself may be overfitted to the specific task: working with ImageNet. Other model architecture may be more flexible in using for something very different from ImageNet.\" But in your case poor performance of B7 may also be resulted by some issues with the training setup, not only because of the model itself.\n\nTo be honest, EfficientNet worked well for me only once, and usually I prefer more traditional models (built without automatic architecture search), like ResNeXt, SWIN, ResNeSt. But given that so many people use EfficientNet everywhere, I may be just a person who doesn't know how to train it properly( or know how to train properly other networks))",
          "votes": 14
        },
        {
          "id": 1398120,
          "postDate": "2021-07-23T18:35:28.750Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1398273,
          "postDate": "2021-07-24T00:55:57.400Z",
          "content": "<p>For others reference, Local KFold CV Score: ResNext50 -&gt; 0.83 and EfficientNet B4 -&gt; 0.86</p>",
          "rawMarkdown": "For others reference, Local KFold CV Score: ResNext50 -> 0.83 and EfficientNet B4 -> 0.86",
          "votes": 1
        },
        {
          "id": 1400169,
          "postDate": "2021-07-26T05:51:15.357Z",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> I was able to train a SEResnet and ResNeXt in much less time with comparable accuracy to a EfficientNet. And from a Deployment perspective the storage size is less as well. </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Local AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNext50</td>\n<td>0.863</td>\n</tr>\n<tr>\n<td>EfficientNetB7</td>\n<td>0.868</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "@iafoss @saurabhbagchi I was able to train a SEResnet and ResNeXt in much less time with comparable accuracy to a EfficientNet. And from a Deployment perspective the storage size is less as well. \n\n| Model | Local AUC |\n| :---: | :---: |\n| ResNext50 | 0.863 |\n| EfficientNetB7 | 0.868 |\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1387464,
      "postDate": "2021-07-14T07:54:23.460Z",
      "content": "<p>Let me give you a direction. As you may have seen in other posts, the preprocessing of this data set is more important. If you handle it well, the LB of the efficientnetv2_b1 model should be close to 0.872 so far.</p>",
      "rawMarkdown": "Let me give you a direction. As you may have seen in other posts, the preprocessing of this data set is more important. If you handle it well, the LB of the efficientnetv2_b1 model should be close to 0.872 so far.",
      "votes": 5,
      "replies": [
        {
          "id": 1387518,
          "postDate": "2021-07-14T08:44:16.240Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/zhangeng\" target=\"_blank\">@zhangeng</a>, will focus more on preprocessing of the data.</p>",
          "rawMarkdown": "Thanks @zhangeng, will focus more on preprocessing of the data."
        }
      ]
    },
    {
      "id": 1397607,
      "postDate": "2021-07-23T10:38:43.317Z",
      "content": "<p>Preprocessing is very important. I'll say it again, but this time I have to add another point, that is, the initial learning rate is also very important. Other models have not been tested yet, but only effv2_ b1.</p>",
      "rawMarkdown": "Preprocessing is very important. I'll say it again, but this time I have to add another point, that is, the initial learning rate is also very important. Other models have not been tested yet, but only effv2_ b1.",
      "votes": 4
    },
    {
      "id": 1397265,
      "postDate": "2021-07-23T02:21:46.763Z",
      "content": "<p>I got a much lower LB score with B0, but I think I botched something (CV ~ 0.9, LB ~ 0.75, weird). I am rather new to these big CNN like Efficientent, so I am wondering, are you using pretrained networks or the \"raw\" version? Normally I would use pretrained, but the spectrograms/Q-transforms are quite different from Imagenet pictures and our dataset here is very large, so maybe it makes more sense to train from scratch here?</p>",
      "rawMarkdown": "I got a much lower LB score with B0, but I think I botched something (CV ~ 0.9, LB ~ 0.75, weird). I am rather new to these big CNN like Efficientent, so I am wondering, are you using pretrained networks or the \"raw\" version? Normally I would use pretrained, but the spectrograms/Q-transforms are quite different from Imagenet pictures and our dataset here is very large, so maybe it makes more sense to train from scratch here?",
      "votes": 1,
      "replies": [
        {
          "id": 1397351,
          "postDate": "2021-07-23T05:36:43.497Z",
          "content": "<p>Yes, we would need to train the models which is why so many GPU hours are going away 😭😭</p>",
          "rawMarkdown": "Yes, we would need to train the models which is why so many GPU hours are going away 😭😭",
          "votes": 1
        },
        {
          "id": 1398116,
          "postDate": "2021-07-23T18:32:32.807Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1398118,
          "postDate": "2021-07-23T18:34:05.890Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1398276,
          "postDate": "2021-07-24T01:05:43.643Z",
          "content": "<p>Thanks, all that makes lots of sense. I found the un-pretrained B0 to learn very fast (overfit after 2 epochs) but with the same CV result as the pre-trained one. Once I get my next GPU hours I'll try smaller learn rates and maybe train some sort of ResNet to compare.</p>",
          "rawMarkdown": "Thanks, all that makes lots of sense. I found the un-pretrained B0 to learn very fast (overfit after 2 epochs) but with the same CV result as the pre-trained one. Once I get my next GPU hours I'll try smaller learn rates and maybe train some sort of ResNet to compare.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1390595,
      "postDate": "2021-07-16T20:15:40.777Z",
      "content": "<p>Your experiments maybe will be helpful, because we have ~75GB of the data and it will be make sense to try deeper model's architecture like B8, because if the model is deeper so it has more probability to be overfiited, if we have low data.  </p>",
      "rawMarkdown": "Your experiments maybe will be helpful, because we have ~75GB of the data and it will be make sense to try deeper model's architecture like B8, because if the model is deeper so it has more probability to be overfiited, if we have low data.  ",
      "votes": 1
    },
    {
      "id": 1388734,
      "postDate": "2021-07-15T07:21:01.213Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a>, can you share the code/metrics, because based on my experiments model size has had no significant improvement (all lie in the ~0.86) range when using q-transformed TFRecords with 4 k-folds.</p>",
      "rawMarkdown": "Hey @saurabhbagchi, can you share the code/metrics, because based on my experiments model size has had no significant improvement (all lie in the ~0.86) range when using q-transformed TFRecords with 4 k-folds.",
      "votes": 1,
      "replies": [
        {
          "id": 1388746,
          "postDate": "2021-07-15T07:35:23.933Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sauravmaheshkar\" target=\"_blank\">@sauravmaheshkar</a>, the difference might be due to usage of mel spectograms versus constant q-transform. </p>",
          "rawMarkdown": "Hi @sauravmaheshkar, the difference might be due to usage of mel spectograms versus constant q-transform. ",
          "votes": 1
        },
        {
          "id": 1392227,
          "postDate": "2021-07-18T13:18:57.967Z",
          "content": "<p>With B7 I'm getting a score 0.868 and with B0 a score of 0.860 with Q-Transformed Dataset. </p>",
          "rawMarkdown": "With B7 I'm getting a score 0.868 and with B0 a score of 0.860 with Q-Transformed Dataset. ",
          "votes": 2
        },
        {
          "id": 1392251,
          "postDate": "2021-07-18T13:40:33.343Z",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sauravmaheshkar\" target=\"_blank\">@sauravmaheshkar</a> , that is great!</p>",
          "rawMarkdown": "Thanks for sharing @sauravmaheshkar , that is great!"
        }
      ]
    },
    {
      "id": 1387662,
      "postDate": "2021-07-14T10:54:52.090Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a>! Do you use mel spectograms or constant q-transform as preprocessing?</p>",
      "rawMarkdown": "Thanks for sharing @saurabhbagchi! Do you use mel spectograms or constant q-transform as preprocessing?",
      "votes": 1,
      "replies": [
        {
          "id": 1387665,
          "postDate": "2021-07-14T10:57:26.723Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a> , have used mel spectograms, hope this helps!</p>",
          "rawMarkdown": "Hi @hannes82 , have used mel spectograms, hope this helps!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1387236,
      "postDate": "2021-07-14T03:51:57.897Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> , looks the bigger backbone the better score for this competition.</p>",
      "rawMarkdown": "Thanks for sharing @saurabhbagchi , looks the bigger backbone the better score for this competition.",
      "votes": 1,
      "replies": [
        {
          "id": 1387288,
          "postDate": "2021-07-14T04:29:01.640Z",
          "content": "<p>Yes, that is what it appears as per the experiments <a href=\"https://www.kaggle.com/superchenhao\" target=\"_blank\">@superchenhao</a> </p>",
          "rawMarkdown": "Yes, that is what it appears as per the experiments @superchenhao "
        }
      ]
    },
    {
      "id": 1387205,
      "postDate": "2021-07-14T03:32:51.653Z",
      "content": "<p>What image size did you use ?</p>",
      "rawMarkdown": "What image size did you use ?",
      "votes": 1,
      "replies": [
        {
          "id": 1387287,
          "postDate": "2021-07-14T04:28:17.650Z",
          "content": "<p>Used 128x123x3 image size as that is what EfficientNet model expects, hope this helps <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> </p>",
          "rawMarkdown": "Used 128x123x3 image size as that is what EfficientNet model expects, hope this helps @mithilsalunkhe "
        },
        {
          "id": 1387290,
          "postDate": "2021-07-14T04:33:32.663Z",
          "content": "<p>Thank you for your insight.May i ask what your training time was with B7</p>",
          "rawMarkdown": "Thank you for your insight.May i ask what your training time was with B7",
          "votes": 1
        },
        {
          "id": 1387294,
          "postDate": "2021-07-14T04:37:45.617Z",
          "content": "<p>For me, it took about 2.5 to 3 hours using Kaggle GPUs <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> </p>",
          "rawMarkdown": "For me, it took about 2.5 to 3 hours using Kaggle GPUs @mithilsalunkhe "
        }
      ]
    },
    {
      "id": 1387174,
      "postDate": "2021-07-14T03:06:56.867Z",
      "content": "<p>Hi Old Monk, could you share how many epochs you trained your models and the learning rate scheduler you used? Thanks.</p>",
      "rawMarkdown": "Hi Old Monk, could you share how many epochs you trained your models and the learning rate scheduler you used? Thanks.",
      "votes": 1,
      "replies": [
        {
          "id": 1387286,
          "postDate": "2021-07-14T04:26:38.733Z",
          "content": "<p>I trained for 5 epochs and used Adam optimizer, hope this helps <a href=\"https://www.kaggle.com/lhkhiem28\" target=\"_blank\">@lhkhiem28</a> </p>",
          "rawMarkdown": "I trained for 5 epochs and used Adam optimizer, hope this helps @lhkhiem28 ",
          "votes": 2
        },
        {
          "id": 1387293,
          "postDate": "2021-07-14T04:36:47.337Z",
          "content": "<p>So useful. Thank you</p>",
          "rawMarkdown": "So useful. Thank you",
          "votes": 2
        },
        {
          "id": 1394057,
          "postDate": "2021-07-20T06:26:37.283Z",
          "content": "<p><a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> could you share what learning rate scheduler you used ?</p>",
          "rawMarkdown": "@saurabhbagchi could you share what learning rate scheduler you used ?"
        }
      ]
    },
    {
      "id": 1561280,
      "postDate": "2021-10-27T13:12:01.580Z",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
    },
    {
      "id": 1386922,
      "postDate": "2021-07-13T19:38:34.560Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1395264,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2021-07-21T04:59:51.923000",
      "content": "<p>I'm not quite sure why do you use so large models… My current LB, 0.875, is just a simple single ResNeXt50 model. I also tried ResNeXt101 that gives only a tiny, ~0.0005 CV boost. Likely for super large models the CV boost could be just ~0.001 with several times longer training…</p>",
      "votes": 15,
      "replies": [
        {
          "id": 1395355,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-07-21T06:49:52.780000",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> , will try out the resnet and resnext series now. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1395682,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2021-07-21T13:00:33.993000",
          "content": "<p>I think the more important thing here may be the proper treatment of the data. It might be a way how some ppl got 0.879.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1397248,
          "author_name": "Saurav Maheshkar ☕️",
          "author_url": "",
          "post_date": "2021-07-23T01:15:42.873000",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>, maybe i'm doing something wrong but my B7 model performs worse than my ResNeXt50. Thoughts ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1397278,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2021-07-23T02:51:41.977000",
          "content": "<p>The thing I usually say in this case: \"EfficientNet is build based on architecture search, and the architecture itself may be overfitted to the specific task: working with ImageNet. Other model architecture may be more flexible in using for something very different from ImageNet.\" But in your case poor performance of B7 may also be resulted by some issues with the training setup, not only because of the model itself.</p>\n<p>To be honest, EfficientNet worked well for me only once, and usually I prefer more traditional models (built without automatic architecture search), like ResNeXt, SWIN, ResNeSt. But given that so many people use EfficientNet everywhere, I may be just a person who doesn't know how to train it properly( or know how to train properly other networks))</p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 1398120,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-23T18:35:28.750000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1398273,
          "author_name": "Saurav Maheshkar ☕️",
          "author_url": "",
          "post_date": "2021-07-24T00:55:57.400000",
          "content": "<p>For others reference, Local KFold CV Score: ResNext50 -&gt; 0.83 and EfficientNet B4 -&gt; 0.86</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1400169,
          "author_name": "Saurav Maheshkar ☕️",
          "author_url": "",
          "post_date": "2021-07-26T05:51:15.357000",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> I was able to train a SEResnet and ResNeXt in much less time with comparable accuracy to a EfficientNet. And from a Deployment perspective the storage size is less as well. </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Local AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNext50</td>\n<td>0.863</td>\n</tr>\n<tr>\n<td>EfficientNetB7</td>\n<td>0.868</td>\n</tr>\n</tbody>\n</table>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1387464,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2021-07-14T07:54:23.460000",
      "content": "<p>Let me give you a direction. As you may have seen in other posts, the preprocessing of this data set is more important. If you handle it well, the LB of the efficientnetv2_b1 model should be close to 0.872 so far.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1387518,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-07-14T08:44:16.240000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/zhangeng\" target=\"_blank\">@zhangeng</a>, will focus more on preprocessing of the data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1397607,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2021-07-23T10:38:43.317000",
      "content": "<p>Preprocessing is very important. I'll say it again, but this time I have to add another point, that is, the initial learning rate is also very important. Other models have not been tested yet, but only effv2_ b1.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1397265,
      "author_name": "Ulrich Ebling",
      "author_url": "",
      "post_date": "2021-07-23T02:21:46.763000",
      "content": "<p>I got a much lower LB score with B0, but I think I botched something (CV ~ 0.9, LB ~ 0.75, weird). I am rather new to these big CNN like Efficientent, so I am wondering, are you using pretrained networks or the \"raw\" version? Normally I would use pretrained, but the spectrograms/Q-transforms are quite different from Imagenet pictures and our dataset here is very large, so maybe it makes more sense to train from scratch here?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1397351,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-07-23T05:36:43.497000",
          "content": "<p>Yes, we would need to train the models which is why so many GPU hours are going away 😭😭</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1398116,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-23T18:32:32.807000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1398118,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-23T18:34:05.890000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1398276,
          "author_name": "Ulrich Ebling",
          "author_url": "",
          "post_date": "2021-07-24T01:05:43.643000",
          "content": "<p>Thanks, all that makes lots of sense. I found the un-pretrained B0 to learn very fast (overfit after 2 epochs) but with the same CV result as the pre-trained one. Once I get my next GPU hours I'll try smaller learn rates and maybe train some sort of ResNet to compare.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1390595,
      "author_name": "Vadim Irtlach",
      "author_url": "",
      "post_date": "2021-07-16T20:15:40.777000",
      "content": "<p>Your experiments maybe will be helpful, because we have ~75GB of the data and it will be make sense to try deeper model's architecture like B8, because if the model is deeper so it has more probability to be overfiited, if we have low data.  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1388734,
      "author_name": "Saurav Maheshkar ☕️",
      "author_url": "",
      "post_date": "2021-07-15T07:21:01.213000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a>, can you share the code/metrics, because based on my experiments model size has had no significant improvement (all lie in the ~0.86) range when using q-transformed TFRecords with 4 k-folds.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1388746,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-07-15T07:35:23.933000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sauravmaheshkar\" target=\"_blank\">@sauravmaheshkar</a>, the difference might be due to usage of mel spectograms versus constant q-transform. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1392227,
          "author_name": "Saurav Maheshkar ☕️",
          "author_url": "",
          "post_date": "2021-07-18T13:18:57.967000",
          "content": "<p>With B7 I'm getting a score 0.868 and with B0 a score of 0.860 with Q-Transformed Dataset. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1392251,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-07-18T13:40:33.343000",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sauravmaheshkar\" target=\"_blank\">@sauravmaheshkar</a> , that is great!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1387662,
      "author_name": "Hannes Öhler",
      "author_url": "",
      "post_date": "2021-07-14T10:54:52.090000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a>! Do you use mel spectograms or constant q-transform as preprocessing?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1387665,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-07-14T10:57:26.723000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a> , have used mel spectograms, hope this helps!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1387236,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2021-07-14T03:51:57.897000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> , looks the bigger backbone the better score for this competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1387288,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-07-14T04:29:01.640000",
          "content": "<p>Yes, that is what it appears as per the experiments <a href=\"https://www.kaggle.com/superchenhao\" target=\"_blank\">@superchenhao</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1387205,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-07-14T03:32:51.653000",
      "content": "<p>What image size did you use ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1387287,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-07-14T04:28:17.650000",
          "content": "<p>Used 128x123x3 image size as that is what EfficientNet model expects, hope this helps <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1387290,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-07-14T04:33:32.663000",
          "content": "<p>Thank you for your insight.May i ask what your training time was with B7</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1387294,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-07-14T04:37:45.617000",
          "content": "<p>For me, it took about 2.5 to 3 hours using Kaggle GPUs <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1387174,
      "author_name": "lhkhiem28",
      "author_url": "",
      "post_date": "2021-07-14T03:06:56.867000",
      "content": "<p>Hi Old Monk, could you share how many epochs you trained your models and the learning rate scheduler you used? Thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1387286,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-07-14T04:26:38.733000",
          "content": "<p>I trained for 5 epochs and used Adam optimizer, hope this helps <a href=\"https://www.kaggle.com/lhkhiem28\" target=\"_blank\">@lhkhiem28</a> </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1387293,
          "author_name": "lhkhiem28",
          "author_url": "",
          "post_date": "2021-07-14T04:36:47.337000",
          "content": "<p>So useful. Thank you</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1394057,
          "author_name": "Atharva Ingle",
          "author_url": "",
          "post_date": "2021-07-20T06:26:37.283000",
          "content": "<p><a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> could you share what learning rate scheduler you used ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1561280,
      "author_name": "ChristopherZerafa",
      "author_url": "",
      "post_date": "2021-10-27T13:12:01.580000",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1386922,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-13T19:38:34.560000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1386690": "Hello Kagglers,\n\nI have tried experimenting with the entire suite of EfficientNet from B0 to B7. We are getting improved LB score for this competition from the higher end of EfficientNet that is B7 compared to B0. For B0 the LB score was in the range of 0.830 while for B7 it was in the range of 0.860. There is also a newer version of EfficientNet B8 which has additional option of adversarial training. Hopefully the LB score and CV score improves with this. The only negative part is a lot of GPU hours went away due to this experimentation, so wanted to share findings in larger forum, in case someone is experimenting on similar lines. Eager to know your thoughts as well.\n\nThanks and Regards,\nOld Monk",
    "1395264": "I'm not quite sure why do you use so large models... My current LB, 0.875, is just a simple single ResNeXt50 model. I also tried ResNeXt101 that gives only a tiny, ~0.0005 CV boost. Likely for super large models the CV boost could be just ~0.001 with several times longer training...",
    "1387464": "Let me give you a direction. As you may have seen in other posts, the preprocessing of this data set is more important. If you handle it well, the LB of the efficientnetv2_b1 model should be close to 0.872 so far.",
    "1397607": "Preprocessing is very important. I'll say it again, but this time I have to add another point, that is, the initial learning rate is also very important. Other models have not been tested yet, but only effv2_ b1.",
    "1397265": "I got a much lower LB score with B0, but I think I botched something (CV ~ 0.9, LB ~ 0.75, weird). I am rather new to these big CNN like Efficientent, so I am wondering, are you using pretrained networks or the \"raw\" version? Normally I would use pretrained, but the spectrograms/Q-transforms are quite different from Imagenet pictures and our dataset here is very large, so maybe it makes more sense to train from scratch here?",
    "1390595": "Your experiments maybe will be helpful, because we have ~75GB of the data and it will be make sense to try deeper model's architecture like B8, because if the model is deeper so it has more probability to be overfiited, if we have low data.  ",
    "1388734": "Hey @saurabhbagchi, can you share the code/metrics, because based on my experiments model size has had no significant improvement (all lie in the ~0.86) range when using q-transformed TFRecords with 4 k-folds.",
    "1387662": "Thanks for sharing @saurabhbagchi! Do you use mel spectograms or constant q-transform as preprocessing?",
    "1387236": "Thanks for sharing @saurabhbagchi , looks the bigger backbone the better score for this competition.",
    "1387205": "What image size did you use ?",
    "1387174": "Hi Old Monk, could you share how many epochs you trained your models and the learning rate scheduler you used? Thanks.",
    "1561280": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
    "1386922": ""
  }
}