{
  "id": 106559,
  "title": "Why doesn't this EfficientNet model work?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/106559",
  "author_name": "",
  "post_date": "2019-08-30T00:10:17.562230600Z",
  "votes": 7,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I have made my EfficientNetB3 training kernel public over here:</p>\n\n<p><a href=\"https://www.kaggle.com/tanlikesmath/aptos-2019-previous-dr-efficientnet-single-model?scriptVersionId=19709106\">https://www.kaggle.com/tanlikesmath/aptos-2019-previous-dr-efficientnet-single-model?scriptVersionId=19709106</a></p>\n\n<p>I pretrained the model on the old dataset, with the validation loss on the 2019 dataset, and got a score of 0.88.</p>\n\n<p>This model gives me a score of 0.750, which is nothing near the high scores of 0.79 or so people are able to get with EfficientNet models.</p>\n\n<p>I have been following most of the tips provided by @DrHB and @heye0507 in the forums. Does the community have any tips for me to improve my EfficientNet models?</p>",
  "messages": [
    {
      "id": "612684",
      "postDate": "08/30/2019 00:10:17",
      "content": "<p>I have made my EfficientNetB3 training kernel public over here:</p>\n\n<p><a href=\"https://www.kaggle.com/tanlikesmath/aptos-2019-previous-dr-efficientnet-single-model?scriptVersionId=19709106\">https://www.kaggle.com/tanlikesmath/aptos-2019-previous-dr-efficientnet-single-model?scriptVersionId=19709106</a></p>\n\n<p>I pretrained the model on the old dataset, with the validation loss on the 2019 dataset, and got a score of 0.88.</p>\n\n<p>This model gives me a score of 0.750, which is nothing near the high scores of 0.79 or so people are able to get with EfficientNet models.</p>\n\n<p>I have been following most of the tips provided by @DrHB and @heye0507 in the forums. Does the community have any tips for me to improve my EfficientNet models?</p>",
      "rawMarkdown": "I have made my EfficientNetB3 training kernel public over here:\n\nhttps://www.kaggle.com/tanlikesmath/aptos-2019-previous-dr-efficientnet-single-model?scriptVersionId=19709106\n\nI pretrained the model on the old dataset, with the validation loss on the 2019 dataset, and got a score of 0.88.\n\nThis model gives me a score of 0.750, which is nothing near the high scores of 0.79 or so people are able to get with EfficientNet models.\n\nI have been following most of the tips provided by @DrHB and @heye0507 in the forums. Does the community have any tips for me to improve my EfficientNet models?",
      "votes": null
    },
    {
      "id": "613004",
      "postDate": "08/30/2019 06:46:31",
      "content": "<p>If I may contribute:\n1. Following tips as close as possible may help reduce things to look out for, as there are a lot of possible \"why\"s (both explainable ones and non-explainable ones)\n2. Since val &amp; test do not match that well, you could just be overfitting to validation.\n3. I see that you save based on best qwk, but that may be enhancing overfitting to val. Try monitoring val_loss instead.\n4. I sometimes found a model to fit to test set better when given no aug (weird...), but you could try that. This also technically would count as \"starting simple\"\n5. A lot of people seem to have different approaches that work/ don't work. Trying out as many things as possible (while annoying) will be worth it 😄 </p>\n\n<p>P.S. Congrats on getting in top 💯 </p>",
      "rawMarkdown": "If I may contribute:\n1. Following tips as close as possible may help reduce things to look out for, as there are a lot of possible \"why\"s (both explainable ones and non-explainable ones)\n2. Since val &amp; test do not match that well, you could just be overfitting to validation.\n3. I see that you save based on best qwk, but that may be enhancing overfitting to val. Try monitoring val_loss instead.\n4. I sometimes found a model to fit to test set better when given no aug (weird...), but you could try that. This also technically would count as \"starting simple\"\n5. A lot of people seem to have different approaches that work/ don't work. Trying out as many things as possible (while annoying) will be worth it 😄 \n\nP.S. Congrats on getting in top 💯",
      "votes": null
    },
    {
      "id": "613220",
      "postDate": "08/30/2019 10:15:15",
      "content": "<p>My original experiments used EfficientNetB3 with image size 300 (correct size for this network), while DrHB used a size of 224, but I did the same kind of pretraining and augmentations etc. Then Dreamdragon discussed how he used freezing and unfreezing for EfficientNetB4, which I applied to my B3 model, but also didn't improve it. So for the most part, I think I following the tips pretty closely. However, one thing which you pointed out was the monitoring of the QWK instead of the valid loss. I will try that out and see if it improves my model. Thanks for pointing that out!</p>\n\n<p>Also thanks for the congrats! I joined another participant, and an ensemble of our current models mutually benefited us, so very excited about that! </p>",
      "rawMarkdown": "My original experiments used EfficientNetB3 with image size 300 (correct size for this network), while DrHB used a size of 224, but I did the same kind of pretraining and augmentations etc. Then Dreamdragon discussed how he used freezing and unfreezing for EfficientNetB4, which I applied to my B3 model, but also didn't improve it. So for the most part, I think I following the tips pretty closely. However, one thing which you pointed out was the monitoring of the QWK instead of the valid loss. I will try that out and see if it improves my model. Thanks for pointing that out!\n\n Also thanks for the congrats! I joined another participant, and an ensemble of our current models mutually benefited us, so very excited about that!",
      "votes": null
    },
    {
      "id": "613900",
      "postDate": "08/30/2019 23:16:51",
      "content": "<p>I remember I said my naive conclusion is freeze / unfreeze can be void if you use default head. The result to me is the same as we just train the whole model together :( </p>\n\n<p>And if you want to use concat_pooling (which I said is fastai head), you have to be very careful with the lr as the lr will just kill the efficient net backbone. I agree with people on the fastai forum that efficient net is very hard to tune... at least for me...</p>\n\n<p>Today I trained a classification model, it is also very interesting... It has the same local cv score as my effb4, which should be in the range of 0.78-0.79 LB. I will wait until I have the new quota to submit, but I guess classification should reach the same level as regression? That's my guess...</p>\n\n<p>Also, gratz on getting into top 100 :)</p>",
      "rawMarkdown": "I remember I said my naive conclusion is freeze / unfreeze can be void if you use default head. The result to me is the same as we just train the whole model together :( \n\nAnd if you want to use concat_pooling (which I said is fastai head), you have to be very careful with the lr as the lr will just kill the efficient net backbone. I agree with people on the fastai forum that efficient net is very hard to tune... at least for me...\n\nToday I trained a classification model, it is also very interesting... It has the same local cv score as my effb4, which should be in the range of 0.78-0.79 LB. I will wait until I have the new quota to submit, but I guess classification should reach the same level as regression? That's my guess...\n\nAlso, gratz on getting into top 100 :)",
      "votes": null
    },
    {
      "id": "613905",
      "postDate": "08/30/2019 23:28:55",
      "content": "<p>By default, you mean the Linear layer that comes with the EfficientNet, right? </p>\n\n<p>I tried training the whole model all at once but got poor results (A 5-fold CV got me 0.778 as I discussed in my previous post). Hence, I tried the freezing and unfreezing, as someone else was telling me it is helpful to rescale the gradients or something along those lines and may improve performance. But it did not improve anything it seems.</p>\n\n<p>I am trying the fastai head right now with an EfficientNetB4 network. Essentially trying to follow what you did. The validation QWK after each epoch seems to be promising, but that's not very reliable. Once it finishes, I will submit and see. Thanks again for sharing your tips and good luck to you!</p>",
      "rawMarkdown": "By default, you mean the Linear layer that comes with the EfficientNet, right? \n\nI tried training the whole model all at once but got poor results (A 5-fold CV got me 0.778 as I discussed in my previous post). Hence, I tried the freezing and unfreezing, as someone else was telling me it is helpful to rescale the gradients or something along those lines and may improve performance. But it did not improve anything it seems.\n\nI am trying the fastai head right now with an EfficientNetB4 network. Essentially trying to follow what you did. The validation QWK after each epoch seems to be promising, but that's not very reliable. Once it finishes, I will submit and see. Thanks again for sharing your tips and good luck to you!",
      "votes": null
    },
    {
      "id": "613908",
      "postDate": "08/30/2019 23:36:36",
      "content": "<p>Yes, you just change the nn.Linear(in_features,1) </p>\n\n<p>And for fastai head, I meant create_head( num_features(body)*2 ....)</p>\n\n<p>The problem is where you want to make the cut, at fc? at conv_head? In my case, cut the conv_head gives slightly better result. But with that much time training the model, I found that you just train the whole thing with single lr works best...</p>\n\n<p>That said, it is very frustrated that discriminative lr didn't work... The idea that give early layer smaller lr and later layer large lr was working like charm for me until this competition, I wonder if that's the case I am treating it as regression problem, but for object detection, discriminative lr was fine... </p>\n\n<p>Also, I remember you are working on classification model, if you have the resource, try resNext101_32x16d, the result is very promising on my local cv. (But the pre-training on 2015 dataset took a whole day... ) </p>",
      "rawMarkdown": "Yes, you just change the nn.Linear(in_features,1) \n\nAnd for fastai head, I meant create_head( num_features(body)*2 ....)\n\nThe problem is where you want to make the cut, at fc? at conv_head? In my case, cut the conv_head gives slightly better result. But with that much time training the model, I found that you just train the whole thing with single lr works best...\n\nThat said, it is very frustrated that discriminative lr didn't work... The idea that give early layer smaller lr and later layer large lr was working like charm for me until this competition, I wonder if that's the case I am treating it as regression problem, but for object detection, discriminative lr was fine... \n\nAlso, I remember you are working on classification model, if you have the resource, try resNext101_32x16d, the result is very promising on my local cv. (But the pre-training on 2015 dataset took a whole day... )",
      "votes": null
    },
    {
      "id": "613910",
      "postDate": "08/30/2019 23:42:56",
      "content": "<p>Yes, I split at the fc. I will look into splitting at convhead. And I will look into ResNext. Thanks for the tips!</p>",
      "rawMarkdown": "Yes, I split at the fc. I will look into splitting at convhead. And I will look into ResNext. Thanks for the tips!",
      "votes": null
    },
    {
      "id": "613981",
      "postDate": "08/31/2019 02:00:49",
      "content": "<p><a href=\"/tanlikesmath\">@tanlikesmath</a>  Dude, you jumped from being in 300s to being in top 100 since last two days i guess. Please share some tips. And whats your single best model score ? </p>",
      "rawMarkdown": "tanlikesmath  Dude, you jumped from being in 300s to being in top 100 since last two days i guess. Please share some tips. And whats your single best model score ?",
      "votes": null
    },
    {
      "id": "613983",
      "postDate": "08/31/2019 02:03:13",
      "content": "<p>I teamed up with someone with 90th place or so. Adding my ResNet50 5-fold CV to his ensemble brought us to 70th place.</p>",
      "rawMarkdown": "I teamed up with someone with 90th place or so. Adding my ResNet50 5-fold CV to his ensemble brought us to 70th place.",
      "votes": null
    },
    {
      "id": "613985",
      "postDate": "08/31/2019 02:04:27",
      "content": "<p>Oh Cool. All the best.  So the tip is to team up xD. Unfortunately, we have crossed the deadline. </p>",
      "rawMarkdown": "Oh Cool. All the best.  So the tip is to team up xD. Unfortunately, we have crossed the deadline.",
      "votes": null
    },
    {
      "id": "613987",
      "postDate": "08/31/2019 02:12:23",
      "content": "<p>Indeed teaming up is beneficial as it seems ensembling is taking people very far. But there are so many other tricks out there you can try to use to improve your score. Good luck!</p>",
      "rawMarkdown": "Indeed teaming up is beneficial as it seems ensembling is taking people very far. But there are so many other tricks out there you can try to use to improve your score. Good luck!",
      "votes": null
    },
    {
      "id": "613991",
      "postDate": "08/31/2019 02:29:55",
      "content": "<p>(slight) Success! I made some changes based on <a href=\"/joonl04\">@joonl04</a> and <a href=\"/heye0507\">@heye0507</a> comments. My EfficientNetB3 model has now achieved single model score of 0.780, which is close to <a href=\"/drhabib\">@drhabib</a>'s starter kernel which achieved a score of 0.783. I will now do a 5-fold CV which will hopefully improve the score further. Thanks again for the help. </p>\n\n<p>EDIT: unfortunately, the EfficientNetB4 only got 0.776, but I will keep trying! </p>",
      "rawMarkdown": "(slight) Success! I made some changes based on @joonl04 and @heye0507 comments. My EfficientNetB3 model has now achieved single model score of 0.780, which is close to @drhabib's starter kernel which achieved a score of 0.783. I will now do a 5-fold CV which will hopefully improve the score further. Thanks again for the help. \n\nEDIT: unfortunately, the EfficientNetB4 only got 0.776, but I will keep trying!",
      "votes": null
    },
    {
      "id": "614441",
      "postDate": "08/31/2019 13:43:37",
      "content": "<p>A lot of people overtrain on new data imo. Try only training on both old and new data combined. Also play around with your optimizer hyperparams</p>",
      "rawMarkdown": "A lot of people overtrain on new data imo. Try only training on both old and new data combined. Also play around with your optimizer hyperparams",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 613004,
      "author_name": "joonl04",
      "author_url": "",
      "post_date": "08/30/2019 06:46:31",
      "content": "<p>If I may contribute:\n1. Following tips as close as possible may help reduce things to look out for, as there are a lot of possible \"why\"s (both explainable ones and non-explainable ones)\n2. Since val &amp; test do not match that well, you could just be overfitting to validation.\n3. I see that you save based on best qwk, but that may be enhancing overfitting to val. Try monitoring val_loss instead.\n4. I sometimes found a model to fit to test set better when given no aug (weird...), but you could try that. This also technically would count as \"starting simple\"\n5. A lot of people seem to have different approaches that work/ don't work. Trying out as many things as possible (while annoying) will be worth it 😄 </p>\n\n<p>P.S. Congrats on getting in top 💯 </p>",
      "votes": null,
      "replies": [
        {
          "id": 613220,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/30/2019 10:15:15",
          "content": "<p>My original experiments used EfficientNetB3 with image size 300 (correct size for this network), while DrHB used a size of 224, but I did the same kind of pretraining and augmentations etc. Then Dreamdragon discussed how he used freezing and unfreezing for EfficientNetB4, which I applied to my B3 model, but also didn't improve it. So for the most part, I think I following the tips pretty closely. However, one thing which you pointed out was the monitoring of the QWK instead of the valid loss. I will try that out and see if it improves my model. Thanks for pointing that out!</p>\n\n<p>Also thanks for the congrats! I joined another participant, and an ensemble of our current models mutually benefited us, so very excited about that! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 613900,
          "author_name": "heye0507",
          "author_url": "",
          "post_date": "08/30/2019 23:16:51",
          "content": "<p>I remember I said my naive conclusion is freeze / unfreeze can be void if you use default head. The result to me is the same as we just train the whole model together :( </p>\n\n<p>And if you want to use concat_pooling (which I said is fastai head), you have to be very careful with the lr as the lr will just kill the efficient net backbone. I agree with people on the fastai forum that efficient net is very hard to tune... at least for me...</p>\n\n<p>Today I trained a classification model, it is also very interesting... It has the same local cv score as my effb4, which should be in the range of 0.78-0.79 LB. I will wait until I have the new quota to submit, but I guess classification should reach the same level as regression? That's my guess...</p>\n\n<p>Also, gratz on getting into top 100 :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 613905,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/30/2019 23:28:55",
          "content": "<p>By default, you mean the Linear layer that comes with the EfficientNet, right? </p>\n\n<p>I tried training the whole model all at once but got poor results (A 5-fold CV got me 0.778 as I discussed in my previous post). Hence, I tried the freezing and unfreezing, as someone else was telling me it is helpful to rescale the gradients or something along those lines and may improve performance. But it did not improve anything it seems.</p>\n\n<p>I am trying the fastai head right now with an EfficientNetB4 network. Essentially trying to follow what you did. The validation QWK after each epoch seems to be promising, but that's not very reliable. Once it finishes, I will submit and see. Thanks again for sharing your tips and good luck to you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 613908,
          "author_name": "heye0507",
          "author_url": "",
          "post_date": "08/30/2019 23:36:36",
          "content": "<p>Yes, you just change the nn.Linear(in_features,1) </p>\n\n<p>And for fastai head, I meant create_head( num_features(body)*2 ....)</p>\n\n<p>The problem is where you want to make the cut, at fc? at conv_head? In my case, cut the conv_head gives slightly better result. But with that much time training the model, I found that you just train the whole thing with single lr works best...</p>\n\n<p>That said, it is very frustrated that discriminative lr didn't work... The idea that give early layer smaller lr and later layer large lr was working like charm for me until this competition, I wonder if that's the case I am treating it as regression problem, but for object detection, discriminative lr was fine... </p>\n\n<p>Also, I remember you are working on classification model, if you have the resource, try resNext101_32x16d, the result is very promising on my local cv. (But the pre-training on 2015 dataset took a whole day... ) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 613910,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/30/2019 23:42:56",
          "content": "<p>Yes, I split at the fc. I will look into splitting at convhead. And I will look into ResNext. Thanks for the tips!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 613981,
          "author_name": "virajbagal",
          "author_url": "",
          "post_date": "08/31/2019 02:00:49",
          "content": "<p><a href=\"/tanlikesmath\">@tanlikesmath</a>  Dude, you jumped from being in 300s to being in top 100 since last two days i guess. Please share some tips. And whats your single best model score ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 613983,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/31/2019 02:03:13",
          "content": "<p>I teamed up with someone with 90th place or so. Adding my ResNet50 5-fold CV to his ensemble brought us to 70th place.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 613985,
          "author_name": "virajbagal",
          "author_url": "",
          "post_date": "08/31/2019 02:04:27",
          "content": "<p>Oh Cool. All the best.  So the tip is to team up xD. Unfortunately, we have crossed the deadline. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 613987,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/31/2019 02:12:23",
          "content": "<p>Indeed teaming up is beneficial as it seems ensembling is taking people very far. But there are so many other tricks out there you can try to use to improve your score. Good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 613991,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "08/31/2019 02:29:55",
      "content": "<p>(slight) Success! I made some changes based on <a href=\"/joonl04\">@joonl04</a> and <a href=\"/heye0507\">@heye0507</a> comments. My EfficientNetB3 model has now achieved single model score of 0.780, which is close to <a href=\"/drhabib\">@drhabib</a>'s starter kernel which achieved a score of 0.783. I will now do a 5-fold CV which will hopefully improve the score further. Thanks again for the help. </p>\n\n<p>EDIT: unfortunately, the EfficientNetB4 only got 0.776, but I will keep trying! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 614441,
      "author_name": "sidhanthholalkere",
      "author_url": "",
      "post_date": "08/31/2019 13:43:37",
      "content": "<p>A lot of people overtrain on new data imo. Try only training on both old and new data combined. Also play around with your optimizer hyperparams</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "612684": "I have made my EfficientNetB3 training kernel public over here:\n\nhttps://www.kaggle.com/tanlikesmath/aptos-2019-previous-dr-efficientnet-single-model?scriptVersionId=19709106\n\nI pretrained the model on the old dataset, with the validation loss on the 2019 dataset, and got a score of 0.88.\n\nThis model gives me a score of 0.750, which is nothing near the high scores of 0.79 or so people are able to get with EfficientNet models.\n\nI have been following most of the tips provided by @DrHB and @heye0507 in the forums. Does the community have any tips for me to improve my EfficientNet models?",
    "613004": "If I may contribute:\n1. Following tips as close as possible may help reduce things to look out for, as there are a lot of possible \"why\"s (both explainable ones and non-explainable ones)\n2. Since val &amp; test do not match that well, you could just be overfitting to validation.\n3. I see that you save based on best qwk, but that may be enhancing overfitting to val. Try monitoring val_loss instead.\n4. I sometimes found a model to fit to test set better when given no aug (weird...), but you could try that. This also technically would count as \"starting simple\"\n5. A lot of people seem to have different approaches that work/ don't work. Trying out as many things as possible (while annoying) will be worth it 😄 \n\nP.S. Congrats on getting in top 💯",
    "613220": "My original experiments used EfficientNetB3 with image size 300 (correct size for this network), while DrHB used a size of 224, but I did the same kind of pretraining and augmentations etc. Then Dreamdragon discussed how he used freezing and unfreezing for EfficientNetB4, which I applied to my B3 model, but also didn't improve it. So for the most part, I think I following the tips pretty closely. However, one thing which you pointed out was the monitoring of the QWK instead of the valid loss. I will try that out and see if it improves my model. Thanks for pointing that out!\n\n Also thanks for the congrats! I joined another participant, and an ensemble of our current models mutually benefited us, so very excited about that!",
    "613900": "I remember I said my naive conclusion is freeze / unfreeze can be void if you use default head. The result to me is the same as we just train the whole model together :( \n\nAnd if you want to use concat_pooling (which I said is fastai head), you have to be very careful with the lr as the lr will just kill the efficient net backbone. I agree with people on the fastai forum that efficient net is very hard to tune... at least for me...\n\nToday I trained a classification model, it is also very interesting... It has the same local cv score as my effb4, which should be in the range of 0.78-0.79 LB. I will wait until I have the new quota to submit, but I guess classification should reach the same level as regression? That's my guess...\n\nAlso, gratz on getting into top 100 :)",
    "613905": "By default, you mean the Linear layer that comes with the EfficientNet, right? \n\nI tried training the whole model all at once but got poor results (A 5-fold CV got me 0.778 as I discussed in my previous post). Hence, I tried the freezing and unfreezing, as someone else was telling me it is helpful to rescale the gradients or something along those lines and may improve performance. But it did not improve anything it seems.\n\nI am trying the fastai head right now with an EfficientNetB4 network. Essentially trying to follow what you did. The validation QWK after each epoch seems to be promising, but that's not very reliable. Once it finishes, I will submit and see. Thanks again for sharing your tips and good luck to you!",
    "613908": "Yes, you just change the nn.Linear(in_features,1) \n\nAnd for fastai head, I meant create_head( num_features(body)*2 ....)\n\nThe problem is where you want to make the cut, at fc? at conv_head? In my case, cut the conv_head gives slightly better result. But with that much time training the model, I found that you just train the whole thing with single lr works best...\n\nThat said, it is very frustrated that discriminative lr didn't work... The idea that give early layer smaller lr and later layer large lr was working like charm for me until this competition, I wonder if that's the case I am treating it as regression problem, but for object detection, discriminative lr was fine... \n\nAlso, I remember you are working on classification model, if you have the resource, try resNext101_32x16d, the result is very promising on my local cv. (But the pre-training on 2015 dataset took a whole day... )",
    "613910": "Yes, I split at the fc. I will look into splitting at convhead. And I will look into ResNext. Thanks for the tips!",
    "613981": "tanlikesmath  Dude, you jumped from being in 300s to being in top 100 since last two days i guess. Please share some tips. And whats your single best model score ?",
    "613983": "I teamed up with someone with 90th place or so. Adding my ResNet50 5-fold CV to his ensemble brought us to 70th place.",
    "613985": "Oh Cool. All the best.  So the tip is to team up xD. Unfortunately, we have crossed the deadline.",
    "613987": "Indeed teaming up is beneficial as it seems ensembling is taking people very far. But there are so many other tricks out there you can try to use to improve your score. Good luck!",
    "613991": "(slight) Success! I made some changes based on @joonl04 and @heye0507 comments. My EfficientNetB3 model has now achieved single model score of 0.780, which is close to @drhabib's starter kernel which achieved a score of 0.783. I will now do a 5-fold CV which will hopefully improve the score further. Thanks again for the help. \n\nEDIT: unfortunately, the EfficientNetB4 only got 0.776, but I will keep trying!",
    "614441": "A lot of people overtrain on new data imo. Try only training on both old and new data combined. Also play around with your optimizer hyperparams"
  },
  "source": "meta"
}