{
  "id": 108209,
  "title": "78th place solution and thoughts (first medal!)",
  "url": "/competitions/aptos2019-blindness-detection/discussion/108209",
  "author_name": "ilovescience",
  "post_date": "2019-09-10T00:00:13.713000",
  "votes": 16,
  "comment_count": 6,
  "views": 0,
  "content": "<h2>My experience</h2>\n\n<p>Congratulations to all the participants! I am very excited as this is my first medal!</p>\n\n<p>Before you read my post, please upvote my teammate's ( <a href=\"/agscin\">@agscin</a> ) post over <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107940#latest-620900\">here</a></p>\n\n<p>Just a little bit of my background before I started this competition. I started on Kaggle 3 years ago. I started working on some competitions seriously, but I was a complete beginner, and basically was forking kernels and playing around with hyperparameters.\nI started taking Kaggle more seriously when I saw the TGS salt competition. I saw some of <a href=\"/iafoss\">@iafoss</a>'s kernels using fastai 0.7 and it introduced me to the fast.ai course. I was not able to compete seriously in TGS salt, but I started taking the fast.ai course. Since I was very busy, I went to about 3 lessons before the end of the year.</p>\n\n<p>But then in January, fast.ai released the third iteration of the course. I was very excited, and started taking the course more seriously. To practice, I started turning to Kaggle. I practiced my image classification skills on Histopathological Cancer Detection competition. I got within the top 14%, which was the best I had done in a Kaggle competition. While studying for the fast.ai course, I tried to do small projects in the form of Kaggle Kernels and competitions. For example, to practice using fast.ai tabular and understanding the usage of NN for tabular data, I tried it out on the LANL competition dataset. I wasn't taking the competition seriously, and it turned out that submission would have gotten me a silver medal, which was sad. To put my skills to the use, I then joined on the Freesound Audio Tagging 2019 competition, working on it seriously. Although it was an audio classification competition, it very quickly became an image classification competition. I was doing well, at one point almost close to a silver medal. However, I reached a dead-end, and my position started dropping rapidly. Eventually, I was only 8 positions away from getting a bronze medal, which was really disappointing to be so close, yet miss the medal.</p>\n\n<p>While working on Kaggle and studying with the fast.ai course, I decided to work on a small practical project to help improve my skills. Coincidentally, I had chosen the previous competition diabetic retinopathy dataset and worked on it during my free time for a couple months before this competition started. When I saw this competition, I was really excited, as I already had some idea regarding the type of data, a little bit about the literature, etc. Because of this, I repurposed my exiting kernels to publish a starter kernel and was able to share the resized version of the previous competition dataset, which I already had.  </p>\n\n<p>At this point, I would like to share my solution and a little bit about what I and my teammate did:</p>\n\n<h2>Solution</h2>\n\n<p>I had worked already with ResNet50 using this dataset so I decided to try that. I used a model I trained on the previous dataset before the competition started as the pretrained model for my experiments. The training for that model used oversampling and mixup.</p>\n\n<p>Apart from the models I submitted from my starter kernel, I did most of my experiments using a 5-fold CV setup. Later I will describe some of the EfficientNet models (which were unsuccessful), in which I submitted single models.</p>\n\n<p>My CV was <em>not</em> stratified. Typically, one would take the out-of-fold predictions, put them all together, and calculate the metric on that. Unfortunately, since I was training one fold per kernel, I took the out-of-fold predictions for each fold, calculated the metric, and averaged them. This is not the way people typically do a CV, but it was the best and fastest way for the constraints I had.</p>\n\n<p>Unless otherwise noted, all models are ResNet50:\n1. Taking the classification model from my starter kernel, adding seed and adding 5-fold CV --&gt; public LB 0.749, private LB 0.893\n2. Adding <a href=\"/ratthachat\">@ratthachat</a> <code>crop_image_from_gray</code> function, I got public LB 0.784, private LB 0.903\n3. Tried adding mixup, only got public LB 0.777, but private LB was 0.906. I did not pursue mixup further\n4. Added label smoothing, got public LB of 0.786, private LB of 0.905.\n5. Trained for more epochs, got public LB of 0.793, private LB of 0.903. <strong>This is very surprising.</strong> Training longer improved public LB, but did not improve private LB. In fact, it was the same!! I am unsure why this is the case!</p>\n\n<p>This was essentially my model for the competition. Further experiments did not succeed on the public LB. But I will review them anyway, as it is interesting to compare the private LB:\n1. Trained label smoothing model longer, got public LB of 0.769, and private LB of 0.905. Once again, <strong>the private LB stayed the same!!</strong>\n2. I tried regression, and the public LB was 0.765, and private LB was 0.903\n3. I tried ResNet152, and I got public LB of 0.775, and private LB of 0.901\n4. I tried oversampling with a single model, with poor results. This is interesting, as oversampling was quite useful for the previous dataset based on previous experiments. </p>\n\n<p>Couple other things I tried that failed:\n- I tried training an XResnet50 from scratch on the previous competition dataset using LSUV initialization and LAMB optimizer.\n- I tried the LAMB optimizer to train a ResNet50 (with Imagenet weights) on the previous competition dataset.</p>\n\n<p>At this point, I was stuck with 0.793 and wanted to see what the hype was about EfficientNets. However, as you can see <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105563#latest-617654\">here</a> and <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106559#latest-614441\">here</a>, I significantly struggled. I finally was able to get 0.780 with an EfficientNetB3 and 0.781 with EfficientNetB4. These did surprisingly well on the private LB, obtaining 0.910, and 0.912. At this time, I teamed up with <a href=\"/agscin\">@agscin</a>. We used these models in a special ensemble we tried out that used Random Forests to take all the predictions from all our models, but surprisingly this ensemble did not do well.</p>\n\n<p>We tried many things that did not work:\n1. Various optimizers: Adam+Lookahead, RAdam, RAdam+Lookahead, Yogi, SGD, SGD+Lookahead\n2. Mish activation\n3. Larger and more complex models like ResNext models, SE-ResNet models, etc.\n4. Even smaller models like MobileNetV2\n5. <code>circular_crop</code>\n6. Probably more I cannot remember right now LOL</p>\n\n<p>My teammate tested some of these on CIFAR10 and obtained interesting results that did not transfer to the APTOS dataset, which was quite unfortunate.</p>\n\n<p>Our final model (all classification) is a simple average ensemble of:</p>\n\n<p>2 x EfficientNetB3 trained purely on 2019 train data (with polar unrolling) for classification.\nImage size: 300x768. Public Score: 0.788, 0.782\n<a href=\"https://www.kaggle.com/agscin/polar-unrolling-preprocessing-84th-place\">https://www.kaggle.com/agscin/polar-unrolling-preprocessing-84th-place</a></p>\n\n<p>1 x EfficientNetB3 pretrained on 2015 data and fine-tuned on 2019 (with polar unrolling) for classification. Public Score: 0.794\nImage size: 300x768</p>\n\n<p>5-fold CV ResNet50 trained for classification.\nImage size: 256x256. Public Score: 0.793</p>\n\n<p>Score: Public LB - 0.822, Private LB - 0.923</p>\n\n<p>I want to thank my teammate, Artur.  I had a great time working with him. I would also like to thank <a href=\"/ratthachat\">@ratthachat</a> for his <code>crop_image_from_gray</code> function. And even though I wasn't able to use the EfficientNets, I would like to thanks <a href=\"/drhabib\">@drhabib</a> for his contributions to the competition. I would finally like to thank the fast.ai team for developing an amazing library and course that helped me jumpstart my deep learning experience. I also would like to thank the fastai community for the support and guidance.</p>\n\n<h2>Thoughts</h2>\n\n<ul>\n<li>If you are hitting a dead-end, ask for help on the forums, and people will be glad to help you out. Also, teaming up and ensembling can also really help you if you are struggling. After all, two minds are better than one!</li>\n<li>In this competition, we had no idea what the private LB would be like. But eventually, it turned out to be like the training data (I was so <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763#latest-609621\">wrong</a> and I am not sure why). But it meant that the age-old adage \"Trust your CV\" once again was proven true!</li>\n<li>Even though some of these new optimizers and activation functions did not turn out to be successful in this competition, some of them did well on CIFAR10 and are worth trying out in other competitions.</li>\n<li>Being honest, the ResNet50 models I developed probably have low generalization ability. Interestingly, the validation loss was slowly increasing while QWK also increased. It is possible that ResNet50 may have some problem generalizing. But I believe some of this may have to do with the optimizer and the hyperparameters used. A few hours before the deadline, I tried training a ResNet50 with SGD and I saw almost no difference between training and validation loss. Unfortunately, there was not enough time to try to submit or add to our ensemble so I had no idea how it did on the private LB. According, to <a href=\"https://www.fast.ai/2018/07/02/adam-weight-decay/\">this article</a> Adam's generalization ability, is bad due to poorly chosen hyperparameters but I am unsure. This is something to look into a little further, but it may be important to use SGD optimizer as a baseline.</li>\n</ul>\n\n<p><strong>In conclusion, I learned a lot from this competition. I am happy to get my first medal, and I am excited to apply what I learned here to other competitions!</strong></p>",
  "messages": [
    {
      "id": 622657,
      "postDate": "2019-09-10T00:00:13.713Z",
      "content": "<h2>My experience</h2>\n\n<p>Congratulations to all the participants! I am very excited as this is my first medal!</p>\n\n<p>Before you read my post, please upvote my teammate's ( <a href=\"/agscin\">@agscin</a> ) post over <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107940#latest-620900\">here</a></p>\n\n<p>Just a little bit of my background before I started this competition. I started on Kaggle 3 years ago. I started working on some competitions seriously, but I was a complete beginner, and basically was forking kernels and playing around with hyperparameters.\nI started taking Kaggle more seriously when I saw the TGS salt competition. I saw some of <a href=\"/iafoss\">@iafoss</a>'s kernels using fastai 0.7 and it introduced me to the fast.ai course. I was not able to compete seriously in TGS salt, but I started taking the fast.ai course. Since I was very busy, I went to about 3 lessons before the end of the year.</p>\n\n<p>But then in January, fast.ai released the third iteration of the course. I was very excited, and started taking the course more seriously. To practice, I started turning to Kaggle. I practiced my image classification skills on Histopathological Cancer Detection competition. I got within the top 14%, which was the best I had done in a Kaggle competition. While studying for the fast.ai course, I tried to do small projects in the form of Kaggle Kernels and competitions. For example, to practice using fast.ai tabular and understanding the usage of NN for tabular data, I tried it out on the LANL competition dataset. I wasn't taking the competition seriously, and it turned out that submission would have gotten me a silver medal, which was sad. To put my skills to the use, I then joined on the Freesound Audio Tagging 2019 competition, working on it seriously. Although it was an audio classification competition, it very quickly became an image classification competition. I was doing well, at one point almost close to a silver medal. However, I reached a dead-end, and my position started dropping rapidly. Eventually, I was only 8 positions away from getting a bronze medal, which was really disappointing to be so close, yet miss the medal.</p>\n\n<p>While working on Kaggle and studying with the fast.ai course, I decided to work on a small practical project to help improve my skills. Coincidentally, I had chosen the previous competition diabetic retinopathy dataset and worked on it during my free time for a couple months before this competition started. When I saw this competition, I was really excited, as I already had some idea regarding the type of data, a little bit about the literature, etc. Because of this, I repurposed my exiting kernels to publish a starter kernel and was able to share the resized version of the previous competition dataset, which I already had.  </p>\n\n<p>At this point, I would like to share my solution and a little bit about what I and my teammate did:</p>\n\n<h2>Solution</h2>\n\n<p>I had worked already with ResNet50 using this dataset so I decided to try that. I used a model I trained on the previous dataset before the competition started as the pretrained model for my experiments. The training for that model used oversampling and mixup.</p>\n\n<p>Apart from the models I submitted from my starter kernel, I did most of my experiments using a 5-fold CV setup. Later I will describe some of the EfficientNet models (which were unsuccessful), in which I submitted single models.</p>\n\n<p>My CV was <em>not</em> stratified. Typically, one would take the out-of-fold predictions, put them all together, and calculate the metric on that. Unfortunately, since I was training one fold per kernel, I took the out-of-fold predictions for each fold, calculated the metric, and averaged them. This is not the way people typically do a CV, but it was the best and fastest way for the constraints I had.</p>\n\n<p>Unless otherwise noted, all models are ResNet50:\n1. Taking the classification model from my starter kernel, adding seed and adding 5-fold CV --&gt; public LB 0.749, private LB 0.893\n2. Adding <a href=\"/ratthachat\">@ratthachat</a> <code>crop_image_from_gray</code> function, I got public LB 0.784, private LB 0.903\n3. Tried adding mixup, only got public LB 0.777, but private LB was 0.906. I did not pursue mixup further\n4. Added label smoothing, got public LB of 0.786, private LB of 0.905.\n5. Trained for more epochs, got public LB of 0.793, private LB of 0.903. <strong>This is very surprising.</strong> Training longer improved public LB, but did not improve private LB. In fact, it was the same!! I am unsure why this is the case!</p>\n\n<p>This was essentially my model for the competition. Further experiments did not succeed on the public LB. But I will review them anyway, as it is interesting to compare the private LB:\n1. Trained label smoothing model longer, got public LB of 0.769, and private LB of 0.905. Once again, <strong>the private LB stayed the same!!</strong>\n2. I tried regression, and the public LB was 0.765, and private LB was 0.903\n3. I tried ResNet152, and I got public LB of 0.775, and private LB of 0.901\n4. I tried oversampling with a single model, with poor results. This is interesting, as oversampling was quite useful for the previous dataset based on previous experiments. </p>\n\n<p>Couple other things I tried that failed:\n- I tried training an XResnet50 from scratch on the previous competition dataset using LSUV initialization and LAMB optimizer.\n- I tried the LAMB optimizer to train a ResNet50 (with Imagenet weights) on the previous competition dataset.</p>\n\n<p>At this point, I was stuck with 0.793 and wanted to see what the hype was about EfficientNets. However, as you can see <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105563#latest-617654\">here</a> and <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106559#latest-614441\">here</a>, I significantly struggled. I finally was able to get 0.780 with an EfficientNetB3 and 0.781 with EfficientNetB4. These did surprisingly well on the private LB, obtaining 0.910, and 0.912. At this time, I teamed up with <a href=\"/agscin\">@agscin</a>. We used these models in a special ensemble we tried out that used Random Forests to take all the predictions from all our models, but surprisingly this ensemble did not do well.</p>\n\n<p>We tried many things that did not work:\n1. Various optimizers: Adam+Lookahead, RAdam, RAdam+Lookahead, Yogi, SGD, SGD+Lookahead\n2. Mish activation\n3. Larger and more complex models like ResNext models, SE-ResNet models, etc.\n4. Even smaller models like MobileNetV2\n5. <code>circular_crop</code>\n6. Probably more I cannot remember right now LOL</p>\n\n<p>My teammate tested some of these on CIFAR10 and obtained interesting results that did not transfer to the APTOS dataset, which was quite unfortunate.</p>\n\n<p>Our final model (all classification) is a simple average ensemble of:</p>\n\n<p>2 x EfficientNetB3 trained purely on 2019 train data (with polar unrolling) for classification.\nImage size: 300x768. Public Score: 0.788, 0.782\n<a href=\"https://www.kaggle.com/agscin/polar-unrolling-preprocessing-84th-place\">https://www.kaggle.com/agscin/polar-unrolling-preprocessing-84th-place</a></p>\n\n<p>1 x EfficientNetB3 pretrained on 2015 data and fine-tuned on 2019 (with polar unrolling) for classification. Public Score: 0.794\nImage size: 300x768</p>\n\n<p>5-fold CV ResNet50 trained for classification.\nImage size: 256x256. Public Score: 0.793</p>\n\n<p>Score: Public LB - 0.822, Private LB - 0.923</p>\n\n<p>I want to thank my teammate, Artur.  I had a great time working with him. I would also like to thank <a href=\"/ratthachat\">@ratthachat</a> for his <code>crop_image_from_gray</code> function. And even though I wasn't able to use the EfficientNets, I would like to thanks <a href=\"/drhabib\">@drhabib</a> for his contributions to the competition. I would finally like to thank the fast.ai team for developing an amazing library and course that helped me jumpstart my deep learning experience. I also would like to thank the fastai community for the support and guidance.</p>\n\n<h2>Thoughts</h2>\n\n<ul>\n<li>If you are hitting a dead-end, ask for help on the forums, and people will be glad to help you out. Also, teaming up and ensembling can also really help you if you are struggling. After all, two minds are better than one!</li>\n<li>In this competition, we had no idea what the private LB would be like. But eventually, it turned out to be like the training data (I was so <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763#latest-609621\">wrong</a> and I am not sure why). But it meant that the age-old adage \"Trust your CV\" once again was proven true!</li>\n<li>Even though some of these new optimizers and activation functions did not turn out to be successful in this competition, some of them did well on CIFAR10 and are worth trying out in other competitions.</li>\n<li>Being honest, the ResNet50 models I developed probably have low generalization ability. Interestingly, the validation loss was slowly increasing while QWK also increased. It is possible that ResNet50 may have some problem generalizing. But I believe some of this may have to do with the optimizer and the hyperparameters used. A few hours before the deadline, I tried training a ResNet50 with SGD and I saw almost no difference between training and validation loss. Unfortunately, there was not enough time to try to submit or add to our ensemble so I had no idea how it did on the private LB. According, to <a href=\"https://www.fast.ai/2018/07/02/adam-weight-decay/\">this article</a> Adam's generalization ability, is bad due to poorly chosen hyperparameters but I am unsure. This is something to look into a little further, but it may be important to use SGD optimizer as a baseline.</li>\n</ul>\n\n<p><strong>In conclusion, I learned a lot from this competition. I am happy to get my first medal, and I am excited to apply what I learned here to other competitions!</strong></p>",
      "rawMarkdown": "## My experience\n\nCongratulations to all the participants! I am very excited as this is my first medal!\n\n\nBefore you read my post, please upvote my teammate's ( @agscin ) post over [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107940#latest-620900)\n\n\nJust a little bit of my background before I started this competition. I started on Kaggle 3 years ago. I started working on some competitions seriously, but I was a complete beginner, and basically was forking kernels and playing around with hyperparameters.\nI started taking Kaggle more seriously when I saw the TGS salt competition. I saw some of @iafoss's kernels using fastai 0.7 and it introduced me to the fast.ai course. I was not able to compete seriously in TGS salt, but I started taking the fast.ai course. Since I was very busy, I went to about 3 lessons before the end of the year.\n\nBut then in January, fast.ai released the third iteration of the course. I was very excited, and started taking the course more seriously. To practice, I started turning to Kaggle. I practiced my image classification skills on Histopathological Cancer Detection competition. I got within the top 14%, which was the best I had done in a Kaggle competition. While studying for the fast.ai course, I tried to do small projects in the form of Kaggle Kernels and competitions. For example, to practice using fast.ai tabular and understanding the usage of NN for tabular data, I tried it out on the LANL competition dataset. I wasn't taking the competition seriously, and it turned out that submission would have gotten me a silver medal, which was sad. To put my skills to the use, I then joined on the Freesound Audio Tagging 2019 competition, working on it seriously. Although it was an audio classification competition, it very quickly became an image classification competition. I was doing well, at one point almost close to a silver medal. However, I reached a dead-end, and my position started dropping rapidly. Eventually, I was only 8 positions away from getting a bronze medal, which was really disappointing to be so close, yet miss the medal.\n\nWhile working on Kaggle and studying with the fast.ai course, I decided to work on a small practical project to help improve my skills. Coincidentally, I had chosen the previous competition diabetic retinopathy dataset and worked on it during my free time for a couple months before this competition started. When I saw this competition, I was really excited, as I already had some idea regarding the type of data, a little bit about the literature, etc. Because of this, I repurposed my exiting kernels to publish a starter kernel and was able to share the resized version of the previous competition dataset, which I already had.  \n\nAt this point, I would like to share my solution and a little bit about what I and my teammate did:\n\n## Solution\n\nI had worked already with ResNet50 using this dataset so I decided to try that. I used a model I trained on the previous dataset before the competition started as the pretrained model for my experiments. The training for that model used oversampling and mixup.\n\nApart from the models I submitted from my starter kernel, I did most of my experiments using a 5-fold CV setup. Later I will describe some of the EfficientNet models (which were unsuccessful), in which I submitted single models.\n\nMy CV was *not* stratified. Typically, one would take the out-of-fold predictions, put them all together, and calculate the metric on that. Unfortunately, since I was training one fold per kernel, I took the out-of-fold predictions for each fold, calculated the metric, and averaged them. This is not the way people typically do a CV, but it was the best and fastest way for the constraints I had.\n\n\nUnless otherwise noted, all models are ResNet50:\n1. Taking the classification model from my starter kernel, adding seed and adding 5-fold CV --&gt; public LB 0.749, private LB 0.893\n2. Adding @ratthachat `crop_image_from_gray` function, I got public LB 0.784, private LB 0.903\n3. Tried adding mixup, only got public LB 0.777, but private LB was 0.906. I did not pursue mixup further\n4. Added label smoothing, got public LB of 0.786, private LB of 0.905.\n5. Trained for more epochs, got public LB of 0.793, private LB of 0.903. **This is very surprising.** Training longer improved public LB, but did not improve private LB. In fact, it was the same!! I am unsure why this is the case!\n\nThis was essentially my model for the competition. Further experiments did not succeed on the public LB. But I will review them anyway, as it is interesting to compare the private LB:\n1. Trained label smoothing model longer, got public LB of 0.769, and private LB of 0.905. Once again, **the private LB stayed the same!!**\n2. I tried regression, and the public LB was 0.765, and private LB was 0.903\n3. I tried ResNet152, and I got public LB of 0.775, and private LB of 0.901\n4. I tried oversampling with a single model, with poor results. This is interesting, as oversampling was quite useful for the previous dataset based on previous experiments. \n\nCouple other things I tried that failed:\n- I tried training an XResnet50 from scratch on the previous competition dataset using LSUV initialization and LAMB optimizer.\n- I tried the LAMB optimizer to train a ResNet50 (with Imagenet weights) on the previous competition dataset.\n\nAt this point, I was stuck with 0.793 and wanted to see what the hype was about EfficientNets. However, as you can see [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105563#latest-617654) and [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106559#latest-614441), I significantly struggled. I finally was able to get 0.780 with an EfficientNetB3 and 0.781 with EfficientNetB4. These did surprisingly well on the private LB, obtaining 0.910, and 0.912. At this time, I teamed up with @agscin. We used these models in a special ensemble we tried out that used Random Forests to take all the predictions from all our models, but surprisingly this ensemble did not do well.\n\nWe tried many things that did not work:\n1. Various optimizers: Adam+Lookahead, RAdam, RAdam+Lookahead, Yogi, SGD, SGD+Lookahead\n2. Mish activation\n3. Larger and more complex models like ResNext models, SE-ResNet models, etc.\n4. Even smaller models like MobileNetV2\n5. `circular_crop`\n6. Probably more I cannot remember right now LOL\n\n\nMy teammate tested some of these on CIFAR10 and obtained interesting results that did not transfer to the APTOS dataset, which was quite unfortunate.\n\nOur final model (all classification) is a simple average ensemble of:\n\n2 x EfficientNetB3 trained purely on 2019 train data (with polar unrolling) for classification.\nImage size: 300x768. Public Score: 0.788, 0.782\nhttps://www.kaggle.com/agscin/polar-unrolling-preprocessing-84th-place\n\n1 x EfficientNetB3 pretrained on 2015 data and fine-tuned on 2019 (with polar unrolling) for classification. Public Score: 0.794\nImage size: 300x768\n\n5-fold CV ResNet50 trained for classification.\nImage size: 256x256. Public Score: 0.793\n\nScore: Public LB - 0.822, Private LB - 0.923\n\nI want to thank my teammate, Artur.  I had a great time working with him. I would also like to thank @ratthachat for his `crop_image_from_gray` function. And even though I wasn't able to use the EfficientNets, I would like to thanks @drhabib for his contributions to the competition. I would finally like to thank the fast.ai team for developing an amazing library and course that helped me jumpstart my deep learning experience. I also would like to thank the [fastai community](forums.fast.ai) for the support and guidance.\n\n## Thoughts\n\n - If you are hitting a dead-end, ask for help on the forums, and people will be glad to help you out. Also, teaming up and ensembling can also really help you if you are struggling. After all, two minds are better than one!\n - In this competition, we had no idea what the private LB would be like. But eventually, it turned out to be like the training data (I was so [wrong](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763#latest-609621) and I am not sure why). But it meant that the age-old adage \"Trust your CV\" once again was proven true!\n - Even though some of these new optimizers and activation functions did not turn out to be successful in this competition, some of them did well on CIFAR10 and are worth trying out in other competitions.\n - Being honest, the ResNet50 models I developed probably have low generalization ability. Interestingly, the validation loss was slowly increasing while QWK also increased. It is possible that ResNet50 may have some problem generalizing. But I believe some of this may have to do with the optimizer and the hyperparameters used. A few hours before the deadline, I tried training a ResNet50 with SGD and I saw almost no difference between training and validation loss. Unfortunately, there was not enough time to try to submit or add to our ensemble so I had no idea how it did on the private LB. According, to [this article](https://www.fast.ai/2018/07/02/adam-weight-decay/) Adam's generalization ability, is bad due to poorly chosen hyperparameters but I am unsure. This is something to look into a little further, but it may be important to use SGD optimizer as a baseline.\n\n**In conclusion, I learned a lot from this competition. I am happy to get my first medal, and I am excited to apply what I learned here to other competitions!**",
      "votes": 16
    },
    {
      "id": 622710,
      "postDate": "2019-09-10T01:49:22.390Z",
      "content": "<p>Thanks for writing and sharing! Your 2015 dataset and starter kernels are also invaluable to all of us! </p>\n\n<p>I have no doubt that you will make many more contributions to our community in the future, and I am looking forward to that!</p>",
      "rawMarkdown": "Thanks for writing and sharing! Your 2015 dataset and starter kernels are also invaluable to all of us! \n\nI have no doubt that you will make many more contributions to our community in the future, and I am looking forward to that!",
      "votes": 3,
      "replies": [
        {
          "id": 622745,
          "postDate": "2019-09-10T03:21:24.647Z",
          "content": "<p>Thank you for your contributions to the competition!</p>",
          "rawMarkdown": "Thank you for your contributions to the competition!",
          "votes": 1
        }
      ]
    },
    {
      "id": 622662,
      "postDate": "2019-09-10T00:16:31.067Z",
      "content": "<p>Congratulations! I really loved your personal journey it’s remind me a bit of my experience:) You are doing great keep learning and practicing:) Congratulations on your first medal!! </p>",
      "rawMarkdown": "Congratulations! I really loved your personal journey it’s remind me a bit of my experience:) You are doing great keep learning and practicing:) Congratulations on your first medal!! ",
      "votes": 4,
      "replies": [
        {
          "id": 622663,
          "postDate": "2019-09-10T00:19:25.590Z",
          "content": "<p>Thanks! congratulations on the gold medal and becoming a master!  Nice to see a fellow fastai user do so well!</p>",
          "rawMarkdown": "Thanks! congratulations on the gold medal and becoming a master!  Nice to see a fellow fastai user do so well!",
          "votes": 1
        }
      ]
    },
    {
      "id": 622724,
      "postDate": "2019-09-10T02:26:32.663Z",
      "content": "<p>Congrats and thanks for sharing your learning journey with us! It is inspiring!</p>",
      "rawMarkdown": "Congrats and thanks for sharing your learning journey with us! It is inspiring!\n",
      "votes": 1
    },
    {
      "id": 622770,
      "postDate": "2019-09-10T04:36:37.193Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 622710,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-09-10T01:49:22.390000",
      "content": "<p>Thanks for writing and sharing! Your 2015 dataset and starter kernels are also invaluable to all of us! </p>\n\n<p>I have no doubt that you will make many more contributions to our community in the future, and I am looking forward to that!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 622745,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2019-09-10T03:21:24.647000",
          "content": "<p>Thank you for your contributions to the competition!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 622662,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2019-09-10T00:16:31.067000",
      "content": "<p>Congratulations! I really loved your personal journey it’s remind me a bit of my experience:) You are doing great keep learning and practicing:) Congratulations on your first medal!! </p>",
      "votes": 4,
      "replies": [
        {
          "id": 622663,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2019-09-10T00:19:25.590000",
          "content": "<p>Thanks! congratulations on the gold medal and becoming a master!  Nice to see a fellow fastai user do so well!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 622724,
      "author_name": "Qile Tan",
      "author_url": "",
      "post_date": "2019-09-10T02:26:32.663000",
      "content": "<p>Congrats and thanks for sharing your learning journey with us! It is inspiring!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 622770,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T04:36:37.193000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "622657": "## My experience\n\nCongratulations to all the participants! I am very excited as this is my first medal!\n\n\nBefore you read my post, please upvote my teammate's ( @agscin ) post over [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107940#latest-620900)\n\n\nJust a little bit of my background before I started this competition. I started on Kaggle 3 years ago. I started working on some competitions seriously, but I was a complete beginner, and basically was forking kernels and playing around with hyperparameters.\nI started taking Kaggle more seriously when I saw the TGS salt competition. I saw some of @iafoss's kernels using fastai 0.7 and it introduced me to the fast.ai course. I was not able to compete seriously in TGS salt, but I started taking the fast.ai course. Since I was very busy, I went to about 3 lessons before the end of the year.\n\nBut then in January, fast.ai released the third iteration of the course. I was very excited, and started taking the course more seriously. To practice, I started turning to Kaggle. I practiced my image classification skills on Histopathological Cancer Detection competition. I got within the top 14%, which was the best I had done in a Kaggle competition. While studying for the fast.ai course, I tried to do small projects in the form of Kaggle Kernels and competitions. For example, to practice using fast.ai tabular and understanding the usage of NN for tabular data, I tried it out on the LANL competition dataset. I wasn't taking the competition seriously, and it turned out that submission would have gotten me a silver medal, which was sad. To put my skills to the use, I then joined on the Freesound Audio Tagging 2019 competition, working on it seriously. Although it was an audio classification competition, it very quickly became an image classification competition. I was doing well, at one point almost close to a silver medal. However, I reached a dead-end, and my position started dropping rapidly. Eventually, I was only 8 positions away from getting a bronze medal, which was really disappointing to be so close, yet miss the medal.\n\nWhile working on Kaggle and studying with the fast.ai course, I decided to work on a small practical project to help improve my skills. Coincidentally, I had chosen the previous competition diabetic retinopathy dataset and worked on it during my free time for a couple months before this competition started. When I saw this competition, I was really excited, as I already had some idea regarding the type of data, a little bit about the literature, etc. Because of this, I repurposed my exiting kernels to publish a starter kernel and was able to share the resized version of the previous competition dataset, which I already had.  \n\nAt this point, I would like to share my solution and a little bit about what I and my teammate did:\n\n## Solution\n\nI had worked already with ResNet50 using this dataset so I decided to try that. I used a model I trained on the previous dataset before the competition started as the pretrained model for my experiments. The training for that model used oversampling and mixup.\n\nApart from the models I submitted from my starter kernel, I did most of my experiments using a 5-fold CV setup. Later I will describe some of the EfficientNet models (which were unsuccessful), in which I submitted single models.\n\nMy CV was *not* stratified. Typically, one would take the out-of-fold predictions, put them all together, and calculate the metric on that. Unfortunately, since I was training one fold per kernel, I took the out-of-fold predictions for each fold, calculated the metric, and averaged them. This is not the way people typically do a CV, but it was the best and fastest way for the constraints I had.\n\n\nUnless otherwise noted, all models are ResNet50:\n1. Taking the classification model from my starter kernel, adding seed and adding 5-fold CV --&gt; public LB 0.749, private LB 0.893\n2. Adding @ratthachat `crop_image_from_gray` function, I got public LB 0.784, private LB 0.903\n3. Tried adding mixup, only got public LB 0.777, but private LB was 0.906. I did not pursue mixup further\n4. Added label smoothing, got public LB of 0.786, private LB of 0.905.\n5. Trained for more epochs, got public LB of 0.793, private LB of 0.903. **This is very surprising.** Training longer improved public LB, but did not improve private LB. In fact, it was the same!! I am unsure why this is the case!\n\nThis was essentially my model for the competition. Further experiments did not succeed on the public LB. But I will review them anyway, as it is interesting to compare the private LB:\n1. Trained label smoothing model longer, got public LB of 0.769, and private LB of 0.905. Once again, **the private LB stayed the same!!**\n2. I tried regression, and the public LB was 0.765, and private LB was 0.903\n3. I tried ResNet152, and I got public LB of 0.775, and private LB of 0.901\n4. I tried oversampling with a single model, with poor results. This is interesting, as oversampling was quite useful for the previous dataset based on previous experiments. \n\nCouple other things I tried that failed:\n- I tried training an XResnet50 from scratch on the previous competition dataset using LSUV initialization and LAMB optimizer.\n- I tried the LAMB optimizer to train a ResNet50 (with Imagenet weights) on the previous competition dataset.\n\nAt this point, I was stuck with 0.793 and wanted to see what the hype was about EfficientNets. However, as you can see [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105563#latest-617654) and [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106559#latest-614441), I significantly struggled. I finally was able to get 0.780 with an EfficientNetB3 and 0.781 with EfficientNetB4. These did surprisingly well on the private LB, obtaining 0.910, and 0.912. At this time, I teamed up with @agscin. We used these models in a special ensemble we tried out that used Random Forests to take all the predictions from all our models, but surprisingly this ensemble did not do well.\n\nWe tried many things that did not work:\n1. Various optimizers: Adam+Lookahead, RAdam, RAdam+Lookahead, Yogi, SGD, SGD+Lookahead\n2. Mish activation\n3. Larger and more complex models like ResNext models, SE-ResNet models, etc.\n4. Even smaller models like MobileNetV2\n5. `circular_crop`\n6. Probably more I cannot remember right now LOL\n\n\nMy teammate tested some of these on CIFAR10 and obtained interesting results that did not transfer to the APTOS dataset, which was quite unfortunate.\n\nOur final model (all classification) is a simple average ensemble of:\n\n2 x EfficientNetB3 trained purely on 2019 train data (with polar unrolling) for classification.\nImage size: 300x768. Public Score: 0.788, 0.782\nhttps://www.kaggle.com/agscin/polar-unrolling-preprocessing-84th-place\n\n1 x EfficientNetB3 pretrained on 2015 data and fine-tuned on 2019 (with polar unrolling) for classification. Public Score: 0.794\nImage size: 300x768\n\n5-fold CV ResNet50 trained for classification.\nImage size: 256x256. Public Score: 0.793\n\nScore: Public LB - 0.822, Private LB - 0.923\n\nI want to thank my teammate, Artur.  I had a great time working with him. I would also like to thank @ratthachat for his `crop_image_from_gray` function. And even though I wasn't able to use the EfficientNets, I would like to thanks @drhabib for his contributions to the competition. I would finally like to thank the fast.ai team for developing an amazing library and course that helped me jumpstart my deep learning experience. I also would like to thank the [fastai community](forums.fast.ai) for the support and guidance.\n\n## Thoughts\n\n - If you are hitting a dead-end, ask for help on the forums, and people will be glad to help you out. Also, teaming up and ensembling can also really help you if you are struggling. After all, two minds are better than one!\n - In this competition, we had no idea what the private LB would be like. But eventually, it turned out to be like the training data (I was so [wrong](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763#latest-609621) and I am not sure why). But it meant that the age-old adage \"Trust your CV\" once again was proven true!\n - Even though some of these new optimizers and activation functions did not turn out to be successful in this competition, some of them did well on CIFAR10 and are worth trying out in other competitions.\n - Being honest, the ResNet50 models I developed probably have low generalization ability. Interestingly, the validation loss was slowly increasing while QWK also increased. It is possible that ResNet50 may have some problem generalizing. But I believe some of this may have to do with the optimizer and the hyperparameters used. A few hours before the deadline, I tried training a ResNet50 with SGD and I saw almost no difference between training and validation loss. Unfortunately, there was not enough time to try to submit or add to our ensemble so I had no idea how it did on the private LB. According, to [this article](https://www.fast.ai/2018/07/02/adam-weight-decay/) Adam's generalization ability, is bad due to poorly chosen hyperparameters but I am unsure. This is something to look into a little further, but it may be important to use SGD optimizer as a baseline.\n\n**In conclusion, I learned a lot from this competition. I am happy to get my first medal, and I am excited to apply what I learned here to other competitions!**",
    "622710": "Thanks for writing and sharing! Your 2015 dataset and starter kernels are also invaluable to all of us! \n\nI have no doubt that you will make many more contributions to our community in the future, and I am looking forward to that!",
    "622662": "Congratulations! I really loved your personal journey it’s remind me a bit of my experience:) You are doing great keep learning and practicing:) Congratulations on your first medal!! ",
    "622724": "Congrats and thanks for sharing your learning journey with us! It is inspiring!\n",
    "622770": ""
  }
}