{
  "id": 107990,
  "title": "9th place solution",
  "url": "/competitions/aptos2019-blindness-detection/discussion/107990",
  "author_name": "Psi",
  "post_date": "2019-09-08T09:42:18.593000",
  "votes": 53,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Thanks to Kaggle for hosting this competition, thanks to my teammates and specifically to <a href=\"/dott1718\">@dott1718</a>  (GM yay! - not to jinx it). I am really, really happy about this gold medal as it was basically the first time for us seriously tackling some computer vision problem meaning we also learned quite a few things throughout this competition. It also closes the circle of our journey over this year where we have done really well in all different kinds of competitions including NLP, tabular data, time series/temporal data, and now image data. We would have never expected to do this well when we started to play around on Kaggle less than a year ago, and I am really proud and humble of our progress. If nothing strange happens, this also means that @dott has reached his GM title. Next, I want to lay out our solution. I like to write a lot, but most likely our solution is very similar to other solutions with shorter explanations 😊.</p>\n\n<p><strong>Preprocessing</strong></p>\n\n<p>Overall, I think there are a few puzzle pieces one must need to put together, to produce a good solution. One is to have robust preprocessing that allows to generalize well and that does not make the models overfit on certain image characteristics, like size of the image or other facets like the markers depicting the device of the recording. This is not that important if you would just look at 2019 training data, but rather when you look at test data, where for example a significant amount of images have been cropped to 640x480. To that end, we have the following rough pre-processing routine which in the end is quite similar to the awesome steps of Ben:</p>\n\n<ol>\n<li>Crop the images to the retina by removing black area</li>\n<li>Scale the image by the radius of the retina keeping the aspect ratio intact. We slightly modified this process for the 640x480 images in test data as we also have each side of the retina cut there, so the actual radius is a bit larger.</li>\n<li>Pad the image to 640x640 meaning we still have the original aspect ratio in place. </li>\n<li>Remove blur</li>\n<li>Circle crop, in training we use 90% crop, in inference 95% crop</li>\n<li>Resize to 320x320</li>\n</ol>\n\n<p><strong>Models</strong></p>\n\n<p>I think the type of model is not that important in this competition. But different models might need slightly different learning rates, epochs, etc. but you should be able to get to similar results with a wide variety of models. In the end, we use two models: Efficientnet-B7 and SE-Resnext 50. We switched to SE-Resnext 50 only later on to have faster experimental possibilities, with quite similar results to the Efnet B7. We did a lot of different variations of those models, will elaborate on that later. </p>\n\n<p><strong>Pretraining</strong></p>\n\n<p>We did pretraining on full 2015 data which was very important to get better results on test data. Unfortunately, we did not figure really out what makes a good prefit, so usually we tried to pick the epochs that have the best val score on 2019. Sometimes you can get “lucky” prefits which score quite better on LB compared to others. We fitted for around 15-30 or so epochs depending on the model, we either did some slight LR scheduling, or just kept it constant and either used Radam or some other Adam variation there.</p>\n\n<p><strong>Finetuning</strong></p>\n\n<p>We did 4-fold CV on 2019 data, usually fitting around 3 bags to embrace the randomness a bit. Here, we used Adam optimizer and did a once cycle policy for overall 4 fixed epochs with learning rate ranging from 0.00001 to 0.0005. We did not do any epoch selection here, and kept the 4 epochs fixed. I am a big fan of fixed epochs as it allows you to better judge CV score, and early stopping can lead to some form of overfitting. But, you might need more models here to balance out “bad” fits you might get.</p>\n\n<p><strong>Inference</strong></p>\n\n<p>Our final best sub has TTA4 (as-is, horizontal flip, vertical flip, horizontal+vertical flip). We analyzed that results are really random. This is mostly due to the fact, that a lot of labels are inconsistent. This became really obvious to us, when we tried to do fixed setting of some test labels with the duplicate image labels from train. Just doing this for around 30 images dropped our score by nearly 0.01 on LB which is quite insane. So what makes a good model apparently, is one that has a diverse set of predictions and then finds a good “average” representation of those. You can also get quite lucky if you by chance take the one or the other label for an image. We did not want to rely on luck, so we decided to try to do a blend that is as robust as possible.</p>\n\n<p><strong>Blending</strong></p>\n\n<p>Our final blend is a combination of 12 different types of models, where each one has been fitted for 4 folds and for each one we do TTA4. So overall, our final solution is an average blend of 192 inference procedures. We also slightly changed the final thresholds based on CV simulations to [0.5, 1.5, 2.43, 3.32] expecting private distribution to be similar to train data. This solution has public LB 0.833 and private LB 0.931. \nWe tried to make our final models as diverse as possible, meaning having different prefits and different types of finetunes. All models have the similar augmentations, doing rotation, horizontal/vertical flip and one of CLAHE or brightness-contrast change. I don’t think image augmentations are that helpful apart from flipping.  I will not go into too much detail here but basically we have a mix of models doing the following things:\n- normal models not doing anything specific on top of the abovementioned\n- at least half of the models have local normalization, meaning that for each image we subtract mean and divide by std of the image itself, then we do not do global normalization\n- most models use MSE, we have a few that have a dualloss with MSE and BCE\n- mostly we just use average pooling, some models have a concat of max/avg pooling and one model has LPP pooling\n- one or two models have some random circle cropping (80-99% cropping)</p>\n\n<p><strong>Further things</strong></p>\n\n<p>We played a lot with finding a good way of blending stuff. In the end simple averaging is what apparently worked best. Our second sub uses extremely randomized trees with different seeds on top of all the predictions. This allowed to introduce a bit more randomness, and worked well on both CV as well as public LB, but a tiny bit worse on private LB. </p>\n\n<p>We did not use any pseudo tagging and apparently that is what is missing to get to top of leader board for us. Of course we also tried it. What we did on the one hand is to test in on CV by pseudo labeling the out-of-folds. This either did not improve much or only a tiny bit. We also tried it with pseudo labeling full test data and refitting the model (public+private). This did not change public LB, but we could not see CV score, so we decided to drop it. We did not try to pseudo label only public LB, and apparently if I am reading solutions correctly, this is what most do. I need to test a bit why this worked, but maybe it helped diversity as public LB is a bit different.</p>\n\n<p>There are many, many more things we tried. A lot with group normalization. A lot with different image sizes, and many other stuff.</p>",
  "messages": [
    {
      "id": 621188,
      "postDate": "2019-09-08T09:42:18.593Z",
      "content": "<p>Thanks to Kaggle for hosting this competition, thanks to my teammates and specifically to <a href=\"/dott1718\">@dott1718</a>  (GM yay! - not to jinx it). I am really, really happy about this gold medal as it was basically the first time for us seriously tackling some computer vision problem meaning we also learned quite a few things throughout this competition. It also closes the circle of our journey over this year where we have done really well in all different kinds of competitions including NLP, tabular data, time series/temporal data, and now image data. We would have never expected to do this well when we started to play around on Kaggle less than a year ago, and I am really proud and humble of our progress. If nothing strange happens, this also means that @dott has reached his GM title. Next, I want to lay out our solution. I like to write a lot, but most likely our solution is very similar to other solutions with shorter explanations 😊.</p>\n\n<p><strong>Preprocessing</strong></p>\n\n<p>Overall, I think there are a few puzzle pieces one must need to put together, to produce a good solution. One is to have robust preprocessing that allows to generalize well and that does not make the models overfit on certain image characteristics, like size of the image or other facets like the markers depicting the device of the recording. This is not that important if you would just look at 2019 training data, but rather when you look at test data, where for example a significant amount of images have been cropped to 640x480. To that end, we have the following rough pre-processing routine which in the end is quite similar to the awesome steps of Ben:</p>\n\n<ol>\n<li>Crop the images to the retina by removing black area</li>\n<li>Scale the image by the radius of the retina keeping the aspect ratio intact. We slightly modified this process for the 640x480 images in test data as we also have each side of the retina cut there, so the actual radius is a bit larger.</li>\n<li>Pad the image to 640x640 meaning we still have the original aspect ratio in place. </li>\n<li>Remove blur</li>\n<li>Circle crop, in training we use 90% crop, in inference 95% crop</li>\n<li>Resize to 320x320</li>\n</ol>\n\n<p><strong>Models</strong></p>\n\n<p>I think the type of model is not that important in this competition. But different models might need slightly different learning rates, epochs, etc. but you should be able to get to similar results with a wide variety of models. In the end, we use two models: Efficientnet-B7 and SE-Resnext 50. We switched to SE-Resnext 50 only later on to have faster experimental possibilities, with quite similar results to the Efnet B7. We did a lot of different variations of those models, will elaborate on that later. </p>\n\n<p><strong>Pretraining</strong></p>\n\n<p>We did pretraining on full 2015 data which was very important to get better results on test data. Unfortunately, we did not figure really out what makes a good prefit, so usually we tried to pick the epochs that have the best val score on 2019. Sometimes you can get “lucky” prefits which score quite better on LB compared to others. We fitted for around 15-30 or so epochs depending on the model, we either did some slight LR scheduling, or just kept it constant and either used Radam or some other Adam variation there.</p>\n\n<p><strong>Finetuning</strong></p>\n\n<p>We did 4-fold CV on 2019 data, usually fitting around 3 bags to embrace the randomness a bit. Here, we used Adam optimizer and did a once cycle policy for overall 4 fixed epochs with learning rate ranging from 0.00001 to 0.0005. We did not do any epoch selection here, and kept the 4 epochs fixed. I am a big fan of fixed epochs as it allows you to better judge CV score, and early stopping can lead to some form of overfitting. But, you might need more models here to balance out “bad” fits you might get.</p>\n\n<p><strong>Inference</strong></p>\n\n<p>Our final best sub has TTA4 (as-is, horizontal flip, vertical flip, horizontal+vertical flip). We analyzed that results are really random. This is mostly due to the fact, that a lot of labels are inconsistent. This became really obvious to us, when we tried to do fixed setting of some test labels with the duplicate image labels from train. Just doing this for around 30 images dropped our score by nearly 0.01 on LB which is quite insane. So what makes a good model apparently, is one that has a diverse set of predictions and then finds a good “average” representation of those. You can also get quite lucky if you by chance take the one or the other label for an image. We did not want to rely on luck, so we decided to try to do a blend that is as robust as possible.</p>\n\n<p><strong>Blending</strong></p>\n\n<p>Our final blend is a combination of 12 different types of models, where each one has been fitted for 4 folds and for each one we do TTA4. So overall, our final solution is an average blend of 192 inference procedures. We also slightly changed the final thresholds based on CV simulations to [0.5, 1.5, 2.43, 3.32] expecting private distribution to be similar to train data. This solution has public LB 0.833 and private LB 0.931. \nWe tried to make our final models as diverse as possible, meaning having different prefits and different types of finetunes. All models have the similar augmentations, doing rotation, horizontal/vertical flip and one of CLAHE or brightness-contrast change. I don’t think image augmentations are that helpful apart from flipping.  I will not go into too much detail here but basically we have a mix of models doing the following things:\n- normal models not doing anything specific on top of the abovementioned\n- at least half of the models have local normalization, meaning that for each image we subtract mean and divide by std of the image itself, then we do not do global normalization\n- most models use MSE, we have a few that have a dualloss with MSE and BCE\n- mostly we just use average pooling, some models have a concat of max/avg pooling and one model has LPP pooling\n- one or two models have some random circle cropping (80-99% cropping)</p>\n\n<p><strong>Further things</strong></p>\n\n<p>We played a lot with finding a good way of blending stuff. In the end simple averaging is what apparently worked best. Our second sub uses extremely randomized trees with different seeds on top of all the predictions. This allowed to introduce a bit more randomness, and worked well on both CV as well as public LB, but a tiny bit worse on private LB. </p>\n\n<p>We did not use any pseudo tagging and apparently that is what is missing to get to top of leader board for us. Of course we also tried it. What we did on the one hand is to test in on CV by pseudo labeling the out-of-folds. This either did not improve much or only a tiny bit. We also tried it with pseudo labeling full test data and refitting the model (public+private). This did not change public LB, but we could not see CV score, so we decided to drop it. We did not try to pseudo label only public LB, and apparently if I am reading solutions correctly, this is what most do. I need to test a bit why this worked, but maybe it helped diversity as public LB is a bit different.</p>\n\n<p>There are many, many more things we tried. A lot with group normalization. A lot with different image sizes, and many other stuff.</p>",
      "rawMarkdown": "Thanks to Kaggle for hosting this competition, thanks to my teammates and specifically to @dott1718  (GM yay! - not to jinx it). I am really, really happy about this gold medal as it was basically the first time for us seriously tackling some computer vision problem meaning we also learned quite a few things throughout this competition. It also closes the circle of our journey over this year where we have done really well in all different kinds of competitions including NLP, tabular data, time series/temporal data, and now image data. We would have never expected to do this well when we started to play around on Kaggle less than a year ago, and I am really proud and humble of our progress. If nothing strange happens, this also means that @dott has reached his GM title. Next, I want to lay out our solution. I like to write a lot, but most likely our solution is very similar to other solutions with shorter explanations 😊.\n\n**Preprocessing**\n\nOverall, I think there are a few puzzle pieces one must need to put together, to produce a good solution. One is to have robust preprocessing that allows to generalize well and that does not make the models overfit on certain image characteristics, like size of the image or other facets like the markers depicting the device of the recording. This is not that important if you would just look at 2019 training data, but rather when you look at test data, where for example a significant amount of images have been cropped to 640x480. To that end, we have the following rough pre-processing routine which in the end is quite similar to the awesome steps of Ben:\n\n1. Crop the images to the retina by removing black area\n2. Scale the image by the radius of the retina keeping the aspect ratio intact. We slightly modified this process for the 640x480 images in test data as we also have each side of the retina cut there, so the actual radius is a bit larger.\n3. Pad the image to 640x640 meaning we still have the original aspect ratio in place. \n4. Remove blur\n5. Circle crop, in training we use 90% crop, in inference 95% crop\n6. Resize to 320x320\n\n**Models**\n\nI think the type of model is not that important in this competition. But different models might need slightly different learning rates, epochs, etc. but you should be able to get to similar results with a wide variety of models. In the end, we use two models: Efficientnet-B7 and SE-Resnext 50. We switched to SE-Resnext 50 only later on to have faster experimental possibilities, with quite similar results to the Efnet B7. We did a lot of different variations of those models, will elaborate on that later. \n\n**Pretraining**\n\nWe did pretraining on full 2015 data which was very important to get better results on test data. Unfortunately, we did not figure really out what makes a good prefit, so usually we tried to pick the epochs that have the best val score on 2019. Sometimes you can get “lucky” prefits which score quite better on LB compared to others. We fitted for around 15-30 or so epochs depending on the model, we either did some slight LR scheduling, or just kept it constant and either used Radam or some other Adam variation there.\n\n**Finetuning**\n\nWe did 4-fold CV on 2019 data, usually fitting around 3 bags to embrace the randomness a bit. Here, we used Adam optimizer and did a once cycle policy for overall 4 fixed epochs with learning rate ranging from 0.00001 to 0.0005. We did not do any epoch selection here, and kept the 4 epochs fixed. I am a big fan of fixed epochs as it allows you to better judge CV score, and early stopping can lead to some form of overfitting. But, you might need more models here to balance out “bad” fits you might get.\n\n**Inference**\n\nOur final best sub has TTA4 (as-is, horizontal flip, vertical flip, horizontal+vertical flip). We analyzed that results are really random. This is mostly due to the fact, that a lot of labels are inconsistent. This became really obvious to us, when we tried to do fixed setting of some test labels with the duplicate image labels from train. Just doing this for around 30 images dropped our score by nearly 0.01 on LB which is quite insane. So what makes a good model apparently, is one that has a diverse set of predictions and then finds a good “average” representation of those. You can also get quite lucky if you by chance take the one or the other label for an image. We did not want to rely on luck, so we decided to try to do a blend that is as robust as possible.\n\n**Blending**\n\nOur final blend is a combination of 12 different types of models, where each one has been fitted for 4 folds and for each one we do TTA4. So overall, our final solution is an average blend of 192 inference procedures. We also slightly changed the final thresholds based on CV simulations to [0.5, 1.5, 2.43, 3.32] expecting private distribution to be similar to train data. This solution has public LB 0.833 and private LB 0.931. \nWe tried to make our final models as diverse as possible, meaning having different prefits and different types of finetunes. All models have the similar augmentations, doing rotation, horizontal/vertical flip and one of CLAHE or brightness-contrast change. I don’t think image augmentations are that helpful apart from flipping.  I will not go into too much detail here but basically we have a mix of models doing the following things:\n- normal models not doing anything specific on top of the abovementioned\n- at least half of the models have local normalization, meaning that for each image we subtract mean and divide by std of the image itself, then we do not do global normalization\n- most models use MSE, we have a few that have a dualloss with MSE and BCE\n- mostly we just use average pooling, some models have a concat of max/avg pooling and one model has LPP pooling\n- one or two models have some random circle cropping (80-99% cropping)\n\n**Further things**\n\nWe played a lot with finding a good way of blending stuff. In the end simple averaging is what apparently worked best. Our second sub uses extremely randomized trees with different seeds on top of all the predictions. This allowed to introduce a bit more randomness, and worked well on both CV as well as public LB, but a tiny bit worse on private LB. \n\nWe did not use any pseudo tagging and apparently that is what is missing to get to top of leader board for us. Of course we also tried it. What we did on the one hand is to test in on CV by pseudo labeling the out-of-folds. This either did not improve much or only a tiny bit. We also tried it with pseudo labeling full test data and refitting the model (public+private). This did not change public LB, but we could not see CV score, so we decided to drop it. We did not try to pseudo label only public LB, and apparently if I am reading solutions correctly, this is what most do. I need to test a bit why this worked, but maybe it helped diversity as public LB is a bit different.\n\nThere are many, many more things we tried. A lot with group normalization. A lot with different image sizes, and many other stuff.\n\n",
      "votes": 53
    },
    {
      "id": 624299,
      "postDate": "2019-09-12T00:08:59.657Z",
      "content": "<p>Congrats. You still have at least 1 competition category to try: non-ML (like annually Christmas optimisation santa competition)</p>",
      "rawMarkdown": "Congrats. You still have at least 1 competition category to try: non-ML (like annually Christmas optimisation santa competition)",
      "votes": 1,
      "replies": [
        {
          "id": 624599,
          "postDate": "2019-09-12T07:48:07.153Z",
          "content": "<p>Will keep it in mind ;)</p>",
          "rawMarkdown": "Will keep it in mind ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 621524,
      "postDate": "2019-09-08T16:10:51.243Z",
      "content": "<p>Congratulations to the team , <a href=\"/philippsinger\">@philippsinger</a>  <a href=\"/dott1718\">@dott1718</a>  !👍 </p>",
      "rawMarkdown": "Congratulations to the team , @philippsinger  @dott1718  !👍 ",
      "votes": 1
    },
    {
      "id": 621514,
      "postDate": "2019-09-08T15:50:25.117Z",
      "content": "<p>Congrats <a href=\"/philippsinger\">@philippsinger</a> and thanks for sharing</p>",
      "rawMarkdown": "Congrats @philippsinger and thanks for sharing",
      "votes": 1
    },
    {
      "id": 621442,
      "postDate": "2019-09-08T14:16:19.893Z",
      "content": "<p>Congratulation and thank you for sharing!</p>",
      "rawMarkdown": "Congratulation and thank you for sharing!",
      "votes": 1
    },
    {
      "id": 621404,
      "postDate": "2019-09-08T13:38:17.553Z",
      "content": "<p>Congrats <a href=\"/philippsinger\">@philippsinger</a>, <a href=\"/dott1718\">@dott1718</a>, <a href=\"/mmotoki\">@mmotoki</a>  et. al. and thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congrats @philippsinger, @dott1718, @mmotoki  et. al. and thanks for sharing your solution overview.",
      "votes": 1
    },
    {
      "id": 621343,
      "postDate": "2019-09-08T12:58:23.403Z",
      "content": "<p>Philipp <a href=\"/philippsinger\">@philippsinger</a>  Not only you are really excellent data scientist, your writing is one of the best in my opinion. In fact, what I learned from your Quora post highly encourage me to totally rethink about machine learning philosophy.</p>\n\n<p>Congrat to <a href=\"/dott1718\">@dott1718</a> to be GM too!</p>",
      "rawMarkdown": "Philipp @philippsinger  Not only you are really excellent data scientist, your writing is one of the best in my opinion. In fact, what I learned from your Quora post highly encourage me to totally rethink about machine learning philosophy.\n\nCongrat to @dott1718 to be GM too!",
      "votes": 1,
      "replies": [
        {
          "id": 621522,
          "postDate": "2019-09-08T16:07:35.863Z",
          "content": "<p>Thanks a lot for the kind words. You did really well on this competition and deserved a better final place, your kernels helped so many people.</p>",
          "rawMarkdown": "Thanks a lot for the kind words. You did really well on this competition and deserved a better final place, your kernels helped so many people.",
          "votes": 1
        }
      ]
    },
    {
      "id": 621285,
      "postDate": "2019-09-08T12:02:06.117Z",
      "content": "<p>Very nice !! congrats for the gold and gontrats to <a href=\"/dott1718\">@dott1718</a>  for the GM, well deserved :)</p>",
      "rawMarkdown": "Very nice !! congrats for the gold and gontrats to @dott1718  for the GM, well deserved :)",
      "votes": 1
    },
    {
      "id": 621895,
      "postDate": "2019-09-09T04:40:20.317Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 624299,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2019-09-12T00:08:59.657000",
      "content": "<p>Congrats. You still have at least 1 competition category to try: non-ML (like annually Christmas optimisation santa competition)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 624599,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-09-12T07:48:07.153000",
          "content": "<p>Will keep it in mind ;)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 621524,
      "author_name": "Noah Weber",
      "author_url": "",
      "post_date": "2019-09-08T16:10:51.243000",
      "content": "<p>Congratulations to the team , <a href=\"/philippsinger\">@philippsinger</a>  <a href=\"/dott1718\">@dott1718</a>  !👍 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621514,
      "author_name": "Carlos Prades K.",
      "author_url": "",
      "post_date": "2019-09-08T15:50:25.117000",
      "content": "<p>Congrats <a href=\"/philippsinger\">@philippsinger</a> and thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621442,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2019-09-08T14:16:19.893000",
      "content": "<p>Congratulation and thank you for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621404,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-09-08T13:38:17.553000",
      "content": "<p>Congrats <a href=\"/philippsinger\">@philippsinger</a>, <a href=\"/dott1718\">@dott1718</a>, <a href=\"/mmotoki\">@mmotoki</a>  et. al. and thanks for sharing your solution overview.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621343,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-09-08T12:58:23.403000",
      "content": "<p>Philipp <a href=\"/philippsinger\">@philippsinger</a>  Not only you are really excellent data scientist, your writing is one of the best in my opinion. In fact, what I learned from your Quora post highly encourage me to totally rethink about machine learning philosophy.</p>\n\n<p>Congrat to <a href=\"/dott1718\">@dott1718</a> to be GM too!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 621522,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-09-08T16:07:35.863000",
          "content": "<p>Thanks a lot for the kind words. You did really well on this competition and deserved a better final place, your kernels helped so many people.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 621285,
      "author_name": "Dimosthenis Karaflos",
      "author_url": "",
      "post_date": "2019-09-08T12:02:06.117000",
      "content": "<p>Very nice !! congrats for the gold and gontrats to <a href=\"/dott1718\">@dott1718</a>  for the GM, well deserved :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621895,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T04:40:20.317000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "621188": "Thanks to Kaggle for hosting this competition, thanks to my teammates and specifically to @dott1718  (GM yay! - not to jinx it). I am really, really happy about this gold medal as it was basically the first time for us seriously tackling some computer vision problem meaning we also learned quite a few things throughout this competition. It also closes the circle of our journey over this year where we have done really well in all different kinds of competitions including NLP, tabular data, time series/temporal data, and now image data. We would have never expected to do this well when we started to play around on Kaggle less than a year ago, and I am really proud and humble of our progress. If nothing strange happens, this also means that @dott has reached his GM title. Next, I want to lay out our solution. I like to write a lot, but most likely our solution is very similar to other solutions with shorter explanations 😊.\n\n**Preprocessing**\n\nOverall, I think there are a few puzzle pieces one must need to put together, to produce a good solution. One is to have robust preprocessing that allows to generalize well and that does not make the models overfit on certain image characteristics, like size of the image or other facets like the markers depicting the device of the recording. This is not that important if you would just look at 2019 training data, but rather when you look at test data, where for example a significant amount of images have been cropped to 640x480. To that end, we have the following rough pre-processing routine which in the end is quite similar to the awesome steps of Ben:\n\n1. Crop the images to the retina by removing black area\n2. Scale the image by the radius of the retina keeping the aspect ratio intact. We slightly modified this process for the 640x480 images in test data as we also have each side of the retina cut there, so the actual radius is a bit larger.\n3. Pad the image to 640x640 meaning we still have the original aspect ratio in place. \n4. Remove blur\n5. Circle crop, in training we use 90% crop, in inference 95% crop\n6. Resize to 320x320\n\n**Models**\n\nI think the type of model is not that important in this competition. But different models might need slightly different learning rates, epochs, etc. but you should be able to get to similar results with a wide variety of models. In the end, we use two models: Efficientnet-B7 and SE-Resnext 50. We switched to SE-Resnext 50 only later on to have faster experimental possibilities, with quite similar results to the Efnet B7. We did a lot of different variations of those models, will elaborate on that later. \n\n**Pretraining**\n\nWe did pretraining on full 2015 data which was very important to get better results on test data. Unfortunately, we did not figure really out what makes a good prefit, so usually we tried to pick the epochs that have the best val score on 2019. Sometimes you can get “lucky” prefits which score quite better on LB compared to others. We fitted for around 15-30 or so epochs depending on the model, we either did some slight LR scheduling, or just kept it constant and either used Radam or some other Adam variation there.\n\n**Finetuning**\n\nWe did 4-fold CV on 2019 data, usually fitting around 3 bags to embrace the randomness a bit. Here, we used Adam optimizer and did a once cycle policy for overall 4 fixed epochs with learning rate ranging from 0.00001 to 0.0005. We did not do any epoch selection here, and kept the 4 epochs fixed. I am a big fan of fixed epochs as it allows you to better judge CV score, and early stopping can lead to some form of overfitting. But, you might need more models here to balance out “bad” fits you might get.\n\n**Inference**\n\nOur final best sub has TTA4 (as-is, horizontal flip, vertical flip, horizontal+vertical flip). We analyzed that results are really random. This is mostly due to the fact, that a lot of labels are inconsistent. This became really obvious to us, when we tried to do fixed setting of some test labels with the duplicate image labels from train. Just doing this for around 30 images dropped our score by nearly 0.01 on LB which is quite insane. So what makes a good model apparently, is one that has a diverse set of predictions and then finds a good “average” representation of those. You can also get quite lucky if you by chance take the one or the other label for an image. We did not want to rely on luck, so we decided to try to do a blend that is as robust as possible.\n\n**Blending**\n\nOur final blend is a combination of 12 different types of models, where each one has been fitted for 4 folds and for each one we do TTA4. So overall, our final solution is an average blend of 192 inference procedures. We also slightly changed the final thresholds based on CV simulations to [0.5, 1.5, 2.43, 3.32] expecting private distribution to be similar to train data. This solution has public LB 0.833 and private LB 0.931. \nWe tried to make our final models as diverse as possible, meaning having different prefits and different types of finetunes. All models have the similar augmentations, doing rotation, horizontal/vertical flip and one of CLAHE or brightness-contrast change. I don’t think image augmentations are that helpful apart from flipping.  I will not go into too much detail here but basically we have a mix of models doing the following things:\n- normal models not doing anything specific on top of the abovementioned\n- at least half of the models have local normalization, meaning that for each image we subtract mean and divide by std of the image itself, then we do not do global normalization\n- most models use MSE, we have a few that have a dualloss with MSE and BCE\n- mostly we just use average pooling, some models have a concat of max/avg pooling and one model has LPP pooling\n- one or two models have some random circle cropping (80-99% cropping)\n\n**Further things**\n\nWe played a lot with finding a good way of blending stuff. In the end simple averaging is what apparently worked best. Our second sub uses extremely randomized trees with different seeds on top of all the predictions. This allowed to introduce a bit more randomness, and worked well on both CV as well as public LB, but a tiny bit worse on private LB. \n\nWe did not use any pseudo tagging and apparently that is what is missing to get to top of leader board for us. Of course we also tried it. What we did on the one hand is to test in on CV by pseudo labeling the out-of-folds. This either did not improve much or only a tiny bit. We also tried it with pseudo labeling full test data and refitting the model (public+private). This did not change public LB, but we could not see CV score, so we decided to drop it. We did not try to pseudo label only public LB, and apparently if I am reading solutions correctly, this is what most do. I need to test a bit why this worked, but maybe it helped diversity as public LB is a bit different.\n\nThere are many, many more things we tried. A lot with group normalization. A lot with different image sizes, and many other stuff.\n\n",
    "624299": "Congrats. You still have at least 1 competition category to try: non-ML (like annually Christmas optimisation santa competition)",
    "621524": "Congratulations to the team , @philippsinger  @dott1718  !👍 ",
    "621514": "Congrats @philippsinger and thanks for sharing",
    "621442": "Congratulation and thank you for sharing!",
    "621404": "Congrats @philippsinger, @dott1718, @mmotoki  et. al. and thanks for sharing your solution overview.",
    "621343": "Philipp @philippsinger  Not only you are really excellent data scientist, your writing is one of the best in my opinion. In fact, what I learned from your Quora post highly encourage me to totally rethink about machine learning philosophy.\n\nCongrat to @dott1718 to be GM too!",
    "621285": "Very nice !! congrats for the gold and gontrats to @dott1718  for the GM, well deserved :)",
    "621895": ""
  }
}