{
  "id": 108711,
  "title": "#16 Solution",
  "url": "/competitions/aptos2019-blindness-detection/writeups/carlospk-16-solution",
  "author_name": "",
  "post_date": "2019-09-13T12:24:24.537147Z",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all. I have been quite busy since the end of the competition so I couldn’t post my solution earlier, but I still want to share with you what I did. \nI hope it is not too late!</p>\n\n<p>My solution was an ensemble of different architectures (Resnet34, Resnet50, EfficientNet-b4 and EfficientNet-b5), different preprocessing methods (Circle+GaussianBlur and CircleToSquare+GaussianBlur) and different input sizes (224, 380 and 456). All they were blended using a weighted average (I manually tuned the weights, which worked better for me than selecting the weights using hyperopt). I used TTA in some models, which improved the score a bit (it was required setting scale=1.0 for TTA to work well using fastai).</p>\n\n<p>A key part of my solution was the preprocessing. At first I tried substracting both median blurred and gaussian blurred images, but I just kept using the latter because I saw much better correlation between CV and LB. The idea was to remove all information that could help the algorithm to recognize the different kind of images, in order to generalize better. However, substracting gaussian blurred image seemed not enough to remove all the original picture information, since the color in the limit of the eye was still a bit different for the different image types:\n<img src=\"https://i.postimg.cc/dVqfCn1s/Gaussian-Blur.png\" alt=\"Circle + Gaussian Blur\"></p>\n\n<p>So the idea was removing the black part before substracting gaussian blurred images. For that purpose I found the <a href=\"https://pypi.org/project/squircle/\">squircle library</a> which has 3 methods to do it. This is how the images look after preprocessing:\n<img src=\"https://i.postimg.cc/HLKH3Fc4/CTS.png\" alt=\"Circle to Square +  Gaussian Blur\"></p>\n\n<p>My best single model was an EfficientNet-b5 based on this CircleToSquare+GaussianBlur preprocessing, input size 380 and the following augmentations parameters (fastai):  get_transforms(max_rotate=0, flip_vert=True, max_warp=0.0, max_zoom=0, max_lighting=0.35). It scored 0.814 in the public LB and 0.927 in the private LB.</p>\n\n<p>I also used pseudolabeling, which helped me a bit, but not so much because of the way I applied it (using a classification model with a threshold of 0.9). I guess if I would have applied pseudolabeling using all test data the way top teams described it my model would have performed better. At least I learned something new ;)</p>\n\n<p>Well, congrats to all winners and thanks to all people who shared in discussions and who shared great kernels. Special thanks to <a href=\"/ratthachat\">@ratthachat</a> for the many preprocessing ideas and to <a href=\"/hmendonca\">@hmendonca</a> for sharing a kernel showing how to apply EfficientNet using fastai.</p>\n\n<p>All the best!</p>",
  "messages": [
    {
      "id": "625757",
      "postDate": "09/13/2019 12:24:24",
      "content": "<p>Hi all. I have been quite busy since the end of the competition so I couldn’t post my solution earlier, but I still want to share with you what I did. \nI hope it is not too late!</p>\n\n<p>My solution was an ensemble of different architectures (Resnet34, Resnet50, EfficientNet-b4 and EfficientNet-b5), different preprocessing methods (Circle+GaussianBlur and CircleToSquare+GaussianBlur) and different input sizes (224, 380 and 456). All they were blended using a weighted average (I manually tuned the weights, which worked better for me than selecting the weights using hyperopt). I used TTA in some models, which improved the score a bit (it was required setting scale=1.0 for TTA to work well using fastai).</p>\n\n<p>A key part of my solution was the preprocessing. At first I tried substracting both median blurred and gaussian blurred images, but I just kept using the latter because I saw much better correlation between CV and LB. The idea was to remove all information that could help the algorithm to recognize the different kind of images, in order to generalize better. However, substracting gaussian blurred image seemed not enough to remove all the original picture information, since the color in the limit of the eye was still a bit different for the different image types:\n<img src=\"https://i.postimg.cc/dVqfCn1s/Gaussian-Blur.png\" alt=\"Circle + Gaussian Blur\"></p>\n\n<p>So the idea was removing the black part before substracting gaussian blurred images. For that purpose I found the <a href=\"https://pypi.org/project/squircle/\">squircle library</a> which has 3 methods to do it. This is how the images look after preprocessing:\n<img src=\"https://i.postimg.cc/HLKH3Fc4/CTS.png\" alt=\"Circle to Square +  Gaussian Blur\"></p>\n\n<p>My best single model was an EfficientNet-b5 based on this CircleToSquare+GaussianBlur preprocessing, input size 380 and the following augmentations parameters (fastai):  get_transforms(max_rotate=0, flip_vert=True, max_warp=0.0, max_zoom=0, max_lighting=0.35). It scored 0.814 in the public LB and 0.927 in the private LB.</p>\n\n<p>I also used pseudolabeling, which helped me a bit, but not so much because of the way I applied it (using a classification model with a threshold of 0.9). I guess if I would have applied pseudolabeling using all test data the way top teams described it my model would have performed better. At least I learned something new ;)</p>\n\n<p>Well, congrats to all winners and thanks to all people who shared in discussions and who shared great kernels. Special thanks to <a href=\"/ratthachat\">@ratthachat</a> for the many preprocessing ideas and to <a href=\"/hmendonca\">@hmendonca</a> for sharing a kernel showing how to apply EfficientNet using fastai.</p>\n\n<p>All the best!</p>",
      "rawMarkdown": "Hi all. I have been quite busy since the end of the competition so I couldn’t post my solution earlier, but I still want to share with you what I did. \nI hope it is not too late!\n\nMy solution was an ensemble of different architectures (Resnet34, Resnet50, EfficientNet-b4 and EfficientNet-b5), different preprocessing methods (Circle+GaussianBlur and CircleToSquare+GaussianBlur) and different input sizes (224, 380 and 456). All they were blended using a weighted average (I manually tuned the weights, which worked better for me than selecting the weights using hyperopt). I used TTA in some models, which improved the score a bit (it was required setting scale=1.0 for TTA to work well using fastai).\n\nA key part of my solution was the preprocessing. At first I tried substracting both median blurred and gaussian blurred images, but I just kept using the latter because I saw much better correlation between CV and LB. The idea was to remove all information that could help the algorithm to recognize the different kind of images, in order to generalize better. However, substracting gaussian blurred image seemed not enough to remove all the original picture information, since the color in the limit of the eye was still a bit different for the different image types:\n![Circle + Gaussian Blur]( https://i.postimg.cc/dVqfCn1s/Gaussian-Blur.png)\n\nSo the idea was removing the black part before substracting gaussian blurred images. For that purpose I found the [squircle library](https://pypi.org/project/squircle/) which has 3 methods to do it. This is how the images look after preprocessing:\n![Circle to Square +  Gaussian Blur]( https://i.postimg.cc/HLKH3Fc4/CTS.png)\n\nMy best single model was an EfficientNet-b5 based on this CircleToSquare+GaussianBlur preprocessing, input size 380 and the following augmentations parameters (fastai):  get_transforms(max_rotate=0, flip_vert=True, max_warp=0.0, max_zoom=0, max_lighting=0.35). It scored 0.814 in the public LB and 0.927 in the private LB.\n\nI also used pseudolabeling, which helped me a bit, but not so much because of the way I applied it (using a classification model with a threshold of 0.9). I guess if I would have applied pseudolabeling using all test data the way top teams described it my model would have performed better. At least I learned something new ;)\n\nWell, congrats to all winners and thanks to all people who shared in discussions and who shared great kernels. Special thanks to @ratthachat for the many preprocessing ideas and to @hmendonca for sharing a kernel showing how to apply EfficientNet using fastai.\n\nAll the best!",
      "votes": null
    },
    {
      "id": "625762",
      "postDate": "09/13/2019 12:31:54",
      "content": "<p>Congrat Carlos, and wish next time you will get +1 place (or rather there are more competitors so +1 gold) ...</p>\n\n<p>BTW, your best model is comparable to mine, when not using pseudo-labelling. Maybe because we use similar preprocessing and the same B5. See you next round!</p>",
      "rawMarkdown": "Congrat Carlos, and wish next time you will get +1 place (or rather there are more competitors so +1 gold) ...\n\nBTW, your best model is comparable to mine, when not using pseudo-labelling. Maybe because we use similar preprocessing and the same B5. See you next round!",
      "votes": null
    },
    {
      "id": "625896",
      "postDate": "09/13/2019 15:11:14",
      "content": "<p>Congratulations Carlos and thanks for sharing.\nMay I ask in which environment you develop your kernel ? In Kaggle directly or on your own machine and then transfer it to Kaggle ?</p>\n\n<p>Xavier.</p>",
      "rawMarkdown": "Congratulations Carlos and thanks for sharing.\nMay I ask in which environment you develop your kernel ? In Kaggle directly or on your own machine and then transfer it to Kaggle ?\n\nXavier.",
      "votes": null
    },
    {
      "id": "625978",
      "postDate": "09/13/2019 16:44:55",
      "content": "<p>I preprocessed the training images in my computer, which speed up the training process a lot, but I have no GPU so I run my experiments in Kaggle ;)</p>",
      "rawMarkdown": "I preprocessed the training images in my computer, which speed up the training process a lot, but I have no GPU so I run my experiments in Kaggle ;)",
      "votes": null
    },
    {
      "id": "626248",
      "postDate": "09/14/2019 04:29:23",
      "content": "<p>Great Work...\nCongratulations\nThank you for Sharing your Approach &amp; Insights....!! <a href=\"/carlospk\">@carlospk</a> </p>",
      "rawMarkdown": "Great Work...\nCongratulations\nThank you for Sharing your Approach &amp; Insights....!! @carlospk",
      "votes": null
    },
    {
      "id": "628484",
      "postDate": "09/17/2019 12:00:21",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 625762,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "09/13/2019 12:31:54",
      "content": "<p>Congrat Carlos, and wish next time you will get +1 place (or rather there are more competitors so +1 gold) ...</p>\n\n<p>BTW, your best model is comparable to mine, when not using pseudo-labelling. Maybe because we use similar preprocessing and the same B5. See you next round!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 625896,
      "author_name": "xavierp",
      "author_url": "",
      "post_date": "09/13/2019 15:11:14",
      "content": "<p>Congratulations Carlos and thanks for sharing.\nMay I ask in which environment you develop your kernel ? In Kaggle directly or on your own machine and then transfer it to Kaggle ?</p>\n\n<p>Xavier.</p>",
      "votes": null,
      "replies": [
        {
          "id": 625978,
          "author_name": "carlospk",
          "author_url": "",
          "post_date": "09/13/2019 16:44:55",
          "content": "<p>I preprocessed the training images in my computer, which speed up the training process a lot, but I have no GPU so I run my experiments in Kaggle ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 628484,
          "author_name": "xavierp",
          "author_url": "",
          "post_date": "09/17/2019 12:00:21",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 626248,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "09/14/2019 04:29:23",
      "content": "<p>Great Work...\nCongratulations\nThank you for Sharing your Approach &amp; Insights....!! <a href=\"/carlospk\">@carlospk</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "625757": "Hi all. I have been quite busy since the end of the competition so I couldn’t post my solution earlier, but I still want to share with you what I did. \nI hope it is not too late!\n\nMy solution was an ensemble of different architectures (Resnet34, Resnet50, EfficientNet-b4 and EfficientNet-b5), different preprocessing methods (Circle+GaussianBlur and CircleToSquare+GaussianBlur) and different input sizes (224, 380 and 456). All they were blended using a weighted average (I manually tuned the weights, which worked better for me than selecting the weights using hyperopt). I used TTA in some models, which improved the score a bit (it was required setting scale=1.0 for TTA to work well using fastai).\n\nA key part of my solution was the preprocessing. At first I tried substracting both median blurred and gaussian blurred images, but I just kept using the latter because I saw much better correlation between CV and LB. The idea was to remove all information that could help the algorithm to recognize the different kind of images, in order to generalize better. However, substracting gaussian blurred image seemed not enough to remove all the original picture information, since the color in the limit of the eye was still a bit different for the different image types:\n![Circle + Gaussian Blur]( https://i.postimg.cc/dVqfCn1s/Gaussian-Blur.png)\n\nSo the idea was removing the black part before substracting gaussian blurred images. For that purpose I found the [squircle library](https://pypi.org/project/squircle/) which has 3 methods to do it. This is how the images look after preprocessing:\n![Circle to Square +  Gaussian Blur]( https://i.postimg.cc/HLKH3Fc4/CTS.png)\n\nMy best single model was an EfficientNet-b5 based on this CircleToSquare+GaussianBlur preprocessing, input size 380 and the following augmentations parameters (fastai):  get_transforms(max_rotate=0, flip_vert=True, max_warp=0.0, max_zoom=0, max_lighting=0.35). It scored 0.814 in the public LB and 0.927 in the private LB.\n\nI also used pseudolabeling, which helped me a bit, but not so much because of the way I applied it (using a classification model with a threshold of 0.9). I guess if I would have applied pseudolabeling using all test data the way top teams described it my model would have performed better. At least I learned something new ;)\n\nWell, congrats to all winners and thanks to all people who shared in discussions and who shared great kernels. Special thanks to @ratthachat for the many preprocessing ideas and to @hmendonca for sharing a kernel showing how to apply EfficientNet using fastai.\n\nAll the best!",
    "625762": "Congrat Carlos, and wish next time you will get +1 place (or rather there are more competitors so +1 gold) ...\n\nBTW, your best model is comparable to mine, when not using pseudo-labelling. Maybe because we use similar preprocessing and the same B5. See you next round!",
    "625896": "Congratulations Carlos and thanks for sharing.\nMay I ask in which environment you develop your kernel ? In Kaggle directly or on your own machine and then transfer it to Kaggle ?\n\nXavier.",
    "625978": "I preprocessed the training images in my computer, which speed up the training process a lot, but I have no GPU so I run my experiments in Kaggle ;)",
    "626248": "Great Work...\nCongratulations\nThank you for Sharing your Approach &amp; Insights....!! @carlospk",
    "628484": "Thanks!"
  },
  "source": "meta"
}