{
  "id": 107987,
  "title": "4th place solution",
  "url": "/competitions/aptos2019-blindness-detection/discussion/107987",
  "author_name": "Qishen Ha",
  "post_date": "2019-09-08T09:08:46.719000",
  "votes": 61,
  "comment_count": 36,
  "views": 0,
  "content": "<h1>My Submissions</h1>\n\n<p>My best public LB:  <code>0.842</code> (with quick submit, so no private LB available)\nMy FINAL submit v1, with public LB <code>0.838</code> , private LB <code>0.932</code> , is based on my 6 models with best local CV.\nMy FINAL submit v2, with public LB <code>0.840</code> , private LB <code>0.933</code> , with more variance in ensemble models. \nMy best single model, with public LB <code>0.840</code> , private LB <code>0.929</code>, which is EfficientNet-B7 with 224x224 input.</p>\n\n<h1>Preprocessing</h1>\n\n<h2>Crop From Gray (Both training and predicting)</h2>\n\n<p>This function is shared in <a href=\"https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping\">here</a> (search keyword <code>crop_image_from_gray</code> in the page, thanks <a href=\"/ratthachat\">@ratthachat</a> for sharing!)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2F0405798d336f82857988977e629f108b%2F_20190908171800.png?generation=1567930710772532&amp;alt=media\" alt=\"\"></p>\n\n<h2>Apply Image Type (Only For Training)</h2>\n\n<p>I apply image type for every image in training set (2015 &amp; 2019) using the brightness of image boundary. \nFrom left to right in the following picture: <code>Type_0</code> , <code>Type_1</code> , <code>Type_2</code> </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2F0fd5e78191a5373707c4faf3c83f1ad6%2F_20190908211528.png?generation=1567944954429547&amp;alt=media\" alt=\"\"></p>\n\n<h2>Transform All Image to Type 2 (Only For Training)</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2F8439f5c3e2a27fd2d548e460eb2da46c%2F_20190908173231.png?generation=1567931590044112&amp;alt=media\" alt=\"\"></p>\n\n<h1>Augmentation</h1>\n\n<p>After crop &amp; resize all images to type_2, I heavily transforming those images by:</p>\n\n<p>Dihedral, RandomCrop, Rotation, Contrast, Brightness, Cutout, PerspectiveTransform, Clahe.</p>\n\n<h1>Model</h1>\n\n<p>I find that large model training on high resolution images is much more likely to be overfitting, due to the small size of training set. So I use small model for high resolution / large model for low resolution.</p>\n\n<ul>\n<li>EfficientNet-B7 (224x224)  LB 0.840</li>\n<li>EfficientNet-B6 (240x240) LB 0.832</li>\n<li>EfficientNet-B5 (256x256) LB 0.831</li>\n<li>EfficientNet-B4 (320x320) LB 0.826</li>\n<li>EfficientNet-B3 (352x352) *LB Don't Know</li>\n<li>EfficientNet-B2 (376x376) LB 0.828</li>\n</ul>\n\n<p><strong>LB is came from 5-fold testing x 8 Dihedral TTA</strong></p>\n\n<p>Despite the difference in Public LB, all of these model have a similar local 5-fold CV around 0.933</p>\n\n<p>↑ My FINAL submit v1 is simple avg these models.</p>\n\n<p>And the FINAL submit v2 is based on v1, replaced B2 &amp; B3 models with a B7 and a B5 that trained on Ben's preprocessing data (image type, along with other tricks are also applied).</p>\n\n<ul>\n<li>EfficientNet-B7 (224x224 Ben's)  LB 0.834</li>\n<li>EfficientNet-B5 (256x256 Ben's) LB 0.833</li>\n</ul>\n\n<p>local 5-fold CV of these 2 models is around 0.927</p>\n\n<h1>Training Process</h1>\n\n<p>Pretrain on 2015 full dataset (both train and test) for 25 epochs without validation. Then do 5-fold CV on 2019 dataset.</p>\n\n<h1>Predicting Process</h1>\n\n<p>Because of the 5-fold CV, I got 5 model files for each experiment.\nFor each experiment do 5 models * 8 Dihedral TTA predicting.\n<strong>I only do image preprocessing once for 5*8=40 test as for saving time.</strong> In this case 40 times test cost ~10mins for predicting 1928 images.\nWith ensemble 6 models, the kernel took 50mins for committing, 6h for submitting.</p>",
  "messages": [
    {
      "id": 621164,
      "postDate": "2019-09-08T09:08:46.720Z",
      "content": "<h1>My Submissions</h1>\n\n<p>My best public LB:  <code>0.842</code> (with quick submit, so no private LB available)\nMy FINAL submit v1, with public LB <code>0.838</code> , private LB <code>0.932</code> , is based on my 6 models with best local CV.\nMy FINAL submit v2, with public LB <code>0.840</code> , private LB <code>0.933</code> , with more variance in ensemble models. \nMy best single model, with public LB <code>0.840</code> , private LB <code>0.929</code>, which is EfficientNet-B7 with 224x224 input.</p>\n\n<h1>Preprocessing</h1>\n\n<h2>Crop From Gray (Both training and predicting)</h2>\n\n<p>This function is shared in <a href=\"https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping\">here</a> (search keyword <code>crop_image_from_gray</code> in the page, thanks <a href=\"/ratthachat\">@ratthachat</a> for sharing!)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2F0405798d336f82857988977e629f108b%2F_20190908171800.png?generation=1567930710772532&amp;alt=media\" alt=\"\"></p>\n\n<h2>Apply Image Type (Only For Training)</h2>\n\n<p>I apply image type for every image in training set (2015 &amp; 2019) using the brightness of image boundary. \nFrom left to right in the following picture: <code>Type_0</code> , <code>Type_1</code> , <code>Type_2</code> </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2F0fd5e78191a5373707c4faf3c83f1ad6%2F_20190908211528.png?generation=1567944954429547&amp;alt=media\" alt=\"\"></p>\n\n<h2>Transform All Image to Type 2 (Only For Training)</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2F8439f5c3e2a27fd2d548e460eb2da46c%2F_20190908173231.png?generation=1567931590044112&amp;alt=media\" alt=\"\"></p>\n\n<h1>Augmentation</h1>\n\n<p>After crop &amp; resize all images to type_2, I heavily transforming those images by:</p>\n\n<p>Dihedral, RandomCrop, Rotation, Contrast, Brightness, Cutout, PerspectiveTransform, Clahe.</p>\n\n<h1>Model</h1>\n\n<p>I find that large model training on high resolution images is much more likely to be overfitting, due to the small size of training set. So I use small model for high resolution / large model for low resolution.</p>\n\n<ul>\n<li>EfficientNet-B7 (224x224)  LB 0.840</li>\n<li>EfficientNet-B6 (240x240) LB 0.832</li>\n<li>EfficientNet-B5 (256x256) LB 0.831</li>\n<li>EfficientNet-B4 (320x320) LB 0.826</li>\n<li>EfficientNet-B3 (352x352) *LB Don't Know</li>\n<li>EfficientNet-B2 (376x376) LB 0.828</li>\n</ul>\n\n<p><strong>LB is came from 5-fold testing x 8 Dihedral TTA</strong></p>\n\n<p>Despite the difference in Public LB, all of these model have a similar local 5-fold CV around 0.933</p>\n\n<p>↑ My FINAL submit v1 is simple avg these models.</p>\n\n<p>And the FINAL submit v2 is based on v1, replaced B2 &amp; B3 models with a B7 and a B5 that trained on Ben's preprocessing data (image type, along with other tricks are also applied).</p>\n\n<ul>\n<li>EfficientNet-B7 (224x224 Ben's)  LB 0.834</li>\n<li>EfficientNet-B5 (256x256 Ben's) LB 0.833</li>\n</ul>\n\n<p>local 5-fold CV of these 2 models is around 0.927</p>\n\n<h1>Training Process</h1>\n\n<p>Pretrain on 2015 full dataset (both train and test) for 25 epochs without validation. Then do 5-fold CV on 2019 dataset.</p>\n\n<h1>Predicting Process</h1>\n\n<p>Because of the 5-fold CV, I got 5 model files for each experiment.\nFor each experiment do 5 models * 8 Dihedral TTA predicting.\n<strong>I only do image preprocessing once for 5*8=40 test as for saving time.</strong> In this case 40 times test cost ~10mins for predicting 1928 images.\nWith ensemble 6 models, the kernel took 50mins for committing, 6h for submitting.</p>",
      "rawMarkdown": "# My Submissions\n\nMy best public LB:  `0.842` (with quick submit, so no private LB available)\nMy FINAL submit v1, with public LB `0.838` , private LB `0.932` , is based on my 6 models with best local CV.\nMy FINAL submit v2, with public LB `0.840` , private LB `0.933` , with more variance in ensemble models. \nMy best single model, with public LB `0.840` , private LB `0.929`, which is EfficientNet-B7 with 224x224 input.\n\n# Preprocessing\n\n## Crop From Gray (Both training and predicting)\n\nThis function is shared in [here](https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping) (search keyword `crop_image_from_gray` in the page, thanks @ratthachat for sharing!)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2F0405798d336f82857988977e629f108b%2F_20190908171800.png?generation=1567930710772532&amp;alt=media)\n\n## Apply Image Type (Only For Training)\n\nI apply image type for every image in training set (2015 &amp; 2019) using the brightness of image boundary. \nFrom left to right in the following picture: `Type_0` , `Type_1` , `Type_2` \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2F0fd5e78191a5373707c4faf3c83f1ad6%2F_20190908211528.png?generation=1567944954429547&amp;alt=media)\n\n\n\n## Transform All Image to Type 2 (Only For Training)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2F8439f5c3e2a27fd2d548e460eb2da46c%2F_20190908173231.png?generation=1567931590044112&amp;alt=media)\n\n\n# Augmentation\n\nAfter crop &amp; resize all images to type\\_2, I heavily transforming those images by:\n\nDihedral, RandomCrop, Rotation, Contrast, Brightness, Cutout, PerspectiveTransform, Clahe.\n\n# Model\n\nI find that large model training on high resolution images is much more likely to be overfitting, due to the small size of training set. So I use small model for high resolution / large model for low resolution.\n\n* EfficientNet-B7 (224x224)  LB 0.840\n* EfficientNet-B6 (240x240) LB 0.832\n* EfficientNet-B5 (256x256) LB 0.831\n* EfficientNet-B4 (320x320) LB 0.826\n* EfficientNet-B3 (352x352) *LB Don't Know\n* EfficientNet-B2 (376x376) LB 0.828\n\n**LB is came from 5-fold testing x 8 Dihedral TTA**\n\nDespite the difference in Public LB, all of these model have a similar local 5-fold CV around 0.933\n\n↑ My FINAL submit v1 is simple avg these models.\n\nAnd the FINAL submit v2 is based on v1, replaced B2 &amp; B3 models with a B7 and a B5 that trained on Ben's preprocessing data (image type, along with other tricks are also applied).\n\n* EfficientNet-B7 (224x224 Ben's)  LB 0.834\n* EfficientNet-B5 (256x256 Ben's) LB 0.833\n\nlocal 5-fold CV of these 2 models is around 0.927\n\n# Training Process\n\nPretrain on 2015 full dataset (both train and test) for 25 epochs without validation. Then do 5-fold CV on 2019 dataset.\n\n# Predicting Process\n\nBecause of the 5-fold CV, I got 5 model files for each experiment.\nFor each experiment do 5 models * 8 Dihedral TTA predicting.\n**I only do image preprocessing once for 5*8=40 test as for saving time.** In this case 40 times test cost ~10mins for predicting 1928 images.\nWith ensemble 6 models, the kernel took 50mins for committing, 6h for submitting.\n\n",
      "votes": 61
    },
    {
      "id": 622743,
      "postDate": "2019-09-10T03:20:00.893Z",
      "content": "<p>Congratulation and thanks for your sharing. It is amazing that 40 times test cost only 10mins, can you give more details about it? Does do image preprocessing only once mean that the image is read from the hard disk once? Which framework do you use? We use fastai, EfficientNet-B7, 224*224 and do 5-fold cv, which takes 700s for predicting 1928 images. Thanks！</p>",
      "rawMarkdown": "Congratulation and thanks for your sharing. It is amazing that 40 times test cost only 10mins, can you give more details about it? Does do image preprocessing only once mean that the image is read from the hard disk once? Which framework do you use? We use fastai, EfficientNet-B7, 224*224 and do 5-fold cv, which takes 700s for predicting 1928 images. Thanks！",
      "votes": 1,
      "replies": [
        {
          "id": 622783,
          "postDate": "2019-09-10T05:05:53.853Z",
          "content": "<p><a href=\"/wenjuhuang\">@wenjuhuang</a>  Congrats to you, too! You got your first competition rank in top 1%!</p>\n\n<p>Sure, I use pytorch and I do something like this:</p>\n\n<p>```\n    models = get_models(enet_type, model_name, out_dim)  # load 5 models into GPU\n    outputs = [[] for x in range(n_fold)]   # n_fold=5</p>\n\n<pre><code>with torch.no_grad():\n    for data in tqdm(data_loader):  # batch_size=16\n        data = data.to(device)\n        for mid, model in enumerate(models):\n            for I in range(n_TTA):  # n_TTA=8\n                if I == 0:\n                    output = model(get_trans(data, I))\n                else:\n                    output += model(get_trans(data, I))\n            outputs[mid].append(output.squeeze() / n_TTA)\n</code></pre>\n\n<p>```</p>\n\n<p>```</p>\n\n<h1>Diheral TTA outside dataloader</h1>\n\n<p>def get_trans(img, I):\n    if I &gt;= 4:\n        img = img.transpose(2,3)\n    if I % 4 == 0:\n        return img\n    elif I % 4 == 1:\n        return img.flip(2)\n    elif I % 4 == 2:\n        return img.flip(3)\n    elif I % 4 == 3:\n        return img.flip(2).flip(3)\n```</p>\n\n<p>Loading and preprocessing is extremely slow in this game due to the high resolution of original images.\nWith the code above, actually the time cost between <code>n_TTA=1</code> and <code>n_TTA=8</code> is quite similar.</p>",
          "rawMarkdown": "@wenjuhuang  Congrats to you, too! You got your first competition rank in top 1%!\n\nSure, I use pytorch and I do something like this:\n\n```\n    models = get_models(enet_type, model_name, out_dim)  # load 5 models into GPU\n    outputs = [[] for x in range(n_fold)]   # n_fold=5\n\n    with torch.no_grad():\n        for data in tqdm(data_loader):  # batch_size=16\n            data = data.to(device)\n            for mid, model in enumerate(models):\n                for I in range(n_TTA):  # n_TTA=8\n                    if I == 0:\n                        output = model(get_trans(data, I))\n                    else:\n                        output += model(get_trans(data, I))\n                outputs[mid].append(output.squeeze() / n_TTA)\n```\n\n\n```\n# Diheral TTA outside dataloader\ndef get_trans(img, I):\n    if I &gt;= 4:\n        img = img.transpose(2,3)\n    if I % 4 == 0:\n        return img\n    elif I % 4 == 1:\n        return img.flip(2)\n    elif I % 4 == 2:\n        return img.flip(3)\n    elif I % 4 == 3:\n        return img.flip(2).flip(3)\n```\n\nLoading and preprocessing is extremely slow in this game due to the high resolution of original images.\nWith the code above, actually the time cost between `n_TTA=1` and `n_TTA=8` is quite similar.",
          "votes": 4
        },
        {
          "id": 622853,
          "postDate": "2019-09-10T06:58:43.367Z",
          "content": "<p>Thank you👍 . yeah, I found fastai loads data at each stage of TTA 😓 </p>",
          "rawMarkdown": "Thank you👍 . yeah, I found fastai loads data at each stage of TTA 😓 "
        }
      ]
    },
    {
      "id": 621554,
      "postDate": "2019-09-08T16:47:44.817Z",
      "content": "<p>Awesome job! Thank you for writing summary! and congratulation on becoming Master =) </p>",
      "rawMarkdown": "Awesome job! Thank you for writing summary! and congratulation on becoming Master =) ",
      "votes": 1,
      "replies": [
        {
          "id": 621560,
          "postDate": "2019-09-08T16:56:19.213Z",
          "content": "<p>Thanks! But I won't be master this time.\nMaybe in the next competition ;)\nAnd thank you for your sharing during this competition!</p>",
          "rawMarkdown": "Thanks! But I won't be master this time.\nMaybe in the next competition ;)\nAnd thank you for your sharing during this competition!",
          "votes": 1
        }
      ]
    },
    {
      "id": 621319,
      "postDate": "2019-09-08T12:37:01.123Z",
      "content": "<p>Very elegant solution. Question: how did you choose the best checkpoint when pre-training without validation? </p>",
      "rawMarkdown": "Very elegant solution. Question: how did you choose the best checkpoint when pre-training without validation? ",
      "votes": 1,
      "replies": [
        {
          "id": 621372,
          "postDate": "2019-09-08T13:19:08.810Z",
          "content": "<p>I find that it's not so different between using final epoch / best checkpoint, so I drop validation when pre-training for time sake.</p>",
          "rawMarkdown": "I find that it's not so different between using final epoch / best checkpoint, so I drop validation when pre-training for time sake.",
          "votes": 3
        }
      ]
    },
    {
      "id": 621174,
      "postDate": "2019-09-08T09:23:06.767Z",
      "content": "<p>Congratulation <a href=\"/haqishen\">@haqishen</a> Qishen !!  It's great to see B6 / B7 in action and great performance with small-size images (interesting!!)</p>",
      "rawMarkdown": "Congratulation @haqishen Qishen !!  It's great to see B6 / B7 in action and great performance with small-size images (interesting!!)",
      "votes": 1,
      "replies": [
        {
          "id": 621314,
          "postDate": "2019-09-08T12:32:40.837Z",
          "content": "<p>Thanks for your sharing in this competition! It is very informatic and inspired me a lot.\nIt's a pity for your team's shake down....\nI'm still curious about how you guys get 0.85+ public LB score.</p>",
          "rawMarkdown": "Thanks for your sharing in this competition! It is very informatic and inspired me a lot.\nIt's a pity for your team's shake down....\nI'm still curious about how you guys get 0.85+ public LB score.",
          "votes": 1
        },
        {
          "id": 621336,
          "postDate": "2019-09-08T12:50:01.190Z",
          "content": "<p>I use very similar model to Gary's (8th place) with 2-heads (regression/classification) and also pseudo labelling. By using iterative pseudo labels very similar to this post by <a href=\"/bibek777\">@bibek777</a> </p>\n\n<p><a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106177#latest-610648\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106177#latest-610648</a></p>\n\n<p>By iterating 3 times and ensemble them I was able to boost from Public 0.82x --&gt; 0.850</p>\n\n<p>I could continue but I was also afraid to overfit ( but I did not see overfitting sign by my CV and my Grad-CAM visualization)</p>\n\n<p>It turns out that this model is not overfit but got a very little boost in private (926 --&gt; 929), this is understandable since public is not like private at all :) Since this is my personal model, we didn't select it and rather try more promising combination to others... I didn't really expect/anticipate that our most promising choice will not fit private ... Choose 2 subs really a difficult job!</p>",
          "rawMarkdown": "I use very similar model to Gary's (8th place) with 2-heads (regression/classification) and also pseudo labelling. By using iterative pseudo labels very similar to this post by @bibek777 \n\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106177#latest-610648\n\nBy iterating 3 times and ensemble them I was able to boost from Public 0.82x --&gt; 0.850\n\nI could continue but I was also afraid to overfit ( but I did not see overfitting sign by my CV and my Grad-CAM visualization)\n\nIt turns out that this model is not overfit but got a very little boost in private (926 --&gt; 929), this is understandable since public is not like private at all :) Since this is my personal model, we didn't select it and rather try more promising combination to others... I didn't really expect/anticipate that our most promising choice will not fit private ... Choose 2 subs really a difficult job!",
          "votes": 2
        },
        {
          "id": 621364,
          "postDate": "2019-09-08T13:16:53.807Z",
          "content": "<p>926 --&gt; 929 is impressive boost in private LB! Thank you for sharing!</p>",
          "rawMarkdown": "926 --&gt; 929 is impressive boost in private LB! Thank you for sharing!"
        }
      ]
    },
    {
      "id": 627489,
      "postDate": "2019-09-16T03:56:59.193Z",
      "content": "<p>Congrats and appreciate the awesome sharing!!!\nI have some questions about the preprocessing applied on the dataset.</p>\n\n<ol>\n<li><p>You said that the image type depend on \"brightness\" in the literature, however that might exist different degree of brightness each type, right? It could looked dark or bright in Type_0, and so on... \nSo I straightly thought you defined each type by the shape of retinopathy like perfectly maintain the circle shape for Type_0, cut out top and bottom for Type_1, slight zoom in for Type_2. Hope that I did not misunderstand that definition!</p></li>\n<li><p>Did you manually define each training data to specific type or training a small model for the task?</p></li>\n<li><p>According the experiments you only took image type normalization for training data, how about validation data in the cross validation and the testing data for submission? Only resizing both of them? </p></li>\n<li><p>How about the augmentation application for validation and testing data?</p></li>\n</ol>\n\n<p>Thank you so much!!!</p>",
      "rawMarkdown": "Congrats and appreciate the awesome sharing!!!\nI have some questions about the preprocessing applied on the dataset.\n\n1. You said that the image type depend on \"brightness\" in the literature, however that might exist different degree of brightness each type, right? It could looked dark or bright in Type_0, and so on... \nSo I straightly thought you defined each type by the shape of retinopathy like perfectly maintain the circle shape for Type_0, cut out top and bottom for Type_1, slight zoom in for Type_2. Hope that I did not misunderstand that definition!\n\n2. Did you manually define each training data to specific type or training a small model for the task?\n\n3. According the experiments you only took image type normalization for training data, how about validation data in the cross validation and the testing data for submission? Only resizing both of them? \n\n4. How about the augmentation application for validation and testing data?\n\nThank you so much!!!",
      "replies": [
        {
          "id": 627520,
          "postDate": "2019-09-16T05:05:43.493Z",
          "content": "<p>Thank you.</p>\n\n<p>&gt; You said that the image type depend on \"brightness\" in the literature, however that might exist different degree of brightness each type, right? It could looked dark or bright in Type0, and so on… So I straightly thought you defined each type by the shape of retinopathy like perfectly maintain the circle shape for Type0, cut out top and bottom for Type1, slight zoom in for Type2. Hope that I did not misunderstand that definition!</p>\n\n<p>&gt; Did you manually define each training data to specific type or training a small model for the task?</p>\n\n<p>To be specifically, I count the number of pixels that brighter than 15 of each boundary for each image, plot the distribution to find out the best threshold, and apply image type base on that threshold. It's not a perfect split but It's good enough for me.</p>\n\n<p>&gt; According the experiments you only took image type normalization for training data, how about validation data in the cross validation and the testing data for submission? Only resizing both of them?</p>\n\n<p>&gt; How about the augmentation application for validation and testing data?</p>\n\n<p>Only crop from gray for validation and prediction. Additionally, 8 times dihedral TTA is applied for prediction.</p>",
          "rawMarkdown": "Thank you.\n\n&gt; You said that the image type depend on \"brightness\" in the literature, however that might exist different degree of brightness each type, right? It could looked dark or bright in Type0, and so on… So I straightly thought you defined each type by the shape of retinopathy like perfectly maintain the circle shape for Type0, cut out top and bottom for Type1, slight zoom in for Type2. Hope that I did not misunderstand that definition!\n\n&gt; Did you manually define each training data to specific type or training a small model for the task?\n\nTo be specifically, I count the number of pixels that brighter than 15 of each boundary for each image, plot the distribution to find out the best threshold, and apply image type base on that threshold. It's not a perfect split but It's good enough for me.\n\n&gt; According the experiments you only took image type normalization for training data, how about validation data in the cross validation and the testing data for submission? Only resizing both of them?\n\n&gt; How about the augmentation application for validation and testing data?\n\nOnly crop from gray for validation and prediction. Additionally, 8 times dihedral TTA is applied for prediction.\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 625605,
      "postDate": "2019-09-13T08:13:07.710Z",
      "content": "<p>I was wandering which of these previously high-scoring teams were cleaned and why....</p>",
      "rawMarkdown": " I was wandering which of these previously high-scoring teams were cleaned and why....",
      "replies": [
        {
          "id": 625667,
          "postDate": "2019-09-13T09:36:10.363Z",
          "content": "<p>Read this post, but we have no evidence at all. Kaggle will not provide any information about that for us.\n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107914#621873\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107914#621873</a></p>",
          "rawMarkdown": "Read this post, but we have no evidence at all. Kaggle will not provide any information about that for us.\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107914#621873"
        }
      ]
    },
    {
      "id": 623949,
      "postDate": "2019-09-11T12:55:10.240Z",
      "content": "<p>Congrats for jumping up to prize zone!!👍 </p>",
      "rawMarkdown": "Congrats for jumping up to prize zone!!👍 ",
      "replies": [
        {
          "id": 623961,
          "postDate": "2019-09-11T13:16:59.787Z",
          "content": "<p>Thanks! My friend shared your tweet 「俺達のAPTOSはまだ終わってない」to me. www</p>",
          "rawMarkdown": "Thanks! My friend shared your tweet 「俺達のAPTOSはまだ終わってない」to me. www"
        }
      ]
    },
    {
      "id": 623777,
      "postDate": "2019-09-11T09:43:07.327Z",
      "content": "<p>Congratulations with the great result! Did you approach this problem as regression? If yes, did you tune the thresholds, and if yes, then how, if no, why not)</p>",
      "rawMarkdown": "Congratulations with the great result! Did you approach this problem as regression? If yes, did you tune the thresholds, and if yes, then how, if no, why not)",
      "replies": [
        {
          "id": 623787,
          "postDate": "2019-09-11T09:51:46.997Z",
          "content": "<p>Thank you.\nI use regression without tuning thresholds, just round().\nI have tried to optimize kappa score by the method shared in a public kernel, but my best model shows that optimized threshold is quite similar to [0.5,1.5,2.5,3.5], something like [0.503, 1.497, 2.508, 3.501]. So I choose not to optimize it.</p>",
          "rawMarkdown": "Thank you.\nI use regression without tuning thresholds, just round().\nI have tried to optimize kappa score by the method shared in a public kernel, but my best model shows that optimized threshold is quite similar to [0.5,1.5,2.5,3.5], something like [0.503, 1.497, 2.508, 3.501]. So I choose not to optimize it."
        }
      ]
    },
    {
      "id": 621542,
      "postDate": "2019-09-08T16:38:01.407Z",
      "content": "<p>Hi thank your for sharing. The normalization idea is impressive. \nI have a question: does your augmentation method (Dihedral, RandomCrop, Rotation, Contrast, Brightness, Cutout, PerspectiveTransform, Clahe.) all interpretable? can you tell why you use this methods?</p>",
      "rawMarkdown": "Hi thank your for sharing. The normalization idea is impressive. \nI have a question: does your augmentation method (Dihedral, RandomCrop, Rotation, Contrast, Brightness, Cutout, PerspectiveTransform, Clahe.) all interpretable? can you tell why you use this methods?",
      "replies": [
        {
          "id": 621559,
          "postDate": "2019-09-08T16:54:47.333Z",
          "content": "<p>The main idea of augmentation is modifying images without changing its label, so that we can enlarge our training set and make our model generalize better.\nWhat I chosen is most likely to do this job for this dataset. Then I'll do experiment on it.\nIt's not an interpretation for each of those augmentation method but I hope it can help you.</p>",
          "rawMarkdown": "The main idea of augmentation is modifying images without changing its label, so that we can enlarge our training set and make our model generalize better.\nWhat I chosen is most likely to do this job for this dataset. Then I'll do experiment on it.\nIt's not an interpretation for each of those augmentation method but I hope it can help you.\n",
          "votes": 1
        },
        {
          "id": 621756,
          "postDate": "2019-09-08T22:34:31.957Z",
          "content": "<p>thank you for your reply. May I ask for 3 more questions:\n- diheral you are using fastai? because i found no package for this besides fastai?\n- randomcrop on raw picture or randomcrop on resized picture?\n- any lr_schedule and how you find the best policy\nThx.</p>",
          "rawMarkdown": "thank you for your reply. May I ask for 3 more questions:\n- diheral you are using fastai? because i found no package for this besides fastai?\n- randomcrop on raw picture or randomcrop on resized picture?\n- any lr_schedule and how you find the best policy\nThx."
        },
        {
          "id": 621883,
          "postDate": "2019-09-09T04:17:51.503Z",
          "content": "<p>Sure.</p>\n\n<p>&gt; diheral you are using fastai? because i found no package for this besides fastai?</p>\n\n<p>I use pytorch. Diheral is not a difficult augmentation to implement. I just used the term that fastai calls it (since I don't know how to call it properly)</p>\n\n<p>&gt; randomcrop on raw picture or randomcrop on resized picture?</p>\n\n<p>Randomcrop on type_2 transformed images. Actually all augmentation process is done on type_2 transformed images.</p>\n\n<p>&gt; any lr_schedule and how you find the best policy</p>\n\n<p>Simple cosine LR schedule, for pre-train I set <code>init_lr = 0.0005</code> , for finetune I set <code>init_lr = 0.00001</code> .\nI've try something like warmup but it seem to be not helpful on my local CV.</p>",
          "rawMarkdown": "Sure.\n\n&gt; diheral you are using fastai? because i found no package for this besides fastai?\n\nI use pytorch. Diheral is not a difficult augmentation to implement. I just used the term that fastai calls it (since I don't know how to call it properly)\n\n&gt; randomcrop on raw picture or randomcrop on resized picture?\n\nRandomcrop on type\\_2 transformed images. Actually all augmentation process is done on type\\_2 transformed images.\n\n&gt; any lr_schedule and how you find the best policy\n\nSimple cosine LR schedule, for pre-train I set `init_lr = 0.0005` , for finetune I set `init_lr = 0.00001` .\nI've try something like warmup but it seem to be not helpful on my local CV.",
          "votes": 1
        },
        {
          "id": 622061,
          "postDate": "2019-09-09T08:22:36.417Z",
          "content": "<p>thank you!</p>",
          "rawMarkdown": "thank you!"
        },
        {
          "id": 626475,
          "postDate": "2019-09-14T11:49:59.840Z",
          "content": "<p>Hi I want to know: is your augmentation: \"Dihedral, RandomCrop, Rotation, Contrast, Brightness, Cutout, PerspectiveTransform, Clahe\" is listing in order? Or their order can be changed and you can achieve similar result? </p>",
          "rawMarkdown": "Hi I want to know: is your augmentation: \"Dihedral, RandomCrop, Rotation, Contrast, Brightness, Cutout, PerspectiveTransform, Clahe\" is listing in order? Or their order can be changed and you can achieve similar result? "
        },
        {
          "id": 626905,
          "postDate": "2019-09-15T05:00:26.840Z",
          "content": "<p>IMO the order is never an important thing in augmentation. You can just visualize images after augmentations to find it out.</p>",
          "rawMarkdown": "IMO the order is never an important thing in augmentation. You can just visualize images after augmentations to find it out."
        }
      ]
    },
    {
      "id": 621392,
      "postDate": "2019-09-08T13:30:06.857Z",
      "content": "<p>Congrats <a href=\"/haqishen\">@haqishen</a>  and thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congrats @haqishen  and thanks for sharing your solution overview."
    },
    {
      "id": 621176,
      "postDate": "2019-09-08T09:27:07.053Z",
      "content": "<p>Congratulations! Using the brightness of image boundary to apply image type is impressive! Why you only use this norm method in training?</p>",
      "rawMarkdown": "Congratulations! Using the brightness of image boundary to apply image type is impressive! Why you only use this norm method in training?",
      "replies": [
        {
          "id": 621308,
          "postDate": "2019-09-08T12:25:49.730Z",
          "content": "<p>First of all, image type normalization can certainly eliminate the effects of meta data.</p>\n\n<p>And,</p>\n\n<p>EXP1. Adding image type normalization all the time dropped my local CV.\nEXP2. Adding image type normalization only for training also dropped my local CV but a little better than exp1.</p>\n\n<p>It means, the information that dropped by image type normalization is actually useful. So I finally didn't apply this trick for testing.</p>",
          "rawMarkdown": "First of all, image type normalization can certainly eliminate the effects of meta data.\n\nAnd,\n\nEXP1. Adding image type normalization all the time dropped my local CV.\nEXP2. Adding image type normalization only for training also dropped my local CV but a little better than exp1.\n\nIt means, the information that dropped by image type normalization is actually useful. So I finally didn't apply this trick for testing.",
          "votes": 1
        }
      ]
    },
    {
      "id": 621172,
      "postDate": "2019-09-08T09:21:49.450Z",
      "content": "<p>Congrats! Image type normalization is impressive.</p>",
      "rawMarkdown": "Congrats! Image type normalization is impressive."
    },
    {
      "id": 621169,
      "postDate": "2019-09-08T09:20:14.047Z",
      "content": "<p>Thanks for sharing.</p>\n\n<p>&gt;My best single model, with public LB 0.840 , private LB 0.929, which is EfficientNet-B7 with 224x224 input.</p>\n\n<p>This also proves the point that you don't necessarily have to use the default image size for the architectures to get the best out of them, My efficientnet-b5 performed best with 256x images instead of 456 (the default). </p>",
      "rawMarkdown": "Thanks for sharing.\n\n&gt;My best single model, with public LB 0.840 , private LB 0.929, which is EfficientNet-B7 with 224x224 input.\n\nThis also proves the point that you don't necessarily have to use the default image size for the architectures to get the best out of them, My efficientnet-b5 performed best with 256x images instead of 456 (the default). \n",
      "replies": [
        {
          "id": 621493,
          "postDate": "2019-09-08T15:29:42.277Z",
          "content": "<p>I think the secret lies on the size of dataset. If we have a ImageNet size diabetic retinopathy dataset, maybe B7 can just fit on higher resolution without being overfitting.</p>",
          "rawMarkdown": "I think the secret lies on the size of dataset. If we have a ImageNet size diabetic retinopathy dataset, maybe B7 can just fit on higher resolution without being overfitting."
        }
      ]
    },
    {
      "id": 621168,
      "postDate": "2019-09-08T09:17:49.063Z",
      "content": "<p>congratulations, a impressive solution, you got a nice score without pseudo labels.</p>",
      "rawMarkdown": "congratulations, a impressive solution, you got a nice score without pseudo labels.",
      "replies": [
        {
          "id": 621173,
          "postDate": "2019-09-08T09:22:01.257Z",
          "content": "<p>Interesting, probably because he is using those type_* pre-processing techniques to make the images resemble to the public test images, majority of which were cropped in these manners.</p>",
          "rawMarkdown": "Interesting, probably because he is using those type_* pre-processing techniques to make the images resemble to the public test images, majority of which were cropped in these manners."
        },
        {
          "id": 621311,
          "postDate": "2019-09-08T12:28:32.033Z",
          "content": "<p><a href=\"/garybios\">@garybios</a> Congrats to you, too!\nI've tried pseudo once but didn't get good result so I gave up it 😂 Maybe I have done something wrong.</p>",
          "rawMarkdown": "@garybios Congrats to you, too!\nI've tried pseudo once but didn't get good result so I gave up it 😂 Maybe I have done something wrong.\n"
        }
      ]
    },
    {
      "id": 621946,
      "postDate": "2019-09-09T05:40:02.287Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 622743,
      "author_name": "Wenju Huang",
      "author_url": "",
      "post_date": "2019-09-10T03:20:00.893000",
      "content": "<p>Congratulation and thanks for your sharing. It is amazing that 40 times test cost only 10mins, can you give more details about it? Does do image preprocessing only once mean that the image is read from the hard disk once? Which framework do you use? We use fastai, EfficientNet-B7, 224*224 and do 5-fold cv, which takes 700s for predicting 1928 images. Thanks！</p>",
      "votes": 1,
      "replies": [
        {
          "id": 622783,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-10T05:05:53.853000",
          "content": "<p><a href=\"/wenjuhuang\">@wenjuhuang</a>  Congrats to you, too! You got your first competition rank in top 1%!</p>\n\n<p>Sure, I use pytorch and I do something like this:</p>\n\n<p>```\n    models = get_models(enet_type, model_name, out_dim)  # load 5 models into GPU\n    outputs = [[] for x in range(n_fold)]   # n_fold=5</p>\n\n<pre><code>with torch.no_grad():\n    for data in tqdm(data_loader):  # batch_size=16\n        data = data.to(device)\n        for mid, model in enumerate(models):\n            for I in range(n_TTA):  # n_TTA=8\n                if I == 0:\n                    output = model(get_trans(data, I))\n                else:\n                    output += model(get_trans(data, I))\n            outputs[mid].append(output.squeeze() / n_TTA)\n</code></pre>\n\n<p>```</p>\n\n<p>```</p>\n\n<h1>Diheral TTA outside dataloader</h1>\n\n<p>def get_trans(img, I):\n    if I &gt;= 4:\n        img = img.transpose(2,3)\n    if I % 4 == 0:\n        return img\n    elif I % 4 == 1:\n        return img.flip(2)\n    elif I % 4 == 2:\n        return img.flip(3)\n    elif I % 4 == 3:\n        return img.flip(2).flip(3)\n```</p>\n\n<p>Loading and preprocessing is extremely slow in this game due to the high resolution of original images.\nWith the code above, actually the time cost between <code>n_TTA=1</code> and <code>n_TTA=8</code> is quite similar.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 622853,
          "author_name": "Wenju Huang",
          "author_url": "",
          "post_date": "2019-09-10T06:58:43.367000",
          "content": "<p>Thank you👍 . yeah, I found fastai loads data at each stage of TTA 😓 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621554,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2019-09-08T16:47:44.817000",
      "content": "<p>Awesome job! Thank you for writing summary! and congratulation on becoming Master =) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 621560,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-08T16:56:19.213000",
          "content": "<p>Thanks! But I won't be master this time.\nMaybe in the next competition ;)\nAnd thank you for your sharing during this competition!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 621319,
      "author_name": "Thomas Yokota",
      "author_url": "",
      "post_date": "2019-09-08T12:37:01.123000",
      "content": "<p>Very elegant solution. Question: how did you choose the best checkpoint when pre-training without validation? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 621372,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-08T13:19:08.810000",
          "content": "<p>I find that it's not so different between using final epoch / best checkpoint, so I drop validation when pre-training for time sake.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 621174,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-09-08T09:23:06.767000",
      "content": "<p>Congratulation <a href=\"/haqishen\">@haqishen</a> Qishen !!  It's great to see B6 / B7 in action and great performance with small-size images (interesting!!)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 621314,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-08T12:32:40.837000",
          "content": "<p>Thanks for your sharing in this competition! It is very informatic and inspired me a lot.\nIt's a pity for your team's shake down....\nI'm still curious about how you guys get 0.85+ public LB score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 621336,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-09-08T12:50:01.190000",
          "content": "<p>I use very similar model to Gary's (8th place) with 2-heads (regression/classification) and also pseudo labelling. By using iterative pseudo labels very similar to this post by <a href=\"/bibek777\">@bibek777</a> </p>\n\n<p><a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106177#latest-610648\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106177#latest-610648</a></p>\n\n<p>By iterating 3 times and ensemble them I was able to boost from Public 0.82x --&gt; 0.850</p>\n\n<p>I could continue but I was also afraid to overfit ( but I did not see overfitting sign by my CV and my Grad-CAM visualization)</p>\n\n<p>It turns out that this model is not overfit but got a very little boost in private (926 --&gt; 929), this is understandable since public is not like private at all :) Since this is my personal model, we didn't select it and rather try more promising combination to others... I didn't really expect/anticipate that our most promising choice will not fit private ... Choose 2 subs really a difficult job!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 621364,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-08T13:16:53.807000",
          "content": "<p>926 --&gt; 929 is impressive boost in private LB! Thank you for sharing!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 627489,
      "author_name": "jayjhlin",
      "author_url": "",
      "post_date": "2019-09-16T03:56:59.193000",
      "content": "<p>Congrats and appreciate the awesome sharing!!!\nI have some questions about the preprocessing applied on the dataset.</p>\n\n<ol>\n<li><p>You said that the image type depend on \"brightness\" in the literature, however that might exist different degree of brightness each type, right? It could looked dark or bright in Type_0, and so on... \nSo I straightly thought you defined each type by the shape of retinopathy like perfectly maintain the circle shape for Type_0, cut out top and bottom for Type_1, slight zoom in for Type_2. Hope that I did not misunderstand that definition!</p></li>\n<li><p>Did you manually define each training data to specific type or training a small model for the task?</p></li>\n<li><p>According the experiments you only took image type normalization for training data, how about validation data in the cross validation and the testing data for submission? Only resizing both of them? </p></li>\n<li><p>How about the augmentation application for validation and testing data?</p></li>\n</ol>\n\n<p>Thank you so much!!!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 627520,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-16T05:05:43.493000",
          "content": "<p>Thank you.</p>\n\n<p>&gt; You said that the image type depend on \"brightness\" in the literature, however that might exist different degree of brightness each type, right? It could looked dark or bright in Type0, and so on… So I straightly thought you defined each type by the shape of retinopathy like perfectly maintain the circle shape for Type0, cut out top and bottom for Type1, slight zoom in for Type2. Hope that I did not misunderstand that definition!</p>\n\n<p>&gt; Did you manually define each training data to specific type or training a small model for the task?</p>\n\n<p>To be specifically, I count the number of pixels that brighter than 15 of each boundary for each image, plot the distribution to find out the best threshold, and apply image type base on that threshold. It's not a perfect split but It's good enough for me.</p>\n\n<p>&gt; According the experiments you only took image type normalization for training data, how about validation data in the cross validation and the testing data for submission? Only resizing both of them?</p>\n\n<p>&gt; How about the augmentation application for validation and testing data?</p>\n\n<p>Only crop from gray for validation and prediction. Additionally, 8 times dihedral TTA is applied for prediction.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 625605,
      "author_name": "Homoalways",
      "author_url": "",
      "post_date": "2019-09-13T08:13:07.710000",
      "content": "<p>I was wandering which of these previously high-scoring teams were cleaned and why....</p>",
      "votes": 0,
      "replies": [
        {
          "id": 625667,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-13T09:36:10.363000",
          "content": "<p>Read this post, but we have no evidence at all. Kaggle will not provide any information about that for us.\n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107914#621873\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107914#621873</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 623949,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2019-09-11T12:55:10.240000",
      "content": "<p>Congrats for jumping up to prize zone!!👍 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 623961,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-11T13:16:59.787000",
          "content": "<p>Thanks! My friend shared your tweet 「俺達のAPTOSはまだ終わってない」to me. www</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 623777,
      "author_name": "Evgeny Kovalev",
      "author_url": "",
      "post_date": "2019-09-11T09:43:07.327000",
      "content": "<p>Congratulations with the great result! Did you approach this problem as regression? If yes, did you tune the thresholds, and if yes, then how, if no, why not)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 623787,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-11T09:51:46.997000",
          "content": "<p>Thank you.\nI use regression without tuning thresholds, just round().\nI have tried to optimize kappa score by the method shared in a public kernel, but my best model shows that optimized threshold is quite similar to [0.5,1.5,2.5,3.5], something like [0.503, 1.497, 2.508, 3.501]. So I choose not to optimize it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621542,
      "author_name": "SchenbergZ",
      "author_url": "",
      "post_date": "2019-09-08T16:38:01.407000",
      "content": "<p>Hi thank your for sharing. The normalization idea is impressive. \nI have a question: does your augmentation method (Dihedral, RandomCrop, Rotation, Contrast, Brightness, Cutout, PerspectiveTransform, Clahe.) all interpretable? can you tell why you use this methods?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 621559,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-08T16:54:47.333000",
          "content": "<p>The main idea of augmentation is modifying images without changing its label, so that we can enlarge our training set and make our model generalize better.\nWhat I chosen is most likely to do this job for this dataset. Then I'll do experiment on it.\nIt's not an interpretation for each of those augmentation method but I hope it can help you.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 621756,
          "author_name": "SchenbergZ",
          "author_url": "",
          "post_date": "2019-09-08T22:34:31.957000",
          "content": "<p>thank you for your reply. May I ask for 3 more questions:\n- diheral you are using fastai? because i found no package for this besides fastai?\n- randomcrop on raw picture or randomcrop on resized picture?\n- any lr_schedule and how you find the best policy\nThx.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 621883,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-09T04:17:51.503000",
          "content": "<p>Sure.</p>\n\n<p>&gt; diheral you are using fastai? because i found no package for this besides fastai?</p>\n\n<p>I use pytorch. Diheral is not a difficult augmentation to implement. I just used the term that fastai calls it (since I don't know how to call it properly)</p>\n\n<p>&gt; randomcrop on raw picture or randomcrop on resized picture?</p>\n\n<p>Randomcrop on type_2 transformed images. Actually all augmentation process is done on type_2 transformed images.</p>\n\n<p>&gt; any lr_schedule and how you find the best policy</p>\n\n<p>Simple cosine LR schedule, for pre-train I set <code>init_lr = 0.0005</code> , for finetune I set <code>init_lr = 0.00001</code> .\nI've try something like warmup but it seem to be not helpful on my local CV.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 622061,
          "author_name": "SchenbergZ",
          "author_url": "",
          "post_date": "2019-09-09T08:22:36.417000",
          "content": "<p>thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 626475,
          "author_name": "SchenbergZ",
          "author_url": "",
          "post_date": "2019-09-14T11:49:59.840000",
          "content": "<p>Hi I want to know: is your augmentation: \"Dihedral, RandomCrop, Rotation, Contrast, Brightness, Cutout, PerspectiveTransform, Clahe\" is listing in order? Or their order can be changed and you can achieve similar result? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 626905,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-15T05:00:26.840000",
          "content": "<p>IMO the order is never an important thing in augmentation. You can just visualize images after augmentations to find it out.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621392,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-09-08T13:30:06.857000",
      "content": "<p>Congrats <a href=\"/haqishen\">@haqishen</a>  and thanks for sharing your solution overview.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621176,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2019-09-08T09:27:07.053000",
      "content": "<p>Congratulations! Using the brightness of image boundary to apply image type is impressive! Why you only use this norm method in training?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 621308,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-08T12:25:49.730000",
          "content": "<p>First of all, image type normalization can certainly eliminate the effects of meta data.</p>\n\n<p>And,</p>\n\n<p>EXP1. Adding image type normalization all the time dropped my local CV.\nEXP2. Adding image type normalization only for training also dropped my local CV but a little better than exp1.</p>\n\n<p>It means, the information that dropped by image type normalization is actually useful. So I finally didn't apply this trick for testing.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 621172,
      "author_name": "Yuanhao",
      "author_url": "",
      "post_date": "2019-09-08T09:21:49.450000",
      "content": "<p>Congrats! Image type normalization is impressive.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621169,
      "author_name": "Rishabh Agrahari",
      "author_url": "",
      "post_date": "2019-09-08T09:20:14.047000",
      "content": "<p>Thanks for sharing.</p>\n\n<p>&gt;My best single model, with public LB 0.840 , private LB 0.929, which is EfficientNet-B7 with 224x224 input.</p>\n\n<p>This also proves the point that you don't necessarily have to use the default image size for the architectures to get the best out of them, My efficientnet-b5 performed best with 256x images instead of 456 (the default). </p>",
      "votes": 0,
      "replies": [
        {
          "id": 621493,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-08T15:29:42.277000",
          "content": "<p>I think the secret lies on the size of dataset. If we have a ImageNet size diabetic retinopathy dataset, maybe B7 can just fit on higher resolution without being overfitting.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621168,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2019-09-08T09:17:49.063000",
      "content": "<p>congratulations, a impressive solution, you got a nice score without pseudo labels.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 621173,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-09-08T09:22:01.257000",
          "content": "<p>Interesting, probably because he is using those type_* pre-processing techniques to make the images resemble to the public test images, majority of which were cropped in these manners.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 621311,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2019-09-08T12:28:32.033000",
          "content": "<p><a href=\"/garybios\">@garybios</a> Congrats to you, too!\nI've tried pseudo once but didn't get good result so I gave up it 😂 Maybe I have done something wrong.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621946,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T05:40:02.287000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "621164": "# My Submissions\n\nMy best public LB:  `0.842` (with quick submit, so no private LB available)\nMy FINAL submit v1, with public LB `0.838` , private LB `0.932` , is based on my 6 models with best local CV.\nMy FINAL submit v2, with public LB `0.840` , private LB `0.933` , with more variance in ensemble models. \nMy best single model, with public LB `0.840` , private LB `0.929`, which is EfficientNet-B7 with 224x224 input.\n\n# Preprocessing\n\n## Crop From Gray (Both training and predicting)\n\nThis function is shared in [here](https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping) (search keyword `crop_image_from_gray` in the page, thanks @ratthachat for sharing!)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2F0405798d336f82857988977e629f108b%2F_20190908171800.png?generation=1567930710772532&amp;alt=media)\n\n## Apply Image Type (Only For Training)\n\nI apply image type for every image in training set (2015 &amp; 2019) using the brightness of image boundary. \nFrom left to right in the following picture: `Type_0` , `Type_1` , `Type_2` \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2F0fd5e78191a5373707c4faf3c83f1ad6%2F_20190908211528.png?generation=1567944954429547&amp;alt=media)\n\n\n\n## Transform All Image to Type 2 (Only For Training)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F448347%2F8439f5c3e2a27fd2d548e460eb2da46c%2F_20190908173231.png?generation=1567931590044112&amp;alt=media)\n\n\n# Augmentation\n\nAfter crop &amp; resize all images to type\\_2, I heavily transforming those images by:\n\nDihedral, RandomCrop, Rotation, Contrast, Brightness, Cutout, PerspectiveTransform, Clahe.\n\n# Model\n\nI find that large model training on high resolution images is much more likely to be overfitting, due to the small size of training set. So I use small model for high resolution / large model for low resolution.\n\n* EfficientNet-B7 (224x224)  LB 0.840\n* EfficientNet-B6 (240x240) LB 0.832\n* EfficientNet-B5 (256x256) LB 0.831\n* EfficientNet-B4 (320x320) LB 0.826\n* EfficientNet-B3 (352x352) *LB Don't Know\n* EfficientNet-B2 (376x376) LB 0.828\n\n**LB is came from 5-fold testing x 8 Dihedral TTA**\n\nDespite the difference in Public LB, all of these model have a similar local 5-fold CV around 0.933\n\n↑ My FINAL submit v1 is simple avg these models.\n\nAnd the FINAL submit v2 is based on v1, replaced B2 &amp; B3 models with a B7 and a B5 that trained on Ben's preprocessing data (image type, along with other tricks are also applied).\n\n* EfficientNet-B7 (224x224 Ben's)  LB 0.834\n* EfficientNet-B5 (256x256 Ben's) LB 0.833\n\nlocal 5-fold CV of these 2 models is around 0.927\n\n# Training Process\n\nPretrain on 2015 full dataset (both train and test) for 25 epochs without validation. Then do 5-fold CV on 2019 dataset.\n\n# Predicting Process\n\nBecause of the 5-fold CV, I got 5 model files for each experiment.\nFor each experiment do 5 models * 8 Dihedral TTA predicting.\n**I only do image preprocessing once for 5*8=40 test as for saving time.** In this case 40 times test cost ~10mins for predicting 1928 images.\nWith ensemble 6 models, the kernel took 50mins for committing, 6h for submitting.\n\n",
    "622743": "Congratulation and thanks for your sharing. It is amazing that 40 times test cost only 10mins, can you give more details about it? Does do image preprocessing only once mean that the image is read from the hard disk once? Which framework do you use? We use fastai, EfficientNet-B7, 224*224 and do 5-fold cv, which takes 700s for predicting 1928 images. Thanks！",
    "621554": "Awesome job! Thank you for writing summary! and congratulation on becoming Master =) ",
    "621319": "Very elegant solution. Question: how did you choose the best checkpoint when pre-training without validation? ",
    "621174": "Congratulation @haqishen Qishen !!  It's great to see B6 / B7 in action and great performance with small-size images (interesting!!)",
    "627489": "Congrats and appreciate the awesome sharing!!!\nI have some questions about the preprocessing applied on the dataset.\n\n1. You said that the image type depend on \"brightness\" in the literature, however that might exist different degree of brightness each type, right? It could looked dark or bright in Type_0, and so on... \nSo I straightly thought you defined each type by the shape of retinopathy like perfectly maintain the circle shape for Type_0, cut out top and bottom for Type_1, slight zoom in for Type_2. Hope that I did not misunderstand that definition!\n\n2. Did you manually define each training data to specific type or training a small model for the task?\n\n3. According the experiments you only took image type normalization for training data, how about validation data in the cross validation and the testing data for submission? Only resizing both of them? \n\n4. How about the augmentation application for validation and testing data?\n\nThank you so much!!!",
    "625605": " I was wandering which of these previously high-scoring teams were cleaned and why....",
    "623949": "Congrats for jumping up to prize zone!!👍 ",
    "623777": "Congratulations with the great result! Did you approach this problem as regression? If yes, did you tune the thresholds, and if yes, then how, if no, why not)",
    "621542": "Hi thank your for sharing. The normalization idea is impressive. \nI have a question: does your augmentation method (Dihedral, RandomCrop, Rotation, Contrast, Brightness, Cutout, PerspectiveTransform, Clahe.) all interpretable? can you tell why you use this methods?",
    "621392": "Congrats @haqishen  and thanks for sharing your solution overview.",
    "621176": "Congratulations! Using the brightness of image boundary to apply image type is impressive! Why you only use this norm method in training?",
    "621172": "Congrats! Image type normalization is impressive.",
    "621169": "Thanks for sharing.\n\n&gt;My best single model, with public LB 0.840 , private LB 0.929, which is EfficientNet-B7 with 224x224 input.\n\nThis also proves the point that you don't necessarily have to use the default image size for the architectures to get the best out of them, My efficientnet-b5 performed best with 256x images instead of 456 (the default). \n",
    "621168": "congratulations, a impressive solution, you got a nice score without pseudo labels.",
    "621946": ""
  }
}