{
  "id": 108030,
  "title": "8th Place Solution ",
  "url": "/competitions/aptos2019-blindness-detection/discussion/108030",
  "author_name": "DrHB",
  "post_date": "2019-09-08T15:17:40.581000",
  "votes": 90,
  "comment_count": 32,
  "views": 0,
  "content": "<p>First of all I want to thank everyone who participated, shared kernels and was part of activate discussion. </p>\n\n<h2>PROBLEM:</h2>\n\n<p>As you have already heard and read many top team solutions have used very simple 2 STEP approach in fact this was heavily discussed here \n(<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815</a>) and I also partially mentinoed in my kernel (<a href=\"https://www.kaggle.com/drhabib/starter-kernel-for-0-79\">https://www.kaggle.com/drhabib/starter-kernel-for-0-79</a>) :\n<code>\nSTEP 1:  pretrain on 2015 data \nSTEP 2:  finetune on 2019 data \n</code>\nThis approach is good. But When I personally started experimenting I was extremely worried that our local CV was around 0.931 but LB was 0.811. This was a huge gap and we were afraid about shake up. Also just pertaining on old data using 2019 as Validation results in LB 0.75. Which indicated that our model was not generalizing good. </p>\n\n<h2>SOLUTION:</h2>\n\n<p><code>Our idea was very simple instead of training in two diffrent stages we will just combine all the data from 2019 and 2015 and try to work from here to improve our generalization.</code></p>\n\n<h3>PREPROCESSING.</h3>\n\n<p><code>ISSUE 1:</code>\nRemove black background.</p>\n\n<p>For Image preprocessing and have used script and info from this awesome kernel <a href=\"https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping\">https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping</a> \nThank you <a href=\"/ratthachat\">@ratthachat</a>  </p>\n\n<p>The script basically removes all the black background. </p>\n\n<p>```\ndef resize_to(img, targ_sz:int, use_min:bool=False):\n    h,w = img.shape[:2]\n    min_sz = (min if use_min else max)(w,h)\n    ratio = targ_sz/min_sz\n    return int(w*ratio),int(h*ratio)</p>\n\n<p>def crop_image_from_gray(img,tol=7):\n    if img.ndim ==2:\n        mask = img&gt;tol\n        return img[np.ix_(mask.any(1),mask.any(0))]\n    elif img.ndim==3:\n        gray_img = cv2.cvtColor(img, cv2.COLOR_RGB2GRAY)\n        mask = gray_img&gt;tol</p>\n\n<pre><code>    check_shape = img[:,:,0][np.ix_(mask.any(1),mask.any(0))].shape[0]\n    if (check_shape == 0): # image is too dark so that we crop out everything,\n        return img # return original image\n    else:\n        img1=img[:,:,0][np.ix_(mask.any(1),mask.any(0))]\n        img2=img[:,:,1][np.ix_(mask.any(1),mask.any(0))]\n        img3=img[:,:,2][np.ix_(mask.any(1),mask.any(0))]\n        img = np.stack([img1,img2,img3],axis=-1)\n\n    return img\n</code></pre>\n\n<p>```</p>\n\n<p>```\n    def load_ben_color(path, size):\n         image = cv2.imread(str(path))\n          image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n          image = crop_image_from_gray(image)\n          target_size = resize_to(image, size, use_min=True)\n           image = cv2.resize(image, target_size)\n           return PIL.Image.fromarray(image)</p>\n\n<p><code>``\nThe only think I modified I added</code>resize_to<code>function to preserve original aspect ratio so the images look natural.  (My function called</code>load_ben_color` it has nothing to do with ben processing i was just too lazy to change name)</p>\n\n<p><code>ISSUE 2:</code>\nRemoving confusing labels and duplications. </p>\n\n<p>Since we were combining 2015 and 2019 in to one dataset. It was really important to remove all the duplicates and confusing labels to prevent any leak. We have used this awesome kernel to do the job (<a href=\"https://www.kaggle.com/h4211819/more-information-about-duplicate\">https://www.kaggle.com/h4211819/more-information-about-duplicate</a>) Thank you <a href=\"/h4211819\">@h4211819</a></p>\n\n<p><code>ISSUE 3.</code></p>\n\n<p>Solving zoom problem. \nAfter all images were preprocessed to remove black background we observe, also heavily discussed on the forums that 2019 data looked zoomed to the center (below is the example). \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F6ce63736925aa4fde27487609d10edce%2FScreen%20Shot%202019-09-08%20at%2010.00.22%20AM.png?generation=1567954302210617&amp;alt=media\" alt=\"\"></p>\n\n<p>Since we did not know how private data was we decided this problem is best solved using Augmentation random zooms to center, in this way we were making sure that we are not overfitting and also generalizing better. </p>\n\n<h3>AUGMENTATION:</h3>\n\n<p>This was very important part our solution (I posted experiments that I have performed with augmentations here: <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815</a>)</p>\n\n<p>Below is our augmentation: </p>\n\n<p><code>Rotation</code> to 360, <code>flips</code>, <code>zoom</code> up to 1.35x , <code>lightning</code>. This is how its looked:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F4405468830b942fdbd0988b087038c7a%2FScreen%20Shot%202019-09-08%20at%2011.02.17%20AM.png?generation=1567954956948361&amp;alt=media\" alt=\"\"></p>\n\n<h3>TRAINING:</h3>\n\n<p>All the training was done using one cycle learning policy (fastai) with gradual increasing of size , we start with low resolution images and gradually go to higher resolution with reusing weights. You can find details on the my Github notebooks  <a href=\"https://github.com/DrHB/APTOS-2019-GOLD-MEDAL-SOLUTION\">here</a> (I will keep updating) .</p>\n\n<p><code>Model 1</code>:</p>\n\n<p>Results:</p>\n\n<p><code>\nModel:          EfficientNet-B5\nLoss :          MSE\nCV Sratagy   :  5 FOLD StratifiedKFold\n</code>\nAfter combining both datasets 2015 + 2019  and applying above mentioned tricks we immediately observed that our Local CV matches very well public scores.</p>\n\n<p>```\nEXP 740:</p>\n\n<p>5 FOLD CV: 0.823\n       LB: 0.821\n       PB: 0.926\n```</p>\n\n<p><code>Model 2</code>:</p>\n\n<p>Results:</p>\n\n<p>```\nModel:          EfficientNet-B4\nLoss :          MSE\nCV Sratagy   :  5 FOLD StratifiedKFold</p>\n\n<p>```\ndata same as Model 1 </p>\n\n<p>```\nEXP 765:</p>\n\n<p>5 FOLD CV: 0.831\n       LB: 0.816</p>\n\n<pre><code>   PB: 0.927\n</code></pre>\n\n<p>```\nThis looked all promising our Local CV was stable and had good correlation with LB score . But stil we wanted to push further. And last think I decided to try is do pseudo labeling: </p>\n\n<p>I took average prediction for 2019 test data from <code>EXP_740</code> and <code>EXP_765</code>. Added pseudo label data to train data and retrain <code>EXP_740</code> and <code>EXP_765</code>  for 3 epoch below are results</p>\n\n<p>```\nEXP_740_PSD:\n5 FOLD CV: 0.831\n       LB: 0.93\n       PB: 0.930</p>\n\n<p><code>\n</code>\nEXP_765_PSD:\n5 FOLD CV: 0.83\n       LB: 0.835\n       PB: 0.929</p>\n\n<p>```</p>\n\n<p>The combined average of this two models results <code>0.931</code> on PB and <code>0.830</code> on LB =)</p>\n\n<h3>CONCLUSION:</h3>\n\n<p>This was very fun competition, as you can see we did not use many tricks just very systematical logical thinking and trusted our local CV =) </p>",
  "messages": [
    {
      "id": 621481,
      "postDate": "2019-09-08T15:17:40.583Z",
      "content": "<p>First of all I want to thank everyone who participated, shared kernels and was part of activate discussion. </p>\n\n<h2>PROBLEM:</h2>\n\n<p>As you have already heard and read many top team solutions have used very simple 2 STEP approach in fact this was heavily discussed here \n(<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815</a>) and I also partially mentinoed in my kernel (<a href=\"https://www.kaggle.com/drhabib/starter-kernel-for-0-79\">https://www.kaggle.com/drhabib/starter-kernel-for-0-79</a>) :\n<code>\nSTEP 1:  pretrain on 2015 data \nSTEP 2:  finetune on 2019 data \n</code>\nThis approach is good. But When I personally started experimenting I was extremely worried that our local CV was around 0.931 but LB was 0.811. This was a huge gap and we were afraid about shake up. Also just pertaining on old data using 2019 as Validation results in LB 0.75. Which indicated that our model was not generalizing good. </p>\n\n<h2>SOLUTION:</h2>\n\n<p><code>Our idea was very simple instead of training in two diffrent stages we will just combine all the data from 2019 and 2015 and try to work from here to improve our generalization.</code></p>\n\n<h3>PREPROCESSING.</h3>\n\n<p><code>ISSUE 1:</code>\nRemove black background.</p>\n\n<p>For Image preprocessing and have used script and info from this awesome kernel <a href=\"https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping\">https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping</a> \nThank you <a href=\"/ratthachat\">@ratthachat</a>  </p>\n\n<p>The script basically removes all the black background. </p>\n\n<p>```\ndef resize_to(img, targ_sz:int, use_min:bool=False):\n    h,w = img.shape[:2]\n    min_sz = (min if use_min else max)(w,h)\n    ratio = targ_sz/min_sz\n    return int(w*ratio),int(h*ratio)</p>\n\n<p>def crop_image_from_gray(img,tol=7):\n    if img.ndim ==2:\n        mask = img&gt;tol\n        return img[np.ix_(mask.any(1),mask.any(0))]\n    elif img.ndim==3:\n        gray_img = cv2.cvtColor(img, cv2.COLOR_RGB2GRAY)\n        mask = gray_img&gt;tol</p>\n\n<pre><code>    check_shape = img[:,:,0][np.ix_(mask.any(1),mask.any(0))].shape[0]\n    if (check_shape == 0): # image is too dark so that we crop out everything,\n        return img # return original image\n    else:\n        img1=img[:,:,0][np.ix_(mask.any(1),mask.any(0))]\n        img2=img[:,:,1][np.ix_(mask.any(1),mask.any(0))]\n        img3=img[:,:,2][np.ix_(mask.any(1),mask.any(0))]\n        img = np.stack([img1,img2,img3],axis=-1)\n\n    return img\n</code></pre>\n\n<p>```</p>\n\n<p>```\n    def load_ben_color(path, size):\n         image = cv2.imread(str(path))\n          image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n          image = crop_image_from_gray(image)\n          target_size = resize_to(image, size, use_min=True)\n           image = cv2.resize(image, target_size)\n           return PIL.Image.fromarray(image)</p>\n\n<p><code>``\nThe only think I modified I added</code>resize_to<code>function to preserve original aspect ratio so the images look natural.  (My function called</code>load_ben_color` it has nothing to do with ben processing i was just too lazy to change name)</p>\n\n<p><code>ISSUE 2:</code>\nRemoving confusing labels and duplications. </p>\n\n<p>Since we were combining 2015 and 2019 in to one dataset. It was really important to remove all the duplicates and confusing labels to prevent any leak. We have used this awesome kernel to do the job (<a href=\"https://www.kaggle.com/h4211819/more-information-about-duplicate\">https://www.kaggle.com/h4211819/more-information-about-duplicate</a>) Thank you <a href=\"/h4211819\">@h4211819</a></p>\n\n<p><code>ISSUE 3.</code></p>\n\n<p>Solving zoom problem. \nAfter all images were preprocessed to remove black background we observe, also heavily discussed on the forums that 2019 data looked zoomed to the center (below is the example). \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F6ce63736925aa4fde27487609d10edce%2FScreen%20Shot%202019-09-08%20at%2010.00.22%20AM.png?generation=1567954302210617&amp;alt=media\" alt=\"\"></p>\n\n<p>Since we did not know how private data was we decided this problem is best solved using Augmentation random zooms to center, in this way we were making sure that we are not overfitting and also generalizing better. </p>\n\n<h3>AUGMENTATION:</h3>\n\n<p>This was very important part our solution (I posted experiments that I have performed with augmentations here: <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815</a>)</p>\n\n<p>Below is our augmentation: </p>\n\n<p><code>Rotation</code> to 360, <code>flips</code>, <code>zoom</code> up to 1.35x , <code>lightning</code>. This is how its looked:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F4405468830b942fdbd0988b087038c7a%2FScreen%20Shot%202019-09-08%20at%2011.02.17%20AM.png?generation=1567954956948361&amp;alt=media\" alt=\"\"></p>\n\n<h3>TRAINING:</h3>\n\n<p>All the training was done using one cycle learning policy (fastai) with gradual increasing of size , we start with low resolution images and gradually go to higher resolution with reusing weights. You can find details on the my Github notebooks  <a href=\"https://github.com/DrHB/APTOS-2019-GOLD-MEDAL-SOLUTION\">here</a> (I will keep updating) .</p>\n\n<p><code>Model 1</code>:</p>\n\n<p>Results:</p>\n\n<p><code>\nModel:          EfficientNet-B5\nLoss :          MSE\nCV Sratagy   :  5 FOLD StratifiedKFold\n</code>\nAfter combining both datasets 2015 + 2019  and applying above mentioned tricks we immediately observed that our Local CV matches very well public scores.</p>\n\n<p>```\nEXP 740:</p>\n\n<p>5 FOLD CV: 0.823\n       LB: 0.821\n       PB: 0.926\n```</p>\n\n<p><code>Model 2</code>:</p>\n\n<p>Results:</p>\n\n<p>```\nModel:          EfficientNet-B4\nLoss :          MSE\nCV Sratagy   :  5 FOLD StratifiedKFold</p>\n\n<p>```\ndata same as Model 1 </p>\n\n<p>```\nEXP 765:</p>\n\n<p>5 FOLD CV: 0.831\n       LB: 0.816</p>\n\n<pre><code>   PB: 0.927\n</code></pre>\n\n<p>```\nThis looked all promising our Local CV was stable and had good correlation with LB score . But stil we wanted to push further. And last think I decided to try is do pseudo labeling: </p>\n\n<p>I took average prediction for 2019 test data from <code>EXP_740</code> and <code>EXP_765</code>. Added pseudo label data to train data and retrain <code>EXP_740</code> and <code>EXP_765</code>  for 3 epoch below are results</p>\n\n<p>```\nEXP_740_PSD:\n5 FOLD CV: 0.831\n       LB: 0.93\n       PB: 0.930</p>\n\n<p><code>\n</code>\nEXP_765_PSD:\n5 FOLD CV: 0.83\n       LB: 0.835\n       PB: 0.929</p>\n\n<p>```</p>\n\n<p>The combined average of this two models results <code>0.931</code> on PB and <code>0.830</code> on LB =)</p>\n\n<h3>CONCLUSION:</h3>\n\n<p>This was very fun competition, as you can see we did not use many tricks just very systematical logical thinking and trusted our local CV =) </p>",
      "rawMarkdown": "First of all I want to thank everyone who participated, shared kernels and was part of activate discussion. \n\n## PROBLEM:\n\nAs you have already heard and read many top team solutions have used very simple 2 STEP approach in fact this was heavily discussed here \n(https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815) and I also partially mentinoed in my kernel (https://www.kaggle.com/drhabib/starter-kernel-for-0-79) :\n```\nSTEP 1:  pretrain on 2015 data \nSTEP 2:  finetune on 2019 data \n```\nThis approach is good. But When I personally started experimenting I was extremely worried that our local CV was around 0.931 but LB was 0.811. This was a huge gap and we were afraid about shake up. Also just pertaining on old data using 2019 as Validation results in LB 0.75. Which indicated that our model was not generalizing good. \n\n## SOLUTION:\n`Our idea was very simple instead of training in two diffrent stages we will just combine all the data from 2019 and 2015 and try to work from here to improve our generalization. `\n\n### PREPROCESSING. \n\n`ISSUE 1:`\nRemove black background.\n\nFor Image preprocessing and have used script and info from this awesome kernel https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping \nThank you @ratthachat  \n\nThe script basically removes all the black background. \n\n```\ndef resize_to(img, targ_sz:int, use_min:bool=False):\n    h,w = img.shape[:2]\n    min_sz = (min if use_min else max)(w,h)\n    ratio = targ_sz/min_sz\n    return int(w*ratio),int(h*ratio)\n\ndef crop_image_from_gray(img,tol=7):\n    if img.ndim ==2:\n        mask = img&gt;tol\n        return img[np.ix_(mask.any(1),mask.any(0))]\n    elif img.ndim==3:\n        gray_img = cv2.cvtColor(img, cv2.COLOR_RGB2GRAY)\n        mask = gray_img&gt;tol\n        \n        check_shape = img[:,:,0][np.ix_(mask.any(1),mask.any(0))].shape[0]\n        if (check_shape == 0): # image is too dark so that we crop out everything,\n            return img # return original image\n        else:\n            img1=img[:,:,0][np.ix_(mask.any(1),mask.any(0))]\n            img2=img[:,:,1][np.ix_(mask.any(1),mask.any(0))]\n            img3=img[:,:,2][np.ix_(mask.any(1),mask.any(0))]\n            img = np.stack([img1,img2,img3],axis=-1)\n\n        return img\n```\n    \n```\n    def load_ben_color(path, size):\n         image = cv2.imread(str(path))\n          image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n          image = crop_image_from_gray(image)\n          target_size = resize_to(image, size, use_min=True)\n           image = cv2.resize(image, target_size)\n           return PIL.Image.fromarray(image)\n\n```\nThe only think I modified I added `resize_to` function to preserve original aspect ratio so the images look natural.  (My function called `load_ben_color` it has nothing to do with ben processing i was just too lazy to change name)\n\n\n`ISSUE 2:`\nRemoving confusing labels and duplications. \n\nSince we were combining 2015 and 2019 in to one dataset. It was really important to remove all the duplicates and confusing labels to prevent any leak. We have used this awesome kernel to do the job (https://www.kaggle.com/h4211819/more-information-about-duplicate) Thank you @h4211819\n\n`ISSUE 3. `\n\nSolving zoom problem. \nAfter all images were preprocessed to remove black background we observe, also heavily discussed on the forums that 2019 data looked zoomed to the center (below is the example). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F6ce63736925aa4fde27487609d10edce%2FScreen%20Shot%202019-09-08%20at%2010.00.22%20AM.png?generation=1567954302210617&amp;alt=media)\n\nSince we did not know how private data was we decided this problem is best solved using Augmentation random zooms to center, in this way we were making sure that we are not overfitting and also generalizing better. \n\n### AUGMENTATION:\n\nThis was very important part our solution (I posted experiments that I have performed with augmentations here: https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815)\n\nBelow is our augmentation: \n\n`Rotation` to 360, `flips`, ` zoom` up to 1.35x , `lightning`. This is how its looked:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F4405468830b942fdbd0988b087038c7a%2FScreen%20Shot%202019-09-08%20at%2011.02.17%20AM.png?generation=1567954956948361&amp;alt=media)\n\n\n### TRAINING:\n\nAll the training was done using one cycle learning policy (fastai) with gradual increasing of size , we start with low resolution images and gradually go to higher resolution with reusing weights. You can find details on the my Github notebooks  [here](https://github.com/DrHB/APTOS-2019-GOLD-MEDAL-SOLUTION) (I will keep updating) .\n\n`Model 1`:\n\nResults:\n\n```\nModel:          EfficientNet-B5\nLoss :          MSE\nCV Sratagy   :  5 FOLD StratifiedKFold\n```\nAfter combining both datasets 2015 + 2019  and applying above mentioned tricks we immediately observed that our Local CV matches very well public scores.\n\n```\nEXP 740:\n\n5 FOLD CV: 0.823\n       LB: 0.821\n       PB: 0.926\n```\n\n`Model 2`:\n\nResults:\n\n```\nModel:          EfficientNet-B4\nLoss :          MSE\nCV Sratagy   :  5 FOLD StratifiedKFold\n\n```\ndata same as Model 1 \n\n```\nEXP 765:\n\n5 FOLD CV: 0.831\n       LB: 0.816\n\n       PB: 0.927\n```\nThis looked all promising our Local CV was stable and had good correlation with LB score . But stil we wanted to push further. And last think I decided to try is do pseudo labeling: \n\nI took average prediction for 2019 test data from `EXP_740` and `EXP_765`. Added pseudo label data to train data and retrain `EXP_740` and `EXP_765`  for 3 epoch below are results\n\n```\nEXP_740_PSD:\n5 FOLD CV: 0.831\n       LB: 0.93\n       PB: 0.930\n\n```\n```\nEXP_765_PSD:\n5 FOLD CV: 0.83\n       LB: 0.835\n       PB: 0.929\n\n```\n\nThe combined average of this two models results `0.931` on PB and `0.830` on LB =)\n\n### CONCLUSION:\nThis was very fun competition, as you can see we did not use many tricks just very systematical logical thinking and trusted our local CV =) ",
      "votes": 90
    },
    {
      "id": 623550,
      "postDate": "2019-09-11T04:35:10.220Z",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> Congratulations〜. During the competition, we learned a lot from your discussions. Thanks.</p>",
      "rawMarkdown": "@drhabib Congratulations〜. During the competition, we learned a lot from your discussions. Thanks.",
      "votes": 1
    },
    {
      "id": 623524,
      "postDate": "2019-09-11T03:54:04.213Z",
      "content": "<p>Wow, I realized that you become master in this game, congratulations!</p>",
      "rawMarkdown": "Wow, I realized that you become master in this game, congratulations!",
      "votes": 1,
      "replies": [
        {
          "id": 623529,
          "postDate": "2019-09-11T04:02:25.423Z",
          "content": "<p>Thank youuuu :) You are next! </p>",
          "rawMarkdown": "Thank youuuu :) You are next! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 623336,
      "postDate": "2019-09-10T19:17:29.787Z",
      "content": "<p>This is such a beautiful solution. And a wonderful write up, thank you so much for sharing.</p>",
      "rawMarkdown": "This is such a beautiful solution. And a wonderful write up, thank you so much for sharing.",
      "votes": 1
    },
    {
      "id": 621766,
      "postDate": "2019-09-08T23:07:18.977Z",
      "content": "<p>Super well written explanation. I hope we will meet new challenges together again very soon. </p>\n\n<p>About this comp, I think everybody all agrees that you are very well deserved. Until next time <a href=\"/drhabib\">@drhabib</a> !!</p>",
      "rawMarkdown": "Super well written explanation. I hope we will meet new challenges together again very soon. \n\nAbout this comp, I think everybody all agrees that you are very well deserved. Until next time @drhabib !!",
      "votes": 1
    },
    {
      "id": 621586,
      "postDate": "2019-09-08T17:20:40.243Z",
      "content": "<p>Nice report, very good to see that you also upload your code to Git!</p>",
      "rawMarkdown": "Nice report, very good to see that you also upload your code to Git!",
      "votes": 1
    },
    {
      "id": 621558,
      "postDate": "2019-09-08T16:52:07.713Z",
      "content": "<p>very well documented and explained..thanks a lot <a href=\"/drhabib\">@drhabib</a> .. learning a lot\nand many many congratulations !!!</p>",
      "rawMarkdown": "very well documented and explained..thanks a lot @drhabib .. learning a lot\nand many many congratulations !!!",
      "votes": 1
    },
    {
      "id": 621513,
      "postDate": "2019-09-08T15:49:26.590Z",
      "content": "<p>Thanks for sharing. As usual, it's informative and well-written. Keep up the good work. :)</p>",
      "rawMarkdown": "Thanks for sharing. As usual, it's informative and well-written. Keep up the good work. :)",
      "votes": 1
    },
    {
      "id": 621502,
      "postDate": "2019-09-08T15:42:21.277Z",
      "content": "<p>Congratz, awesome job. I am still puzzled why public lb pseudo helped private lb because you add mostly 2s and 0s there. We tried private LB pseudo and it did not do too much, maybe because dist is more similar to train.</p>",
      "rawMarkdown": "Congratz, awesome job. I am still puzzled why public lb pseudo helped private lb because you add mostly 2s and 0s there. We tried private LB pseudo and it did not do too much, maybe because dist is more similar to train.",
      "votes": 1,
      "replies": [
        {
          "id": 621517,
          "postDate": "2019-09-08T16:01:36.257Z",
          "content": "<p>Thank you very much..  I dont  have yet explanation for pseudo labels.  Congratulation to your team as well for gold medal!</p>",
          "rawMarkdown": "Thank you very much..  I dont  have yet explanation for pseudo labels.  Congratulation to your team as well for gold medal!",
          "votes": 1
        }
      ]
    },
    {
      "id": 621501,
      "postDate": "2019-09-08T15:41:42.717Z",
      "content": "<p>Thank you for the post. Am I right that pseudo labeling pushed your score from 0.816 to 0.93 (on public) ?</p>",
      "rawMarkdown": "Thank you for the post. Am I right that pseudo labeling pushed your score from 0.816 to 0.93 (on public) ?",
      "votes": 1,
      "replies": [
        {
          "id": 621508,
          "postDate": "2019-09-08T15:45:14.817Z",
          "content": "<p>Have you used for pseudo label only public test?</p>",
          "rawMarkdown": "Have you used for pseudo label only public test?"
        },
        {
          "id": 621520,
          "postDate": "2019-09-08T16:03:33.330Z",
          "content": "<p>We did psuedo labeling for 2019 Test only. It helped us to push from 0.926 to 0.931 on Private =) </p>",
          "rawMarkdown": "We did psuedo labeling for 2019 Test only. It helped us to push from 0.926 to 0.931 on Private =) ",
          "votes": 1
        }
      ]
    },
    {
      "id": 621496,
      "postDate": "2019-09-08T15:34:20.753Z",
      "content": "<p>Congrats and thanks for sharing this solution and your insights during the competition. That helped me a lot!</p>",
      "rawMarkdown": "Congrats and thanks for sharing this solution and your insights during the competition. That helped me a lot!",
      "votes": 1,
      "replies": [
        {
          "id": 621497,
          "postDate": "2019-09-08T15:35:01.867Z",
          "content": "<p>congrats on your gold! =) </p>",
          "rawMarkdown": "congrats on your gold! =) "
        }
      ]
    },
    {
      "id": 621484,
      "postDate": "2019-09-08T15:19:11.883Z",
      "content": "<p>Congrats <a href=\"/drhabib\">@drhabib</a> and thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congrats @drhabib and thanks for sharing your solution overview.",
      "votes": 1,
      "replies": [
        {
          "id": 621487,
          "postDate": "2019-09-08T15:23:23.117Z",
          "content": "<p>thank you =) and you are welcome! =) </p>",
          "rawMarkdown": "thank you =) and you are welcome! =) ",
          "votes": 1
        }
      ]
    },
    {
      "id": 624050,
      "postDate": "2019-09-11T15:13:27.053Z",
      "content": "<p>wow you have 700+ experiments, which I only have 150+😂 </p>",
      "rawMarkdown": "wow you have 700+ experiments, which I only have 150+😂 ",
      "votes": 2,
      "replies": [
        {
          "id": 624051,
          "postDate": "2019-09-11T15:19:24.587Z",
          "content": "<p>this includes experiments in my mind as well 😂 =) </p>",
          "rawMarkdown": "this includes experiments in my mind as well 😂 =) ",
          "votes": 2
        }
      ]
    },
    {
      "id": 622317,
      "postDate": "2019-09-09T13:53:15.560Z",
      "content": "<p>Thank you for sharing this solution. \nI got to learn something very unique from your solution and lots of congratulations on your ranking on LB.</p>",
      "rawMarkdown": "Thank you for sharing this solution. \nI got to learn something very unique from your solution and lots of congratulations on your ranking on LB.",
      "votes": 2
    },
    {
      "id": 622046,
      "postDate": "2019-09-09T08:01:15.813Z",
      "content": "<p>Congratulations ! That shows basics are very important :)</p>",
      "rawMarkdown": "Congratulations ! That shows basics are very important :)",
      "votes": 2
    },
    {
      "id": 621905,
      "postDate": "2019-09-09T04:48:44.503Z",
      "content": "<p>Congratulations...\nGreat Work...\nThanks for Sharing your Approach &amp; Insights... !! <a href=\"/drhabib\">@drhabib</a> </p>",
      "rawMarkdown": "Congratulations...\nGreat Work...\nThanks for Sharing your Approach &amp; Insights... !! @drhabib ",
      "votes": 2
    },
    {
      "id": 621832,
      "postDate": "2019-09-09T02:10:33.473Z",
      "content": "<p>A big THANK YOU to your awesome posts and kernels, I have definitely learned a lot! And congratulations !!!</p>",
      "rawMarkdown": "A big THANK YOU to your awesome posts and kernels, I have definitely learned a lot! And congratulations !!!",
      "votes": 2
    },
    {
      "id": 621796,
      "postDate": "2019-09-09T00:45:03.693Z",
      "content": "<p>Thank you very much. I learned a lot from your post, and really appreciated your insight for this competition. </p>\n\n<p>One thing I am still curious is if you have ever cut your efficient net model, and if you have ever tired concat pooling instead of the final efficient net adaptive pooling.</p>\n\n<p>personally speaking, I did some experiments but didn't push too far, but cutting layer groups and gradual unfreezing seems help but not a lot.</p>\n\n<p>And discriminative lr didn't help much. </p>\n\n<p>Thanks and congratz! </p>",
      "rawMarkdown": "Thank you very much. I learned a lot from your post, and really appreciated your insight for this competition. \n\nOne thing I am still curious is if you have ever cut your efficient net model, and if you have ever tired concat pooling instead of the final efficient net adaptive pooling.\n\npersonally speaking, I did some experiments but didn't push too far, but cutting layer groups and gradual unfreezing seems help but not a lot.\n\nAnd discriminative lr didn't help much. \n\nThanks and congratz! ",
      "votes": 2,
      "replies": [
        {
          "id": 622240,
          "postDate": "2019-09-09T12:24:43.217Z",
          "content": "<p>Hi congrats on your silver =) After some experimenting I found out in my case just replacing last linear layer works the best =) </p>",
          "rawMarkdown": "Hi congrats on your silver =) After some experimenting I found out in my case just replacing last linear layer works the best =) "
        }
      ]
    },
    {
      "id": 621784,
      "postDate": "2019-09-09T00:19:02.807Z",
      "content": "<p>Congrats <a href=\"/drhabib\">@drhabib</a> I learnt a lot from your posts on the discussion boards, congrats on a great result!</p>",
      "rawMarkdown": "Congrats @drhabib I learnt a lot from your posts on the discussion boards, congrats on a great result!",
      "votes": 2
    },
    {
      "id": 1180103,
      "postDate": "2021-02-01T05:13:48.880Z",
      "content": "<p>Sir, can u explain what is local CV, LB and PB??</p>",
      "rawMarkdown": "Sir, can u explain what is local CV, LB and PB??",
      "replies": [
        {
          "id": 1180942,
          "postDate": "2021-02-01T14:45:18.320Z",
          "content": "<p>CV - <code>cross validation</code> (which means my local score on validation data)<br>\nLB - <code>Leaderboard</code> score when I submit to leaderboard<br>\nPB - <code>Privateboard</code> score, after competition ends we can view score on private (which determines final standings) </p>",
          "rawMarkdown": "CV - `cross validation` (which means my local score on validation data)\nLB - `Leaderboard` score when I submit to leaderboard\nPB - `Privateboard` score, after competition ends we can view score on private (which determines final standings) "
        }
      ]
    },
    {
      "id": 788978,
      "postDate": "2020-03-28T08:19:03.697Z",
      "content": "<p>@DrHB\nThanks a lot for your amazing solution. Could you please help me with one observation that is troubling me.\n[1] When i look at your results, as the image size increases the kappa score is decreasing (and val/train loss is increasing). But the overall position on the leaderboard is increasing. What might be the reason for this.</p>\n\n<p>[2] In 4th place solution of this competitions, it was mentioned that the larger efficientnets do well on the lower resolution images to prevent overfitting. So, doesn't the approach of progressively growing image sizes during training violate this notion. Does it make sense to go in opposite direction ?</p>",
      "rawMarkdown": "@DrHB\nThanks a lot for your amazing solution. Could you please help me with one observation that is troubling me.\n[1] When i look at your results, as the image size increases the kappa score is decreasing (and val/train loss is increasing). But the overall position on the leaderboard is increasing. What might be the reason for this.\n\n[2] In 4th place solution of this competitions, it was mentioned that the larger efficientnets do well on the lower resolution images to prevent overfitting. So, doesn't the approach of progressively growing image sizes during training violate this notion. Does it make sense to go in opposite direction ?\n"
    },
    {
      "id": 624139,
      "postDate": "2019-09-11T18:12:33.360Z",
      "content": "<p>Great work <a href=\"/drhabib\">@drhabib</a>. I am not sure if Kaggle will share the solution file for the public test set. Can you please share your outputs for test set, which got you the highest private LB. Would like to compare it against what I have. Thanks.</p>",
      "rawMarkdown": "Great work @drhabib. I am not sure if Kaggle will share the solution file for the public test set. Can you please share your outputs for test set, which got you the highest private LB. Would like to compare it against what I have. Thanks."
    },
    {
      "id": 623204,
      "postDate": "2019-09-10T15:40:21.677Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 621801,
      "postDate": "2019-09-09T00:53:57.553Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 623550,
      "author_name": "krsw",
      "author_url": "",
      "post_date": "2019-09-11T04:35:10.220000",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> Congratulations〜. During the competition, we learned a lot from your discussions. Thanks.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 623524,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2019-09-11T03:54:04.213000",
      "content": "<p>Wow, I realized that you become master in this game, congratulations!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 623529,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-11T04:02:25.423000",
          "content": "<p>Thank youuuu :) You are next! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 623336,
      "author_name": "Hilal Shaath",
      "author_url": "",
      "post_date": "2019-09-10T19:17:29.787000",
      "content": "<p>This is such a beautiful solution. And a wonderful write up, thank you so much for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621766,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-09-08T23:07:18.977000",
      "content": "<p>Super well written explanation. I hope we will meet new challenges together again very soon. </p>\n\n<p>About this comp, I think everybody all agrees that you are very well deserved. Until next time <a href=\"/drhabib\">@drhabib</a> !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621586,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2019-09-08T17:20:40.243000",
      "content": "<p>Nice report, very good to see that you also upload your code to Git!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621558,
      "author_name": "Anurag Trivedi",
      "author_url": "",
      "post_date": "2019-09-08T16:52:07.713000",
      "content": "<p>very well documented and explained..thanks a lot <a href=\"/drhabib\">@drhabib</a> .. learning a lot\nand many many congratulations !!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621513,
      "author_name": "Tahsin Mostafiz",
      "author_url": "",
      "post_date": "2019-09-08T15:49:26.590000",
      "content": "<p>Thanks for sharing. As usual, it's informative and well-written. Keep up the good work. :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621502,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2019-09-08T15:42:21.277000",
      "content": "<p>Congratz, awesome job. I am still puzzled why public lb pseudo helped private lb because you add mostly 2s and 0s there. We tried private LB pseudo and it did not do too much, maybe because dist is more similar to train.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 621517,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-08T16:01:36.257000",
          "content": "<p>Thank you very much..  I dont  have yet explanation for pseudo labels.  Congratulation to your team as well for gold medal!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 621501,
      "author_name": "Insaf Ashrapov",
      "author_url": "",
      "post_date": "2019-09-08T15:41:42.717000",
      "content": "<p>Thank you for the post. Am I right that pseudo labeling pushed your score from 0.816 to 0.93 (on public) ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 621508,
          "author_name": "Insaf Ashrapov",
          "author_url": "",
          "post_date": "2019-09-08T15:45:14.817000",
          "content": "<p>Have you used for pseudo label only public test?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 621520,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-08T16:03:33.330000",
          "content": "<p>We did psuedo labeling for 2019 Test only. It helped us to push from 0.926 to 0.931 on Private =) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 621496,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2019-09-08T15:34:20.753000",
      "content": "<p>Congrats and thanks for sharing this solution and your insights during the competition. That helped me a lot!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 621497,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-08T15:35:01.867000",
          "content": "<p>congrats on your gold! =) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621484,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-09-08T15:19:11.883000",
      "content": "<p>Congrats <a href=\"/drhabib\">@drhabib</a> and thanks for sharing your solution overview.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 621487,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-08T15:23:23.117000",
          "content": "<p>thank you =) and you are welcome! =) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 624050,
      "author_name": "SchenbergZ",
      "author_url": "",
      "post_date": "2019-09-11T15:13:27.053000",
      "content": "<p>wow you have 700+ experiments, which I only have 150+😂 </p>",
      "votes": 2,
      "replies": [
        {
          "id": 624051,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-11T15:19:24.587000",
          "content": "<p>this includes experiments in my mind as well 😂 =) </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 622317,
      "author_name": "Saikat Biswas",
      "author_url": "",
      "post_date": "2019-09-09T13:53:15.560000",
      "content": "<p>Thank you for sharing this solution. \nI got to learn something very unique from your solution and lots of congratulations on your ranking on LB.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 622046,
      "author_name": "Filemon",
      "author_url": "",
      "post_date": "2019-09-09T08:01:15.813000",
      "content": "<p>Congratulations ! That shows basics are very important :)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 621905,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2019-09-09T04:48:44.503000",
      "content": "<p>Congratulations...\nGreat Work...\nThanks for Sharing your Approach &amp; Insights... !! <a href=\"/drhabib\">@drhabib</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 621832,
      "author_name": "Howard Hsu",
      "author_url": "",
      "post_date": "2019-09-09T02:10:33.473000",
      "content": "<p>A big THANK YOU to your awesome posts and kernels, I have definitely learned a lot! And congratulations !!!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 621796,
      "author_name": "Hao He",
      "author_url": "",
      "post_date": "2019-09-09T00:45:03.693000",
      "content": "<p>Thank you very much. I learned a lot from your post, and really appreciated your insight for this competition. </p>\n\n<p>One thing I am still curious is if you have ever cut your efficient net model, and if you have ever tired concat pooling instead of the final efficient net adaptive pooling.</p>\n\n<p>personally speaking, I did some experiments but didn't push too far, but cutting layer groups and gradual unfreezing seems help but not a lot.</p>\n\n<p>And discriminative lr didn't help much. </p>\n\n<p>Thanks and congratz! </p>",
      "votes": 2,
      "replies": [
        {
          "id": 622240,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-09T12:24:43.217000",
          "content": "<p>Hi congrats on your silver =) After some experimenting I found out in my case just replacing last linear layer works the best =) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621784,
      "author_name": "Antonio De Perio",
      "author_url": "",
      "post_date": "2019-09-09T00:19:02.807000",
      "content": "<p>Congrats <a href=\"/drhabib\">@drhabib</a> I learnt a lot from your posts on the discussion boards, congrats on a great result!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1180103,
      "author_name": "dhaval",
      "author_url": "",
      "post_date": "2021-02-01T05:13:48.880000",
      "content": "<p>Sir, can u explain what is local CV, LB and PB??</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1180942,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2021-02-01T14:45:18.320000",
          "content": "<p>CV - <code>cross validation</code> (which means my local score on validation data)<br>\nLB - <code>Leaderboard</code> score when I submit to leaderboard<br>\nPB - <code>Privateboard</code> score, after competition ends we can view score on private (which determines final standings) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 788978,
      "author_name": "overfitting",
      "author_url": "",
      "post_date": "2020-03-28T08:19:03.697000",
      "content": "<p>@DrHB\nThanks a lot for your amazing solution. Could you please help me with one observation that is troubling me.\n[1] When i look at your results, as the image size increases the kappa score is decreasing (and val/train loss is increasing). But the overall position on the leaderboard is increasing. What might be the reason for this.</p>\n\n<p>[2] In 4th place solution of this competitions, it was mentioned that the larger efficientnets do well on the lower resolution images to prevent overfitting. So, doesn't the approach of progressively growing image sizes during training violate this notion. Does it make sense to go in opposite direction ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 624139,
      "author_name": "Basit Riaz Sheikh",
      "author_url": "",
      "post_date": "2019-09-11T18:12:33.360000",
      "content": "<p>Great work <a href=\"/drhabib\">@drhabib</a>. I am not sure if Kaggle will share the solution file for the public test set. Can you please share your outputs for test set, which got you the highest private LB. Would like to compare it against what I have. Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 623204,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T15:40:21.677000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 621801,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T00:53:57.553000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "621481": "First of all I want to thank everyone who participated, shared kernels and was part of activate discussion. \n\n## PROBLEM:\n\nAs you have already heard and read many top team solutions have used very simple 2 STEP approach in fact this was heavily discussed here \n(https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815) and I also partially mentinoed in my kernel (https://www.kaggle.com/drhabib/starter-kernel-for-0-79) :\n```\nSTEP 1:  pretrain on 2015 data \nSTEP 2:  finetune on 2019 data \n```\nThis approach is good. But When I personally started experimenting I was extremely worried that our local CV was around 0.931 but LB was 0.811. This was a huge gap and we were afraid about shake up. Also just pertaining on old data using 2019 as Validation results in LB 0.75. Which indicated that our model was not generalizing good. \n\n## SOLUTION:\n`Our idea was very simple instead of training in two diffrent stages we will just combine all the data from 2019 and 2015 and try to work from here to improve our generalization. `\n\n### PREPROCESSING. \n\n`ISSUE 1:`\nRemove black background.\n\nFor Image preprocessing and have used script and info from this awesome kernel https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping \nThank you @ratthachat  \n\nThe script basically removes all the black background. \n\n```\ndef resize_to(img, targ_sz:int, use_min:bool=False):\n    h,w = img.shape[:2]\n    min_sz = (min if use_min else max)(w,h)\n    ratio = targ_sz/min_sz\n    return int(w*ratio),int(h*ratio)\n\ndef crop_image_from_gray(img,tol=7):\n    if img.ndim ==2:\n        mask = img&gt;tol\n        return img[np.ix_(mask.any(1),mask.any(0))]\n    elif img.ndim==3:\n        gray_img = cv2.cvtColor(img, cv2.COLOR_RGB2GRAY)\n        mask = gray_img&gt;tol\n        \n        check_shape = img[:,:,0][np.ix_(mask.any(1),mask.any(0))].shape[0]\n        if (check_shape == 0): # image is too dark so that we crop out everything,\n            return img # return original image\n        else:\n            img1=img[:,:,0][np.ix_(mask.any(1),mask.any(0))]\n            img2=img[:,:,1][np.ix_(mask.any(1),mask.any(0))]\n            img3=img[:,:,2][np.ix_(mask.any(1),mask.any(0))]\n            img = np.stack([img1,img2,img3],axis=-1)\n\n        return img\n```\n    \n```\n    def load_ben_color(path, size):\n         image = cv2.imread(str(path))\n          image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n          image = crop_image_from_gray(image)\n          target_size = resize_to(image, size, use_min=True)\n           image = cv2.resize(image, target_size)\n           return PIL.Image.fromarray(image)\n\n```\nThe only think I modified I added `resize_to` function to preserve original aspect ratio so the images look natural.  (My function called `load_ben_color` it has nothing to do with ben processing i was just too lazy to change name)\n\n\n`ISSUE 2:`\nRemoving confusing labels and duplications. \n\nSince we were combining 2015 and 2019 in to one dataset. It was really important to remove all the duplicates and confusing labels to prevent any leak. We have used this awesome kernel to do the job (https://www.kaggle.com/h4211819/more-information-about-duplicate) Thank you @h4211819\n\n`ISSUE 3. `\n\nSolving zoom problem. \nAfter all images were preprocessed to remove black background we observe, also heavily discussed on the forums that 2019 data looked zoomed to the center (below is the example). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F6ce63736925aa4fde27487609d10edce%2FScreen%20Shot%202019-09-08%20at%2010.00.22%20AM.png?generation=1567954302210617&amp;alt=media)\n\nSince we did not know how private data was we decided this problem is best solved using Augmentation random zooms to center, in this way we were making sure that we are not overfitting and also generalizing better. \n\n### AUGMENTATION:\n\nThis was very important part our solution (I posted experiments that I have performed with augmentations here: https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815)\n\nBelow is our augmentation: \n\n`Rotation` to 360, `flips`, ` zoom` up to 1.35x , `lightning`. This is how its looked:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F4405468830b942fdbd0988b087038c7a%2FScreen%20Shot%202019-09-08%20at%2011.02.17%20AM.png?generation=1567954956948361&amp;alt=media)\n\n\n### TRAINING:\n\nAll the training was done using one cycle learning policy (fastai) with gradual increasing of size , we start with low resolution images and gradually go to higher resolution with reusing weights. You can find details on the my Github notebooks  [here](https://github.com/DrHB/APTOS-2019-GOLD-MEDAL-SOLUTION) (I will keep updating) .\n\n`Model 1`:\n\nResults:\n\n```\nModel:          EfficientNet-B5\nLoss :          MSE\nCV Sratagy   :  5 FOLD StratifiedKFold\n```\nAfter combining both datasets 2015 + 2019  and applying above mentioned tricks we immediately observed that our Local CV matches very well public scores.\n\n```\nEXP 740:\n\n5 FOLD CV: 0.823\n       LB: 0.821\n       PB: 0.926\n```\n\n`Model 2`:\n\nResults:\n\n```\nModel:          EfficientNet-B4\nLoss :          MSE\nCV Sratagy   :  5 FOLD StratifiedKFold\n\n```\ndata same as Model 1 \n\n```\nEXP 765:\n\n5 FOLD CV: 0.831\n       LB: 0.816\n\n       PB: 0.927\n```\nThis looked all promising our Local CV was stable and had good correlation with LB score . But stil we wanted to push further. And last think I decided to try is do pseudo labeling: \n\nI took average prediction for 2019 test data from `EXP_740` and `EXP_765`. Added pseudo label data to train data and retrain `EXP_740` and `EXP_765`  for 3 epoch below are results\n\n```\nEXP_740_PSD:\n5 FOLD CV: 0.831\n       LB: 0.93\n       PB: 0.930\n\n```\n```\nEXP_765_PSD:\n5 FOLD CV: 0.83\n       LB: 0.835\n       PB: 0.929\n\n```\n\nThe combined average of this two models results `0.931` on PB and `0.830` on LB =)\n\n### CONCLUSION:\nThis was very fun competition, as you can see we did not use many tricks just very systematical logical thinking and trusted our local CV =) ",
    "623550": "@drhabib Congratulations〜. During the competition, we learned a lot from your discussions. Thanks.",
    "623524": "Wow, I realized that you become master in this game, congratulations!",
    "623336": "This is such a beautiful solution. And a wonderful write up, thank you so much for sharing.",
    "621766": "Super well written explanation. I hope we will meet new challenges together again very soon. \n\nAbout this comp, I think everybody all agrees that you are very well deserved. Until next time @drhabib !!",
    "621586": "Nice report, very good to see that you also upload your code to Git!",
    "621558": "very well documented and explained..thanks a lot @drhabib .. learning a lot\nand many many congratulations !!!",
    "621513": "Thanks for sharing. As usual, it's informative and well-written. Keep up the good work. :)",
    "621502": "Congratz, awesome job. I am still puzzled why public lb pseudo helped private lb because you add mostly 2s and 0s there. We tried private LB pseudo and it did not do too much, maybe because dist is more similar to train.",
    "621501": "Thank you for the post. Am I right that pseudo labeling pushed your score from 0.816 to 0.93 (on public) ?",
    "621496": "Congrats and thanks for sharing this solution and your insights during the competition. That helped me a lot!",
    "621484": "Congrats @drhabib and thanks for sharing your solution overview.",
    "624050": "wow you have 700+ experiments, which I only have 150+😂 ",
    "622317": "Thank you for sharing this solution. \nI got to learn something very unique from your solution and lots of congratulations on your ranking on LB.",
    "622046": "Congratulations ! That shows basics are very important :)",
    "621905": "Congratulations...\nGreat Work...\nThanks for Sharing your Approach &amp; Insights... !! @drhabib ",
    "621832": "A big THANK YOU to your awesome posts and kernels, I have definitely learned a lot! And congratulations !!!",
    "621796": "Thank you very much. I learned a lot from your post, and really appreciated your insight for this competition. \n\nOne thing I am still curious is if you have ever cut your efficient net model, and if you have ever tired concat pooling instead of the final efficient net adaptive pooling.\n\npersonally speaking, I did some experiments but didn't push too far, but cutting layer groups and gradual unfreezing seems help but not a lot.\n\nAnd discriminative lr didn't help much. \n\nThanks and congratz! ",
    "621784": "Congrats @drhabib I learnt a lot from your posts on the discussion boards, congrats on a great result!",
    "1180103": "Sir, can u explain what is local CV, LB and PB??",
    "788978": "@DrHB\nThanks a lot for your amazing solution. Could you please help me with one observation that is troubling me.\n[1] When i look at your results, as the image size increases the kappa score is decreasing (and val/train loss is increasing). But the overall position on the leaderboard is increasing. What might be the reason for this.\n\n[2] In 4th place solution of this competitions, it was mentioned that the larger efficientnets do well on the lower resolution images to prevent overfitting. So, doesn't the approach of progressively growing image sizes during training violate this notion. Does it make sense to go in opposite direction ?\n",
    "624139": "Great work @drhabib. I am not sure if Kaggle will share the solution file for the public test set. Can you please share your outputs for test set, which got you the highest private LB. Would like to compare it against what I have. Thanks.",
    "623204": "",
    "621801": ""
  }
}