{
  "id": 108013,
  "title": "99th Place Solution",
  "url": "/competitions/aptos2019-blindness-detection/discussion/108013",
  "author_name": "",
  "post_date": "2019-09-08T13:29:29.870624500Z",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Congratatulations to the winners, and thanks to kaggle, competition sponsor and kernel contributors who shared their insights, it helped us a lot. </p>\n\n<p>Models\nWe tried many different models including effecient-nets, renets, resnexts etc. However the delusional Public LeaderBoard made us to believe that Efficient-net B3 gives the best result (However our B5 submission also got some great results).</p>\n\n<p>Preprocessing\nWe tried Ben's color preprocessing, but it didn't boost score, also changing colour-space to HSV, did no magic so we stopped trying in the middle of the competition. Only circular cropping and resizing was doone</p>\n\n<p>Augmentations\nWe used a lot of augmentations, Flipping, Zooming(1.5x - as test data was found to be (or atleast our hypothesis was) center cropped version of train data), Gaussian Blur, FastAI Lighting.</p>\n\n<p>Training\nRegression Based Task\nPretrain on old data - \nWe Combined the old train and test data and used new train data as validation for pretraining our models. Stopped the training when validation loss stopped decreasing (overfitted for 1-2 epochs). LR started from 1e-03 and used ReduceLRonPlateau.\nTrain on current competition data - \n5 fold stratified Cross Validation was used initially. The training procedure was divided into 3 stages.\nStage 1 - Freezing all except last layer - 10 epochs\nStage 2 - Unfreezing all - 50 epochs\nStage 3-1 - Freezing all but last 3 layers - 20 epochs\nStage 3-2 - Freezing all but last layer - 10 epochs\nThe same procedure was followed for 10 folds. </p>\n\n<p>Post - Training \nFinal stage's 2 models of every fold (total being 30 models) were average ensembled.\nWe did try other ensembling methods but again the delusiuonal Public LB misguided us.\nThe predictions were sorted and 0.2 * 30 (6+6=12 predictions were trimmed from both side) then average ensembling of 22 median predictions.</p>",
  "messages": [
    {
      "id": "621390",
      "postDate": "09/08/2019 13:29:29",
      "content": "<p>Congratatulations to the winners, and thanks to kaggle, competition sponsor and kernel contributors who shared their insights, it helped us a lot. </p>\n\n<p>Models\nWe tried many different models including effecient-nets, renets, resnexts etc. However the delusional Public LeaderBoard made us to believe that Efficient-net B3 gives the best result (However our B5 submission also got some great results).</p>\n\n<p>Preprocessing\nWe tried Ben's color preprocessing, but it didn't boost score, also changing colour-space to HSV, did no magic so we stopped trying in the middle of the competition. Only circular cropping and resizing was doone</p>\n\n<p>Augmentations\nWe used a lot of augmentations, Flipping, Zooming(1.5x - as test data was found to be (or atleast our hypothesis was) center cropped version of train data), Gaussian Blur, FastAI Lighting.</p>\n\n<p>Training\nRegression Based Task\nPretrain on old data - \nWe Combined the old train and test data and used new train data as validation for pretraining our models. Stopped the training when validation loss stopped decreasing (overfitted for 1-2 epochs). LR started from 1e-03 and used ReduceLRonPlateau.\nTrain on current competition data - \n5 fold stratified Cross Validation was used initially. The training procedure was divided into 3 stages.\nStage 1 - Freezing all except last layer - 10 epochs\nStage 2 - Unfreezing all - 50 epochs\nStage 3-1 - Freezing all but last 3 layers - 20 epochs\nStage 3-2 - Freezing all but last layer - 10 epochs\nThe same procedure was followed for 10 folds. </p>\n\n<p>Post - Training \nFinal stage's 2 models of every fold (total being 30 models) were average ensembled.\nWe did try other ensembling methods but again the delusiuonal Public LB misguided us.\nThe predictions were sorted and 0.2 * 30 (6+6=12 predictions were trimmed from both side) then average ensembling of 22 median predictions.</p>",
      "rawMarkdown": "Congratatulations to the winners, and thanks to kaggle, competition sponsor and kernel contributors who shared their insights, it helped us a lot. \n\nModels\nWe tried many different models including effecient-nets, renets, resnexts etc. However the delusional Public LeaderBoard made us to believe that Efficient-net B3 gives the best result (However our B5 submission also got some great results).\n\nPreprocessing\nWe tried Ben's color preprocessing, but it didn't boost score, also changing colour-space to HSV, did no magic so we stopped trying in the middle of the competition. Only circular cropping and resizing was doone\n\nAugmentations\nWe used a lot of augmentations, Flipping, Zooming(1.5x - as test data was found to be (or atleast our hypothesis was) center cropped version of train data), Gaussian Blur, FastAI Lighting.\n\nTraining\nRegression Based Task\nPretrain on old data - \nWe Combined the old train and test data and used new train data as validation for pretraining our models. Stopped the training when validation loss stopped decreasing (overfitted for 1-2 epochs). LR started from 1e-03 and used ReduceLRonPlateau.\nTrain on current competition data - \n5 fold stratified Cross Validation was used initially. The training procedure was divided into 3 stages.\nStage 1 - Freezing all except last layer - 10 epochs\nStage 2 - Unfreezing all - 50 epochs\nStage 3-1 - Freezing all but last 3 layers - 20 epochs\nStage 3-2 - Freezing all but last layer - 10 epochs\nThe same procedure was followed for 10 folds. \n\nPost - Training \nFinal stage's 2 models of every fold (total being 30 models) were average ensembled.\nWe did try other ensembling methods but again the delusiuonal Public LB misguided us.\nThe predictions were sorted and 0.2 * 30 (6+6=12 predictions were trimmed from both side) then average ensembling of 22 median predictions.",
      "votes": null
    },
    {
      "id": "621420",
      "postDate": "09/08/2019 13:53:03",
      "content": "<p>Can you tell why you trained the fine tuning model on new data for so many epochs? Does that not count as overfitting? And it also seems it might had a long running time.</p>",
      "rawMarkdown": "Can you tell why you trained the fine tuning model on new data for so many epochs? Does that not count as overfitting? And it also seems it might had a long running time.",
      "votes": null
    },
    {
      "id": "621423",
      "postDate": "09/08/2019 13:57:40",
      "content": "<p>Initially we were training for 20 epochs instead of 50...later we reduced the learning rate and trained for 30 more epochs and selected the best CV loss models</p>",
      "rawMarkdown": "Initially we were training for 20 epochs instead of 50...later we reduced the learning rate and trained for 30 more epochs and selected the best CV loss models",
      "votes": null
    },
    {
      "id": "621916",
      "postDate": "09/09/2019 04:59:52",
      "content": "<p>Congratulations...\nGreat Work...\nThanks for Sharing your Approach &amp; Insights... !! <a href=\"/carnav0400\">@carnav0400</a> </p>",
      "rawMarkdown": "Congratulations...\nGreat Work...\nThanks for Sharing your Approach &amp; Insights... !! @carnav0400",
      "votes": null
    },
    {
      "id": "621976",
      "postDate": "09/09/2019 06:13:08",
      "content": "<p>We will update our final code soon on this post! We will keep you posted!</p>",
      "rawMarkdown": "We will update our final code soon on this post! We will keep you posted!",
      "votes": null
    },
    {
      "id": "622326",
      "postDate": "09/09/2019 14:02:36",
      "content": "<p>Congratulations.\nThanks for Sharing your Approach</p>",
      "rawMarkdown": "Congratulations.\nThanks for Sharing your Approach",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 621420,
      "author_name": "imnitishng",
      "author_url": "",
      "post_date": "09/08/2019 13:53:03",
      "content": "<p>Can you tell why you trained the fine tuning model on new data for so many epochs? Does that not count as overfitting? And it also seems it might had a long running time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 621423,
          "author_name": "carnav0400",
          "author_url": "",
          "post_date": "09/08/2019 13:57:40",
          "content": "<p>Initially we were training for 20 epochs instead of 50...later we reduced the learning rate and trained for 30 more epochs and selected the best CV loss models</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 621916,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "09/09/2019 04:59:52",
      "content": "<p>Congratulations...\nGreat Work...\nThanks for Sharing your Approach &amp; Insights... !! <a href=\"/carnav0400\">@carnav0400</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 621976,
          "author_name": "ubamba98",
          "author_url": "",
          "post_date": "09/09/2019 06:13:08",
          "content": "<p>We will update our final code soon on this post! We will keep you posted!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 622326,
      "author_name": "jmourad100",
      "author_url": "",
      "post_date": "09/09/2019 14:02:36",
      "content": "<p>Congratulations.\nThanks for Sharing your Approach</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "621390": "Congratatulations to the winners, and thanks to kaggle, competition sponsor and kernel contributors who shared their insights, it helped us a lot. \n\nModels\nWe tried many different models including effecient-nets, renets, resnexts etc. However the delusional Public LeaderBoard made us to believe that Efficient-net B3 gives the best result (However our B5 submission also got some great results).\n\nPreprocessing\nWe tried Ben's color preprocessing, but it didn't boost score, also changing colour-space to HSV, did no magic so we stopped trying in the middle of the competition. Only circular cropping and resizing was doone\n\nAugmentations\nWe used a lot of augmentations, Flipping, Zooming(1.5x - as test data was found to be (or atleast our hypothesis was) center cropped version of train data), Gaussian Blur, FastAI Lighting.\n\nTraining\nRegression Based Task\nPretrain on old data - \nWe Combined the old train and test data and used new train data as validation for pretraining our models. Stopped the training when validation loss stopped decreasing (overfitted for 1-2 epochs). LR started from 1e-03 and used ReduceLRonPlateau.\nTrain on current competition data - \n5 fold stratified Cross Validation was used initially. The training procedure was divided into 3 stages.\nStage 1 - Freezing all except last layer - 10 epochs\nStage 2 - Unfreezing all - 50 epochs\nStage 3-1 - Freezing all but last 3 layers - 20 epochs\nStage 3-2 - Freezing all but last layer - 10 epochs\nThe same procedure was followed for 10 folds. \n\nPost - Training \nFinal stage's 2 models of every fold (total being 30 models) were average ensembled.\nWe did try other ensembling methods but again the delusiuonal Public LB misguided us.\nThe predictions were sorted and 0.2 * 30 (6+6=12 predictions were trimmed from both side) then average ensembling of 22 median predictions.",
    "621420": "Can you tell why you trained the fine tuning model on new data for so many epochs? Does that not count as overfitting? And it also seems it might had a long running time.",
    "621423": "Initially we were training for 20 epochs instead of 50...later we reduced the learning rate and trained for 30 more epochs and selected the best CV loss models",
    "621916": "Congratulations...\nGreat Work...\nThanks for Sharing your Approach &amp; Insights... !! @carnav0400",
    "621976": "We will update our final code soon on this post! We will keep you posted!",
    "622326": "Congratulations.\nThanks for Sharing your Approach"
  },
  "source": "meta"
}