{
  "id": 109401,
  "title": "112th place solution (Fastai student)",
  "url": "/competitions/aptos2019-blindness-detection/discussion/109401",
  "author_name": "Hao He",
  "post_date": "2019-09-18T22:38:42.629000",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I*<em>Introduction</em>*: It's my first competition after finishing the fastai class. As there are lots of great solutions posted, I will keep my simple and short.</p>\n\n<p><strong>Validation Construction (didn't work out)</strong></p>\n\n<p>Model used: Resnet18, Resnet34, Resnet50\nValidation set: Random split, Stratify split, Mixup test and train to carefully select those closed (Threshold 0.65, for prediction that lower than 0.65 confidence, consider similar to test)</p>\n\n<p>The idea is based on Fastai ML class, use three different models, create different validation set, use validation score and LB score to plot to see if it makes sense.</p>\n\n<p>The goal is if there is one validation set that can have consistent local / and LB score (linearly as model complexity went up)</p>\n\n<p>The result is not pretty, the gap between validation set and LB is still huge. But besides later random seed selection, I also trained some model base on the validation set I create. </p>\n\n<p><strong>Models</strong>:\nUsing the regression idea, as if the target is 4, predicting 1 should be punished more for predicting 3. I think regression + clipping works fine for the case. Also thanks for <a href=\"/abhishek\">@abhishek</a> for his great kernel of clipping. \n<a href=\"https://www.kaggle.com/abhishek/very-simple-pytorch-training-0-59\">https://www.kaggle.com/abhishek/very-simple-pytorch-training-0-59</a></p>\n\n<p>Model used are:\nResnet 50, 101\nResNext 101 32*16d\nEfficient Net B0-B4</p>\n\n<p>Base on my test results, Efficient Net B2 and B4 are giving the best results. I am able to have LB around 0.79+ with baseline model (fastai default argumentation, lr=2e-4, seed=42, split = 0.1)</p>\n\n<p>I start to apply progressive image resizing. Surprisingly, it didn't improve much score (size-&gt;128 to size-&gt;244). Therefore, based on <a href=\"/drhabib\">@drhabib</a> and <a href=\"/taindow\">@taindow</a> discussion thread, I set my image size = (256,256). In fastai, if you pass tuple for image size, it will use Squish instead of Crop</p>\n\n<p>I started to pre-train 2015 data, used ImageNet and Instragram pretrained weights, use 2019 data as validation set to monitor model performance. Also, used Ben's cropping method base on the great kernel <a href=\"https://www.kaggle.com/ratthachat/aptos-eye-preprocessing-in-diabetic-retinopathy\">https://www.kaggle.com/ratthachat/aptos-eye-preprocessing-in-diabetic-retinopathy</a></p>\n\n<p>However, I kept two pretrained models, one with cropping, one without cropping. </p>\n\n<p>From here, I can have single fold efficient net score LB 0.801 in 8 epochs. </p>\n\n<p>I started to test rAdam, different with other posted in the discussion thread, rAdam with fastai lr_find () use suggested lr 10 times smaller gives me single pre-trained model 0.800 LB score (and it is very solid)</p>\n\n<p>I also tried splitting efficient net model at ._fc layer, ._conv_head layer, and find the middle filter change layer split at middle, ._conv_head to make it 3 layer groups then apply discriminative lr. </p>\n\n<p>It didn't work out. </p>\n\n<p>The final thing I tried is remove ._conv_head, and stick fastai concat pooling on top of the efficient net body. The model performed in a very strange way. I don't need to train the body, only the concat pooling head giving me 0.78+ LB, but if I fine-tune the entire model, the LB dropped to 0.74+. Therefore I decided to use consistent lr for all layers</p>\n\n<p><strong>Ensemble</strong></p>\n\n<p>5 fold Efficient Net b4 gives me 0.804 LB score.</p>\n\n<p>I finally ensembled all my 0.8+ LB score model, boosted to 0.809.\nCarefully select some high scoring models base on different cropping, complexity, finally got me to 0.814 LB.</p>\n\n<p>It's a fun competition overall, have been studied a lot from the community. And it is my first competition after 9 months of study deep learning. I can't image that I started at 2018 December with 0 experience of ML / DL, 0 experience of python. 9 months later got my first silver at Kaggle.</p>\n\n<p>Special thanks for the kaggle / fastai community! </p>\n\n<p>All the best, </p>",
  "messages": [
    {
      "id": 629543,
      "postDate": "2019-09-18T22:38:42.630Z",
      "content": "<p>I*<em>Introduction</em>*: It's my first competition after finishing the fastai class. As there are lots of great solutions posted, I will keep my simple and short.</p>\n\n<p><strong>Validation Construction (didn't work out)</strong></p>\n\n<p>Model used: Resnet18, Resnet34, Resnet50\nValidation set: Random split, Stratify split, Mixup test and train to carefully select those closed (Threshold 0.65, for prediction that lower than 0.65 confidence, consider similar to test)</p>\n\n<p>The idea is based on Fastai ML class, use three different models, create different validation set, use validation score and LB score to plot to see if it makes sense.</p>\n\n<p>The goal is if there is one validation set that can have consistent local / and LB score (linearly as model complexity went up)</p>\n\n<p>The result is not pretty, the gap between validation set and LB is still huge. But besides later random seed selection, I also trained some model base on the validation set I create. </p>\n\n<p><strong>Models</strong>:\nUsing the regression idea, as if the target is 4, predicting 1 should be punished more for predicting 3. I think regression + clipping works fine for the case. Also thanks for <a href=\"/abhishek\">@abhishek</a> for his great kernel of clipping. \n<a href=\"https://www.kaggle.com/abhishek/very-simple-pytorch-training-0-59\">https://www.kaggle.com/abhishek/very-simple-pytorch-training-0-59</a></p>\n\n<p>Model used are:\nResnet 50, 101\nResNext 101 32*16d\nEfficient Net B0-B4</p>\n\n<p>Base on my test results, Efficient Net B2 and B4 are giving the best results. I am able to have LB around 0.79+ with baseline model (fastai default argumentation, lr=2e-4, seed=42, split = 0.1)</p>\n\n<p>I start to apply progressive image resizing. Surprisingly, it didn't improve much score (size-&gt;128 to size-&gt;244). Therefore, based on <a href=\"/drhabib\">@drhabib</a> and <a href=\"/taindow\">@taindow</a> discussion thread, I set my image size = (256,256). In fastai, if you pass tuple for image size, it will use Squish instead of Crop</p>\n\n<p>I started to pre-train 2015 data, used ImageNet and Instragram pretrained weights, use 2019 data as validation set to monitor model performance. Also, used Ben's cropping method base on the great kernel <a href=\"https://www.kaggle.com/ratthachat/aptos-eye-preprocessing-in-diabetic-retinopathy\">https://www.kaggle.com/ratthachat/aptos-eye-preprocessing-in-diabetic-retinopathy</a></p>\n\n<p>However, I kept two pretrained models, one with cropping, one without cropping. </p>\n\n<p>From here, I can have single fold efficient net score LB 0.801 in 8 epochs. </p>\n\n<p>I started to test rAdam, different with other posted in the discussion thread, rAdam with fastai lr_find () use suggested lr 10 times smaller gives me single pre-trained model 0.800 LB score (and it is very solid)</p>\n\n<p>I also tried splitting efficient net model at ._fc layer, ._conv_head layer, and find the middle filter change layer split at middle, ._conv_head to make it 3 layer groups then apply discriminative lr. </p>\n\n<p>It didn't work out. </p>\n\n<p>The final thing I tried is remove ._conv_head, and stick fastai concat pooling on top of the efficient net body. The model performed in a very strange way. I don't need to train the body, only the concat pooling head giving me 0.78+ LB, but if I fine-tune the entire model, the LB dropped to 0.74+. Therefore I decided to use consistent lr for all layers</p>\n\n<p><strong>Ensemble</strong></p>\n\n<p>5 fold Efficient Net b4 gives me 0.804 LB score.</p>\n\n<p>I finally ensembled all my 0.8+ LB score model, boosted to 0.809.\nCarefully select some high scoring models base on different cropping, complexity, finally got me to 0.814 LB.</p>\n\n<p>It's a fun competition overall, have been studied a lot from the community. And it is my first competition after 9 months of study deep learning. I can't image that I started at 2018 December with 0 experience of ML / DL, 0 experience of python. 9 months later got my first silver at Kaggle.</p>\n\n<p>Special thanks for the kaggle / fastai community! </p>\n\n<p>All the best, </p>",
      "rawMarkdown": "I**Introduction**: It's my first competition after finishing the fastai class. As there are lots of great solutions posted, I will keep my simple and short.\n\n**Validation Construction (didn't work out)**\n\nModel used: Resnet18, Resnet34, Resnet50\nValidation set: Random split, Stratify split, Mixup test and train to carefully select those closed (Threshold 0.65, for prediction that lower than 0.65 confidence, consider similar to test)\n\nThe idea is based on Fastai ML class, use three different models, create different validation set, use validation score and LB score to plot to see if it makes sense.\n\nThe goal is if there is one validation set that can have consistent local / and LB score (linearly as model complexity went up)\n\nThe result is not pretty, the gap between validation set and LB is still huge. But besides later random seed selection, I also trained some model base on the validation set I create. \n\n**Models**:\nUsing the regression idea, as if the target is 4, predicting 1 should be punished more for predicting 3. I think regression + clipping works fine for the case. Also thanks for @abhishek for his great kernel of clipping. \nhttps://www.kaggle.com/abhishek/very-simple-pytorch-training-0-59\n\nModel used are:\nResnet 50, 101\nResNext 101 32*16d\nEfficient Net B0-B4\n\nBase on my test results, Efficient Net B2 and B4 are giving the best results. I am able to have LB around 0.79+ with baseline model (fastai default argumentation, lr=2e-4, seed=42, split = 0.1)\n\nI start to apply progressive image resizing. Surprisingly, it didn't improve much score (size-&gt;128 to size-&gt;244). Therefore, based on @drhabib and @taindow discussion thread, I set my image size = (256,256). In fastai, if you pass tuple for image size, it will use Squish instead of Crop\n\nI started to pre-train 2015 data, used ImageNet and Instragram pretrained weights, use 2019 data as validation set to monitor model performance. Also, used Ben's cropping method base on the great kernel https://www.kaggle.com/ratthachat/aptos-eye-preprocessing-in-diabetic-retinopathy\n\nHowever, I kept two pretrained models, one with cropping, one without cropping. \n\nFrom here, I can have single fold efficient net score LB 0.801 in 8 epochs. \n\nI started to test rAdam, different with other posted in the discussion thread, rAdam with fastai lr_find () use suggested lr 10 times smaller gives me single pre-trained model 0.800 LB score (and it is very solid)\n\nI also tried splitting efficient net model at ._fc layer, ._conv_head layer, and find the middle filter change layer split at middle, ._conv_head to make it 3 layer groups then apply discriminative lr. \n\nIt didn't work out. \n\nThe final thing I tried is remove ._conv_head, and stick fastai concat pooling on top of the efficient net body. The model performed in a very strange way. I don't need to train the body, only the concat pooling head giving me 0.78+ LB, but if I fine-tune the entire model, the LB dropped to 0.74+. Therefore I decided to use consistent lr for all layers\n\n**Ensemble**\n\n5 fold Efficient Net b4 gives me 0.804 LB score.\n\nI finally ensembled all my 0.8+ LB score model, boosted to 0.809.\nCarefully select some high scoring models base on different cropping, complexity, finally got me to 0.814 LB.\n\nIt's a fun competition overall, have been studied a lot from the community. And it is my first competition after 9 months of study deep learning. I can't image that I started at 2018 December with 0 experience of ML / DL, 0 experience of python. 9 months later got my first silver at Kaggle.\n\nSpecial thanks for the kaggle / fastai community! \n\nAll the best, ",
      "votes": 6
    },
    {
      "id": 1000992,
      "postDate": "2020-09-07T01:44:30.477Z",
      "content": "<p>Thanks for sharing your approach. I learned quite much from you as a beginner in this field. </p>",
      "rawMarkdown": "Thanks for sharing your approach. I learned quite much from you as a beginner in this field. "
    },
    {
      "id": 939078,
      "postDate": "2020-07-22T02:41:09.200Z",
      "content": "<p>Do you have a label for the test set?</p>",
      "rawMarkdown": "Do you have a label for the test set?"
    },
    {
      "id": 629706,
      "postDate": "2019-09-19T04:55:38.403Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1000992,
      "author_name": "Tan Pei Seng",
      "author_url": "",
      "post_date": "2020-09-07T01:44:30.477000",
      "content": "<p>Thanks for sharing your approach. I learned quite much from you as a beginner in this field. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 939078,
      "author_name": "YUANdong1018",
      "author_url": "",
      "post_date": "2020-07-22T02:41:09.200000",
      "content": "<p>Do you have a label for the test set?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 629706,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-19T04:55:38.403000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "629543": "I**Introduction**: It's my first competition after finishing the fastai class. As there are lots of great solutions posted, I will keep my simple and short.\n\n**Validation Construction (didn't work out)**\n\nModel used: Resnet18, Resnet34, Resnet50\nValidation set: Random split, Stratify split, Mixup test and train to carefully select those closed (Threshold 0.65, for prediction that lower than 0.65 confidence, consider similar to test)\n\nThe idea is based on Fastai ML class, use three different models, create different validation set, use validation score and LB score to plot to see if it makes sense.\n\nThe goal is if there is one validation set that can have consistent local / and LB score (linearly as model complexity went up)\n\nThe result is not pretty, the gap between validation set and LB is still huge. But besides later random seed selection, I also trained some model base on the validation set I create. \n\n**Models**:\nUsing the regression idea, as if the target is 4, predicting 1 should be punished more for predicting 3. I think regression + clipping works fine for the case. Also thanks for @abhishek for his great kernel of clipping. \nhttps://www.kaggle.com/abhishek/very-simple-pytorch-training-0-59\n\nModel used are:\nResnet 50, 101\nResNext 101 32*16d\nEfficient Net B0-B4\n\nBase on my test results, Efficient Net B2 and B4 are giving the best results. I am able to have LB around 0.79+ with baseline model (fastai default argumentation, lr=2e-4, seed=42, split = 0.1)\n\nI start to apply progressive image resizing. Surprisingly, it didn't improve much score (size-&gt;128 to size-&gt;244). Therefore, based on @drhabib and @taindow discussion thread, I set my image size = (256,256). In fastai, if you pass tuple for image size, it will use Squish instead of Crop\n\nI started to pre-train 2015 data, used ImageNet and Instragram pretrained weights, use 2019 data as validation set to monitor model performance. Also, used Ben's cropping method base on the great kernel https://www.kaggle.com/ratthachat/aptos-eye-preprocessing-in-diabetic-retinopathy\n\nHowever, I kept two pretrained models, one with cropping, one without cropping. \n\nFrom here, I can have single fold efficient net score LB 0.801 in 8 epochs. \n\nI started to test rAdam, different with other posted in the discussion thread, rAdam with fastai lr_find () use suggested lr 10 times smaller gives me single pre-trained model 0.800 LB score (and it is very solid)\n\nI also tried splitting efficient net model at ._fc layer, ._conv_head layer, and find the middle filter change layer split at middle, ._conv_head to make it 3 layer groups then apply discriminative lr. \n\nIt didn't work out. \n\nThe final thing I tried is remove ._conv_head, and stick fastai concat pooling on top of the efficient net body. The model performed in a very strange way. I don't need to train the body, only the concat pooling head giving me 0.78+ LB, but if I fine-tune the entire model, the LB dropped to 0.74+. Therefore I decided to use consistent lr for all layers\n\n**Ensemble**\n\n5 fold Efficient Net b4 gives me 0.804 LB score.\n\nI finally ensembled all my 0.8+ LB score model, boosted to 0.809.\nCarefully select some high scoring models base on different cropping, complexity, finally got me to 0.814 LB.\n\nIt's a fun competition overall, have been studied a lot from the community. And it is my first competition after 9 months of study deep learning. I can't image that I started at 2018 December with 0 experience of ML / DL, 0 experience of python. 9 months later got my first silver at Kaggle.\n\nSpecial thanks for the kaggle / fastai community! \n\nAll the best, ",
    "1000992": "Thanks for sharing your approach. I learned quite much from you as a beginner in this field. ",
    "939078": "Do you have a label for the test set?",
    "629706": ""
  }
}