{
  "id": 107995,
  "title": "12th place solution",
  "url": "/competitions/aptos2019-blindness-detection/discussion/107995",
  "author_name": "JIANJIAN",
  "post_date": "2019-09-08T10:50:49.998000",
  "votes": 27,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Congrats to everyone, I am very happy to get my first gold medal! Here is my solution.</p>\n\n<h2>models</h2>\n\n<ul>\n<li>efficientnet-b5</li>\n<li>efficientnet-b6\nMy final submission includes 80% regression models and 20%  classification models.\n<h2>Data preprocessing and Augmentations</h2></li>\n</ul>\n\n<p>I used two sets of data_crop augmentations:\n- Cropped black background  and keep the aspect ratio resize to 320x320.\n- Random add black background to  image and resize to 288x288.</p>\n\n<p>Both crop methods use: fliplr, flipud, random_angle_rotate, random_erase, random_blur, brightness_shift, color_shift,hue_shift and gaussian_noise.</p>\n\n<h2>Training strategy</h2>\n\n<ul>\n<li>pretrain on 2015 data and finetune on 2019 data, Thaks to DrHB <a href=\"/drhabib\">@drhabib</a> , I get a start from your methods. </li>\n<li>5 fold CV（Select the two or three highest accuracy folds to submit）</li>\n<li>Finetune should preserve the optimizer's parameters from pretrain, it made a big difference to me.</li>\n<li>TTA*2</li>\n<li>threshold: [0.5, 1.5, 2.5, 3.5]</li>\n<li>Classification models using the voting method to ensemble.\n<h2>Pseudo labels</h2></li>\n</ul>\n\n<p>Using pseudo labels give me LB 0.840-&gt;0.842, but I'm worried about overfitting, so I only give a very low weight to the pseudo-tag model. As it turned out, I was wrong.</p>\n\n<h2>Conclusion</h2>\n\n<ul>\n<li>A good pretrain is helpful. </li>\n<li>Pseudo labels is attractive and dangerous, but it's worth trying.</li>\n<li>Hold out until the last minute!!!💪 </li>\n</ul>",
  "messages": [
    {
      "id": 621227,
      "postDate": "2019-09-08T10:50:49.997Z",
      "content": "<p>Congrats to everyone, I am very happy to get my first gold medal! Here is my solution.</p>\n\n<h2>models</h2>\n\n<ul>\n<li>efficientnet-b5</li>\n<li>efficientnet-b6\nMy final submission includes 80% regression models and 20%  classification models.\n<h2>Data preprocessing and Augmentations</h2></li>\n</ul>\n\n<p>I used two sets of data_crop augmentations:\n- Cropped black background  and keep the aspect ratio resize to 320x320.\n- Random add black background to  image and resize to 288x288.</p>\n\n<p>Both crop methods use: fliplr, flipud, random_angle_rotate, random_erase, random_blur, brightness_shift, color_shift,hue_shift and gaussian_noise.</p>\n\n<h2>Training strategy</h2>\n\n<ul>\n<li>pretrain on 2015 data and finetune on 2019 data, Thaks to DrHB <a href=\"/drhabib\">@drhabib</a> , I get a start from your methods. </li>\n<li>5 fold CV（Select the two or three highest accuracy folds to submit）</li>\n<li>Finetune should preserve the optimizer's parameters from pretrain, it made a big difference to me.</li>\n<li>TTA*2</li>\n<li>threshold: [0.5, 1.5, 2.5, 3.5]</li>\n<li>Classification models using the voting method to ensemble.\n<h2>Pseudo labels</h2></li>\n</ul>\n\n<p>Using pseudo labels give me LB 0.840-&gt;0.842, but I'm worried about overfitting, so I only give a very low weight to the pseudo-tag model. As it turned out, I was wrong.</p>\n\n<h2>Conclusion</h2>\n\n<ul>\n<li>A good pretrain is helpful. </li>\n<li>Pseudo labels is attractive and dangerous, but it's worth trying.</li>\n<li>Hold out until the last minute!!!💪 </li>\n</ul>",
      "rawMarkdown": "Congrats to everyone, I am very happy to get my first gold medal! Here is my solution.\n## models\n- efficientnet-b5\n- efficientnet-b6\nMy final submission includes 80% regression models and 20%  classification models.\n## Data preprocessing and Augmentations\nI used two sets of data_crop augmentations:\n- Cropped black background  and keep the aspect ratio resize to 320x320.\n- Random add black background to  image and resize to 288x288.\n\nBoth crop methods use: fliplr, flipud, random_angle_rotate, random_erase, random_blur, brightness_shift, color_shift,hue_shift and gaussian_noise.\n## Training strategy\n- pretrain on 2015 data and finetune on 2019 data, Thaks to DrHB @drhabib , I get a start from your methods. \n- 5 fold CV（Select the two or three highest accuracy folds to submit）\n- Finetune should preserve the optimizer's parameters from pretrain, it made a big difference to me.\n- TTA*2\n- threshold: [0.5, 1.5, 2.5, 3.5]\n-  Classification models using the voting method to ensemble.\n## Pseudo labels\nUsing pseudo labels give me LB 0.840-&gt;0.842, but I'm worried about overfitting, so I only give a very low weight to the pseudo-tag model. As it turned out, I was wrong.\n##Conclusion\n- A good pretrain is helpful. \n-  Pseudo labels is attractive and dangerous, but it's worth trying.\n- Hold out until the last minute!!!💪 \n",
      "votes": 27
    },
    {
      "id": 621379,
      "postDate": "2019-09-08T13:23:30.997Z",
      "content": "<p>Congratulations! New Kaggle Master! 北航吴彦祖666</p>",
      "rawMarkdown": "Congratulations! New Kaggle Master! 北航吴彦祖666",
      "votes": 1
    },
    {
      "id": 621368,
      "postDate": "2019-09-08T13:17:30.587Z",
      "content": "<p>Congrats!! Finetune should preserve the optimizer's parameters from pretrain, it made a big difference to me. 精髓找到了。 And how much does classification model help in ensemble?</p>",
      "rawMarkdown": "Congrats!! Finetune should preserve the optimizer's parameters from pretrain, it made a big difference to me. 精髓找到了。 And how much does classification model help in ensemble?",
      "votes": 1
    },
    {
      "id": 623082,
      "postDate": "2019-09-10T13:17:53.710Z",
      "content": "<p>Congratulations,  北航吴彦祖. Could you tell me more details about how to do transform learning between 2015 data and 2019 data</p>",
      "rawMarkdown": "Congratulations,  北航吴彦祖. Could you tell me more details about how to do transform learning between 2015 data and 2019 data"
    },
    {
      "id": 621551,
      "postDate": "2019-09-08T16:43:50.853Z",
      "content": "<p>Congratulations! </p>",
      "rawMarkdown": "Congratulations! "
    },
    {
      "id": 621512,
      "postDate": "2019-09-08T15:49:12.773Z",
      "content": "<p>Congrats and thanks for sharing your solution!\nMay I ask 2 questions below??\n1. You used default threshold, does that mean it's better than use threshold optimizer?\n2. Can you please tell me how to ensemble regression model and classification model?</p>\n\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Congrats and thanks for sharing your solution!\nMay I ask 2 questions below??\n1. You used default threshold, does that mean it's better than use threshold optimizer?\n2. Can you please tell me how to ensemble regression model and classification model?\n\nThanks in advance!",
      "replies": [
        {
          "id": 621811,
          "postDate": "2019-09-09T01:16:12.977Z",
          "content": "<ol>\n<li>I tried [0.57, 1.27, 2.67, 3.57] to make there more 0 and 2, it gave me a higher LB(0.843) and a lower LP(0.928). For fear of overfitting, I don't choose it as my final submission. I try threshold optimizer on valid set too, it does't help any thing.</li>\n<li>Use blending  to ensemble  regression model, as for classification model,  I used a method similar to voting, then you can ensemble them just like regression model, code:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2156075%2Fceec6ea463266f4e20bcdf03857d7641%2F2019-09-09%209.17.18.png?generation=1567991954111290&amp;alt=media\" alt=\"\"></li>\n</ol>",
          "rawMarkdown": "1. I tried [0.57, 1.27, 2.67, 3.57] to make there more 0 and 2, it gave me a higher LB(0.843) and a lower LP(0.928). For fear of overfitting, I don't choose it as my final submission. I try threshold optimizer on valid set too, it does't help any thing.\n2. Use blending  to ensemble  regression model, as for classification model,  I used a method similar to voting, then you can ensemble them just like regression model, code:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2156075%2Fceec6ea463266f4e20bcdf03857d7641%2F2019-09-09%209.17.18.png?generation=1567991954111290&amp;alt=media)\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 621452,
      "postDate": "2019-09-08T14:27:45.960Z",
      "content": "<p>北航吴彦祖！！Congratulations!</p>",
      "rawMarkdown": "北航吴彦祖！！Congratulations!"
    },
    {
      "id": 621409,
      "postDate": "2019-09-08T13:43:52.070Z",
      "content": "<p>Congrats <a href=\"/buaazijian\">@buaazijian</a> and thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congrats @buaazijian and thanks for sharing your solution overview."
    },
    {
      "id": 621402,
      "postDate": "2019-09-08T13:35:39.077Z",
      "content": "<p>Congratulations! 北航吴彦祖！！！</p>",
      "rawMarkdown": "Congratulations! 北航吴彦祖！！！"
    },
    {
      "id": 621330,
      "postDate": "2019-09-08T12:45:04.450Z",
      "content": "<p>Congratulation JianJian!! You are deserved to be \"Master\" :)</p>",
      "rawMarkdown": "Congratulation JianJian!! You are deserved to be \"Master\" :)",
      "replies": [
        {
          "id": 621354,
          "postDate": "2019-09-08T13:09:36.917Z",
          "content": "<p>Thank you.</p>",
          "rawMarkdown": "Thank you."
        }
      ]
    },
    {
      "id": 621307,
      "postDate": "2019-09-08T12:25:32.090Z",
      "content": "<p>Congratulations! 北航吴彦祖！！</p>",
      "rawMarkdown": "Congratulations! 北航吴彦祖！！",
      "replies": [
        {
          "id": 621321,
          "postDate": "2019-09-08T12:37:36.580Z",
          "content": "<p>谢谢，谢谢</p>",
          "rawMarkdown": "谢谢，谢谢"
        }
      ]
    },
    {
      "id": 621272,
      "postDate": "2019-09-08T11:50:12.297Z",
      "content": "<p>Congrats. What do you mean by <strong>Pseudo labels</strong>?</p>",
      "rawMarkdown": "Congrats. What do you mean by **Pseudo labels**?",
      "replies": [
        {
          "id": 621289,
          "postDate": "2019-09-08T12:08:27.040Z",
          "content": "<p>One method of semi-Supervised Learning,.\n- Use train set to train a model\n- Use there model to lable the test set, and add these data to train set\n- Repeat the above steps </p>",
          "rawMarkdown": " One method of semi-Supervised Learning,.\n- Use train set to train a model\n- Use there model to lable the test set, and add these data to train set\n- Repeat the above steps ",
          "votes": 3
        },
        {
          "id": 621388,
          "postDate": "2019-09-08T13:27:37.630Z",
          "content": "<p>Great explanation. Thank you for responding.</p>",
          "rawMarkdown": "Great explanation. Thank you for responding."
        }
      ]
    },
    {
      "id": 621258,
      "postDate": "2019-09-08T11:35:10.450Z",
      "content": "<p>😍 北航吴彦祖！！</p>",
      "rawMarkdown": "😍 北航吴彦祖！！",
      "replies": [
        {
          "id": 621267,
          "postDate": "2019-09-08T11:42:04.850Z",
          "content": "<p>哈哈哈</p>",
          "rawMarkdown": "哈哈哈"
        }
      ]
    },
    {
      "id": 621943,
      "postDate": "2019-09-09T05:37:08.093Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 621383,
      "postDate": "2019-09-08T13:25:24.980Z",
      "content": "<p>Thank you for sharing! Congrats!</p>",
      "rawMarkdown": "Thank you for sharing! Congrats!"
    }
  ],
  "comments": [
    {
      "id": 621379,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2019-09-08T13:23:30.997000",
      "content": "<p>Congratulations! New Kaggle Master! 北航吴彦祖666</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621368,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "2019-09-08T13:17:30.587000",
      "content": "<p>Congrats!! Finetune should preserve the optimizer's parameters from pretrain, it made a big difference to me. 精髓找到了。 And how much does classification model help in ensemble?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 623082,
      "author_name": "dnfys",
      "author_url": "",
      "post_date": "2019-09-10T13:17:53.710000",
      "content": "<p>Congratulations,  北航吴彦祖. Could you tell me more details about how to do transform learning between 2015 data and 2019 data</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621551,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2019-09-08T16:43:50.853000",
      "content": "<p>Congratulations! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621512,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2019-09-08T15:49:12.773000",
      "content": "<p>Congrats and thanks for sharing your solution!\nMay I ask 2 questions below??\n1. You used default threshold, does that mean it's better than use threshold optimizer?\n2. Can you please tell me how to ensemble regression model and classification model?</p>\n\n<p>Thanks in advance!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 621811,
          "author_name": "JIANJIAN",
          "author_url": "",
          "post_date": "2019-09-09T01:16:12.977000",
          "content": "<ol>\n<li>I tried [0.57, 1.27, 2.67, 3.57] to make there more 0 and 2, it gave me a higher LB(0.843) and a lower LP(0.928). For fear of overfitting, I don't choose it as my final submission. I try threshold optimizer on valid set too, it does't help any thing.</li>\n<li>Use blending  to ensemble  regression model, as for classification model,  I used a method similar to voting, then you can ensemble them just like regression model, code:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2156075%2Fceec6ea463266f4e20bcdf03857d7641%2F2019-09-09%209.17.18.png?generation=1567991954111290&amp;alt=media\" alt=\"\"></li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 621452,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2019-09-08T14:27:45.960000",
      "content": "<p>北航吴彦祖！！Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621409,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-09-08T13:43:52.070000",
      "content": "<p>Congrats <a href=\"/buaazijian\">@buaazijian</a> and thanks for sharing your solution overview.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621402,
      "author_name": "Wang Xinliang",
      "author_url": "",
      "post_date": "2019-09-08T13:35:39.077000",
      "content": "<p>Congratulations! 北航吴彦祖！！！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621330,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-09-08T12:45:04.450000",
      "content": "<p>Congratulation JianJian!! You are deserved to be \"Master\" :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 621354,
          "author_name": "JIANJIAN",
          "author_url": "",
          "post_date": "2019-09-08T13:09:36.917000",
          "content": "<p>Thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621307,
      "author_name": "Chanhu",
      "author_url": "",
      "post_date": "2019-09-08T12:25:32.090000",
      "content": "<p>Congratulations! 北航吴彦祖！！</p>",
      "votes": 0,
      "replies": [
        {
          "id": 621321,
          "author_name": "JIANJIAN",
          "author_url": "",
          "post_date": "2019-09-08T12:37:36.580000",
          "content": "<p>谢谢，谢谢</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621272,
      "author_name": "JM100",
      "author_url": "",
      "post_date": "2019-09-08T11:50:12.297000",
      "content": "<p>Congrats. What do you mean by <strong>Pseudo labels</strong>?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 621289,
          "author_name": "JIANJIAN",
          "author_url": "",
          "post_date": "2019-09-08T12:08:27.040000",
          "content": "<p>One method of semi-Supervised Learning,.\n- Use train set to train a model\n- Use there model to lable the test set, and add these data to train set\n- Repeat the above steps </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 621388,
          "author_name": "JM100",
          "author_url": "",
          "post_date": "2019-09-08T13:27:37.630000",
          "content": "<p>Great explanation. Thank you for responding.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621258,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2019-09-08T11:35:10.450000",
      "content": "<p>😍 北航吴彦祖！！</p>",
      "votes": 0,
      "replies": [
        {
          "id": 621267,
          "author_name": "JIANJIAN",
          "author_url": "",
          "post_date": "2019-09-08T11:42:04.850000",
          "content": "<p>哈哈哈</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621943,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T05:37:08.093000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621383,
      "author_name": "JM100",
      "author_url": "",
      "post_date": "2019-09-08T13:25:24.980000",
      "content": "<p>Thank you for sharing! Congrats!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "621227": "Congrats to everyone, I am very happy to get my first gold medal! Here is my solution.\n## models\n- efficientnet-b5\n- efficientnet-b6\nMy final submission includes 80% regression models and 20%  classification models.\n## Data preprocessing and Augmentations\nI used two sets of data_crop augmentations:\n- Cropped black background  and keep the aspect ratio resize to 320x320.\n- Random add black background to  image and resize to 288x288.\n\nBoth crop methods use: fliplr, flipud, random_angle_rotate, random_erase, random_blur, brightness_shift, color_shift,hue_shift and gaussian_noise.\n## Training strategy\n- pretrain on 2015 data and finetune on 2019 data, Thaks to DrHB @drhabib , I get a start from your methods. \n- 5 fold CV（Select the two or three highest accuracy folds to submit）\n- Finetune should preserve the optimizer's parameters from pretrain, it made a big difference to me.\n- TTA*2\n- threshold: [0.5, 1.5, 2.5, 3.5]\n-  Classification models using the voting method to ensemble.\n## Pseudo labels\nUsing pseudo labels give me LB 0.840-&gt;0.842, but I'm worried about overfitting, so I only give a very low weight to the pseudo-tag model. As it turned out, I was wrong.\n##Conclusion\n- A good pretrain is helpful. \n-  Pseudo labels is attractive and dangerous, but it's worth trying.\n- Hold out until the last minute!!!💪 \n",
    "621379": "Congratulations! New Kaggle Master! 北航吴彦祖666",
    "621368": "Congrats!! Finetune should preserve the optimizer's parameters from pretrain, it made a big difference to me. 精髓找到了。 And how much does classification model help in ensemble?",
    "623082": "Congratulations,  北航吴彦祖. Could you tell me more details about how to do transform learning between 2015 data and 2019 data",
    "621551": "Congratulations! ",
    "621512": "Congrats and thanks for sharing your solution!\nMay I ask 2 questions below??\n1. You used default threshold, does that mean it's better than use threshold optimizer?\n2. Can you please tell me how to ensemble regression model and classification model?\n\nThanks in advance!",
    "621452": "北航吴彦祖！！Congratulations!",
    "621409": "Congrats @buaazijian and thanks for sharing your solution overview.",
    "621402": "Congratulations! 北航吴彦祖！！！",
    "621330": "Congratulation JianJian!! You are deserved to be \"Master\" :)",
    "621307": "Congratulations! 北航吴彦祖！！",
    "621272": "Congrats. What do you mean by **Pseudo labels**?",
    "621258": "😍 北航吴彦祖！！",
    "621943": "",
    "621383": "Thank you for sharing! Congrats!"
  }
}