{
  "id": 107916,
  "title": "17th place solution",
  "url": "/competitions/aptos2019-blindness-detection/discussion/107916",
  "author_name": "Xuan Cao",
  "post_date": "2019-09-08T00:07:16.958000",
  "votes": 40,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Congrats to all the winners. Here is my solution. </p>\n\n<p>My solution is pretty straight forward. I treat it as a regression problem. It is a linear stacking of \n- efficientnet-b0\n- efficientnet-b3\n- efficientnet-b4\n- efficientnet-b5</p>\n\n<p>Other important features: <br>\n- Image size: 380\n- Pretrain Data: 2015 train data, valid on 2015 public test and test on 2015 private test. \n- Finetune Data: 2019 train data (5-fold CV). \n- Augmentation: H/V flip, rotation(0-360), zoom(0 - 0.1), contrast(0-0.2) and brightness(0-0.2). \n- TTA: 8x(H/V flip + zoom(0, 0.1)</p>\n\n<p>The only thing can be considered as trick is the cropping applied: \n- 2015 data is of high quality, so I decide to do center square crop on all of them. \n- For 2019 data, all the 1:1 images are zoom crop (randomly zoom in 0 - 0.18) using 4:3 aspect ratio. </p>\n\n<p>Why different treatments? This is because I found there are only two aspect ratios (1:1 and 4:3) in 2019 data after removing all the black region, and ~99% of 1:1 images are Class 1. When zoomed in at 0.18, the remaining black region in the 1:1 images look exactly the same as the rest of 4:3 images. I hope after this treatment, the model would classify Class 1 images based on the information in the center region, rather than the black region ratio/shape or aspect ratio. </p>\n\n<p>I think the 4:3 images are enlarged images from 1:1 images because when the doctors found potential markers for the illness in one image, they would zoom in to verify their findings. After zooming, they simply saved the enlarged images. Why 4:3? I guess this is the way most software UI are designed. </p>\n\n<p>It is kind of hard for me to accept that I am so close to gold zone, but this is how kaggle works. I guess I need to work harder for my solo gold. </p>\n\n<p>BTW, CV qwk is aligned with private LB perfectly. For my best solution (QWK 0.930), the local CV (QWK) is 0.93456. The golden rule \"Trust local CV\" applies again! </p>",
  "messages": [
    {
      "id": 620749,
      "postDate": "2019-09-08T00:07:16.957Z",
      "content": "<p>Congrats to all the winners. Here is my solution. </p>\n\n<p>My solution is pretty straight forward. I treat it as a regression problem. It is a linear stacking of \n- efficientnet-b0\n- efficientnet-b3\n- efficientnet-b4\n- efficientnet-b5</p>\n\n<p>Other important features: <br>\n- Image size: 380\n- Pretrain Data: 2015 train data, valid on 2015 public test and test on 2015 private test. \n- Finetune Data: 2019 train data (5-fold CV). \n- Augmentation: H/V flip, rotation(0-360), zoom(0 - 0.1), contrast(0-0.2) and brightness(0-0.2). \n- TTA: 8x(H/V flip + zoom(0, 0.1)</p>\n\n<p>The only thing can be considered as trick is the cropping applied: \n- 2015 data is of high quality, so I decide to do center square crop on all of them. \n- For 2019 data, all the 1:1 images are zoom crop (randomly zoom in 0 - 0.18) using 4:3 aspect ratio. </p>\n\n<p>Why different treatments? This is because I found there are only two aspect ratios (1:1 and 4:3) in 2019 data after removing all the black region, and ~99% of 1:1 images are Class 1. When zoomed in at 0.18, the remaining black region in the 1:1 images look exactly the same as the rest of 4:3 images. I hope after this treatment, the model would classify Class 1 images based on the information in the center region, rather than the black region ratio/shape or aspect ratio. </p>\n\n<p>I think the 4:3 images are enlarged images from 1:1 images because when the doctors found potential markers for the illness in one image, they would zoom in to verify their findings. After zooming, they simply saved the enlarged images. Why 4:3? I guess this is the way most software UI are designed. </p>\n\n<p>It is kind of hard for me to accept that I am so close to gold zone, but this is how kaggle works. I guess I need to work harder for my solo gold. </p>\n\n<p>BTW, CV qwk is aligned with private LB perfectly. For my best solution (QWK 0.930), the local CV (QWK) is 0.93456. The golden rule \"Trust local CV\" applies again! </p>",
      "rawMarkdown": "Congrats to all the winners. Here is my solution. \n\nMy solution is pretty straight forward. I treat it as a regression problem. It is a linear stacking of \n- efficientnet-b0\n- efficientnet-b3\n- efficientnet-b4\n- efficientnet-b5\n\nOther important features:  \n- Image size: 380\n- Pretrain Data: 2015 train data, valid on 2015 public test and test on 2015 private test. \n- Finetune Data: 2019 train data (5-fold CV). \n- Augmentation: H/V flip, rotation(0-360), zoom(0 - 0.1), contrast(0-0.2) and brightness(0-0.2). \n- TTA: 8x(H/V flip + zoom(0, 0.1)\n\n\nThe only thing can be considered as trick is the cropping applied: \n- 2015 data is of high quality, so I decide to do center square crop on all of them. \n- For 2019 data, all the 1:1 images are zoom crop (randomly zoom in 0 - 0.18) using 4:3 aspect ratio. \n\n\nWhy different treatments? This is because I found there are only two aspect ratios (1:1 and 4:3) in 2019 data after removing all the black region, and ~99% of 1:1 images are Class 1. When zoomed in at 0.18, the remaining black region in the 1:1 images look exactly the same as the rest of 4:3 images. I hope after this treatment, the model would classify Class 1 images based on the information in the center region, rather than the black region ratio/shape or aspect ratio. \n\nI think the 4:3 images are enlarged images from 1:1 images because when the doctors found potential markers for the illness in one image, they would zoom in to verify their findings. After zooming, they simply saved the enlarged images. Why 4:3? I guess this is the way most software UI are designed. \n\n\nIt is kind of hard for me to accept that I am so close to gold zone, but this is how kaggle works. I guess I need to work harder for my solo gold. \n\n\nBTW, CV qwk is aligned with private LB perfectly. For my best solution (QWK 0.930), the local CV (QWK) is 0.93456. The golden rule \"Trust local CV\" applies again! ",
      "votes": 40
    },
    {
      "id": 620782,
      "postDate": "2019-09-08T00:51:22.943Z",
      "content": "<p>The diagnosis class distribution of 1:1 images and 4:3 images for 2015 and 2019 data respectively. This is another reason why there is almost no shake in 2015 competition besides the public test size. . </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Fad42ba2dcd55d29a1a6383eb3ef1ea2a%2FWeChat%20Image_20190907174956.png?generation=1567903823182280&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The diagnosis class distribution of 1:1 images and 4:3 images for 2015 and 2019 data respectively. This is another reason why there is almost no shake in 2015 competition besides the public test size. . \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Fad42ba2dcd55d29a1a6383eb3ef1ea2a%2FWeChat%20Image_20190907174956.png?generation=1567903823182280&amp;alt=media)\n",
      "votes": 3
    },
    {
      "id": 620916,
      "postDate": "2019-09-08T04:11:16.637Z",
      "content": "<p><a href=\"/naivelamb\">@naivelamb</a> , Thanks for sharing your approach!</p>",
      "rawMarkdown": "@naivelamb , Thanks for sharing your approach!",
      "votes": 1
    },
    {
      "id": 620767,
      "postDate": "2019-09-08T00:22:34.190Z",
      "content": "<p>Hi <a href=\"/naivelamb\">@naivelamb</a> !</p>\n\n<p>What kind of image preprocessing have you used?</p>",
      "rawMarkdown": "Hi @naivelamb !\n\nWhat kind of image preprocessing have you used?",
      "votes": 1,
      "replies": [
        {
          "id": 620773,
          "postDate": "2019-09-08T00:30:58.130Z",
          "content": "<p>Removing the black pixels + different cropping methods for 2015 and 2019 data respectively. </p>",
          "rawMarkdown": "Removing the black pixels + different cropping methods for 2015 and 2019 data respectively. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 620753,
      "postDate": "2019-09-08T00:11:27.927Z",
      "content": "<p>Thank you！Your analysis of shake gives me confidence to continue!</p>",
      "rawMarkdown": "Thank you！Your analysis of shake gives me confidence to continue!",
      "votes": 2,
      "replies": [
        {
          "id": 620784,
          "postDate": "2019-09-08T00:53:17.860Z",
          "content": "<p>The shake turns out to be more severe than I expected! </p>\n\n<p>Congrats on your gold medal! </p>",
          "rawMarkdown": "The shake turns out to be more severe than I expected! \n\nCongrats on your gold medal! "
        }
      ]
    },
    {
      "id": 621737,
      "postDate": "2019-09-08T21:35:28.247Z",
      "content": "<p>When you mentioned linear stacking - do you mean taking mean of the 4 model predictions? Also may I know the thresholds you use to classify into each class? </p>",
      "rawMarkdown": "When you mentioned linear stacking - do you mean taking mean of the 4 model predictions? Also may I know the thresholds you use to classify into each class? ",
      "replies": [
        {
          "id": 621789,
          "postDate": "2019-09-09T00:33:58.847Z",
          "content": "<p><code>\nlr = LinearRegrassion()\nlr.fit(Y_pred_valid, Y_true)\nY_pred = lr.predict(Y_pred_test)\n</code></p>",
          "rawMarkdown": "```\nlr = LinearRegrassion()\nlr.fit(Y_pred_valid, Y_true)\nY_pred = lr.predict(Y_pred_test)\n```"
        },
        {
          "id": 622616,
          "postDate": "2019-09-09T22:17:03.250Z",
          "content": "<p>Thanks for the reply! Are Y_pred_valid and Y_true from out of fold or just one single validation set here? After which did you use [0.5, 1.5, 2.5, 3.5) to classify each into respective classes or other forms of cutoff mechanism?</p>",
          "rawMarkdown": "Thanks for the reply! Are Y_pred_valid and Y_true from out of fold or just one single validation set here? After which did you use [0.5, 1.5, 2.5, 3.5) to classify each into respective classes or other forms of cutoff mechanism?"
        },
        {
          "id": 622638,
          "postDate": "2019-09-09T23:19:05.953Z",
          "content": "<p>Y_pred_valid are the out-of-fold predictions. If you have only one fold then just use one fold, otherwise combine them to build the oof over the whole train set. </p>\n\n<p>I only used [0.5, 1.5, 2.5, 3.5] to gives the hard label. I think this can prevent overfitting. </p>",
          "rawMarkdown": "Y_pred_valid are the out-of-fold predictions. If you have only one fold then just use one fold, otherwise combine them to build the oof over the whole train set. \n\nI only used [0.5, 1.5, 2.5, 3.5] to gives the hard label. I think this can prevent overfitting. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 620845,
      "postDate": "2019-09-08T02:18:24.233Z",
      "content": "<p>Congratulations! and thanks for sharing. I like your training strategy based on image analysis :)</p>",
      "rawMarkdown": "Congratulations! and thanks for sharing. I like your training strategy based on image analysis :)"
    },
    {
      "id": 620820,
      "postDate": "2019-09-08T01:39:52.683Z",
      "content": "<p>How much do you think tta helped?</p>",
      "rawMarkdown": "How much do you think tta helped?",
      "replies": [
        {
          "id": 620832,
          "postDate": "2019-09-08T02:02:29.560Z",
          "content": "<p>I don't have results on 2019. My 2015 experiments show that 4xTTA gives ~0.04 boost on private LB (2019). </p>\n\n<p>For my pipeline, I think TTA is a necessary, since I have a random zoom-crop when processing the 1:1 images in 2019. </p>",
          "rawMarkdown": "I don't have results on 2019. My 2015 experiments show that 4xTTA gives ~0.04 boost on private LB (2019). \n\nFor my pipeline, I think TTA is a necessary, since I have a random zoom-crop when processing the 1:1 images in 2019. "
        }
      ]
    },
    {
      "id": 620912,
      "postDate": "2019-09-08T04:07:59.483Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 620782,
      "author_name": "Xuan Cao",
      "author_url": "",
      "post_date": "2019-09-08T00:51:22.943000",
      "content": "<p>The diagnosis class distribution of 1:1 images and 4:3 images for 2015 and 2019 data respectively. This is another reason why there is almost no shake in 2015 competition besides the public test size. . </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Fad42ba2dcd55d29a1a6383eb3ef1ea2a%2FWeChat%20Image_20190907174956.png?generation=1567903823182280&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 620916,
      "author_name": "Kranthi Kumar",
      "author_url": "",
      "post_date": "2019-09-08T04:11:16.637000",
      "content": "<p><a href=\"/naivelamb\">@naivelamb</a> , Thanks for sharing your approach!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 620767,
      "author_name": "Cyr1ll",
      "author_url": "",
      "post_date": "2019-09-08T00:22:34.190000",
      "content": "<p>Hi <a href=\"/naivelamb\">@naivelamb</a> !</p>\n\n<p>What kind of image preprocessing have you used?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 620773,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-09-08T00:30:58.130000",
          "content": "<p>Removing the black pixels + different cropping methods for 2015 and 2019 data respectively. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 620753,
      "author_name": "JIANJIAN",
      "author_url": "",
      "post_date": "2019-09-08T00:11:27.927000",
      "content": "<p>Thank you！Your analysis of shake gives me confidence to continue!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 620784,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-09-08T00:53:17.860000",
          "content": "<p>The shake turns out to be more severe than I expected! </p>\n\n<p>Congrats on your gold medal! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621737,
      "author_name": "ChinHuiC",
      "author_url": "",
      "post_date": "2019-09-08T21:35:28.247000",
      "content": "<p>When you mentioned linear stacking - do you mean taking mean of the 4 model predictions? Also may I know the thresholds you use to classify into each class? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 621789,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-09-09T00:33:58.847000",
          "content": "<p><code>\nlr = LinearRegrassion()\nlr.fit(Y_pred_valid, Y_true)\nY_pred = lr.predict(Y_pred_test)\n</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 622616,
          "author_name": "ChinHuiC",
          "author_url": "",
          "post_date": "2019-09-09T22:17:03.250000",
          "content": "<p>Thanks for the reply! Are Y_pred_valid and Y_true from out of fold or just one single validation set here? After which did you use [0.5, 1.5, 2.5, 3.5) to classify each into respective classes or other forms of cutoff mechanism?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 622638,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-09-09T23:19:05.953000",
          "content": "<p>Y_pred_valid are the out-of-fold predictions. If you have only one fold then just use one fold, otherwise combine them to build the oof over the whole train set. </p>\n\n<p>I only used [0.5, 1.5, 2.5, 3.5] to gives the hard label. I think this can prevent overfitting. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 620845,
      "author_name": "Rishabh Agrahari",
      "author_url": "",
      "post_date": "2019-09-08T02:18:24.233000",
      "content": "<p>Congratulations! and thanks for sharing. I like your training strategy based on image analysis :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 620820,
      "author_name": "sh",
      "author_url": "",
      "post_date": "2019-09-08T01:39:52.683000",
      "content": "<p>How much do you think tta helped?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 620832,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-09-08T02:02:29.560000",
          "content": "<p>I don't have results on 2019. My 2015 experiments show that 4xTTA gives ~0.04 boost on private LB (2019). </p>\n\n<p>For my pipeline, I think TTA is a necessary, since I have a random zoom-crop when processing the 1:1 images in 2019. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 620912,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-08T04:07:59.483000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "620749": "Congrats to all the winners. Here is my solution. \n\nMy solution is pretty straight forward. I treat it as a regression problem. It is a linear stacking of \n- efficientnet-b0\n- efficientnet-b3\n- efficientnet-b4\n- efficientnet-b5\n\nOther important features:  \n- Image size: 380\n- Pretrain Data: 2015 train data, valid on 2015 public test and test on 2015 private test. \n- Finetune Data: 2019 train data (5-fold CV). \n- Augmentation: H/V flip, rotation(0-360), zoom(0 - 0.1), contrast(0-0.2) and brightness(0-0.2). \n- TTA: 8x(H/V flip + zoom(0, 0.1)\n\n\nThe only thing can be considered as trick is the cropping applied: \n- 2015 data is of high quality, so I decide to do center square crop on all of them. \n- For 2019 data, all the 1:1 images are zoom crop (randomly zoom in 0 - 0.18) using 4:3 aspect ratio. \n\n\nWhy different treatments? This is because I found there are only two aspect ratios (1:1 and 4:3) in 2019 data after removing all the black region, and ~99% of 1:1 images are Class 1. When zoomed in at 0.18, the remaining black region in the 1:1 images look exactly the same as the rest of 4:3 images. I hope after this treatment, the model would classify Class 1 images based on the information in the center region, rather than the black region ratio/shape or aspect ratio. \n\nI think the 4:3 images are enlarged images from 1:1 images because when the doctors found potential markers for the illness in one image, they would zoom in to verify their findings. After zooming, they simply saved the enlarged images. Why 4:3? I guess this is the way most software UI are designed. \n\n\nIt is kind of hard for me to accept that I am so close to gold zone, but this is how kaggle works. I guess I need to work harder for my solo gold. \n\n\nBTW, CV qwk is aligned with private LB perfectly. For my best solution (QWK 0.930), the local CV (QWK) is 0.93456. The golden rule \"Trust local CV\" applies again! ",
    "620782": "The diagnosis class distribution of 1:1 images and 4:3 images for 2015 and 2019 data respectively. This is another reason why there is almost no shake in 2015 competition besides the public test size. . \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F999074%2Fad42ba2dcd55d29a1a6383eb3ef1ea2a%2FWeChat%20Image_20190907174956.png?generation=1567903823182280&amp;alt=media)\n",
    "620916": "@naivelamb , Thanks for sharing your approach!",
    "620767": "Hi @naivelamb !\n\nWhat kind of image preprocessing have you used?",
    "620753": "Thank you！Your analysis of shake gives me confidence to continue!",
    "621737": "When you mentioned linear stacking - do you mean taking mean of the 4 model predictions? Also may I know the thresholds you use to classify into each class? ",
    "620845": "Congratulations! and thanks for sharing. I like your training strategy based on image analysis :)",
    "620820": "How much do you think tta helped?",
    "620912": ""
  }
}