{
  "id": 103023,
  "title": "A paper from Nature (24 July 2019) (2095 × 2095 image size)",
  "url": "/competitions/aptos2019-blindness-detection/discussion/103023",
  "author_name": "",
  "post_date": "2019-08-06T15:22:43.217018700Z",
  "votes": 30,
  "comment_count": 7,
  "views": 0,
  "content": "<p><a href=\"https://www.nature.com/articles/s41598-019-47181-w\">https://www.nature.com/articles/s41598-019-47181-w</a>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F654629%2F82379f6745973ced564ca1a9ad9a5238%2F2019-08-06%2018_21_42-Clipboard.png?generation=1565104928179689&amp;alt=media\" alt=\"\"></p>\n\n<p>\"Ltd provided a non-open, anonymized retinal image dataset of patients with diabetes, including **41122 **graded retinal color images from 14624 patients.\"</p>\n\n<p>\"Datasets used in model training, tuning and primary validation were provided by Digifundus Ltd. This dataset is not publicly available and restriction apply to their use. The Messidor dataset may be requested from <a href=\"http://www.adcis.net/en/Download-Third-Party/Messidor.htm\">http://www.adcis.net/en/Download-Third-Party/Messidor.htm</a>\"</p>",
  "messages": [
    {
      "id": "593422",
      "postDate": "08/06/2019 15:22:43",
      "content": "<p><a href=\"https://www.nature.com/articles/s41598-019-47181-w\">https://www.nature.com/articles/s41598-019-47181-w</a>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F654629%2F82379f6745973ced564ca1a9ad9a5238%2F2019-08-06%2018_21_42-Clipboard.png?generation=1565104928179689&amp;alt=media\" alt=\"\"></p>\n\n<p>\"Ltd provided a non-open, anonymized retinal image dataset of patients with diabetes, including **41122 **graded retinal color images from 14624 patients.\"</p>\n\n<p>\"Datasets used in model training, tuning and primary validation were provided by Digifundus Ltd. This dataset is not publicly available and restriction apply to their use. The Messidor dataset may be requested from <a href=\"http://www.adcis.net/en/Download-Third-Party/Messidor.htm\">http://www.adcis.net/en/Download-Third-Party/Messidor.htm</a>\"</p>",
      "rawMarkdown": "https://www.nature.com/articles/s41598-019-47181-w\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F654629%2F82379f6745973ced564ca1a9ad9a5238%2F2019-08-06%2018_21_42-Clipboard.png?generation=1565104928179689&amp;alt=media)\n\n\n\n\"Ltd provided a non-open, anonymized retinal image dataset of patients with diabetes, including **41122 **graded retinal color images from 14624 patients.\"\n\n\"Datasets used in model training, tuning and primary validation were provided by Digifundus Ltd. This dataset is not publicly available and restriction apply to their use. The Messidor dataset may be requested from http://www.adcis.net/en/Download-Third-Party/Messidor.htm\"",
      "votes": null
    },
    {
      "id": "593465",
      "postDate": "08/06/2019 16:30:47",
      "content": "<p>TL DR:</p>\n\n<p><strong>The results show that the ensemble of six models outperforms the corresponding single model in every experiment with 512 × 512 image size. In addition, the ensemble model results for 512 × 512 image size turn out to be slightly better than our results for larger image sizes (i.e. 1024 × 1024 and 2095 × 2095)</strong></p>\n\n<p>*<em>Other boring stuff *</em>\nImage processing:</p>\n\n<p><code>In the model training and subsequent primary validation, we used preprocessed versions of the original images. The preprocessing consisted of image cropping followed by resizing. Each image was cropped to a square shape which included the most tightly contained circular area of fundus. The procedure removed most of the black borders and all of the patient related annotations from the image data. Each of the cropped images were then resized to five different standard input image sizes of 256 × 256, 299 × 299, 512 × 512, 1024 × 1024, and 2095 × 2095 pixels. The largest image size was the smallest native resolution of the retinal cameras after the preprocessing steps. Here the creation of multiple resolutions was done for the purposes of analyzing the effect of the input image resolution on the classification performance.</code></p>\n\n<p>Model:</p>\n\n<p><code>The network architecture, we selected, is based on the Inception-v3 architecture13 that was pretrained on ImageNet dataset14</code></p>\n\n<p>Image Augmentation</p>\n\n<p><code>We employ dropout regularization method, as described in the Experimental setup of this Supplement, which causes randomness in the training of the models. We also feed the training images in different (randomized) order and with different random augmentations for each of the models trained for the ensemble to encourage the networks to be dissimilar. Different random behavior was ensured by augmenting the random number generator for each model training.</code></p>\n\n<p>Training:</p>\n\n<p><code>Due to the memory constraints, the models trained on 2095 × 2095 pixels input images were trained using mini-batch size of 1, their batch normalization layers were replaced by instance normalization layers, and parameter updates were accumulated across 15 mini-batches.</code></p>",
      "rawMarkdown": "TL DR:\n\n\n**The results show that the ensemble of six models outperforms the corresponding single model in every experiment with 512 × 512 image size. In addition, the ensemble model results for 512 × 512 image size turn out to be slightly better than our results for larger image sizes (i.e. 1024 × 1024 and 2095 × 2095)**\n\n\n**Other boring stuff **\nImage processing:\n\n`In the model training and subsequent primary validation, we used preprocessed versions of the original images. The preprocessing consisted of image cropping followed by resizing. Each image was cropped to a square shape which included the most tightly contained circular area of fundus. The procedure removed most of the black borders and all of the patient related annotations from the image data. Each of the cropped images were then resized to five different standard input image sizes of 256 × 256, 299 × 299, 512 × 512, 1024 × 1024, and 2095 × 2095 pixels. The largest image size was the smallest native resolution of the retinal cameras after the preprocessing steps. Here the creation of multiple resolutions was done for the purposes of analyzing the effect of the input image resolution on the classification performance.`\n\n\n\nModel:\n\n`The network architecture, we selected, is based on the Inception-v3 architecture13 that was pretrained on ImageNet dataset14`\n\nImage Augmentation\n\n`We employ dropout regularization method, as described in the Experimental setup of this Supplement, which causes randomness in the training of the models. We also feed the training images in different (randomized) order and with different random augmentations for each of the models trained for the ensemble to encourage the networks to be dissimilar. Different random behavior was ensured by augmenting the random number generator for each model training.`\n\nTraining:\n\n`Due to the memory constraints, the models trained on 2095 × 2095 pixels input images were trained using mini-batch size of 1, their batch normalization layers were replaced by instance normalization layers, and parameter updates were accumulated across 15 mini-batches.`",
      "votes": null
    },
    {
      "id": "593484",
      "postDate": "08/06/2019 16:56:12",
      "content": "<p>Thanks for this. Sent you a private message. </p>",
      "rawMarkdown": "Thanks for this. Sent you a private message.",
      "votes": null
    },
    {
      "id": "593492",
      "postDate": "08/06/2019 17:07:11",
      "content": "<p>If I see this right by quickly skimming through it: They do a bunch of different tasks like binary classification differentiating between two different types of diseases: (i) macular edema and (ii) retinopathy.</p>",
      "rawMarkdown": "If I see this right by quickly skimming through it: They do a bunch of different tasks like binary classification differentiating between two different types of diseases: (i) macular edema and (ii) retinopathy.",
      "votes": null
    },
    {
      "id": "593510",
      "postDate": "08/06/2019 17:39:50",
      "content": "<p>Indeed, the only merit I see is the existence of the dataset itself. </p>",
      "rawMarkdown": "Indeed, the only merit I see is the existence of the dataset itself.",
      "votes": null
    },
    {
      "id": "593837",
      "postDate": "08/07/2019 07:13:52",
      "content": "<p>By the way, it looks like they have different labels scheme : </p>\n\n<ul>\n<li>0 (Normal): (μA = 0) AND (H = 0)</li>\n<li>1: (0 &lt; μA &lt;= 5) AND (H = 0)</li>\n<li>2: ((5 &lt; μA &lt; 15) OR (0 &lt; H &lt; 5)) AND (NV = 0)</li>\n<li>3: (μA &gt;= 15) OR (H &gt;=5) OR (NV = 1)</li>\n</ul>\n\n<p>μA: number of microaneurysms\nH: number of hemorrhages\nNV = 1: neovascularization\nNV = 0: no neovascularization</p>\n\n<p>Unlike what we have here (5 levels)</p>\n\n<p>Still, a very interesting paper and I am sure we can learn from it. So thanks for sharing.  </p>",
      "rawMarkdown": "By the way, it looks like they have different labels scheme : \n\n- 0 (Normal): (μA = 0) AND (H = 0)\n- 1: (0 &lt; μA &lt;= 5) AND (H = 0)\n- 2: ((5 &lt; μA &lt; 15) OR (0 &lt; H &lt; 5)) AND (NV = 0)\n- 3: (μA &gt;= 15) OR (H &gt;=5) OR (NV = 1)\n\nμA: number of microaneurysms\nH: number of hemorrhages\nNV = 1: neovascularization\nNV = 0: no neovascularization\n\nUnlike what we have here (5 levels)\n\nStill, a very interesting paper and I am sure we can learn from it. So thanks for sharing.",
      "votes": null
    },
    {
      "id": "594916",
      "postDate": "08/08/2019 16:17:28",
      "content": "<p>Thank you very much for sharing this paper!!</p>",
      "rawMarkdown": "Thank you very much for sharing this paper!!",
      "votes": null
    },
    {
      "id": "596390",
      "postDate": "08/10/2019 15:25:26",
      "content": "<p>Nice! Thanks for sharing! Interesting approach to use mini-batches of 1. What learning rate did they use for the 2095x2095 approach? Wouldn't it need to be insanely small to avoid noisy training?</p>",
      "rawMarkdown": "Nice! Thanks for sharing! Interesting approach to use mini-batches of 1. What learning rate did they use for the 2095x2095 approach? Wouldn't it need to be insanely small to avoid noisy training?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 593465,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "08/06/2019 16:30:47",
      "content": "<p>TL DR:</p>\n\n<p><strong>The results show that the ensemble of six models outperforms the corresponding single model in every experiment with 512 × 512 image size. In addition, the ensemble model results for 512 × 512 image size turn out to be slightly better than our results for larger image sizes (i.e. 1024 × 1024 and 2095 × 2095)</strong></p>\n\n<p>*<em>Other boring stuff *</em>\nImage processing:</p>\n\n<p><code>In the model training and subsequent primary validation, we used preprocessed versions of the original images. The preprocessing consisted of image cropping followed by resizing. Each image was cropped to a square shape which included the most tightly contained circular area of fundus. The procedure removed most of the black borders and all of the patient related annotations from the image data. Each of the cropped images were then resized to five different standard input image sizes of 256 × 256, 299 × 299, 512 × 512, 1024 × 1024, and 2095 × 2095 pixels. The largest image size was the smallest native resolution of the retinal cameras after the preprocessing steps. Here the creation of multiple resolutions was done for the purposes of analyzing the effect of the input image resolution on the classification performance.</code></p>\n\n<p>Model:</p>\n\n<p><code>The network architecture, we selected, is based on the Inception-v3 architecture13 that was pretrained on ImageNet dataset14</code></p>\n\n<p>Image Augmentation</p>\n\n<p><code>We employ dropout regularization method, as described in the Experimental setup of this Supplement, which causes randomness in the training of the models. We also feed the training images in different (randomized) order and with different random augmentations for each of the models trained for the ensemble to encourage the networks to be dissimilar. Different random behavior was ensured by augmenting the random number generator for each model training.</code></p>\n\n<p>Training:</p>\n\n<p><code>Due to the memory constraints, the models trained on 2095 × 2095 pixels input images were trained using mini-batch size of 1, their batch normalization layers were replaced by instance normalization layers, and parameter updates were accumulated across 15 mini-batches.</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 593484,
          "author_name": "solomonk",
          "author_url": "",
          "post_date": "08/06/2019 16:56:12",
          "content": "<p>Thanks for this. Sent you a private message. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 593492,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "08/06/2019 17:07:11",
      "content": "<p>If I see this right by quickly skimming through it: They do a bunch of different tasks like binary classification differentiating between two different types of diseases: (i) macular edema and (ii) retinopathy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 593510,
          "author_name": "solomonk",
          "author_url": "",
          "post_date": "08/06/2019 17:39:50",
          "content": "<p>Indeed, the only merit I see is the existence of the dataset itself. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 593837,
      "author_name": "omershect",
      "author_url": "",
      "post_date": "08/07/2019 07:13:52",
      "content": "<p>By the way, it looks like they have different labels scheme : </p>\n\n<ul>\n<li>0 (Normal): (μA = 0) AND (H = 0)</li>\n<li>1: (0 &lt; μA &lt;= 5) AND (H = 0)</li>\n<li>2: ((5 &lt; μA &lt; 15) OR (0 &lt; H &lt; 5)) AND (NV = 0)</li>\n<li>3: (μA &gt;= 15) OR (H &gt;=5) OR (NV = 1)</li>\n</ul>\n\n<p>μA: number of microaneurysms\nH: number of hemorrhages\nNV = 1: neovascularization\nNV = 0: no neovascularization</p>\n\n<p>Unlike what we have here (5 levels)</p>\n\n<p>Still, a very interesting paper and I am sure we can learn from it. So thanks for sharing.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 594916,
      "author_name": "nanditab35",
      "author_url": "",
      "post_date": "08/08/2019 16:17:28",
      "content": "<p>Thank you very much for sharing this paper!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 596390,
      "author_name": "carlolepelaars",
      "author_url": "",
      "post_date": "08/10/2019 15:25:26",
      "content": "<p>Nice! Thanks for sharing! Interesting approach to use mini-batches of 1. What learning rate did they use for the 2095x2095 approach? Wouldn't it need to be insanely small to avoid noisy training?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "593422": "https://www.nature.com/articles/s41598-019-47181-w\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F654629%2F82379f6745973ced564ca1a9ad9a5238%2F2019-08-06%2018_21_42-Clipboard.png?generation=1565104928179689&amp;alt=media)\n\n\n\n\"Ltd provided a non-open, anonymized retinal image dataset of patients with diabetes, including **41122 **graded retinal color images from 14624 patients.\"\n\n\"Datasets used in model training, tuning and primary validation were provided by Digifundus Ltd. This dataset is not publicly available and restriction apply to their use. The Messidor dataset may be requested from http://www.adcis.net/en/Download-Third-Party/Messidor.htm\"",
    "593465": "TL DR:\n\n\n**The results show that the ensemble of six models outperforms the corresponding single model in every experiment with 512 × 512 image size. In addition, the ensemble model results for 512 × 512 image size turn out to be slightly better than our results for larger image sizes (i.e. 1024 × 1024 and 2095 × 2095)**\n\n\n**Other boring stuff **\nImage processing:\n\n`In the model training and subsequent primary validation, we used preprocessed versions of the original images. The preprocessing consisted of image cropping followed by resizing. Each image was cropped to a square shape which included the most tightly contained circular area of fundus. The procedure removed most of the black borders and all of the patient related annotations from the image data. Each of the cropped images were then resized to five different standard input image sizes of 256 × 256, 299 × 299, 512 × 512, 1024 × 1024, and 2095 × 2095 pixels. The largest image size was the smallest native resolution of the retinal cameras after the preprocessing steps. Here the creation of multiple resolutions was done for the purposes of analyzing the effect of the input image resolution on the classification performance.`\n\n\n\nModel:\n\n`The network architecture, we selected, is based on the Inception-v3 architecture13 that was pretrained on ImageNet dataset14`\n\nImage Augmentation\n\n`We employ dropout regularization method, as described in the Experimental setup of this Supplement, which causes randomness in the training of the models. We also feed the training images in different (randomized) order and with different random augmentations for each of the models trained for the ensemble to encourage the networks to be dissimilar. Different random behavior was ensured by augmenting the random number generator for each model training.`\n\nTraining:\n\n`Due to the memory constraints, the models trained on 2095 × 2095 pixels input images were trained using mini-batch size of 1, their batch normalization layers were replaced by instance normalization layers, and parameter updates were accumulated across 15 mini-batches.`",
    "593484": "Thanks for this. Sent you a private message.",
    "593492": "If I see this right by quickly skimming through it: They do a bunch of different tasks like binary classification differentiating between two different types of diseases: (i) macular edema and (ii) retinopathy.",
    "593510": "Indeed, the only merit I see is the existence of the dataset itself.",
    "593837": "By the way, it looks like they have different labels scheme : \n\n- 0 (Normal): (μA = 0) AND (H = 0)\n- 1: (0 &lt; μA &lt;= 5) AND (H = 0)\n- 2: ((5 &lt; μA &lt; 15) OR (0 &lt; H &lt; 5)) AND (NV = 0)\n- 3: (μA &gt;= 15) OR (H &gt;=5) OR (NV = 1)\n\nμA: number of microaneurysms\nH: number of hemorrhages\nNV = 1: neovascularization\nNV = 0: no neovascularization\n\nUnlike what we have here (5 levels)\n\nStill, a very interesting paper and I am sure we can learn from it. So thanks for sharing.",
    "594916": "Thank you very much for sharing this paper!!",
    "596390": "Nice! Thanks for sharing! Interesting approach to use mini-batches of 1. What learning rate did they use for the 2095x2095 approach? Wouldn't it need to be insanely small to avoid noisy training?"
  },
  "source": "meta"
}