{
  "id": 98856,
  "title": "Unreliable Labels ?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/98856",
  "author_name": "Maxwell",
  "post_date": "2019-07-07T02:40:02.707000",
  "votes": 53,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Although I could not yet manged to build a good model to predict the retinopathic severity, guessing my model seems to be more accurate in some cases ( I'm not an ophthalmologist, it's my intuition ;) ). Here is the cases I observed.  </p>\n\n<p><strong>Fundus images</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F4e4289f3e3482027732e1b8ef9e2d201%2Fimg1.png?generation=1562465525301620&amp;alt=media\" alt=\"fundus images\"></p>\n\n<p>After taking a quick check on the Internet, I found out that the label inaccuracy has been an issue in this field. For example, the label inconsistency among ophthalmologists is introduced in the following movie.  </p>\n\n<p>TensorFlow Dev Summit 2017 Retinal Imaging\n<a href=\"https://youtu.be/oOeZ7IgEN4o?t=156\">https://youtu.be/oOeZ7IgEN4o?t=156</a></p>\n\n<p>I suppose, to tackle this issue, we must implement a robust loss function considering abnormal losses or other remedies as seen in the previous competition (DRD). <br>\nAny opinions are welcome.</p>\n\n<hr>\n\n<p>Updated Jul.10.2019  </p>\n\n<p><strong>The paper of GoogleAI related to above movie</strong> <br>\nDeep Learning vs. Human Graders for Classifying Severity Levels of Diabetic Retinopathy in a Real-World Nationwide Screening Program</p>\n\n<p><a href=\"https://arxiv.org/abs/1810.08290\">https://arxiv.org/abs/1810.08290</a></p>",
  "messages": [
    {
      "id": 569611,
      "postDate": "2019-07-07T02:40:02.707Z",
      "content": "<p>Although I could not yet manged to build a good model to predict the retinopathic severity, guessing my model seems to be more accurate in some cases ( I'm not an ophthalmologist, it's my intuition ;) ). Here is the cases I observed.  </p>\n\n<p><strong>Fundus images</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F4e4289f3e3482027732e1b8ef9e2d201%2Fimg1.png?generation=1562465525301620&amp;alt=media\" alt=\"fundus images\"></p>\n\n<p>After taking a quick check on the Internet, I found out that the label inaccuracy has been an issue in this field. For example, the label inconsistency among ophthalmologists is introduced in the following movie.  </p>\n\n<p>TensorFlow Dev Summit 2017 Retinal Imaging\n<a href=\"https://youtu.be/oOeZ7IgEN4o?t=156\">https://youtu.be/oOeZ7IgEN4o?t=156</a></p>\n\n<p>I suppose, to tackle this issue, we must implement a robust loss function considering abnormal losses or other remedies as seen in the previous competition (DRD). <br>\nAny opinions are welcome.</p>\n\n<hr>\n\n<p>Updated Jul.10.2019  </p>\n\n<p><strong>The paper of GoogleAI related to above movie</strong> <br>\nDeep Learning vs. Human Graders for Classifying Severity Levels of Diabetic Retinopathy in a Real-World Nationwide Screening Program</p>\n\n<p><a href=\"https://arxiv.org/abs/1810.08290\">https://arxiv.org/abs/1810.08290</a></p>",
      "rawMarkdown": "Although I could not yet manged to build a good model to predict the retinopathic severity, guessing my model seems to be more accurate in some cases ( I'm not an ophthalmologist, it's my intuition ;) ). Here is the cases I observed.  \n\n**Fundus images**\n\n![fundus images](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F4e4289f3e3482027732e1b8ef9e2d201%2Fimg1.png?generation=1562465525301620&amp;alt=media)\n\n  \nAfter taking a quick check on the Internet, I found out that the label inaccuracy has been an issue in this field. For example, the label inconsistency among ophthalmologists is introduced in the following movie.  \n\nTensorFlow Dev Summit 2017 Retinal Imaging\n[https://youtu.be/oOeZ7IgEN4o?t=156](https://youtu.be/oOeZ7IgEN4o?t=156)\n\nI suppose, to tackle this issue, we must implement a robust loss function considering abnormal losses or other remedies as seen in the previous competition (DRD).  \nAny opinions are welcome.\n\n---\nUpdated Jul.10.2019  \n\n**The paper of GoogleAI related to above movie**  \nDeep Learning vs. Human Graders for Classifying Severity Levels of Diabetic Retinopathy in a Real-World Nationwide Screening Program\n\nhttps://arxiv.org/abs/1810.08290",
      "votes": 53
    },
    {
      "id": 569759,
      "postDate": "2019-07-07T09:14:44.780Z",
      "content": "<p>As Maxwell said, this is the specific slide showing inconsistency of label estimation:\nNote that each row illustrates a single patient, and each column represents a single doctor's severity estimation of that patient case.</p>\n\n<p><img src=\"https://i.ibb.co/6rQ2sFG/inconsistent-estimation.png\" alt=\"inconsistent  estimation in diabetic retinophaty\"></p>",
      "rawMarkdown": "As Maxwell said, this is the specific slide showing inconsistency of label estimation:\nNote that each row illustrates a single patient, and each column represents a single doctor's severity estimation of that patient case.\n\n![inconsistent  estimation in diabetic retinophaty](https://i.ibb.co/6rQ2sFG/inconsistent-estimation.png)",
      "votes": 16
    },
    {
      "id": 569636,
      "postDate": "2019-07-07T04:29:17.710Z",
      "content": "<p>I feel the same. \nI found 268 duplicated images in train data, and 64 among them have different labels (e.g. 14e3f84445f7 and f0f89314e860). \nDuplication also occurs between train and test, and I guess some of them have different labels.</p>",
      "rawMarkdown": "I feel the same. \nI found 268 duplicated images in train data, and 64 among them have different labels (e.g. 14e3f84445f7 and f0f89314e860). \nDuplication also occurs between train and test, and I guess some of them have different labels.",
      "votes": 13,
      "replies": [
        {
          "id": 569648,
          "postDate": "2019-07-07T05:13:22.670Z",
          "content": "<p><a href=\"/naka2ka\">@naka2ka</a> \nGreat insight.\nSo our models are forced to learn the different labels on the same image😣... </p>\n\n<p>In this kernel, as you did, the duplication was checked by <a href=\"/miklgr500\">@miklgr500</a> .\n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98856#latest-569611\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98856#latest-569611</a></p>\n\n<p>I felt the models tend to overfit, one of the reason of that will be this...</p>",
          "rawMarkdown": "@naka2ka \nGreat insight.\nSo our models are forced to learn the different labels on the same image😣... \n\nIn this kernel, as you did, the duplication was checked by @miklgr500 .\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98856#latest-569611\n\nI felt the models tend to overfit, one of the reason of that will be this...",
          "votes": 2
        },
        {
          "id": 569662,
          "postDate": "2019-07-07T05:48:37.777Z",
          "content": "<p>Yeah. <a href=\"/sohier\">@sohier</a>\nid_code: 14e3f84445f7 diagnosis: 3</p>\n\n<p>id_code: f0f89314e860 diagnosis: 4</p>\n\n<p>These two photos are the same but have different labels.</p>",
          "rawMarkdown": "Yeah. @sohier\nid_code: 14e3f84445f7 diagnosis: 3\n\nid_code: f0f89314e860 diagnosis: 4\n\nThese two photos are the same but have different labels.",
          "votes": 3
        },
        {
          "id": 569666,
          "postDate": "2019-07-07T05:55:16.280Z",
          "content": "<p>Are any ophthalmologists on board?\nAre any ophthalmologists on board?</p>",
          "rawMarkdown": "Are any ophthalmologists on board?\nAre any ophthalmologists on board?",
          "votes": 3
        }
      ]
    },
    {
      "id": 570171,
      "postDate": "2019-07-08T00:10:26.410Z",
      "content": "<p>Updated: <br>\nI published a little bit modified version of below kernel. <br>\nYou can download duplicated list of train data.  </p>\n\n<p><a href=\"https://www.kaggle.com/maxwell110/duplicated-list-csv-file\">https://www.kaggle.com/maxwell110/duplicated-list-csv-file</a></p>\n\n<hr>\n\n<p>To detect the duplicated images, this awesome kernel will be a help. <br>\n<a href=\"https://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data\">https://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data</a></p>\n\n<p>Thanks <a href=\"/manojprabhaakr\">@manojprabhaakr</a> !</p>",
      "rawMarkdown": "Updated:  \nI published a little bit modified version of below kernel.  \nYou can download duplicated list of train data.  \n\nhttps://www.kaggle.com/maxwell110/duplicated-list-csv-file\n\n---\n\nTo detect the duplicated images, this awesome kernel will be a help.  \nhttps://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data\n\nThanks @manojprabhaakr !",
      "votes": 5
    },
    {
      "id": 598054,
      "postDate": "2019-08-13T04:01:59Z",
      "content": "<p><a href=\"/maxwell110\">@maxwell110</a>,  I also see there is label inaccuracy.  Below severity stage differences could help. I feel Data Cleaning is must but how would that be accounted on the hidden test dataset?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F43269%2Ff37410a72422e57d78e9ea5c7eb2d024%2Fretinopathy.png?generation=1565668605750117&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "@maxwell110,  I also see there is label inaccuracy.  Below severity stage differences could help. I feel Data Cleaning is must but how would that be accounted on the hidden test dataset?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F43269%2Ff37410a72422e57d78e9ea5c7eb2d024%2Fretinopathy.png?generation=1565668605750117&amp;alt=media)\n\n",
      "votes": 3,
      "replies": [
        {
          "id": 598667,
          "postDate": "2019-08-13T21:23:06.887Z",
          "content": "<p><a href=\"/hmnshu\">@hmnshu</a> \nI'm not sure Data Cleansing works well or not for private test dataset. But I think private test dataset will also include inaccurate labels.  </p>\n\n<p>For example, DRD2015(previous compettition) dataset seems to include ambiguous labels. <br>\nIn this paper, the author said that 75% images are gradable. <br>\n<a href=\"https://www.biorxiv.org/content/biorxiv/early/2018/06/19/225508.full.pdf\">https://www.biorxiv.org/content/biorxiv/early/2018/06/19/225508.full.pdf</a></p>",
          "rawMarkdown": "@hmnshu \nI'm not sure Data Cleansing works well or not for private test dataset. But I think private test dataset will also include inaccurate labels.  \n\nFor example, DRD2015(previous compettition) dataset seems to include ambiguous labels.  \nIn this paper, the author said that 75% images are gradable.  \nhttps://www.biorxiv.org/content/biorxiv/early/2018/06/19/225508.full.pdf",
          "votes": 1
        }
      ]
    },
    {
      "id": 570766,
      "postDate": "2019-07-08T18:39:38.153Z",
      "content": "<p>I have used the cleaned dataset to train, better cv : 0.9189-&gt;0.9284, but lower lb: 0.782-&gt;0.769; maybe those duplicate images could work like regularization, I'm a noob 😯 </p>",
      "rawMarkdown": "I have used the cleaned dataset to train, better cv : 0.9189-&gt;0.9284, but lower lb: 0.782-&gt;0.769; maybe those duplicate images could work like regularization, I'm a noob 😯 ",
      "votes": 3,
      "replies": [
        {
          "id": 571019,
          "postDate": "2019-07-09T04:17:27.060Z",
          "content": "<p>I recommend you forget  current test data with public lb.</p>",
          "rawMarkdown": "I recommend you forget  current test data with public lb.",
          "votes": 3
        },
        {
          "id": 571035,
          "postDate": "2019-07-09T04:52:57.397Z",
          "content": "<p>Thanks for your recommendation, I have read your kernel about the data distribution, seems the training data has a different distribution with testing data, so maybe cv is not compromising. How about the relationship between public test data and private test data? Based on your kernel, cv should be not promising right?</p>",
          "rawMarkdown": "Thanks for your recommendation, I have read your kernel about the data distribution, seems the training data has a different distribution with testing data, so maybe cv is not compromising. How about the relationship between public test data and private test data? Based on your kernel, cv should be not promising right?",
          "votes": 1
        },
        {
          "id": 571068,
          "postDate": "2019-07-09T06:00:43.243Z",
          "content": "<p>Current public lb calculated on 15%(~300 img) of test data, which have 8% of leaks. Also we know that final decision must have possibility run on 20GB test dataset, that meaning that test dataset changed in the end of competition. So all above make this competition without real public lb.</p>",
          "rawMarkdown": "Current public lb calculated on 15%(~300 img) of test data, which have 8% of leaks. Also we know that final decision must have possibility run on 20GB test dataset, that meaning that test dataset changed in the end of competition. So all above make this competition without real public lb.",
          "votes": 2
        },
        {
          "id": 571074,
          "postDate": "2019-07-09T06:12:27.400Z",
          "content": "<p>But since we have stage 2 running (a new test set), I think the public lb may contain whole 1928 test images that we can see in test.csv.</p>",
          "rawMarkdown": "But since we have stage 2 running (a new test set), I think the public lb may contain whole 1928 test images that we can see in test.csv.",
          "votes": 1
        },
        {
          "id": 571082,
          "postDate": "2019-07-09T06:23:25.180Z",
          "content": "<p>Ok, if it's true, then we have 300/13000=3/130~1/40 -&gt; 2.5%. Are you really believe that 2.5% of final test dataset can influence on your choice better decision?:)</p>",
          "rawMarkdown": "Ok, if it's true, then we have 300/13000=3/130~1/40 -&gt; 2.5%. Are you really believe that 2.5% of final test dataset can influence on your choice better decision?:)"
        },
        {
          "id": 572069,
          "postDate": "2019-07-10T12:10:28.573Z",
          "content": "<p>Do  you  use   external  dataset   or    you    just    use    the   official  dataset   to   achieve    7th?</p>",
          "rawMarkdown": "Do  you  use   external  dataset   or    you    just    use    the   official  dataset   to   achieve    7th?"
        },
        {
          "id": 598799,
          "postDate": "2019-08-14T03:34:21.933Z",
          "content": "<p>Hi <a href=\"/jionie\">@jionie</a> , by \"cleaned dataset\" did you mean you had also removed images in training set that seem to have unreliable labels or you just removed duplicates?</p>",
          "rawMarkdown": "Hi @jionie , by \"cleaned dataset\" did you mean you had also removed images in training set that seem to have unreliable labels or you just removed duplicates?",
          "votes": 1
        },
        {
          "id": 598800,
          "postDate": "2019-08-14T03:38:27.607Z",
          "content": "<p><a href=\"/miklgr500\">@miklgr500</a> Isn't the public lb calculated based on 1928 images(out of 1928 / 15% test images)? They won't use only 300 images for the public lb... Also if the 15% are randomly sampled(instead of not sampling evenly on purpose) from the whole test set, then public lb would be much more reliable than cv.</p>",
          "rawMarkdown": "@miklgr500 Isn't the public lb calculated based on 1928 images(out of 1928 / 15% test images)? They won't use only 300 images for the public lb... Also if the 15% are randomly sampled(instead of not sampling evenly on purpose) from the whole test set, then public lb would be much more reliable than cv."
        },
        {
          "id": 599436,
          "postDate": "2019-08-15T00:50:13.553Z",
          "content": "<p>I removed all samples based on the public csv, but now I just keep all.</p>",
          "rawMarkdown": "I removed all samples based on the public csv, but now I just keep all."
        }
      ]
    },
    {
      "id": 570168,
      "postDate": "2019-07-08T00:06:07.260Z",
      "content": "<p>Well, I guess this is another evidence suggesting that some hard work needs to be done on data cleaning. </p>",
      "rawMarkdown": "Well, I guess this is another evidence suggesting that some hard work needs to be done on data cleaning. ",
      "votes": 1,
      "replies": [
        {
          "id": 570350,
          "postDate": "2019-07-08T07:05:07.760Z",
          "content": "<p>But what if test lables are also mislabled and cleaning the training set would result in poor performance in model?</p>\n\n<p>This manual cleaning can also result in overfitting of training set, right?</p>",
          "rawMarkdown": "But what if test lables are also mislabled and cleaning the training set would result in poor performance in model?\n\nThis manual cleaning can also result in overfitting of training set, right?",
          "votes": 1
        }
      ]
    },
    {
      "id": 570098,
      "postDate": "2019-07-07T19:13:32.523Z",
      "content": "<p>Emm, may need to remove those images with different labels in the training set, just view them as unlabeled data.</p>",
      "rawMarkdown": "Emm, may need to remove those images with different labels in the training set, just view them as unlabeled data.",
      "votes": 2,
      "replies": [
        {
          "id": 570170,
          "postDate": "2019-07-08T00:08:26.007Z",
          "content": "<p>Thanks to <a href=\"/manojprabhaakr\">@manojprabhaakr</a> ,\n<a href=\"https://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data\">https://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data</a>\nwe can detect the duplicated images.  </p>\n\n<p>According to this kernel, there are <strong>143</strong> duplicated images in train data 😩 . <br>\nI will exclude these extra images not to overfit when training models.</p>",
          "rawMarkdown": "Thanks to @manojprabhaakr ,\nhttps://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data\nwe can detect the duplicated images.  \n\nAccording to this kernel, there are **143** duplicated images in train data 😩 .  \nI will exclude these extra images not to overfit when training models.",
          "votes": 5
        },
        {
          "id": 598354,
          "postDate": "2019-08-13T14:07:33.233Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 571308,
      "postDate": "2019-07-09T13:36:10.477Z",
      "content": "<p>I believe there may be manual errors leading to this result</p>",
      "rawMarkdown": "I believe there may be manual errors leading to this result"
    },
    {
      "id": 571264,
      "postDate": "2019-07-09T12:08:49.540Z",
      "content": "<p>agree with you</p>",
      "rawMarkdown": "agree with you"
    },
    {
      "id": 571046,
      "postDate": "2019-07-09T05:25:30.813Z",
      "content": "<p>This is cool. Did you consider not using accuracy at all, but something like precision or recall? f-score (a little of both)?</p>",
      "rawMarkdown": "This is cool. Did you consider not using accuracy at all, but something like precision or recall? f-score (a little of both)?",
      "replies": [
        {
          "id": 571159,
          "postDate": "2019-07-09T08:49:24.773Z",
          "content": "<p><a href=\"/kbridge14\">@kbridge14</a> \nActually, I'm using MSE. I have tried to solve this task as a regression problem.</p>",
          "rawMarkdown": "@kbridge14 \nActually, I'm using MSE. I have tried to solve this task as a regression problem.",
          "votes": 2
        },
        {
          "id": 598143,
          "postDate": "2019-08-13T07:39:13.720Z",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> Given you are solving this task as a regression problem, What are your thoughts on Mixup Data augmentation? </p>",
          "rawMarkdown": "@maxwell110 Given you are solving this task as a regression problem, What are your thoughts on Mixup Data augmentation? "
        },
        {
          "id": 598658,
          "postDate": "2019-08-13T21:05:14.580Z",
          "content": "<p><a href=\"/monsterspy\">@monsterspy</a> \nIMO, Mixup will not work well due to the nonlinearity of diabetic retinopathy grading. I mean, ophthalmologists grade images with some non-linear thresholds. <br>\nFor example, imagine when you mix grade 1 and 3 images with equivalent probability(= 0.5), the target value would be 2(= 1 * 0.5 + 3 * 0.5). But I suppose ophthalmologists will make a diagnosis as grade 3 or higher. Because they will think the image has a lot of blurred but serious findings.</p>\n\n<p>For more detail about diagnosis criteria, check the following material. <br>\n(Sorry, I forgot who introduced this in the discussion thread.)\n<a href=\"http://eyesteve.com/diabetic-retinopathy-grading/\">http://eyesteve.com/diabetic-retinopathy-grading/</a></p>",
          "rawMarkdown": "@monsterspy \nIMO, Mixup will not work well due to the nonlinearity of diabetic retinopathy grading. I mean, ophthalmologists grade images with some non-linear thresholds.  \nFor example, imagine when you mix grade 1 and 3 images with equivalent probability(= 0.5), the target value would be 2(= 1 * 0.5 + 3 * 0.5). But I suppose ophthalmologists will make a diagnosis as grade 3 or higher. Because they will think the image has a lot of blurred but serious findings.\n  \nFor more detail about diagnosis criteria, check the following material.  \n(Sorry, I forgot who introduced this in the discussion thread.)\nhttp://eyesteve.com/diabetic-retinopathy-grading/",
          "votes": 2
        },
        {
          "id": 600502,
          "postDate": "2019-08-16T07:31:05.570Z",
          "content": "<p>Thanks <a href=\"/maxwell110\">@maxwell110</a> for sharing your thoughts. </p>",
          "rawMarkdown": "Thanks @maxwell110 for sharing your thoughts. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 570270,
      "postDate": "2019-07-08T04:28:42.027Z",
      "content": "<p>up</p>",
      "rawMarkdown": "up"
    }
  ],
  "comments": [
    {
      "id": 569759,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-07-07T09:14:44.780000",
      "content": "<p>As Maxwell said, this is the specific slide showing inconsistency of label estimation:\nNote that each row illustrates a single patient, and each column represents a single doctor's severity estimation of that patient case.</p>\n\n<p><img src=\"https://i.ibb.co/6rQ2sFG/inconsistent-estimation.png\" alt=\"inconsistent  estimation in diabetic retinophaty\"></p>",
      "votes": 16,
      "replies": []
    },
    {
      "id": 569636,
      "author_name": "ynktk",
      "author_url": "",
      "post_date": "2019-07-07T04:29:17.710000",
      "content": "<p>I feel the same. \nI found 268 duplicated images in train data, and 64 among them have different labels (e.g. 14e3f84445f7 and f0f89314e860). \nDuplication also occurs between train and test, and I guess some of them have different labels.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 569648,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-07-07T05:13:22.670000",
          "content": "<p><a href=\"/naka2ka\">@naka2ka</a> \nGreat insight.\nSo our models are forced to learn the different labels on the same image😣... </p>\n\n<p>In this kernel, as you did, the duplication was checked by <a href=\"/miklgr500\">@miklgr500</a> .\n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98856#latest-569611\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98856#latest-569611</a></p>\n\n<p>I felt the models tend to overfit, one of the reason of that will be this...</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 569662,
          "author_name": "Zheng Li",
          "author_url": "",
          "post_date": "2019-07-07T05:48:37.777000",
          "content": "<p>Yeah. <a href=\"/sohier\">@sohier</a>\nid_code: 14e3f84445f7 diagnosis: 3</p>\n\n<p>id_code: f0f89314e860 diagnosis: 4</p>\n\n<p>These two photos are the same but have different labels.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 569666,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-07-07T05:55:16.280000",
          "content": "<p>Are any ophthalmologists on board?\nAre any ophthalmologists on board?</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 570171,
      "author_name": "Maxwell",
      "author_url": "",
      "post_date": "2019-07-08T00:10:26.410000",
      "content": "<p>Updated: <br>\nI published a little bit modified version of below kernel. <br>\nYou can download duplicated list of train data.  </p>\n\n<p><a href=\"https://www.kaggle.com/maxwell110/duplicated-list-csv-file\">https://www.kaggle.com/maxwell110/duplicated-list-csv-file</a></p>\n\n<hr>\n\n<p>To detect the duplicated images, this awesome kernel will be a help. <br>\n<a href=\"https://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data\">https://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data</a></p>\n\n<p>Thanks <a href=\"/manojprabhaakr\">@manojprabhaakr</a> !</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 598054,
      "author_name": "Himanshu Gamit",
      "author_url": "",
      "post_date": "2019-08-13T04:01:59",
      "content": "<p><a href=\"/maxwell110\">@maxwell110</a>,  I also see there is label inaccuracy.  Below severity stage differences could help. I feel Data Cleaning is must but how would that be accounted on the hidden test dataset?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F43269%2Ff37410a72422e57d78e9ea5c7eb2d024%2Fretinopathy.png?generation=1565668605750117&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 598667,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-08-13T21:23:06.887000",
          "content": "<p><a href=\"/hmnshu\">@hmnshu</a> \nI'm not sure Data Cleansing works well or not for private test dataset. But I think private test dataset will also include inaccurate labels.  </p>\n\n<p>For example, DRD2015(previous compettition) dataset seems to include ambiguous labels. <br>\nIn this paper, the author said that 75% images are gradable. <br>\n<a href=\"https://www.biorxiv.org/content/biorxiv/early/2018/06/19/225508.full.pdf\">https://www.biorxiv.org/content/biorxiv/early/2018/06/19/225508.full.pdf</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 570766,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "2019-07-08T18:39:38.153000",
      "content": "<p>I have used the cleaned dataset to train, better cv : 0.9189-&gt;0.9284, but lower lb: 0.782-&gt;0.769; maybe those duplicate images could work like regularization, I'm a noob 😯 </p>",
      "votes": 3,
      "replies": [
        {
          "id": 571019,
          "author_name": "Welf Crozzo",
          "author_url": "",
          "post_date": "2019-07-09T04:17:27.060000",
          "content": "<p>I recommend you forget  current test data with public lb.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 571035,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "2019-07-09T04:52:57.397000",
          "content": "<p>Thanks for your recommendation, I have read your kernel about the data distribution, seems the training data has a different distribution with testing data, so maybe cv is not compromising. How about the relationship between public test data and private test data? Based on your kernel, cv should be not promising right?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 571068,
          "author_name": "Welf Crozzo",
          "author_url": "",
          "post_date": "2019-07-09T06:00:43.243000",
          "content": "<p>Current public lb calculated on 15%(~300 img) of test data, which have 8% of leaks. Also we know that final decision must have possibility run on 20GB test dataset, that meaning that test dataset changed in the end of competition. So all above make this competition without real public lb.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 571074,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "2019-07-09T06:12:27.400000",
          "content": "<p>But since we have stage 2 running (a new test set), I think the public lb may contain whole 1928 test images that we can see in test.csv.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 571082,
          "author_name": "Welf Crozzo",
          "author_url": "",
          "post_date": "2019-07-09T06:23:25.180000",
          "content": "<p>Ok, if it's true, then we have 300/13000=3/130~1/40 -&gt; 2.5%. Are you really believe that 2.5% of final test dataset can influence on your choice better decision?:)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 572069,
          "author_name": "哈尔的移动城堡",
          "author_url": "",
          "post_date": "2019-07-10T12:10:28.573000",
          "content": "<p>Do  you  use   external  dataset   or    you    just    use    the   official  dataset   to   achieve    7th?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 598799,
          "author_name": "Homoalways",
          "author_url": "",
          "post_date": "2019-08-14T03:34:21.933000",
          "content": "<p>Hi <a href=\"/jionie\">@jionie</a> , by \"cleaned dataset\" did you mean you had also removed images in training set that seem to have unreliable labels or you just removed duplicates?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 598800,
          "author_name": "Homoalways",
          "author_url": "",
          "post_date": "2019-08-14T03:38:27.607000",
          "content": "<p><a href=\"/miklgr500\">@miklgr500</a> Isn't the public lb calculated based on 1928 images(out of 1928 / 15% test images)? They won't use only 300 images for the public lb... Also if the 15% are randomly sampled(instead of not sampling evenly on purpose) from the whole test set, then public lb would be much more reliable than cv.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 599436,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "2019-08-15T00:50:13.553000",
          "content": "<p>I removed all samples based on the public csv, but now I just keep all.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 570168,
      "author_name": "Xuan Cao",
      "author_url": "",
      "post_date": "2019-07-08T00:06:07.260000",
      "content": "<p>Well, I guess this is another evidence suggesting that some hard work needs to be done on data cleaning. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 570350,
          "author_name": "Rahul",
          "author_url": "",
          "post_date": "2019-07-08T07:05:07.760000",
          "content": "<p>But what if test lables are also mislabled and cleaning the training set would result in poor performance in model?</p>\n\n<p>This manual cleaning can also result in overfitting of training set, right?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 570098,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "2019-07-07T19:13:32.523000",
      "content": "<p>Emm, may need to remove those images with different labels in the training set, just view them as unlabeled data.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 570170,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-07-08T00:08:26.007000",
          "content": "<p>Thanks to <a href=\"/manojprabhaakr\">@manojprabhaakr</a> ,\n<a href=\"https://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data\">https://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data</a>\nwe can detect the duplicated images.  </p>\n\n<p>According to this kernel, there are <strong>143</strong> duplicated images in train data 😩 . <br>\nI will exclude these extra images not to overfit when training models.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 598354,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-08-13T14:07:33.233000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 571308,
      "author_name": "Bibyutatsu",
      "author_url": "",
      "post_date": "2019-07-09T13:36:10.477000",
      "content": "<p>I believe there may be manual errors leading to this result</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 571264,
      "author_name": "Wangchentong",
      "author_url": "",
      "post_date": "2019-07-09T12:08:49.540000",
      "content": "<p>agree with you</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 571046,
      "author_name": "Kerry",
      "author_url": "",
      "post_date": "2019-07-09T05:25:30.813000",
      "content": "<p>This is cool. Did you consider not using accuracy at all, but something like precision or recall? f-score (a little of both)?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 571159,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-07-09T08:49:24.773000",
          "content": "<p><a href=\"/kbridge14\">@kbridge14</a> \nActually, I'm using MSE. I have tried to solve this task as a regression problem.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 598143,
          "author_name": "Arjun",
          "author_url": "",
          "post_date": "2019-08-13T07:39:13.720000",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> Given you are solving this task as a regression problem, What are your thoughts on Mixup Data augmentation? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 598658,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-08-13T21:05:14.580000",
          "content": "<p><a href=\"/monsterspy\">@monsterspy</a> \nIMO, Mixup will not work well due to the nonlinearity of diabetic retinopathy grading. I mean, ophthalmologists grade images with some non-linear thresholds. <br>\nFor example, imagine when you mix grade 1 and 3 images with equivalent probability(= 0.5), the target value would be 2(= 1 * 0.5 + 3 * 0.5). But I suppose ophthalmologists will make a diagnosis as grade 3 or higher. Because they will think the image has a lot of blurred but serious findings.</p>\n\n<p>For more detail about diagnosis criteria, check the following material. <br>\n(Sorry, I forgot who introduced this in the discussion thread.)\n<a href=\"http://eyesteve.com/diabetic-retinopathy-grading/\">http://eyesteve.com/diabetic-retinopathy-grading/</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 600502,
          "author_name": "Arjun",
          "author_url": "",
          "post_date": "2019-08-16T07:31:05.570000",
          "content": "<p>Thanks <a href=\"/maxwell110\">@maxwell110</a> for sharing your thoughts. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 570270,
      "author_name": "aisyah depalan",
      "author_url": "",
      "post_date": "2019-07-08T04:28:42.027000",
      "content": "<p>up</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "569611": "Although I could not yet manged to build a good model to predict the retinopathic severity, guessing my model seems to be more accurate in some cases ( I'm not an ophthalmologist, it's my intuition ;) ). Here is the cases I observed.  \n\n**Fundus images**\n\n![fundus images](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F4e4289f3e3482027732e1b8ef9e2d201%2Fimg1.png?generation=1562465525301620&amp;alt=media)\n\n  \nAfter taking a quick check on the Internet, I found out that the label inaccuracy has been an issue in this field. For example, the label inconsistency among ophthalmologists is introduced in the following movie.  \n\nTensorFlow Dev Summit 2017 Retinal Imaging\n[https://youtu.be/oOeZ7IgEN4o?t=156](https://youtu.be/oOeZ7IgEN4o?t=156)\n\nI suppose, to tackle this issue, we must implement a robust loss function considering abnormal losses or other remedies as seen in the previous competition (DRD).  \nAny opinions are welcome.\n\n---\nUpdated Jul.10.2019  \n\n**The paper of GoogleAI related to above movie**  \nDeep Learning vs. Human Graders for Classifying Severity Levels of Diabetic Retinopathy in a Real-World Nationwide Screening Program\n\nhttps://arxiv.org/abs/1810.08290",
    "569759": "As Maxwell said, this is the specific slide showing inconsistency of label estimation:\nNote that each row illustrates a single patient, and each column represents a single doctor's severity estimation of that patient case.\n\n![inconsistent  estimation in diabetic retinophaty](https://i.ibb.co/6rQ2sFG/inconsistent-estimation.png)",
    "569636": "I feel the same. \nI found 268 duplicated images in train data, and 64 among them have different labels (e.g. 14e3f84445f7 and f0f89314e860). \nDuplication also occurs between train and test, and I guess some of them have different labels.",
    "570171": "Updated:  \nI published a little bit modified version of below kernel.  \nYou can download duplicated list of train data.  \n\nhttps://www.kaggle.com/maxwell110/duplicated-list-csv-file\n\n---\n\nTo detect the duplicated images, this awesome kernel will be a help.  \nhttps://www.kaggle.com/manojprabhaakr/similar-duplicate-images-in-aptos-data\n\nThanks @manojprabhaakr !",
    "598054": "@maxwell110,  I also see there is label inaccuracy.  Below severity stage differences could help. I feel Data Cleaning is must but how would that be accounted on the hidden test dataset?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F43269%2Ff37410a72422e57d78e9ea5c7eb2d024%2Fretinopathy.png?generation=1565668605750117&amp;alt=media)\n\n",
    "570766": "I have used the cleaned dataset to train, better cv : 0.9189-&gt;0.9284, but lower lb: 0.782-&gt;0.769; maybe those duplicate images could work like regularization, I'm a noob 😯 ",
    "570168": "Well, I guess this is another evidence suggesting that some hard work needs to be done on data cleaning. ",
    "570098": "Emm, may need to remove those images with different labels in the training set, just view them as unlabeled data.",
    "571308": "I believe there may be manual errors leading to this result",
    "571264": "agree with you",
    "571046": "This is cool. Did you consider not using accuracy at all, but something like precision or recall? f-score (a little of both)?",
    "570270": "up"
  }
}