{
  "id": 209318,
  "title": "Is this competition a lottery?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/209318",
  "author_name": "arutema47",
  "post_date": "2021-01-07T05:57:16.990000",
  "votes": 65,
  "comment_count": 36,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201471\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201471</a></p>\n<p>Regarding the noise in training, public and private data, I'm feeling that this competition is a lottery. If there is noise in private (I'm sure there is due to the PB leak), what are we predicting in the models? Are we predicting the annotator's habits..?😷 <br>\nThe competition will be far interesting if hosts intended to fix the noise in the private set (and the competition will be so much more meaningful!)</p>",
  "messages": [
    {
      "id": 1142071,
      "postDate": "2021-01-07T05:57:16.990Z",
      "content": "<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201471\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201471</a></p>\n<p>Regarding the noise in training, public and private data, I'm feeling that this competition is a lottery. If there is noise in private (I'm sure there is due to the PB leak), what are we predicting in the models? Are we predicting the annotator's habits..?😷 <br>\nThe competition will be far interesting if hosts intended to fix the noise in the private set (and the competition will be so much more meaningful!)</p>",
      "rawMarkdown": "https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201471\n\nRegarding the noise in training, public and private data, I'm feeling that this competition is a lottery. If there is noise in private (I'm sure there is due to the PB leak), what are we predicting in the models? Are we predicting the annotator's habits..?😷 \nThe competition will be far interesting if hosts intended to fix the noise in the private set (and the competition will be so much more meaningful!)",
      "votes": 65
    },
    {
      "id": 1142756,
      "postDate": "2021-01-07T15:21:30.803Z",
      "content": "<p>This paper might be addressing few concerns around picking models/validating in the presence of noise: <a href=\"https://arxiv.org/pdf/2012.04193.pdf\" target=\"_blank\">Robustness of Accuracy Metric and its Inspirations in Learning with Noisy Labels</a>.</p>\n<pre><code>Under diagonally-dominant class-conditional label noise, the main conclusions are as follows.\n\nA classifier maximizing its accuracy on the noisy distribution is guaranteed to maximize the accuracy on clean distribution.\nWe can obtain an optimal classifier by maximizing training accuracy on sufficiently many noisy samples.\nA noisy validation set is reliable.\n</code></pre>\n<p>My initial guess would be, if label noise is systematic, e.g class conditional noise when most of the leaves are healthy annotators tending to label an image as healthy even though there exists a single disease leaf at the corner, then models would be able to learn it without needing to apply any sophisticated robust training. But if this is not the case, e.g. noise is not predictable by the models than training a robust model with methods studied in literature should be our safe bet for private leaderboard. </p>",
      "rawMarkdown": "This paper might be addressing few concerns around picking models/validating in the presence of noise: [Robustness of Accuracy Metric and its Inspirations in Learning with Noisy Labels](https://arxiv.org/pdf/2012.04193.pdf).\n\n```\nUnder diagonally-dominant class-conditional label noise, the main conclusions are as follows.\n\nA classifier maximizing its accuracy on the noisy distribution is guaranteed to maximize the accuracy on clean distribution.\nWe can obtain an optimal classifier by maximizing training accuracy on sufficiently many noisy samples.\nA noisy validation set is reliable.\n```\n\nMy initial guess would be, if label noise is systematic, e.g class conditional noise when most of the leaves are healthy annotators tending to label an image as healthy even though there exists a single disease leaf at the corner, then models would be able to learn it without needing to apply any sophisticated robust training. But if this is not the case, e.g. noise is not predictable by the models than training a robust model with methods studied in literature should be our safe bet for private leaderboard. ",
      "votes": 24,
      "replies": [
        {
          "id": 1142805,
          "postDate": "2021-01-07T15:42:31.890Z",
          "content": "<p>This is a great paper I read few days ago. </p>\n<p>It clearly prove that methodology is  all we need. </p>",
          "rawMarkdown": "This is a great paper I read few days ago. \n\nIt clearly prove that methodology is  all we need. \n\n",
          "votes": 4
        },
        {
          "id": 1142815,
          "postDate": "2021-01-07T15:52:13.713Z",
          "content": "<p>Yes, it released a lot of stress out of me too :)</p>",
          "rawMarkdown": "Yes, it released a lot of stress out of me too :)",
          "votes": 1
        },
        {
          "id": 1142839,
          "postDate": "2021-01-07T16:12:58.150Z",
          "content": "<p>And it's the first paper I read to validate on noisy data.  While taking care to distiguish and explain systematic and non-systematic noise. </p>",
          "rawMarkdown": "And it's the first paper I read to validate on noisy data.  While taking care to distiguish and explain systematic and non-systematic noise. ",
          "votes": 1
        },
        {
          "id": 1142893,
          "postDate": "2021-01-07T16:44:33.753Z",
          "content": "<p>Exactly! Before coming to this paper I was always wondering if a robust learning paper reports their results on noisy validation or not. They usually use clean validation set, which is not very suitable for real world.</p>",
          "rawMarkdown": "Exactly! Before coming to this paper I was always wondering if a robust learning paper reports their results on noisy validation or not. They usually use clean validation set, which is not very suitable for real world.",
          "votes": 1
        },
        {
          "id": 1142920,
          "postDate": "2021-01-07T17:07:23.400Z",
          "content": "<p>Yep !  And as said by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> in this   <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201471#1106726\" target=\"_blank\">topic</a> .  It may be impossible in real life to collect  only clean data to validate your model. </p>",
          "rawMarkdown": "Yep !  And as said by @hengck23 in this   [topic](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201471#1106726) .  It may be impossible in real life to collect  only clean data to validate your model. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1142305,
      "postDate": "2021-01-07T09:57:42.737Z",
      "content": "<p>Self denoising works . I mean denoising done in Self supervised manner during training (and not manually)<br>\nIt improved both my CV and LB.  I dont know what will happen on the private LB though. </p>\n<p>Anyway , these noise are more likely annotation errors (inherent to humans) than general and biased pattern of the whole data (the latter would make non sense).Thus,  a good and well trained model should give better generalisation and better accuracy in both clean data and clean+noisy data, than a model which overfitted on noisy training data.</p>\n<p>But the difficulty is to train a model which maximize both Robustness and Accuracy in the presence of noise  and most of our models tend to quickly overfit on noise by just <strong>memorizing them</strong>.   That why it's really difficult to get further improvement at some point. This is total contrary of lottery IMHO</p>",
      "rawMarkdown": "Self denoising works . I mean denoising done in Self supervised manner during training (and not manually)\nIt improved both my CV and LB.  I dont know what will happen on the private LB though. \n \nAnyway , these noise are more likely annotation errors (inherent to humans) than general and biased pattern of the whole data (the latter would make non sense).Thus,  a good and well trained model should give better generalisation and better accuracy in both clean data and clean+noisy data, than a model which overfitted on noisy training data.\n\nBut the difficulty is to train a model which maximize both Robustness and Accuracy in the presence of noise  and most of our models tend to quickly overfit on noise by just **memorizing them**.   That why it's really difficult to get further improvement at some point. This is total contrary of lottery IMHO\n",
      "votes": 12,
      "replies": [
        {
          "id": 1142324,
          "postDate": "2021-01-07T10:11:20.490Z",
          "content": "<p>Thanks for the deep insights. I tried self denoising and that's my best model right now but quite far from the top team's scores. Guess I need to dig deeper.</p>\n<p>I liked the previous PANDA competition where the private set noise was guaranteed to be much lower than the train/public sets, this competition would have been interesting for me if the randomness were comforted.</p>",
          "rawMarkdown": "Thanks for the deep insights. I tried self denoising and that's my best model right now but quite far from the top team's scores. Guess I need to dig deeper.\n\nI liked the previous PANDA competition where the private set noise was guaranteed to be much lower than the train/public sets, this competition would have been interesting for me if the randomness were comforted."
        },
        {
          "id": 1142373,
          "postDate": "2021-01-07T10:42:31.873Z",
          "content": "<blockquote>\n  <p>I liked the previous PANDA competition where the private set noise was guaranteed to be much lower than the train/public sets</p>\n</blockquote>\n<p>I didn't participate on PANDA competition .  But did Kaggle confirm this while the competition was still ongoing ? </p>\n<p>I hope there will be much less noise on private LB here too. </p>",
          "rawMarkdown": "> I liked the previous PANDA competition where the private set noise was guaranteed to be much lower than the train/public sets\n\n\nI didn't participate on PANDA competition .  But did Kaggle confirm this while the competition was still ongoing ? \n\nI hope there will be much less noise on private LB here too. "
        },
        {
          "id": 1142391,
          "postDate": "2021-01-07T10:56:01.003Z",
          "content": "<p>Yes, the hosts released a paper on how the data were labeled before the competition.<br>\nFor public, the labeling was done by one expert and for private, the labeling was a voting between three experts.</p>",
          "rawMarkdown": "Yes, the hosts released a paper on how the data were labeled before the competition.\nFor public, the labeling was done by one expert and for private, the labeling was a voting between three experts.",
          "votes": 2
        },
        {
          "id": 1143704,
          "postDate": "2021-01-08T02:36:30.867Z",
          "content": "<p>Do you mind sharing what are some of the methods available for self-supervised denoising?</p>",
          "rawMarkdown": "Do you mind sharing what are some of the methods available for self-supervised denoising?"
        },
        {
          "id": 1143882,
          "postDate": "2021-01-08T05:46:04.640Z",
          "content": "<p>You can read our solution in PANDA<br>\n<a href=\"https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution\" target=\"_blank\">https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution</a></p>",
          "rawMarkdown": "You can read our solution in PANDA\nhttps://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution",
          "votes": 2
        },
        {
          "id": 1143913,
          "postDate": "2021-01-08T06:04:35.327Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1143914,
          "postDate": "2021-01-08T06:08:29.883Z",
          "content": "<p>Thanks a lot :)</p>",
          "rawMarkdown": "Thanks a lot :)"
        },
        {
          "id": 1144188,
          "postDate": "2021-01-08T09:48:52.457Z",
          "content": "<p><a href=\"https://www.kaggle.com/kyoshioka47\" target=\"_blank\">@kyoshioka47</a> could you share the paper on how the data were labeled :D thanks</p>",
          "rawMarkdown": "@kyoshioka47 could you share the paper on how the data were labeled :D thanks"
        },
        {
          "id": 1144350,
          "postDate": "2021-01-08T12:14:44.330Z",
          "content": "<p>that paper I mentioned was for PANDA, you can find it in the competition site :)</p>",
          "rawMarkdown": "that paper I mentioned was for PANDA, you can find it in the competition site :)",
          "votes": 1
        },
        {
          "id": 1145352,
          "postDate": "2021-01-09T04:52:51.780Z",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Can you kindly link some resource on the self supervised data denoising? Coming across image denoising task when searching on google.</p>",
          "rawMarkdown": "@serigne Can you kindly link some resource on the self supervised data denoising? Coming across image denoising task when searching on google."
        }
      ]
    },
    {
      "id": 1144896,
      "postDate": "2021-01-08T18:27:45.633Z",
      "content": "<p>Does anyone know why the best public notebooks say they score <code>LB 0.900+</code> when in list view, but when we go inside none score over 0.900?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F24faacd71988322f89f808ff36e102e5%2Fone.png?generation=1610130442382506&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F2e9febb1bf3d4910ee2af18711d7a55d%2Ftwo.png?generation=1610130452419433&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Does anyone know why the best public notebooks say they score `LB 0.900+` when in list view, but when we go inside none score over 0.900?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F24faacd71988322f89f808ff36e102e5%2Fone.png?generation=1610130442382506&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F2e9febb1bf3d4910ee2af18711d7a55d%2Ftwo.png?generation=1610130452419433&alt=media)",
      "votes": 10,
      "replies": [
        {
          "id": 1144904,
          "postDate": "2021-01-08T18:32:45.843Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I think its because some images where removed from the public/private test set and they did a leaderboard rescore</p>",
          "rawMarkdown": "@cdeotte I think its because some images where removed from the public/private test set and they did a leaderboard rescore",
          "votes": 5
        },
        {
          "id": 1144931,
          "postDate": "2021-01-08T18:50:37.533Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200803\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200803</a></p>\n<p>If you check the discussion above, you can see that public lb has been rescored.<br>\n(But the criteria for sorting notebooks have not been reset.)<br>\n<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
          "rawMarkdown": "https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200803\n\nIf you check the discussion above, you can see that public lb has been rescored.\n(But the criteria for sorting notebooks have not been reset.)\n@cdeotte ",
          "votes": 4
        },
        {
          "id": 1145041,
          "postDate": "2021-01-08T20:31:33.730Z",
          "content": "<p>thanks Yann and Heroseo</p>",
          "rawMarkdown": "thanks Yann and Heroseo",
          "votes": 2
        }
      ]
    },
    {
      "id": 1144182,
      "postDate": "2021-01-08T09:43:49.423Z",
      "content": "<p>In the wheat detection competition, the data set also had noise, but the host responded positively. But the host seldom uttered this game, which is very bad for us. I hope the host can come out and answer some questions.<br>\ne.g. <br>\n1.Is the label really wrong or because we are not experts and we can’t make a correct judgment<br>\n2.division criteria for the coexistence of multiple diseases</p>",
      "rawMarkdown": "In the wheat detection competition, the data set also had noise, but the host responded positively. But the host seldom uttered this game, which is very bad for us. I hope the host can come out and answer some questions.\ne.g. \n1.Is the label really wrong or because we are not experts and we can’t make a correct judgment\n2.division criteria for the coexistence of multiple diseases",
      "votes": 6,
      "replies": [
        {
          "id": 1149468,
          "postDate": "2021-01-11T21:17:44.057Z",
          "content": "<p>This is very good question and should be answered. Is there a way we can tag host of the competition?</p>",
          "rawMarkdown": "This is very good question and should be answered. Is there a way we can tag host of the competition?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1150915,
      "postDate": "2021-01-13T00:43:08.947Z",
      "content": "<p>So far I feel like my experiments have suggested:</p>\n<ul>\n<li><p>filtering out images which look suspiciously labeled from the training has not really had a big positive effect (though am already using some label smoothing). i think with original loss i managed to get around +0.3 or similar compared to a lower starting point.</p></li>\n<li><p>making minor changes to inference…some of which surely should not be negative e.g. more tta cycles…has in some cases lowered LB score by 0.1-0.3. </p></li>\n</ul>\n<p>i dont know if im making the correct interpretation here but it seems to me that this suggests there is some level of 'noise' or blind luck at least in terms of predicting a small portion of borderline images. im not sure if this necessarily says anything for sure about errors in test labeling. i guess it would be surprising if the host could say with confidence that there cannot be a single error in test image labels as i figure it would usually be difficult to diagnose with complete accuracy from photos?</p>\n<p>hopefully there are still plenty of things people will figure out (with or without labeling noise) to make models more accurate overall.</p>",
      "rawMarkdown": "So far I feel like my experiments have suggested:\n\n- filtering out images which look suspiciously labeled from the training has not really had a big positive effect (though am already using some label smoothing). i think with original loss i managed to get around +0.3 or similar compared to a lower starting point.\n\n- making minor changes to inference...some of which surely should not be negative e.g. more tta cycles...has in some cases lowered LB score by 0.1-0.3. \n\ni dont know if im making the correct interpretation here but it seems to me that this suggests there is some level of 'noise' or blind luck at least in terms of predicting a small portion of borderline images. im not sure if this necessarily says anything for sure about errors in test labeling. i guess it would be surprising if the host could say with confidence that there cannot be a single error in test image labels as i figure it would usually be difficult to diagnose with complete accuracy from photos?\n\nhopefully there are still plenty of things people will figure out (with or without labeling noise) to make models more accurate overall.",
      "votes": 3
    },
    {
      "id": 1144121,
      "postDate": "2021-01-08T08:45:35.813Z",
      "content": "<p>It will be a lottery in terms of the leaderboard rank</p>\n<p>1) categorical accuracy<br>\n2) imbalance classes<br>\n3) noisy data<br>\n4) too narrow score between people</p>",
      "rawMarkdown": "It will be a lottery in terms of the leaderboard rank\n\n1) categorical accuracy\n2) imbalance classes\n3) noisy data\n4) too narrow score between people",
      "votes": 3
    },
    {
      "id": 1142497,
      "postDate": "2021-01-07T12:28:22.227Z",
      "content": "<p>What we need is clarification from the hosts about LB and private LB data. If it similar PANDA competition, we need change our model and Pretreatment🤥🤥🤥<br>\nThe gold miner competition will lose the meaning of competition. We want to defeat our opponents with wisdom and skill, not luck.(Although I like gold miners🙈</p>",
      "rawMarkdown": "What we need is clarification from the hosts about LB and private LB data. If it similar PANDA competition, we need change our model and Pretreatment🤥🤥🤥\nThe gold miner competition will lose the meaning of competition. We want to defeat our opponents with wisdom and skill, not luck.(Although I like gold miners🙈",
      "votes": 1
    },
    {
      "id": 1142262,
      "postDate": "2021-01-07T09:03:51.317Z",
      "content": "<p>How do you know that public and private labels are noisy?</p>",
      "rawMarkdown": "How do you know that public and private labels are noisy?",
      "votes": 2,
      "replies": [
        {
          "id": 1142283,
          "postDate": "2021-01-07T09:32:04.537Z",
          "content": "<p>I'm speculating since the LB and PB were consistent when monitored during the data bleach</p>",
          "rawMarkdown": "I'm speculating since the LB and PB were consistent when monitored during the data bleach"
        },
        {
          "id": 1142286,
          "postDate": "2021-01-07T09:37:21.290Z",
          "content": "<p>It only tells that public and private are probably random split. Why do you think that labels are noisy?</p>",
          "rawMarkdown": "It only tells that public and private are probably random split. Why do you think that labels are noisy?",
          "votes": 2
        },
        {
          "id": 1142292,
          "postDate": "2021-01-07T09:46:41.380Z",
          "content": "<p>I feel that labels are somewhat \"random\", when analyzing the confusion between disease and healthy.</p>\n<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673</a><br>\nAs pointed out in this thread, the label annotation strategy may not be consistent especially when there is both disease and healthy leaf inside the image.</p>\n<p>But it's just my opinion, would love to hear your views too.</p>",
          "rawMarkdown": "I feel that labels are somewhat \"random\", when analyzing the confusion between disease and healthy.\n\nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673\nAs pointed out in this thread, the label annotation strategy may not be consistent especially when there is both disease and healthy leaf inside the image.\n\nBut it's just my opinion, would love to hear your views too.",
          "votes": 4
        },
        {
          "id": 1142297,
          "postDate": "2021-01-07T09:50:55.513Z",
          "content": "<p>Ok, so training set labels are noisy, we still don't know if test set labels are noisy. Even if they are, it is the nature of the problem unless the test set noise is biased.</p>",
          "rawMarkdown": "Ok, so training set labels are noisy, we still don't know if test set labels are noisy. Even if they are, it is the nature of the problem unless the test set noise is biased.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1142186,
      "postDate": "2021-01-07T08:11:09.043Z",
      "content": "<p>Well said. Because of this, I lost motivation after achieving 0.898 ☹️</p>",
      "rawMarkdown": "Well said. Because of this, I lost motivation after achieving 0.898 ☹️",
      "votes": 2
    },
    {
      "id": 1142133,
      "postDate": "2021-01-07T07:12:50.080Z",
      "content": "<p>yes, it is.</p>",
      "rawMarkdown": "yes, it is."
    },
    {
      "id": 1145777,
      "postDate": "2021-01-09T10:27:28.587Z",
      "content": "<p>It would be nice if we could get some information on the quality of the test set. Eg. is it as noisy as the training set?</p>",
      "rawMarkdown": "It would be nice if we could get some information on the quality of the test set. Eg. is it as noisy as the training set?"
    },
    {
      "id": 1142580,
      "postDate": "2021-01-07T13:33:57.483Z",
      "content": "<p>Same feelings here. Especially after the PB leak.</p>",
      "rawMarkdown": "Same feelings here. Especially after the PB leak."
    },
    {
      "id": 1142107,
      "postDate": "2021-01-07T06:44:39.303Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1142756,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2021-01-07T15:21:30.803000",
      "content": "<p>This paper might be addressing few concerns around picking models/validating in the presence of noise: <a href=\"https://arxiv.org/pdf/2012.04193.pdf\" target=\"_blank\">Robustness of Accuracy Metric and its Inspirations in Learning with Noisy Labels</a>.</p>\n<pre><code>Under diagonally-dominant class-conditional label noise, the main conclusions are as follows.\n\nA classifier maximizing its accuracy on the noisy distribution is guaranteed to maximize the accuracy on clean distribution.\nWe can obtain an optimal classifier by maximizing training accuracy on sufficiently many noisy samples.\nA noisy validation set is reliable.\n</code></pre>\n<p>My initial guess would be, if label noise is systematic, e.g class conditional noise when most of the leaves are healthy annotators tending to label an image as healthy even though there exists a single disease leaf at the corner, then models would be able to learn it without needing to apply any sophisticated robust training. But if this is not the case, e.g. noise is not predictable by the models than training a robust model with methods studied in literature should be our safe bet for private leaderboard. </p>",
      "votes": 24,
      "replies": [
        {
          "id": 1142805,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-01-07T15:42:31.890000",
          "content": "<p>This is a great paper I read few days ago. </p>\n<p>It clearly prove that methodology is  all we need. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1142815,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2021-01-07T15:52:13.713000",
          "content": "<p>Yes, it released a lot of stress out of me too :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142839,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-01-07T16:12:58.150000",
          "content": "<p>And it's the first paper I read to validate on noisy data.  While taking care to distiguish and explain systematic and non-systematic noise. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142893,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2021-01-07T16:44:33.753000",
          "content": "<p>Exactly! Before coming to this paper I was always wondering if a robust learning paper reports their results on noisy validation or not. They usually use clean validation set, which is not very suitable for real world.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142920,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-01-07T17:07:23.400000",
          "content": "<p>Yep !  And as said by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> in this   <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201471#1106726\" target=\"_blank\">topic</a> .  It may be impossible in real life to collect  only clean data to validate your model. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1142305,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2021-01-07T09:57:42.737000",
      "content": "<p>Self denoising works . I mean denoising done in Self supervised manner during training (and not manually)<br>\nIt improved both my CV and LB.  I dont know what will happen on the private LB though. </p>\n<p>Anyway , these noise are more likely annotation errors (inherent to humans) than general and biased pattern of the whole data (the latter would make non sense).Thus,  a good and well trained model should give better generalisation and better accuracy in both clean data and clean+noisy data, than a model which overfitted on noisy training data.</p>\n<p>But the difficulty is to train a model which maximize both Robustness and Accuracy in the presence of noise  and most of our models tend to quickly overfit on noise by just <strong>memorizing them</strong>.   That why it's really difficult to get further improvement at some point. This is total contrary of lottery IMHO</p>",
      "votes": 12,
      "replies": [
        {
          "id": 1142324,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2021-01-07T10:11:20.490000",
          "content": "<p>Thanks for the deep insights. I tried self denoising and that's my best model right now but quite far from the top team's scores. Guess I need to dig deeper.</p>\n<p>I liked the previous PANDA competition where the private set noise was guaranteed to be much lower than the train/public sets, this competition would have been interesting for me if the randomness were comforted.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142373,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-01-07T10:42:31.873000",
          "content": "<blockquote>\n  <p>I liked the previous PANDA competition where the private set noise was guaranteed to be much lower than the train/public sets</p>\n</blockquote>\n<p>I didn't participate on PANDA competition .  But did Kaggle confirm this while the competition was still ongoing ? </p>\n<p>I hope there will be much less noise on private LB here too. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142391,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2021-01-07T10:56:01.003000",
          "content": "<p>Yes, the hosts released a paper on how the data were labeled before the competition.<br>\nFor public, the labeling was done by one expert and for private, the labeling was a voting between three experts.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1143704,
          "author_name": "Sally80",
          "author_url": "",
          "post_date": "2021-01-08T02:36:30.867000",
          "content": "<p>Do you mind sharing what are some of the methods available for self-supervised denoising?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1143882,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2021-01-08T05:46:04.640000",
          "content": "<p>You can read our solution in PANDA<br>\n<a href=\"https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution\" target=\"_blank\">https://github.com/kentaroy47/Kaggle-PANDA-1st-place-solution</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1143913,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-08T06:04:35.327000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1143914,
          "author_name": "Sally80",
          "author_url": "",
          "post_date": "2021-01-08T06:08:29.883000",
          "content": "<p>Thanks a lot :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1144188,
          "author_name": "Benlei Cui",
          "author_url": "",
          "post_date": "2021-01-08T09:48:52.457000",
          "content": "<p><a href=\"https://www.kaggle.com/kyoshioka47\" target=\"_blank\">@kyoshioka47</a> could you share the paper on how the data were labeled :D thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1144350,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2021-01-08T12:14:44.330000",
          "content": "<p>that paper I mentioned was for PANDA, you can find it in the competition site :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1145352,
          "author_name": "Rahul Deora",
          "author_url": "",
          "post_date": "2021-01-09T04:52:51.780000",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Can you kindly link some resource on the self supervised data denoising? Coming across image denoising task when searching on google.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1144896,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-01-08T18:27:45.633000",
      "content": "<p>Does anyone know why the best public notebooks say they score <code>LB 0.900+</code> when in list view, but when we go inside none score over 0.900?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F24faacd71988322f89f808ff36e102e5%2Fone.png?generation=1610130442382506&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F2e9febb1bf3d4910ee2af18711d7a55d%2Ftwo.png?generation=1610130452419433&amp;alt=media\" alt=\"\"></p>",
      "votes": 10,
      "replies": [
        {
          "id": 1144904,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2021-01-08T18:32:45.843000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I think its because some images where removed from the public/private test set and they did a leaderboard rescore</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1144931,
          "author_name": "Heroseo",
          "author_url": "",
          "post_date": "2021-01-08T18:50:37.533000",
          "content": "<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200803\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200803</a></p>\n<p>If you check the discussion above, you can see that public lb has been rescored.<br>\n(But the criteria for sorting notebooks have not been reset.)<br>\n<a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1145041,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-01-08T20:31:33.730000",
          "content": "<p>thanks Yann and Heroseo</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1144182,
      "author_name": "Benlei Cui",
      "author_url": "",
      "post_date": "2021-01-08T09:43:49.423000",
      "content": "<p>In the wheat detection competition, the data set also had noise, but the host responded positively. But the host seldom uttered this game, which is very bad for us. I hope the host can come out and answer some questions.<br>\ne.g. <br>\n1.Is the label really wrong or because we are not experts and we can’t make a correct judgment<br>\n2.division criteria for the coexistence of multiple diseases</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1149468,
          "author_name": "Sandeep",
          "author_url": "",
          "post_date": "2021-01-11T21:17:44.057000",
          "content": "<p>This is very good question and should be answered. Is there a way we can tag host of the competition?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1150915,
      "author_name": "Dave E",
      "author_url": "",
      "post_date": "2021-01-13T00:43:08.947000",
      "content": "<p>So far I feel like my experiments have suggested:</p>\n<ul>\n<li><p>filtering out images which look suspiciously labeled from the training has not really had a big positive effect (though am already using some label smoothing). i think with original loss i managed to get around +0.3 or similar compared to a lower starting point.</p></li>\n<li><p>making minor changes to inference…some of which surely should not be negative e.g. more tta cycles…has in some cases lowered LB score by 0.1-0.3. </p></li>\n</ul>\n<p>i dont know if im making the correct interpretation here but it seems to me that this suggests there is some level of 'noise' or blind luck at least in terms of predicting a small portion of borderline images. im not sure if this necessarily says anything for sure about errors in test labeling. i guess it would be surprising if the host could say with confidence that there cannot be a single error in test image labels as i figure it would usually be difficult to diagnose with complete accuracy from photos?</p>\n<p>hopefully there are still plenty of things people will figure out (with or without labeling noise) to make models more accurate overall.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1144121,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2021-01-08T08:45:35.813000",
      "content": "<p>It will be a lottery in terms of the leaderboard rank</p>\n<p>1) categorical accuracy<br>\n2) imbalance classes<br>\n3) noisy data<br>\n4) too narrow score between people</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1142497,
      "author_name": "Mingjie Wang",
      "author_url": "",
      "post_date": "2021-01-07T12:28:22.227000",
      "content": "<p>What we need is clarification from the hosts about LB and private LB data. If it similar PANDA competition, we need change our model and Pretreatment🤥🤥🤥<br>\nThe gold miner competition will lose the meaning of competition. We want to defeat our opponents with wisdom and skill, not luck.(Although I like gold miners🙈</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1142262,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2021-01-07T09:03:51.317000",
      "content": "<p>How do you know that public and private labels are noisy?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1142283,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2021-01-07T09:32:04.537000",
          "content": "<p>I'm speculating since the LB and PB were consistent when monitored during the data bleach</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142286,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2021-01-07T09:37:21.290000",
          "content": "<p>It only tells that public and private are probably random split. Why do you think that labels are noisy?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1142292,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2021-01-07T09:46:41.380000",
          "content": "<p>I feel that labels are somewhat \"random\", when analyzing the confusion between disease and healthy.</p>\n<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202673</a><br>\nAs pointed out in this thread, the label annotation strategy may not be consistent especially when there is both disease and healthy leaf inside the image.</p>\n<p>But it's just my opinion, would love to hear your views too.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1142297,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2021-01-07T09:50:55.513000",
          "content": "<p>Ok, so training set labels are noisy, we still don't know if test set labels are noisy. Even if they are, it is the nature of the problem unless the test set noise is biased.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1142186,
      "author_name": "Kaushal Shah",
      "author_url": "",
      "post_date": "2021-01-07T08:11:09.043000",
      "content": "<p>Well said. Because of this, I lost motivation after achieving 0.898 ☹️</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1142133,
      "author_name": "Tian",
      "author_url": "",
      "post_date": "2021-01-07T07:12:50.080000",
      "content": "<p>yes, it is.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1145777,
      "author_name": "SohamTamba",
      "author_url": "",
      "post_date": "2021-01-09T10:27:28.587000",
      "content": "<p>It would be nice if we could get some information on the quality of the test set. Eg. is it as noisy as the training set?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1142580,
      "author_name": "Tsai29",
      "author_url": "",
      "post_date": "2021-01-07T13:33:57.483000",
      "content": "<p>Same feelings here. Especially after the PB leak.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1142107,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-07T06:44:39.303000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1142071": "https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/201471\n\nRegarding the noise in training, public and private data, I'm feeling that this competition is a lottery. If there is noise in private (I'm sure there is due to the PB leak), what are we predicting in the models? Are we predicting the annotator's habits..?😷 \nThe competition will be far interesting if hosts intended to fix the noise in the private set (and the competition will be so much more meaningful!)",
    "1142756": "This paper might be addressing few concerns around picking models/validating in the presence of noise: [Robustness of Accuracy Metric and its Inspirations in Learning with Noisy Labels](https://arxiv.org/pdf/2012.04193.pdf).\n\n```\nUnder diagonally-dominant class-conditional label noise, the main conclusions are as follows.\n\nA classifier maximizing its accuracy on the noisy distribution is guaranteed to maximize the accuracy on clean distribution.\nWe can obtain an optimal classifier by maximizing training accuracy on sufficiently many noisy samples.\nA noisy validation set is reliable.\n```\n\nMy initial guess would be, if label noise is systematic, e.g class conditional noise when most of the leaves are healthy annotators tending to label an image as healthy even though there exists a single disease leaf at the corner, then models would be able to learn it without needing to apply any sophisticated robust training. But if this is not the case, e.g. noise is not predictable by the models than training a robust model with methods studied in literature should be our safe bet for private leaderboard. ",
    "1142305": "Self denoising works . I mean denoising done in Self supervised manner during training (and not manually)\nIt improved both my CV and LB.  I dont know what will happen on the private LB though. \n \nAnyway , these noise are more likely annotation errors (inherent to humans) than general and biased pattern of the whole data (the latter would make non sense).Thus,  a good and well trained model should give better generalisation and better accuracy in both clean data and clean+noisy data, than a model which overfitted on noisy training data.\n\nBut the difficulty is to train a model which maximize both Robustness and Accuracy in the presence of noise  and most of our models tend to quickly overfit on noise by just **memorizing them**.   That why it's really difficult to get further improvement at some point. This is total contrary of lottery IMHO\n",
    "1144896": "Does anyone know why the best public notebooks say they score `LB 0.900+` when in list view, but when we go inside none score over 0.900?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F24faacd71988322f89f808ff36e102e5%2Fone.png?generation=1610130442382506&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F2e9febb1bf3d4910ee2af18711d7a55d%2Ftwo.png?generation=1610130452419433&alt=media)",
    "1144182": "In the wheat detection competition, the data set also had noise, but the host responded positively. But the host seldom uttered this game, which is very bad for us. I hope the host can come out and answer some questions.\ne.g. \n1.Is the label really wrong or because we are not experts and we can’t make a correct judgment\n2.division criteria for the coexistence of multiple diseases",
    "1150915": "So far I feel like my experiments have suggested:\n\n- filtering out images which look suspiciously labeled from the training has not really had a big positive effect (though am already using some label smoothing). i think with original loss i managed to get around +0.3 or similar compared to a lower starting point.\n\n- making minor changes to inference...some of which surely should not be negative e.g. more tta cycles...has in some cases lowered LB score by 0.1-0.3. \n\ni dont know if im making the correct interpretation here but it seems to me that this suggests there is some level of 'noise' or blind luck at least in terms of predicting a small portion of borderline images. im not sure if this necessarily says anything for sure about errors in test labeling. i guess it would be surprising if the host could say with confidence that there cannot be a single error in test image labels as i figure it would usually be difficult to diagnose with complete accuracy from photos?\n\nhopefully there are still plenty of things people will figure out (with or without labeling noise) to make models more accurate overall.",
    "1144121": "It will be a lottery in terms of the leaderboard rank\n\n1) categorical accuracy\n2) imbalance classes\n3) noisy data\n4) too narrow score between people",
    "1142497": "What we need is clarification from the hosts about LB and private LB data. If it similar PANDA competition, we need change our model and Pretreatment🤥🤥🤥\nThe gold miner competition will lose the meaning of competition. We want to defeat our opponents with wisdom and skill, not luck.(Although I like gold miners🙈",
    "1142262": "How do you know that public and private labels are noisy?",
    "1142186": "Well said. Because of this, I lost motivation after achieving 0.898 ☹️",
    "1142133": "yes, it is.",
    "1145777": "It would be nice if we could get some information on the quality of the test set. Eg. is it as noisy as the training set?",
    "1142580": "Same feelings here. Especially after the PB leak.",
    "1142107": ""
  }
}