{
  "id": 219621,
  "title": "Cleaner Dataset = worse performance?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/219621",
  "author_name": "",
  "post_date": "2021-02-15T18:53:11.291340100Z",
  "votes": 11,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I'm quite baffled by some of the results I'm getting. Maybe some of you have an idea of what’s causing this. </p>\n<p>I've started with just the 2020 dataset and got a CV score of ~90.3% and 89.4% LB.<br>\nAfter that I started cleaning the merged (2020 and 2019) dataset by removing a few hundred mislabeled images. At this point I have done a few iterations of this and have removed about 1500 images from the ds. Now the part that I don’t understand. Every iteration my CV score improved but my LB score dropped by about 0.1 - 0.3%. Now I am at 94.7% CV (so the cleaning is obviously helping) and yet my LB score is down at 88.7%…<br>\nDoes anybody have an idea why this could be happening? Any input is appreciated!</p>\n<p>I’m using EfficientNetB4 with CutMix, GridMask and some other augs (flip, crop, etc.)</p>",
  "messages": [
    {
      "id": "1201934",
      "postDate": "02/15/2021 18:53:11",
      "content": "<p>I'm quite baffled by some of the results I'm getting. Maybe some of you have an idea of what’s causing this. </p>\n<p>I've started with just the 2020 dataset and got a CV score of ~90.3% and 89.4% LB.<br>\nAfter that I started cleaning the merged (2020 and 2019) dataset by removing a few hundred mislabeled images. At this point I have done a few iterations of this and have removed about 1500 images from the ds. Now the part that I don’t understand. Every iteration my CV score improved but my LB score dropped by about 0.1 - 0.3%. Now I am at 94.7% CV (so the cleaning is obviously helping) and yet my LB score is down at 88.7%…<br>\nDoes anybody have an idea why this could be happening? Any input is appreciated!</p>\n<p>I’m using EfficientNetB4 with CutMix, GridMask and some other augs (flip, crop, etc.)</p>",
      "rawMarkdown": "I'm quite baffled by some of the results I'm getting. Maybe some of you have an idea of what’s causing this. \n\nI've started with just the 2020 dataset and got a CV score of ~90.3% and 89.4% LB.\nAfter that I started cleaning the merged (2020 and 2019) dataset by removing a few hundred mislabeled images. At this point I have done a few iterations of this and have removed about 1500 images from the ds. Now the part that I don’t understand. Every iteration my CV score improved but my LB score dropped by about 0.1 - 0.3%. Now I am at 94.7% CV (so the cleaning is obviously helping) and yet my LB score is down at 88.7%...\nDoes anybody have an idea why this could be happening? Any input is appreciated!\n\nI’m using EfficientNetB4 with CutMix, GridMask and some other augs (flip, crop, etc.)",
      "votes": null
    },
    {
      "id": "1201964",
      "postDate": "02/15/2021 19:17:56",
      "content": "<p>I have similar results with EfficientNet when using 2019 dataset. I think it's something to do with image dimensions because 2020 dataset you don't have to resize first but in 2019 dataset. There are many images with dimensions smaller than 512 ( 600x500 or even with 387 or something). </p>\n<p>I tried to play around with augmentation and rescale in the inference submission but got no luck. I just gave up :(</p>",
      "rawMarkdown": "I have similar results with EfficientNet when using 2019 dataset. I think it's something to do with image dimensions because 2020 dataset you don't have to resize first but in 2019 dataset. There are many images with dimensions smaller than 512 ( 600x500 or even with 387 or something). \n\nI tried to play around with augmentation and rescale in the inference submission but got no luck. I just gave up :(",
      "votes": null
    },
    {
      "id": "1201991",
      "postDate": "02/15/2021 19:34:43",
      "content": "<p>It is probably because the test images also have many mislabeled images. When you models learn from noisy labels, they may also learn how to be prone to label the mistaken ones.</p>",
      "rawMarkdown": "It is probably because the test images also have many mislabeled images. When you models learn from noisy labels, they may also learn how to be prone to label the mistaken ones.",
      "votes": null
    },
    {
      "id": "1202013",
      "postDate": "02/15/2021 19:52:26",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a>. I also suggest performing CV on the entire dataset (and/or make predictions for filtered out labels specifically), not only the cleaned dataset to observe the impact of cleaning, doing vice versa is just deceiving yourself. </p>\n<p>As for me, filtering out heavily mislabeled predictions led to worse CV and LB.</p>",
      "rawMarkdown": "I agree with @woshifym. I also suggest performing CV on the entire dataset (and/or make predictions for filtered out labels specifically), not only the cleaned dataset to observe the impact of cleaning, doing vice versa is just deceiving yourself. \n\nAs for me, filtering out heavily mislabeled predictions led to worse CV and LB.",
      "votes": null
    },
    {
      "id": "1202081",
      "postDate": "02/15/2021 20:30:28",
      "content": "<p>I tried it myself and spent a week to remove mislabelled identical images but my LB decreased because of it. The reason why it happened is because the test data is noisy like train data, this approach will not work. We need to train with the noisy data.</p>",
      "rawMarkdown": "I tried it myself and spent a week to remove mislabelled identical images but my LB decreased because of it. The reason why it happened is because the test data is noisy like train data, this approach will not work. We need to train with the noisy data.",
      "votes": null
    },
    {
      "id": "1202111",
      "postDate": "02/15/2021 21:07:55",
      "content": "<p>Thanks for your answers! That makes sence. So I guess if the model learns in noise it may be able to find a (slight)structure in the mislabled images? Now another question would be which of the models will perform better at the real world task? The one that learned in noise or the one that learned on \"clean\" data? In my head it should be the clean data one. Any thoughts on that?</p>",
      "rawMarkdown": "Thanks for your answers! That makes sence. So I guess if the model learns in noise it may be able to find a (slight)structure in the mislabled images? Now another question would be which of the models will perform better at the real world task? The one that learned in noise or the one that learned on \"clean\" data? In my head it should be the clean data one. Any thoughts on that?",
      "votes": null
    },
    {
      "id": "1202112",
      "postDate": "02/15/2021 21:09:20",
      "content": "<p>Thats a great point! I should validate the models on the original data since thats how the test set will look like. Never thought about that thanks!</p>",
      "rawMarkdown": "Thats a great point! I should validate the models on the original data since thats how the test set will look like. Never thought about that thanks!",
      "votes": null
    },
    {
      "id": "1204193",
      "postDate": "02/15/2021 23:00:36",
      "content": "<p>In the real world, I think the model trained on \"clean\" data would definitely perform the best. But, in real world, it is hard to make a \"clean\" dataset because experts would also make mistakes when labeling the images. So, dealing with noisy labels is always the challenge. </p>\n<p>For this competition, we deal with the unseen \"labeled\" images. We do not know how much noisy the test data is, so different experiments should be done to explore the various possibilities. Probably we are overfitting this competition.</p>",
      "rawMarkdown": "In the real world, I think the model trained on \"clean\" data would definitely perform the best. But, in real world, it is hard to make a \"clean\" dataset because experts would also make mistakes when labeling the images. So, dealing with noisy labels is always the challenge. \n\nFor this competition, we deal with the unseen \"labeled\" images. We do not know how much noisy the test data is, so different experiments should be done to explore the various possibilities. Probably we are overfitting this competition.",
      "votes": null
    },
    {
      "id": "1204299",
      "postDate": "02/16/2021 03:20:19",
      "content": "<p>I observed the same. After removing 1k images, I get a local CV of 0.93 over the full train dataset. Which leads me to believe that the visible portion of the test dataset (public LB) has a noisy distribution. If the rest of the private LB is similar, this model will perform pretty bad. On the contrary, if the private LB is cleaner this would perform pretty well. I expect a shakeup if the private LB is Iess noisy and models that were overfit for the public LB would move to lower rankings. </p>",
      "rawMarkdown": "I observed the same. After removing 1k images, I get a local CV of 0.93 over the full train dataset. Which leads me to believe that the visible portion of the test dataset (public LB) has a noisy distribution. If the rest of the private LB is similar, this model will perform pretty bad. On the contrary, if the private LB is cleaner this would perform pretty well. I expect a shakeup if the private LB is Iess noisy and models that were overfit for the public LB would move to lower rankings.",
      "votes": null
    },
    {
      "id": "1204392",
      "postDate": "02/16/2021 05:39:56",
      "content": "<p>You should use the original data for validation and cleaner data for training to confirm the improvement of your models.<br>\nIf you use the relabeled images for validation, the score of validation must be improved because cleaner data is easy to predict.<br>\nActually I tried that, CV were slightly improved, but public LB was not changed.</p>",
      "rawMarkdown": "You should use the original data for validation and cleaner data for training to confirm the improvement of your models.\nIf you use the relabeled images for validation, the score of validation must be improved because cleaner data is easy to predict.\nActually I tried that, CV were slightly improved, but public LB was not changed.",
      "votes": null
    },
    {
      "id": "1204395",
      "postDate": "02/16/2021 05:48:51",
      "content": "<p>You can try bi tempered loss which is used to train in noisy labels <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215911\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "You can try bi tempered loss which is used to train in noisy labels [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215911)",
      "votes": null
    },
    {
      "id": "1204550",
      "postDate": "02/16/2021 07:54:08",
      "content": "<p>I met similar situations.<br>\nI denoise 2019merged data with full data trained models.<br>\nIn denoise training, fold validation is not denoised, and local CV was not improved.<br>\nPublic LB worse with denoised data model,,,,  I think why this happen because noise has any pattern. </p>",
      "rawMarkdown": "I met similar situations.\nI denoise 2019merged data with full data trained models.\nIn denoise training, fold validation is not denoised, and local CV was not improved.\nPublic LB worse with denoised data model,,,,  I think why this happen because noise has any pattern.",
      "votes": null
    },
    {
      "id": "1204899",
      "postDate": "02/16/2021 12:10:24",
      "content": "<p>+1. After relableling the dataset my CV has been improved from 0.900 to 0.920, but LB scores only 0.895 (5 folds), when single fold trained on the noisy data scores 0.900. I think the people who use knowledge distillation will get they real places in the private LB, where should be much less noisy data. Otherwise, it'll be just a kind of lottery and we'll see random solutions in the green and gold zones.</p>",
      "rawMarkdown": "1. After relableling the dataset my CV has been improved from 0.900 to 0.920, but LB scores only 0.895 (5 folds), when single fold trained on the noisy data scores 0.900. I think the people who use knowledge distillation will get they real places in the private LB, where should be much less noisy data. Otherwise, it'll be just a kind of lottery and we'll see random solutions in the green and gold zones.",
      "votes": null
    },
    {
      "id": "1205136",
      "postDate": "02/16/2021 14:55:57",
      "content": "<p>Manual relabelling introduces confirmation bias  and you should keep the same original validation set. </p>",
      "rawMarkdown": "Manual relabelling introduces confirmation bias  and you should keep the same original validation set.",
      "votes": null
    },
    {
      "id": "1205544",
      "postDate": "02/16/2021 18:52:51",
      "content": "<blockquote>\n  <p>I think the people who use knowledge distillation will get they real places in the private LB</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/vadimtimakin\" target=\"_blank\">@vadimtimakin</a> What do you mean by this? You think that those using knowledge distillation will be shaken up or down in the private LB?</p>\n<p>According to <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/207410\" target=\"_blank\">the private score breach from 2 months ago</a>, private score is very close to the public LB score, it's very likely that the private set is as noisy as the public one, I think that whatever effect knowledge distillation has on the public will also apply for private.</p>",
      "rawMarkdown": "> I think the people who use knowledge distillation will get they real places in the private LB\n\n@vadimtimakin What do you mean by this? You think that those using knowledge distillation will be shaken up or down in the private LB?\n\nAccording to [the private score breach from 2 months ago](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/207410), private score is very close to the public LB score, it's very likely that the private set is as noisy as the public one, I think that whatever effect knowledge distillation has on the public will also apply for private.",
      "votes": null
    },
    {
      "id": "1205625",
      "postDate": "02/16/2021 20:28:46",
      "content": "<p>Everyone understands that shake-up will be huge if the private data is clear. However, some people say that it's false and try to prove it remember this post and the fact that according to those data the difference between public and private parts is small. I can explain by saying that it was only a beginning competition and most of the people hadn't relabeled the data yet by that moment. I'm sure that shake-up will affect at least  people who trained their models on the cleaner data (in their favor).</p>",
      "rawMarkdown": "Everyone understands that shake-up will be huge if the private data is clear. However, some people say that it's false and try to prove it remember this post and the fact that according to those data the difference between public and private parts is small. I can explain by saying that it was only a beginning competition and most of the people hadn't relabeled the data yet by that moment. I'm sure that shake-up will affect at least  people who trained their models on the cleaner data (in their favor).",
      "votes": null
    },
    {
      "id": "1207218",
      "postDate": "02/17/2021 18:48:09",
      "content": "<p>Sorry simple question here. How do you manually relabel? I started working through the images and had a hard time determining which disease was which?</p>",
      "rawMarkdown": "Sorry simple question here. How do you manually relabel? I started working through the images and had a hard time determining which disease was which?",
      "votes": null
    },
    {
      "id": "1208315",
      "postDate": "02/18/2021 08:41:31",
      "content": "<p>I don't manually relabel anything</p>",
      "rawMarkdown": "I don't manually relabel anything",
      "votes": null
    },
    {
      "id": "1208703",
      "postDate": "02/18/2021 12:54:19",
      "content": "<p>Hi Max, newbie question here. Did you create a process to clean the data, so you did not do it manually?<br>\nWhich would be quite a bit of work.<br>\nThanks<br>\n_Bob</p>",
      "rawMarkdown": "Hi Max, newbie question here. Did you create a process to clean the data, so you did not do it manually?\nWhich would be quite a bit of work.\nThanks\n_Bob",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1201964,
      "author_name": "tom88jerry",
      "author_url": "",
      "post_date": "02/15/2021 19:17:56",
      "content": "<p>I have similar results with EfficientNet when using 2019 dataset. I think it's something to do with image dimensions because 2020 dataset you don't have to resize first but in 2019 dataset. There are many images with dimensions smaller than 512 ( 600x500 or even with 387 or something). </p>\n<p>I tried to play around with augmentation and rescale in the inference submission but got no luck. I just gave up :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1201991,
      "author_name": "woshifym",
      "author_url": "",
      "post_date": "02/15/2021 19:34:43",
      "content": "<p>It is probably because the test images also have many mislabeled images. When you models learn from noisy labels, they may also learn how to be prone to label the mistaken ones.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1202013,
      "author_name": "nickuzmenkov",
      "author_url": "",
      "post_date": "02/15/2021 19:52:26",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a>. I also suggest performing CV on the entire dataset (and/or make predictions for filtered out labels specifically), not only the cleaned dataset to observe the impact of cleaning, doing vice versa is just deceiving yourself. </p>\n<p>As for me, filtering out heavily mislabeled predictions led to worse CV and LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1202112,
          "author_name": "maxneumann",
          "author_url": "",
          "post_date": "02/15/2021 21:09:20",
          "content": "<p>Thats a great point! I should validate the models on the original data since thats how the test set will look like. Never thought about that thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1202081,
      "author_name": "vickygoyal",
      "author_url": "",
      "post_date": "02/15/2021 20:30:28",
      "content": "<p>I tried it myself and spent a week to remove mislabelled identical images but my LB decreased because of it. The reason why it happened is because the test data is noisy like train data, this approach will not work. We need to train with the noisy data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1202111,
      "author_name": "maxneumann",
      "author_url": "",
      "post_date": "02/15/2021 21:07:55",
      "content": "<p>Thanks for your answers! That makes sence. So I guess if the model learns in noise it may be able to find a (slight)structure in the mislabled images? Now another question would be which of the models will perform better at the real world task? The one that learned in noise or the one that learned on \"clean\" data? In my head it should be the clean data one. Any thoughts on that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1204193,
          "author_name": "woshifym",
          "author_url": "",
          "post_date": "02/15/2021 23:00:36",
          "content": "<p>In the real world, I think the model trained on \"clean\" data would definitely perform the best. But, in real world, it is hard to make a \"clean\" dataset because experts would also make mistakes when labeling the images. So, dealing with noisy labels is always the challenge. </p>\n<p>For this competition, we deal with the unseen \"labeled\" images. We do not know how much noisy the test data is, so different experiments should be done to explore the various possibilities. Probably we are overfitting this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1204299,
      "author_name": "trushk",
      "author_url": "",
      "post_date": "02/16/2021 03:20:19",
      "content": "<p>I observed the same. After removing 1k images, I get a local CV of 0.93 over the full train dataset. Which leads me to believe that the visible portion of the test dataset (public LB) has a noisy distribution. If the rest of the private LB is similar, this model will perform pretty bad. On the contrary, if the private LB is cleaner this would perform pretty well. I expect a shakeup if the private LB is Iess noisy and models that were overfit for the public LB would move to lower rankings. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1204392,
      "author_name": "yosukeyama",
      "author_url": "",
      "post_date": "02/16/2021 05:39:56",
      "content": "<p>You should use the original data for validation and cleaner data for training to confirm the improvement of your models.<br>\nIf you use the relabeled images for validation, the score of validation must be improved because cleaner data is easy to predict.<br>\nActually I tried that, CV were slightly improved, but public LB was not changed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1204395,
      "author_name": "vickygoyal",
      "author_url": "",
      "post_date": "02/16/2021 05:48:51",
      "content": "<p>You can try bi tempered loss which is used to train in noisy labels <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215911\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1204550,
      "author_name": "yoshito",
      "author_url": "",
      "post_date": "02/16/2021 07:54:08",
      "content": "<p>I met similar situations.<br>\nI denoise 2019merged data with full data trained models.<br>\nIn denoise training, fold validation is not denoised, and local CV was not improved.<br>\nPublic LB worse with denoised data model,,,,  I think why this happen because noise has any pattern. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1204899,
      "author_name": "vadimtimakin",
      "author_url": "",
      "post_date": "02/16/2021 12:10:24",
      "content": "<p>+1. After relableling the dataset my CV has been improved from 0.900 to 0.920, but LB scores only 0.895 (5 folds), when single fold trained on the noisy data scores 0.900. I think the people who use knowledge distillation will get they real places in the private LB, where should be much less noisy data. Otherwise, it'll be just a kind of lottery and we'll see random solutions in the green and gold zones.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1205544,
          "author_name": "amiiiney",
          "author_url": "",
          "post_date": "02/16/2021 18:52:51",
          "content": "<blockquote>\n  <p>I think the people who use knowledge distillation will get they real places in the private LB</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/vadimtimakin\" target=\"_blank\">@vadimtimakin</a> What do you mean by this? You think that those using knowledge distillation will be shaken up or down in the private LB?</p>\n<p>According to <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/207410\" target=\"_blank\">the private score breach from 2 months ago</a>, private score is very close to the public LB score, it's very likely that the private set is as noisy as the public one, I think that whatever effect knowledge distillation has on the public will also apply for private.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1205625,
          "author_name": "vadimtimakin",
          "author_url": "",
          "post_date": "02/16/2021 20:28:46",
          "content": "<p>Everyone understands that shake-up will be huge if the private data is clear. However, some people say that it's false and try to prove it remember this post and the fact that according to those data the difference between public and private parts is small. I can explain by saying that it was only a beginning competition and most of the people hadn't relabeled the data yet by that moment. I'm sure that shake-up will affect at least  people who trained their models on the cleaner data (in their favor).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1205136,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "02/16/2021 14:55:57",
      "content": "<p>Manual relabelling introduces confirmation bias  and you should keep the same original validation set. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1207218,
          "author_name": "bobbybobsterbutwhy",
          "author_url": "",
          "post_date": "02/17/2021 18:48:09",
          "content": "<p>Sorry simple question here. How do you manually relabel? I started working through the images and had a hard time determining which disease was which?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208315,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "02/18/2021 08:41:31",
          "content": "<p>I don't manually relabel anything</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208703,
      "author_name": "bobbybobsterbutwhy",
      "author_url": "",
      "post_date": "02/18/2021 12:54:19",
      "content": "<p>Hi Max, newbie question here. Did you create a process to clean the data, so you did not do it manually?<br>\nWhich would be quite a bit of work.<br>\nThanks<br>\n_Bob</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1201934": "I'm quite baffled by some of the results I'm getting. Maybe some of you have an idea of what’s causing this. \n\nI've started with just the 2020 dataset and got a CV score of ~90.3% and 89.4% LB.\nAfter that I started cleaning the merged (2020 and 2019) dataset by removing a few hundred mislabeled images. At this point I have done a few iterations of this and have removed about 1500 images from the ds. Now the part that I don’t understand. Every iteration my CV score improved but my LB score dropped by about 0.1 - 0.3%. Now I am at 94.7% CV (so the cleaning is obviously helping) and yet my LB score is down at 88.7%...\nDoes anybody have an idea why this could be happening? Any input is appreciated!\n\nI’m using EfficientNetB4 with CutMix, GridMask and some other augs (flip, crop, etc.)",
    "1201964": "I have similar results with EfficientNet when using 2019 dataset. I think it's something to do with image dimensions because 2020 dataset you don't have to resize first but in 2019 dataset. There are many images with dimensions smaller than 512 ( 600x500 or even with 387 or something). \n\nI tried to play around with augmentation and rescale in the inference submission but got no luck. I just gave up :(",
    "1201991": "It is probably because the test images also have many mislabeled images. When you models learn from noisy labels, they may also learn how to be prone to label the mistaken ones.",
    "1202013": "I agree with @woshifym. I also suggest performing CV on the entire dataset (and/or make predictions for filtered out labels specifically), not only the cleaned dataset to observe the impact of cleaning, doing vice versa is just deceiving yourself. \n\nAs for me, filtering out heavily mislabeled predictions led to worse CV and LB.",
    "1202081": "I tried it myself and spent a week to remove mislabelled identical images but my LB decreased because of it. The reason why it happened is because the test data is noisy like train data, this approach will not work. We need to train with the noisy data.",
    "1202111": "Thanks for your answers! That makes sence. So I guess if the model learns in noise it may be able to find a (slight)structure in the mislabled images? Now another question would be which of the models will perform better at the real world task? The one that learned in noise or the one that learned on \"clean\" data? In my head it should be the clean data one. Any thoughts on that?",
    "1202112": "Thats a great point! I should validate the models on the original data since thats how the test set will look like. Never thought about that thanks!",
    "1204193": "In the real world, I think the model trained on \"clean\" data would definitely perform the best. But, in real world, it is hard to make a \"clean\" dataset because experts would also make mistakes when labeling the images. So, dealing with noisy labels is always the challenge. \n\nFor this competition, we deal with the unseen \"labeled\" images. We do not know how much noisy the test data is, so different experiments should be done to explore the various possibilities. Probably we are overfitting this competition.",
    "1204299": "I observed the same. After removing 1k images, I get a local CV of 0.93 over the full train dataset. Which leads me to believe that the visible portion of the test dataset (public LB) has a noisy distribution. If the rest of the private LB is similar, this model will perform pretty bad. On the contrary, if the private LB is cleaner this would perform pretty well. I expect a shakeup if the private LB is Iess noisy and models that were overfit for the public LB would move to lower rankings.",
    "1204392": "You should use the original data for validation and cleaner data for training to confirm the improvement of your models.\nIf you use the relabeled images for validation, the score of validation must be improved because cleaner data is easy to predict.\nActually I tried that, CV were slightly improved, but public LB was not changed.",
    "1204395": "You can try bi tempered loss which is used to train in noisy labels [here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215911)",
    "1204550": "I met similar situations.\nI denoise 2019merged data with full data trained models.\nIn denoise training, fold validation is not denoised, and local CV was not improved.\nPublic LB worse with denoised data model,,,,  I think why this happen because noise has any pattern.",
    "1204899": "1. After relableling the dataset my CV has been improved from 0.900 to 0.920, but LB scores only 0.895 (5 folds), when single fold trained on the noisy data scores 0.900. I think the people who use knowledge distillation will get they real places in the private LB, where should be much less noisy data. Otherwise, it'll be just a kind of lottery and we'll see random solutions in the green and gold zones.",
    "1205136": "Manual relabelling introduces confirmation bias  and you should keep the same original validation set.",
    "1205544": "> I think the people who use knowledge distillation will get they real places in the private LB\n\n@vadimtimakin What do you mean by this? You think that those using knowledge distillation will be shaken up or down in the private LB?\n\nAccording to [the private score breach from 2 months ago](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/207410), private score is very close to the public LB score, it's very likely that the private set is as noisy as the public one, I think that whatever effect knowledge distillation has on the public will also apply for private.",
    "1205625": "Everyone understands that shake-up will be huge if the private data is clear. However, some people say that it's false and try to prove it remember this post and the fact that according to those data the difference between public and private parts is small. I can explain by saying that it was only a beginning competition and most of the people hadn't relabeled the data yet by that moment. I'm sure that shake-up will affect at least  people who trained their models on the cleaner data (in their favor).",
    "1207218": "Sorry simple question here. How do you manually relabel? I started working through the images and had a hard time determining which disease was which?",
    "1208315": "I don't manually relabel anything",
    "1208703": "Hi Max, newbie question here. Did you create a process to clean the data, so you did not do it manually?\nWhich would be quite a bit of work.\nThanks\n_Bob"
  },
  "source": "meta"
}