{
  "id": 202206,
  "title": "Why we have a noisy dataset (CutMix vs MixUp)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/202206",
  "author_name": "JGar",
  "post_date": "2020-12-08T23:26:48.650000",
  "votes": 12,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I have seen kernels using both image augmentation via <strong>CutMix</strong> and <strong>MixUp</strong>. Nevertheless I would like to open a discussion on whether to choose and why.</p>\n<ul>\n<li><strong>CutMix</strong> is adding a cropped image in the corner of another image, and change the training label giving both different weights in function of the respective sizes . </li>\n<li><strong>MixUp</strong> is supperposing the images also with corresponding weighted labels. </li>\n</ul>\n<p>They have both proved to give better results on <strong>ImageNet</strong>.</p>\n<p>However, with the <strong>Cassava Dataset</strong>, I'm afraid it would add up statistically too much noise from the other image by showing different scaled images (some are showing close up leaves, some are more like a global view). So I doubt it would be a good choice to use them. I tried both and am not convinced. In this kernel <a href=\"https://www.kaggle.com/tanlikesmath/cassava-classification-eda-fastai-starter\" target=\"_blank\">this kernel</a> , the performance improved when adding it at the same time with smooth labels. Smooth labels have given improvement for us, what about MixUp/CutMix?</p>\n<p>Also, <a href=\"https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations\" target=\"_blank\">this (other) kernel</a> also showed an interesting feature when clustering the data in different groups. It seems that different clusters partly meant different distance to the photographer. It could mean that the noise in our dataset is somehow partly due to the distance at which each photo were taken. I would be happy to hear your thoughts on this.</p>",
  "messages": [
    {
      "id": 1106566,
      "postDate": "2020-12-08T23:26:48.650Z",
      "content": "<p>I have seen kernels using both image augmentation via <strong>CutMix</strong> and <strong>MixUp</strong>. Nevertheless I would like to open a discussion on whether to choose and why.</p>\n<ul>\n<li><strong>CutMix</strong> is adding a cropped image in the corner of another image, and change the training label giving both different weights in function of the respective sizes . </li>\n<li><strong>MixUp</strong> is supperposing the images also with corresponding weighted labels. </li>\n</ul>\n<p>They have both proved to give better results on <strong>ImageNet</strong>.</p>\n<p>However, with the <strong>Cassava Dataset</strong>, I'm afraid it would add up statistically too much noise from the other image by showing different scaled images (some are showing close up leaves, some are more like a global view). So I doubt it would be a good choice to use them. I tried both and am not convinced. In this kernel <a href=\"https://www.kaggle.com/tanlikesmath/cassava-classification-eda-fastai-starter\" target=\"_blank\">this kernel</a> , the performance improved when adding it at the same time with smooth labels. Smooth labels have given improvement for us, what about MixUp/CutMix?</p>\n<p>Also, <a href=\"https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations\" target=\"_blank\">this (other) kernel</a> also showed an interesting feature when clustering the data in different groups. It seems that different clusters partly meant different distance to the photographer. It could mean that the noise in our dataset is somehow partly due to the distance at which each photo were taken. I would be happy to hear your thoughts on this.</p>",
      "rawMarkdown": "I have seen kernels using both image augmentation via **CutMix** and **MixUp**. Nevertheless I would like to open a discussion on whether to choose and why.\n\n- **CutMix** is adding a cropped image in the corner of another image, and change the training label giving both different weights in function of the respective sizes . \n- **MixUp** is supperposing the images also with corresponding weighted labels. \n\nThey have both proved to give better results on **ImageNet**.\n\nHowever, with the **Cassava Dataset**, I'm afraid it would add up statistically too much noise from the other image by showing different scaled images (some are showing close up leaves, some are more like a global view). So I doubt it would be a good choice to use them. I tried both and am not convinced. In this kernel [this kernel](https://www.kaggle.com/tanlikesmath/cassava-classification-eda-fastai-starter) , the performance improved when adding it at the same time with smooth labels. Smooth labels have given improvement for us, what about MixUp/CutMix?\n\n\nAlso, [this (other) kernel](https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations) also showed an interesting feature when clustering the data in different groups. It seems that different clusters partly meant different distance to the photographer. It could mean that the noise in our dataset is somehow partly due to the distance at which each photo were taken. I would be happy to hear your thoughts on this.\n\n\n\n\n\n",
      "votes": 11
    },
    {
      "id": 1115464,
      "postDate": "2020-12-16T10:07:04.823Z",
      "content": "<p>Mixup didn't improve my LB score, even harmful</p>",
      "rawMarkdown": "Mixup didn't improve my LB score, even harmful",
      "votes": 1
    },
    {
      "id": 1115414,
      "postDate": "2020-12-16T09:20:49.143Z",
      "content": "<p>I think it's risky to use them, even some said they got imporved both in CV and LB. <br>\nthink of the ROI, you cant guarantee the label consistency. </p>",
      "rawMarkdown": "I think it's risky to use them, even some said they got imporved both in CV and LB. \nthink of the ROI, you cant guarantee the label consistency. ",
      "votes": 1
    },
    {
      "id": 1106625,
      "postDate": "2020-12-09T01:15:06.283Z",
      "content": "<p>Also, the same kernel  <a href=\"https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations\" target=\"_blank\">https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations</a> displays a cluster where the images represent roots of the cassava plant. So what is happening in our models?</p>\n<p>The model is actually trying to learn the same label from roots and leaves image : images that show different features. I am not sure this is a good thing for our model to perform best. Maybe our CNNs are simply creating both filters for leaves and roots. What are your thoughts on that ? </p>",
      "rawMarkdown": "Also, the same kernel  https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations displays a cluster where the images represent roots of the cassava plant. So what is happening in our models?\n\nThe model is actually trying to learn the same label from roots and leaves image : images that show different features. I am not sure this is a good thing for our model to perform best. Maybe our CNNs are simply creating both filters for leaves and roots. What are your thoughts on that ? \n",
      "votes": 1
    },
    {
      "id": 1106623,
      "postDate": "2020-12-09T01:10:37.140Z",
      "content": "<p>Solving the noise problem is not easy because we have to take into account the noise in the test set. For me it's probable that long distance images are harder to label as they contain the disease information in only few pixels. They could partly refer to what we call \"noise\".</p>\n<p>I actually think between 2019 and 2020 they tried to change data creation method. They possibly wanted to add more rigor by ensuring the distance was a more or less constant parameter. It's possible they populated the public LB with  more \"noisy\" images as our scores went down, and kept cleaner data for private LB.</p>",
      "rawMarkdown": "Solving the noise problem is not easy because we have to take into account the noise in the test set. For me it's probable that long distance images are harder to label as they contain the disease information in only few pixels. They could partly refer to what we call \"noise\".\n\nI actually think between 2019 and 2020 they tried to change data creation method. They possibly wanted to add more rigor by ensuring the distance was a more or less constant parameter. It's possible they populated the public LB with  more \"noisy\" images as our scores went down, and kept cleaner data for private LB.",
      "votes": 1,
      "replies": [
        {
          "id": 1115268,
          "postDate": "2020-12-16T06:37:26.850Z",
          "content": "<p>I don't believe that the private data set has cleaner data. In my opinion the number of images in class will be some what equal or inverse of training data. The distribution of of classes is somewhat similar in public test data and training data. I think the noise will be persisting in all three: training , private and public data. The class distribution my lead to shakeup.</p>",
          "rawMarkdown": "I don't believe that the private data set has cleaner data. In my opinion the number of images in class will be some what equal or inverse of training data. The distribution of of classes is somewhat similar in public test data and training data. I think the noise will be persisting in all three: training , private and public data. The class distribution my lead to shakeup.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1108015,
      "postDate": "2020-12-10T07:27:11.997Z",
      "content": "<p>CutMix didn't work for me.</p>",
      "rawMarkdown": "CutMix didn't work for me.",
      "votes": 2
    },
    {
      "id": 1185690,
      "postDate": "2021-02-04T09:39:39.877Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1115464,
      "author_name": "Matrix",
      "author_url": "",
      "post_date": "2020-12-16T10:07:04.823000",
      "content": "<p>Mixup didn't improve my LB score, even harmful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1115414,
      "author_name": "Tian",
      "author_url": "",
      "post_date": "2020-12-16T09:20:49.143000",
      "content": "<p>I think it's risky to use them, even some said they got imporved both in CV and LB. <br>\nthink of the ROI, you cant guarantee the label consistency. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1106625,
      "author_name": "JGar",
      "author_url": "",
      "post_date": "2020-12-09T01:15:06.283000",
      "content": "<p>Also, the same kernel  <a href=\"https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations\" target=\"_blank\">https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations</a> displays a cluster where the images represent roots of the cassava plant. So what is happening in our models?</p>\n<p>The model is actually trying to learn the same label from roots and leaves image : images that show different features. I am not sure this is a good thing for our model to perform best. Maybe our CNNs are simply creating both filters for leaves and roots. What are your thoughts on that ? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1106623,
      "author_name": "JGar",
      "author_url": "",
      "post_date": "2020-12-09T01:10:37.140000",
      "content": "<p>Solving the noise problem is not easy because we have to take into account the noise in the test set. For me it's probable that long distance images are harder to label as they contain the disease information in only few pixels. They could partly refer to what we call \"noise\".</p>\n<p>I actually think between 2019 and 2020 they tried to change data creation method. They possibly wanted to add more rigor by ensuring the distance was a more or less constant parameter. It's possible they populated the public LB with  more \"noisy\" images as our scores went down, and kept cleaner data for private LB.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1115268,
          "author_name": "Sagar",
          "author_url": "",
          "post_date": "2020-12-16T06:37:26.850000",
          "content": "<p>I don't believe that the private data set has cleaner data. In my opinion the number of images in class will be some what equal or inverse of training data. The distribution of of classes is somewhat similar in public test data and training data. I think the noise will be persisting in all three: training , private and public data. The class distribution my lead to shakeup.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1108015,
      "author_name": "Roman",
      "author_url": "",
      "post_date": "2020-12-10T07:27:11.997000",
      "content": "<p>CutMix didn't work for me.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1185690,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-04T09:39:39.877000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1106566": "I have seen kernels using both image augmentation via **CutMix** and **MixUp**. Nevertheless I would like to open a discussion on whether to choose and why.\n\n- **CutMix** is adding a cropped image in the corner of another image, and change the training label giving both different weights in function of the respective sizes . \n- **MixUp** is supperposing the images also with corresponding weighted labels. \n\nThey have both proved to give better results on **ImageNet**.\n\nHowever, with the **Cassava Dataset**, I'm afraid it would add up statistically too much noise from the other image by showing different scaled images (some are showing close up leaves, some are more like a global view). So I doubt it would be a good choice to use them. I tried both and am not convinced. In this kernel [this kernel](https://www.kaggle.com/tanlikesmath/cassava-classification-eda-fastai-starter) , the performance improved when adding it at the same time with smooth labels. Smooth labels have given improvement for us, what about MixUp/CutMix?\n\n\nAlso, [this (other) kernel](https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations) also showed an interesting feature when clustering the data in different groups. It seems that different clusters partly meant different distance to the photographer. It could mean that the noise in our dataset is somehow partly due to the distance at which each photo were taken. I would be happy to hear your thoughts on this.\n\n\n\n\n\n",
    "1115464": "Mixup didn't improve my LB score, even harmful",
    "1115414": "I think it's risky to use them, even some said they got imporved both in CV and LB. \nthink of the ROI, you cant guarantee the label consistency. ",
    "1106625": "Also, the same kernel  https://www.kaggle.com/tanulsingh077/rethinking-cv-strategy-with-visualizations displays a cluster where the images represent roots of the cassava plant. So what is happening in our models?\n\nThe model is actually trying to learn the same label from roots and leaves image : images that show different features. I am not sure this is a good thing for our model to perform best. Maybe our CNNs are simply creating both filters for leaves and roots. What are your thoughts on that ? \n",
    "1106623": "Solving the noise problem is not easy because we have to take into account the noise in the test set. For me it's probable that long distance images are harder to label as they contain the disease information in only few pixels. They could partly refer to what we call \"noise\".\n\nI actually think between 2019 and 2020 they tried to change data creation method. They possibly wanted to add more rigor by ensuring the distance was a more or less constant parameter. It's possible they populated the public LB with  more \"noisy\" images as our scores went down, and kept cleaner data for private LB.",
    "1108015": "CutMix didn't work for me.",
    "1185690": ""
  }
}