{
  "id": 211020,
  "title": "Should ‘HEALTHY’ become a certain class？",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/211020",
  "author_name": "",
  "post_date": "2021-01-13T09:48:10.883067700Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I notice that healthy leaf exists in most images with disease.But we define it as a certain class’HEALTHY’.It seems an unwise action.So I tried decreasing images labeled’4’ in my training process from 2000 to 200，my cv boost from 0.89to 0.95,it's amazing！And my LB score nearly has no change😂.Does this mean we don't need this class?</p>",
  "messages": [
    {
      "id": "1151398",
      "postDate": "01/13/2021 09:48:10",
      "content": "<p>I notice that healthy leaf exists in most images with disease.But we define it as a certain class’HEALTHY’.It seems an unwise action.So I tried decreasing images labeled’4’ in my training process from 2000 to 200，my cv boost from 0.89to 0.95,it's amazing！And my LB score nearly has no change😂.Does this mean we don't need this class?</p>",
      "rawMarkdown": "I notice that healthy leaf exists in most images with disease.But we define it as a certain class’HEALTHY’.It seems an unwise action.So I tried decreasing images labeled’4’ in my training process from 2000 to 200，my cv boost from 0.89to 0.95,it's amazing！And my LB score nearly has no change😂.Does this mean we don't need this class?",
      "votes": null
    },
    {
      "id": "1151490",
      "postDate": "01/13/2021 11:16:13",
      "content": "<p>Since you cleaned up too many images with noise, the training set looks cleaner and more accurate, but this result confirms that the test set also has a lot of noise, so a cleaner training set cannot identify the test set more accurately Category, or that sentence, the test set may also contain a lot of noisy images, or even incorrectly labeled images. When encountering such problems, our job is to predict the test set (including noisy or incorrectly labeled images). So even if your model predicts the correct category, it may not be suitable for a messy test set. It may be that the test set has wrong labels. Then you must predict the test set according to the wrong labels of the training set…<br>\n<strong>In a word, adapting to the noise environment of the test set is the key to this game.</strong></p>",
      "rawMarkdown": "Since you cleaned up too many images with noise, the training set looks cleaner and more accurate, but this result confirms that the test set also has a lot of noise, so a cleaner training set cannot identify the test set more accurately Category, or that sentence, the test set may also contain a lot of noisy images, or even incorrectly labeled images. When encountering such problems, our job is to predict the test set (including noisy or incorrectly labeled images). So even if your model predicts the correct category, it may not be suitable for a messy test set. It may be that the test set has wrong labels. Then you must predict the test set according to the wrong labels of the training set...\n**In a word, adapting to the noise environment of the test set is the key to this game.**",
      "votes": null
    },
    {
      "id": "1151516",
      "postDate": "01/13/2021 11:36:58",
      "content": "<blockquote>\n  <p>adapting to the noise environment of the test set is the key to this game.<br>\n  Do you mean filteration will not work in this game?<br>\n  But I even find some bulk and root pictures in the dataset,I have other dataset with lightly filtering,but they also perform badly…😧</p>\n</blockquote>",
      "rawMarkdown": "> adapting to the noise environment of the test set is the key to this game.\nDo you mean filteration will not work in this game?\nBut I even find some bulk and root pictures in the dataset,I have other dataset with lightly filtering,but they also perform badly...😧",
      "votes": null
    },
    {
      "id": "1151680",
      "postDate": "01/13/2021 13:47:33",
      "content": "<p>Data is <a href=\"https://ieeexplore.ieee.org/document/4804817\" target=\"_blank\">unreasonably Effective</a></p>\n<p>By removing such huge data without any shot learning, you're keeping your model from learning good representations and rich features from the data, specially from hard samples. Thus you model is about  learning and validating mostly from easy samples. That's why you observe huge CV increase but no LB improvement</p>",
      "rawMarkdown": "Data is [unreasonably Effective](https://ieeexplore.ieee.org/document/4804817)\n\nBy removing such huge data without any shot learning, you're keeping your model from learning good representations and rich features from the data, specially from hard samples. Thus you model is about  learning and validating mostly from easy samples. That's why you observe huge CV increase but no LB improvement",
      "votes": null
    },
    {
      "id": "1151786",
      "postDate": "01/13/2021 14:52:57",
      "content": "<p>Thanks,I got it.<br>\nIf my model always learn clean data,then it isn't robust enough for testing new data.</p>",
      "rawMarkdown": "Thanks,I got it.\nIf my model always learn clean data,then it isn't robust enough for testing new data.",
      "votes": null
    },
    {
      "id": "1151857",
      "postDate": "01/13/2021 16:03:02",
      "content": "<p>I was going to say the same as <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a>, you need to make sure you keep the same validation folds in order to compare CV, otherwise removing images will make the validation easier if hard examples are removed</p>",
      "rawMarkdown": "I was going to say the same as @serigne, you need to make sure you keep the same validation folds in order to compare CV, otherwise removing images will make the validation easier if hard examples are removed",
      "votes": null
    },
    {
      "id": "1152039",
      "postDate": "01/13/2021 18:12:54",
      "content": "<blockquote>\n  <p>my cv boost from 0.89to 0.95,it's amazing！</p>\n</blockquote>\n<p>Are you altering your validation set too? If so, then you can't compare experiments.</p>\n<p>The best way to determine if changing the train data via upsample, downsample, external images etc, is to keep the same validation set in all experiments. So first use all train data with your validation set and record the val score. Next alter the train data and use the same validation set and record the val score. Finally compare those val scores.</p>",
      "rawMarkdown": ">my cv boost from 0.89to 0.95,it's amazing！\n\nAre you altering your validation set too? If so, then you can't compare experiments.\n\nThe best way to determine if changing the train data via upsample, downsample, external images etc, is to keep the same validation set in all experiments. So first use all train data with your validation set and record the val score. Next alter the train data and use the same validation set and record the val score. Finally compare those val scores.",
      "votes": null
    },
    {
      "id": "1152264",
      "postDate": "01/14/2021 01:53:18",
      "content": "<p>I don't think the test set is clean and noise-free, so I need to adapt to the noise of the test set.</p>",
      "rawMarkdown": "I don't think the test set is clean and noise-free, so I need to adapt to the noise of the test set.",
      "votes": null
    },
    {
      "id": "1152269",
      "postDate": "01/14/2021 02:02:48",
      "content": "<p>Thanks for all suggestions! I think i made a noob mistake: EVERYTIME my val. datasets change randomly.<br>\nBut you can also try reducing class'4' from your training dataset,you would find that your LB just change a little😂<br>\nDoes it mean there are very little healthy data or <strong>our model just can not distinguish them from the data with disease?</strong><br>\nI think the biggest problem exists in class'4'😳</p>",
      "rawMarkdown": "Thanks for all suggestions! I think i made a noob mistake: EVERYTIME my val. datasets change randomly.\nBut you can also try reducing class'4' from your training dataset,you would find that your LB just change a little😂\nDoes it mean there are very little healthy data or **our model just can not distinguish them from the data with disease?**\nI think the biggest problem exists in class'4'😳",
      "votes": null
    },
    {
      "id": "1152271",
      "postDate": "01/14/2021 02:04:58",
      "content": "<p>I'm also not sure if there is little healthy data in the private test dataset😭.</p>",
      "rawMarkdown": "I'm also not sure if there is little healthy data in the private test dataset😭.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1151490,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "01/13/2021 11:16:13",
      "content": "<p>Since you cleaned up too many images with noise, the training set looks cleaner and more accurate, but this result confirms that the test set also has a lot of noise, so a cleaner training set cannot identify the test set more accurately Category, or that sentence, the test set may also contain a lot of noisy images, or even incorrectly labeled images. When encountering such problems, our job is to predict the test set (including noisy or incorrectly labeled images). So even if your model predicts the correct category, it may not be suitable for a messy test set. It may be that the test set has wrong labels. Then you must predict the test set according to the wrong labels of the training set…<br>\n<strong>In a word, adapting to the noise environment of the test set is the key to this game.</strong></p>",
      "votes": null,
      "replies": [
        {
          "id": 1151516,
          "author_name": "xzy777",
          "author_url": "",
          "post_date": "01/13/2021 11:36:58",
          "content": "<blockquote>\n  <p>adapting to the noise environment of the test set is the key to this game.<br>\n  Do you mean filteration will not work in this game?<br>\n  But I even find some bulk and root pictures in the dataset,I have other dataset with lightly filtering,but they also perform badly…😧</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1152264,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "01/14/2021 01:53:18",
          "content": "<p>I don't think the test set is clean and noise-free, so I need to adapt to the noise of the test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1151680,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "01/13/2021 13:47:33",
      "content": "<p>Data is <a href=\"https://ieeexplore.ieee.org/document/4804817\" target=\"_blank\">unreasonably Effective</a></p>\n<p>By removing such huge data without any shot learning, you're keeping your model from learning good representations and rich features from the data, specially from hard samples. Thus you model is about  learning and validating mostly from easy samples. That's why you observe huge CV increase but no LB improvement</p>",
      "votes": null,
      "replies": [
        {
          "id": 1151786,
          "author_name": "xzy777",
          "author_url": "",
          "post_date": "01/13/2021 14:52:57",
          "content": "<p>Thanks,I got it.<br>\nIf my model always learn clean data,then it isn't robust enough for testing new data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1151857,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "01/13/2021 16:03:02",
          "content": "<p>I was going to say the same as <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a>, you need to make sure you keep the same validation folds in order to compare CV, otherwise removing images will make the validation easier if hard examples are removed</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1152039,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/13/2021 18:12:54",
      "content": "<blockquote>\n  <p>my cv boost from 0.89to 0.95,it's amazing！</p>\n</blockquote>\n<p>Are you altering your validation set too? If so, then you can't compare experiments.</p>\n<p>The best way to determine if changing the train data via upsample, downsample, external images etc, is to keep the same validation set in all experiments. So first use all train data with your validation set and record the val score. Next alter the train data and use the same validation set and record the val score. Finally compare those val scores.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1152269,
      "author_name": "xzy777",
      "author_url": "",
      "post_date": "01/14/2021 02:02:48",
      "content": "<p>Thanks for all suggestions! I think i made a noob mistake: EVERYTIME my val. datasets change randomly.<br>\nBut you can also try reducing class'4' from your training dataset,you would find that your LB just change a little😂<br>\nDoes it mean there are very little healthy data or <strong>our model just can not distinguish them from the data with disease?</strong><br>\nI think the biggest problem exists in class'4'😳</p>",
      "votes": null,
      "replies": [
        {
          "id": 1152271,
          "author_name": "xzy777",
          "author_url": "",
          "post_date": "01/14/2021 02:04:58",
          "content": "<p>I'm also not sure if there is little healthy data in the private test dataset😭.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1151398": "I notice that healthy leaf exists in most images with disease.But we define it as a certain class’HEALTHY’.It seems an unwise action.So I tried decreasing images labeled’4’ in my training process from 2000 to 200，my cv boost from 0.89to 0.95,it's amazing！And my LB score nearly has no change😂.Does this mean we don't need this class?",
    "1151490": "Since you cleaned up too many images with noise, the training set looks cleaner and more accurate, but this result confirms that the test set also has a lot of noise, so a cleaner training set cannot identify the test set more accurately Category, or that sentence, the test set may also contain a lot of noisy images, or even incorrectly labeled images. When encountering such problems, our job is to predict the test set (including noisy or incorrectly labeled images). So even if your model predicts the correct category, it may not be suitable for a messy test set. It may be that the test set has wrong labels. Then you must predict the test set according to the wrong labels of the training set...\n**In a word, adapting to the noise environment of the test set is the key to this game.**",
    "1151516": "> adapting to the noise environment of the test set is the key to this game.\nDo you mean filteration will not work in this game?\nBut I even find some bulk and root pictures in the dataset,I have other dataset with lightly filtering,but they also perform badly...😧",
    "1151680": "Data is [unreasonably Effective](https://ieeexplore.ieee.org/document/4804817)\n\nBy removing such huge data without any shot learning, you're keeping your model from learning good representations and rich features from the data, specially from hard samples. Thus you model is about  learning and validating mostly from easy samples. That's why you observe huge CV increase but no LB improvement",
    "1151786": "Thanks,I got it.\nIf my model always learn clean data,then it isn't robust enough for testing new data.",
    "1151857": "I was going to say the same as @serigne, you need to make sure you keep the same validation folds in order to compare CV, otherwise removing images will make the validation easier if hard examples are removed",
    "1152039": ">my cv boost from 0.89to 0.95,it's amazing！\n\nAre you altering your validation set too? If so, then you can't compare experiments.\n\nThe best way to determine if changing the train data via upsample, downsample, external images etc, is to keep the same validation set in all experiments. So first use all train data with your validation set and record the val score. Next alter the train data and use the same validation set and record the val score. Finally compare those val scores.",
    "1152264": "I don't think the test set is clean and noise-free, so I need to adapt to the noise of the test set.",
    "1152269": "Thanks for all suggestions! I think i made a noob mistake: EVERYTIME my val. datasets change randomly.\nBut you can also try reducing class'4' from your training dataset,you would find that your LB just change a little😂\nDoes it mean there are very little healthy data or **our model just can not distinguish them from the data with disease?**\nI think the biggest problem exists in class'4'😳",
    "1152271": "I'm also not sure if there is little healthy data in the private test dataset😭."
  },
  "source": "meta"
}