{
  "id": 214291,
  "title": "How to build a reliable local cv?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/214291",
  "author_name": "",
  "post_date": "2021-01-26T03:42:56.177703800Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Now the train set has too much noise, which obviously makes CV unreliable. LB only uses 31% of the data, which may be quite different from PB, so now reconstructing a reliable local CV score will be the key to our survival in the coming huge shake.</p>\n<p>The 2019 data set is much cleaner than 2020. Is it possible to use the 2019 data set as reliable CV ?</p>",
  "messages": [
    {
      "id": "1170188",
      "postDate": "01/26/2021 03:42:56",
      "content": "<p>Now the train set has too much noise, which obviously makes CV unreliable. LB only uses 31% of the data, which may be quite different from PB, so now reconstructing a reliable local CV score will be the key to our survival in the coming huge shake.</p>\n<p>The 2019 data set is much cleaner than 2020. Is it possible to use the 2019 data set as reliable CV ?</p>",
      "rawMarkdown": "Now the train set has too much noise, which obviously makes CV unreliable. LB only uses 31% of the data, which may be quite different from PB, so now reconstructing a reliable local CV score will be the key to our survival in the coming huge shake.\n\nThe 2019 data set is much cleaner than 2020. Is it possible to use the 2019 data set as reliable CV ?",
      "votes": null
    },
    {
      "id": "1170221",
      "postDate": "01/26/2021 04:31:20",
      "content": "<ol>\n<li>IMHO, the presence of noisy labels does not necessarily mean that the CV score is not reliable. It all depends on how well your learning algorithm can deal with noisy labels. </li>\n<li>I am curious why do you think that the 2019 data set is cleaner that the 2020 one? I recall as a few weeks ago, a member of the Kaggle team said that the 2020 data received much more scrutiny than the 2019 data during the data preparation phase. </li>\n</ol>",
      "rawMarkdown": "1. IMHO, the presence of noisy labels does not necessarily mean that the CV score is not reliable. It all depends on how well your learning algorithm can deal with noisy labels. \n2. I am curious why do you think that the 2019 data set is cleaner that the 2020 one? I recall as a few weeks ago, a member of the Kaggle team said that the 2020 data received much more scrutiny than the 2019 data during the data preparation phase.",
      "votes": null
    },
    {
      "id": "1170247",
      "postDate": "01/26/2021 05:01:19",
      "content": "<p>How did you know 2019 data was cleaner than 2020?</p>",
      "rawMarkdown": "How did you know 2019 data was cleaner than 2020?",
      "votes": null
    },
    {
      "id": "1170329",
      "postDate": "01/26/2021 06:29:33",
      "content": "<p><a href=\"https://www.kaggle.com/graf10a\" target=\"_blank\">@graf10a</a> <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> I checked the health class, and in 2020, there is about 20% noise (more than 500 images have been cleaned and there are significant errors), and almost no healthy image in 2019 is wrong.</p>",
      "rawMarkdown": "graf10a @mrinath I checked the health class, and in 2020, there is about 20% noise (more than 500 images have been cleaned and there are significant errors), and almost no healthy image in 2019 is wrong.",
      "votes": null
    },
    {
      "id": "1170340",
      "postDate": "01/26/2021 06:43:41",
      "content": "<p>what's more, LB is noisy too. I tried some clean op to the dataset, resulting in the lower LB </p>",
      "rawMarkdown": "what's more, LB is noisy too. I tried some clean op to the dataset, resulting in the lower LB",
      "votes": null
    },
    {
      "id": "1173243",
      "postDate": "01/27/2021 18:41:43",
      "content": "<p>Adding the data of 2019 to 2020 did result in lower CV and LB, my first guess was that 2019 dataset has an equal amount of noise compared to 2020 dataset. </p>\n<p>Did it increase your CV/LB in any way?</p>",
      "rawMarkdown": "Adding the data of 2019 to 2020 did result in lower CV and LB, my first guess was that 2019 dataset has an equal amount of noise compared to 2020 dataset. \n\nDid it increase your CV/LB in any way?",
      "votes": null
    },
    {
      "id": "1175213",
      "postDate": "01/29/2021 02:30:34",
      "content": "<p>It did imporve, train it with unqiued 2019 dataset</p>",
      "rawMarkdown": "It did imporve, train it with unqiued 2019 dataset",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1170221,
      "author_name": "graf10a",
      "author_url": "",
      "post_date": "01/26/2021 04:31:20",
      "content": "<ol>\n<li>IMHO, the presence of noisy labels does not necessarily mean that the CV score is not reliable. It all depends on how well your learning algorithm can deal with noisy labels. </li>\n<li>I am curious why do you think that the 2019 data set is cleaner that the 2020 one? I recall as a few weeks ago, a member of the Kaggle team said that the 2020 data received much more scrutiny than the 2019 data during the data preparation phase. </li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1170329,
          "author_name": "raininbox",
          "author_url": "",
          "post_date": "01/26/2021 06:29:33",
          "content": "<p><a href=\"https://www.kaggle.com/graf10a\" target=\"_blank\">@graf10a</a> <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> I checked the health class, and in 2020, there is about 20% noise (more than 500 images have been cleaned and there are significant errors), and almost no healthy image in 2019 is wrong.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1170247,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "01/26/2021 05:01:19",
      "content": "<p>How did you know 2019 data was cleaner than 2020?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1170340,
          "author_name": "raininbox",
          "author_url": "",
          "post_date": "01/26/2021 06:43:41",
          "content": "<p>what's more, LB is noisy too. I tried some clean op to the dataset, resulting in the lower LB </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1173243,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/27/2021 18:41:43",
          "content": "<p>Adding the data of 2019 to 2020 did result in lower CV and LB, my first guess was that 2019 dataset has an equal amount of noise compared to 2020 dataset. </p>\n<p>Did it increase your CV/LB in any way?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1175213,
          "author_name": "raininbox",
          "author_url": "",
          "post_date": "01/29/2021 02:30:34",
          "content": "<p>It did imporve, train it with unqiued 2019 dataset</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1170188": "Now the train set has too much noise, which obviously makes CV unreliable. LB only uses 31% of the data, which may be quite different from PB, so now reconstructing a reliable local CV score will be the key to our survival in the coming huge shake.\n\nThe 2019 data set is much cleaner than 2020. Is it possible to use the 2019 data set as reliable CV ?",
    "1170221": "1. IMHO, the presence of noisy labels does not necessarily mean that the CV score is not reliable. It all depends on how well your learning algorithm can deal with noisy labels. \n2. I am curious why do you think that the 2019 data set is cleaner that the 2020 one? I recall as a few weeks ago, a member of the Kaggle team said that the 2020 data received much more scrutiny than the 2019 data during the data preparation phase.",
    "1170247": "How did you know 2019 data was cleaner than 2020?",
    "1170329": "graf10a @mrinath I checked the health class, and in 2020, there is about 20% noise (more than 500 images have been cleaned and there are significant errors), and almost no healthy image in 2019 is wrong.",
    "1170340": "what's more, LB is noisy too. I tried some clean op to the dataset, resulting in the lower LB",
    "1173243": "Adding the data of 2019 to 2020 did result in lower CV and LB, my first guess was that 2019 dataset has an equal amount of noise compared to 2020 dataset. \n\nDid it increase your CV/LB in any way?",
    "1175213": "It did imporve, train it with unqiued 2019 dataset"
  },
  "source": "meta"
}