{
  "id": 210198,
  "title": "Relabelling and 2020-2019 Compound dataset",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/210198",
  "author_name": "Yi Wu",
  "post_date": "2021-01-10T03:08:21.947000",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I am trying relabelling and training model with <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200201\" target=\"_blank\">2020-2019 compound dataset</a>, but it turned out to be no help here, can anyone else share what his/her experiment results went when using relabelling or both dataset?</p>",
  "messages": [
    {
      "id": 1146764,
      "postDate": "2021-01-10T03:08:21.947Z",
      "content": "<p>I am trying relabelling and training model with <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200201\" target=\"_blank\">2020-2019 compound dataset</a>, but it turned out to be no help here, can anyone else share what his/her experiment results went when using relabelling or both dataset?</p>",
      "rawMarkdown": "I am trying relabelling and training model with [2020-2019 compound dataset](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200201), but it turned out to be no help here, can anyone else share what his/her experiment results went when using relabelling or both dataset?",
      "votes": 2
    },
    {
      "id": 1146785,
      "postDate": "2021-01-10T03:42:21.523Z",
      "content": "<p>I think there might be multiple possible reasons that why relabeling didn't work in your experiments.</p>\n<ol>\n<li>Your relabeling label might not be very correct, since the condition of clean label is quite hard to tell even by the experts.</li>\n<li>The public testset is also noisy. So let the model learn the noise distribution of data might help you gain better result than clean data(well, relatively clean).</li>\n</ol>",
      "rawMarkdown": "I think there might be multiple possible reasons that why relabeling didn't work in your experiments.\n1. Your relabeling label might not be very correct, since the condition of clean label is quite hard to tell even by the experts.\n2. The public testset is also noisy. So let the model learn the noise distribution of data might help you gain better result than clean data(well, relatively clean).",
      "replies": [
        {
          "id": 1146894,
          "postDate": "2021-01-10T06:29:29.633Z",
          "content": "<p>same as the old dataset, people's opinions vary a lot. some say even if noise present in LB and PB, it may still be important to relabel some obviously incorrect labels. hard to tell which is better and how to achieve the better</p>",
          "rawMarkdown": "same as the old dataset, people's opinions vary a lot. some say even if noise present in LB and PB, it may still be important to relabel some obviously incorrect labels. hard to tell which is better and how to achieve the better"
        },
        {
          "id": 1146900,
          "postDate": "2021-01-10T06:36:16.813Z",
          "content": "<p>I would like to recommend this paper which shared by other kaggler in another discussion.<br>\nMaybe it can give you some insights about this.<br>\n<a href=\"https://arxiv.org/pdf/2012.04193.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.04193.pdf</a></p>",
          "rawMarkdown": "I would like to recommend this paper which shared by other kaggler in another discussion.\nMaybe it can give you some insights about this.\nhttps://arxiv.org/pdf/2012.04193.pdf"
        },
        {
          "id": 1146923,
          "postDate": "2021-01-10T07:06:36.403Z",
          "content": "<p>will study it, thanks!</p>",
          "rawMarkdown": "will study it, thanks!"
        }
      ]
    },
    {
      "id": 1146779,
      "postDate": "2021-01-10T03:28:13.837Z",
      "content": "<p>A placeholder :).</p>",
      "rawMarkdown": "A placeholder :)."
    },
    {
      "id": 1146857,
      "postDate": "2021-01-10T05:27:44.477Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1146891,
          "postDate": "2021-01-10T06:26:36.727Z",
          "content": "<p>thanks, but some ideas are vague, say 2019 dataset</p>\n<blockquote>\n  <p>I know that this worked for some of you, so it is not completely out of my todo list</p>\n</blockquote>",
          "rawMarkdown": "thanks, but some ideas are vague, say 2019 dataset\n> I know that this worked for some of you, so it is not completely out of my todo list"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1146785,
      "author_name": "Tsai29",
      "author_url": "",
      "post_date": "2021-01-10T03:42:21.523000",
      "content": "<p>I think there might be multiple possible reasons that why relabeling didn't work in your experiments.</p>\n<ol>\n<li>Your relabeling label might not be very correct, since the condition of clean label is quite hard to tell even by the experts.</li>\n<li>The public testset is also noisy. So let the model learn the noise distribution of data might help you gain better result than clean data(well, relatively clean).</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 1146894,
          "author_name": "Yi Wu",
          "author_url": "",
          "post_date": "2021-01-10T06:29:29.633000",
          "content": "<p>same as the old dataset, people's opinions vary a lot. some say even if noise present in LB and PB, it may still be important to relabel some obviously incorrect labels. hard to tell which is better and how to achieve the better</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1146900,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2021-01-10T06:36:16.813000",
          "content": "<p>I would like to recommend this paper which shared by other kaggler in another discussion.<br>\nMaybe it can give you some insights about this.<br>\n<a href=\"https://arxiv.org/pdf/2012.04193.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.04193.pdf</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1146923,
          "author_name": "Yi Wu",
          "author_url": "",
          "post_date": "2021-01-10T07:06:36.403000",
          "content": "<p>will study it, thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1146779,
      "author_name": "哈尔的移动城堡",
      "author_url": "",
      "post_date": "2021-01-10T03:28:13.837000",
      "content": "<p>A placeholder :).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1146857,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-10T05:27:44.477000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1146891,
          "author_name": "Yi Wu",
          "author_url": "",
          "post_date": "2021-01-10T06:26:36.727000",
          "content": "<p>thanks, but some ideas are vague, say 2019 dataset</p>\n<blockquote>\n  <p>I know that this worked for some of you, so it is not completely out of my todo list</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1146764": "I am trying relabelling and training model with [2020-2019 compound dataset](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200201), but it turned out to be no help here, can anyone else share what his/her experiment results went when using relabelling or both dataset?",
    "1146785": "I think there might be multiple possible reasons that why relabeling didn't work in your experiments.\n1. Your relabeling label might not be very correct, since the condition of clean label is quite hard to tell even by the experts.\n2. The public testset is also noisy. So let the model learn the noise distribution of data might help you gain better result than clean data(well, relatively clean).",
    "1146779": "A placeholder :).",
    "1146857": ""
  }
}