{
  "id": 220762,
  "title": "My Noisy label countermeasure",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/220762",
  "author_name": "sinchir0",
  "post_date": "2021-02-19T13:13:45.627000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello, Kaggler! </p>\n<p>In this competition, I wanted to establish a countermeasure against the noisy label, even in the test! But my hopes were not fulfilled‥</p>\n<p>Below is a list I have taken to counter the noisy label.<br>\nPlease let me know what you are doing about noisy label.</p>\n<p><strong>- Using cleanlab to detect noisy labels in train data.</strong><br>\na  joint probability distribution is created from the predictions and the original labels, and noisy labels of the train are detected based on this distribution. I found this notebook very helpful.<br>\n<a href=\"https://www.kaggle.com/telljoy/noisy-label-eda-with-cleanlab\" target=\"_blank\">https://www.kaggle.com/telljoy/noisy-label-eda-with-cleanlab</a></p>\n<p><strong>- Building a Noise Classifier</strong><br>\nDefine noisy(1) where LB0.89 model mis-predict. And I've built a model to classify 1 or 0.<br>\nBut, my Noise Classifier precision is 22%‥, too low. So I gave up.</p>\n<p><strong>- Define the probability of a label that is likely to be wrong, and multiply it by prediction.</strong><br>\nWhen people make mistakes in labeling, there is some kind of rule. For example, people might mistake the image of a dog for a wolf, but not a dog for a cat.  So, I calculated the probability distribution of making a mistake. </p>\n<p>For example, 0 is the next most likely to be mistaken → 4 : 0.12%, 1 : 0.06%, 3 : 0.03%, 2 : 0.02%</p>\n<p>I use this Probability distribution(Mistake Probability distribution) for postprocess. <br>\nBut the accuracy got worse(CV:0.8924→0.8335)</p>\n<p><strong>- I use Mistake Probability distribution only noisy label in test.</strong>  <br>\nI defined noisy data in the train from cleanlab , and defined noisy data in test  where the clean and normal model predict different labels. And I use Mistake Probability distribution to it.<br>\nBut the accuracy got worse(CV:0.8924→0.8870)</p>\n<p><strong>- Give the is_noisy flag and give it as a feature of lightgbm(max_depth5) when stacking.</strong><br>\nCV increased slightly from 0.8989 to 0.8991. However, the CV of max_depth1 was the best at 0.8997, so it was not used.</p>\n<p>All my challenge didn't work. How was your challenge?</p>",
  "messages": [
    {
      "id": 1210472,
      "postDate": "2021-02-19T13:13:45.627Z",
      "content": "<p>Hello, Kaggler! </p>\n<p>In this competition, I wanted to establish a countermeasure against the noisy label, even in the test! But my hopes were not fulfilled‥</p>\n<p>Below is a list I have taken to counter the noisy label.<br>\nPlease let me know what you are doing about noisy label.</p>\n<p><strong>- Using cleanlab to detect noisy labels in train data.</strong><br>\na  joint probability distribution is created from the predictions and the original labels, and noisy labels of the train are detected based on this distribution. I found this notebook very helpful.<br>\n<a href=\"https://www.kaggle.com/telljoy/noisy-label-eda-with-cleanlab\" target=\"_blank\">https://www.kaggle.com/telljoy/noisy-label-eda-with-cleanlab</a></p>\n<p><strong>- Building a Noise Classifier</strong><br>\nDefine noisy(1) where LB0.89 model mis-predict. And I've built a model to classify 1 or 0.<br>\nBut, my Noise Classifier precision is 22%‥, too low. So I gave up.</p>\n<p><strong>- Define the probability of a label that is likely to be wrong, and multiply it by prediction.</strong><br>\nWhen people make mistakes in labeling, there is some kind of rule. For example, people might mistake the image of a dog for a wolf, but not a dog for a cat.  So, I calculated the probability distribution of making a mistake. </p>\n<p>For example, 0 is the next most likely to be mistaken → 4 : 0.12%, 1 : 0.06%, 3 : 0.03%, 2 : 0.02%</p>\n<p>I use this Probability distribution(Mistake Probability distribution) for postprocess. <br>\nBut the accuracy got worse(CV:0.8924→0.8335)</p>\n<p><strong>- I use Mistake Probability distribution only noisy label in test.</strong>  <br>\nI defined noisy data in the train from cleanlab , and defined noisy data in test  where the clean and normal model predict different labels. And I use Mistake Probability distribution to it.<br>\nBut the accuracy got worse(CV:0.8924→0.8870)</p>\n<p><strong>- Give the is_noisy flag and give it as a feature of lightgbm(max_depth5) when stacking.</strong><br>\nCV increased slightly from 0.8989 to 0.8991. However, the CV of max_depth1 was the best at 0.8997, so it was not used.</p>\n<p>All my challenge didn't work. How was your challenge?</p>",
      "rawMarkdown": "Hello, Kaggler! \n\nIn this competition, I wanted to establish a countermeasure against the noisy label, even in the test! But my hopes were not fulfilled‥\n\nBelow is a list I have taken to counter the noisy label.\nPlease let me know what you are doing about noisy label.\n\n**- Using cleanlab to detect noisy labels in train data.**\na  joint probability distribution is created from the predictions and the original labels, and noisy labels of the train are detected based on this distribution. I found this notebook very helpful.\nhttps://www.kaggle.com/telljoy/noisy-label-eda-with-cleanlab\n\n**- Building a Noise Classifier**\nDefine noisy(1) where LB0.89 model mis-predict. And I've built a model to classify 1 or 0.\nBut, my Noise Classifier precision is 22%‥, too low. So I gave up.\n\n**- Define the probability of a label that is likely to be wrong, and multiply it by prediction.**\nWhen people make mistakes in labeling, there is some kind of rule. For example, people might mistake the image of a dog for a wolf, but not a dog for a cat.  So, I calculated the probability distribution of making a mistake. \n\nFor example, 0 is the next most likely to be mistaken → 4 : 0.12%, 1 : 0.06%, 3 : 0.03%, 2 : 0.02%\n\nI use this Probability distribution(Mistake Probability distribution) for postprocess. \nBut the accuracy got worse(CV:0.8924→0.8335)\n\n**- I use Mistake Probability distribution only noisy label in test.**  \nI defined noisy data in the train from cleanlab , and defined noisy data in test  where the clean and normal model predict different labels. And I use Mistake Probability distribution to it.\nBut the accuracy got worse(CV:0.8924→0.8870)\n\n**- Give the is_noisy flag and give it as a feature of lightgbm(max_depth5) when stacking.**\nCV increased slightly from 0.8989 to 0.8991. However, the CV of max_depth1 was the best at 0.8997, so it was not used.\n\nAll my challenge didn't work. How was your challenge?",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1210472": "Hello, Kaggler! \n\nIn this competition, I wanted to establish a countermeasure against the noisy label, even in the test! But my hopes were not fulfilled‥\n\nBelow is a list I have taken to counter the noisy label.\nPlease let me know what you are doing about noisy label.\n\n**- Using cleanlab to detect noisy labels in train data.**\na  joint probability distribution is created from the predictions and the original labels, and noisy labels of the train are detected based on this distribution. I found this notebook very helpful.\nhttps://www.kaggle.com/telljoy/noisy-label-eda-with-cleanlab\n\n**- Building a Noise Classifier**\nDefine noisy(1) where LB0.89 model mis-predict. And I've built a model to classify 1 or 0.\nBut, my Noise Classifier precision is 22%‥, too low. So I gave up.\n\n**- Define the probability of a label that is likely to be wrong, and multiply it by prediction.**\nWhen people make mistakes in labeling, there is some kind of rule. For example, people might mistake the image of a dog for a wolf, but not a dog for a cat.  So, I calculated the probability distribution of making a mistake. \n\nFor example, 0 is the next most likely to be mistaken → 4 : 0.12%, 1 : 0.06%, 3 : 0.03%, 2 : 0.02%\n\nI use this Probability distribution(Mistake Probability distribution) for postprocess. \nBut the accuracy got worse(CV:0.8924→0.8335)\n\n**- I use Mistake Probability distribution only noisy label in test.**  \nI defined noisy data in the train from cleanlab , and defined noisy data in test  where the clean and normal model predict different labels. And I use Mistake Probability distribution to it.\nBut the accuracy got worse(CV:0.8924→0.8870)\n\n**- Give the is_noisy flag and give it as a feature of lightgbm(max_depth5) when stacking.**\nCV increased slightly from 0.8989 to 0.8991. However, the CV of max_depth1 was the best at 0.8997, so it was not used.\n\nAll my challenge didn't work. How was your challenge?"
  }
}