{
  "id": 210557,
  "title": "Cleanlab: The standard package for machine learning with noisy labels and finding mislabeled data in Python.",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/210557",
  "author_name": "ktm",
  "post_date": "2021-01-11T09:33:36.265000",
  "votes": 36,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I found this package while searching ways to deal with noisy data.  <br>\ngithub: <a href=\"https://github.com/cgnorthcutt/cleanlab\" target=\"_blank\">https://github.com/cgnorthcutt/cleanlab</a>  <br>\ndetail: <a href=\"https://l7.curtisnorthcutt.com/confident-learning\" target=\"_blank\">https://l7.curtisnorthcutt.com/confident-learning</a>  <br>\npaper: <a href=\"https://arxiv.org/abs/1911.00068\" target=\"_blank\">https://arxiv.org/abs/1911.00068</a>  </p>\n<p><code>cleanlab</code> is a machine learning python package for <strong>learning with noisy labels</strong> and <strong>finding label errors in datasets</strong>. <code>cleanlab</code> CLEANs LABels. It is powered by the theory of <strong>confident learning</strong>, published in this <a href=\"https://arxiv.org/abs/1911.00068\" target=\"_blank\">paper</a> | <a href=\"https://l7.curtisnorthcutt.com/confident-learning\" target=\"_blank\">blog</a>.  </p>\n<p>I have made a dataset and demo notebook.  <br>\ndataset: <a href=\"https://www.kaggle.com/tmhrkt/cleanlab\" target=\"_blank\">https://www.kaggle.com/tmhrkt/cleanlab</a>  <br>\nnotebook: <a href=\"https://www.kaggle.com/tmhrkt/cassava-cleanlab-with-efficientnet-b0\" target=\"_blank\">https://www.kaggle.com/tmhrkt/cassava-cleanlab-with-efficientnet-b0</a>  <br>\nI hope it helps you :)</p>",
  "messages": [
    {
      "id": 1148645,
      "postDate": "2021-01-11T09:33:36.267Z",
      "content": "<p>I found this package while searching ways to deal with noisy data.  <br>\ngithub: <a href=\"https://github.com/cgnorthcutt/cleanlab\" target=\"_blank\">https://github.com/cgnorthcutt/cleanlab</a>  <br>\ndetail: <a href=\"https://l7.curtisnorthcutt.com/confident-learning\" target=\"_blank\">https://l7.curtisnorthcutt.com/confident-learning</a>  <br>\npaper: <a href=\"https://arxiv.org/abs/1911.00068\" target=\"_blank\">https://arxiv.org/abs/1911.00068</a>  </p>\n<p><code>cleanlab</code> is a machine learning python package for <strong>learning with noisy labels</strong> and <strong>finding label errors in datasets</strong>. <code>cleanlab</code> CLEANs LABels. It is powered by the theory of <strong>confident learning</strong>, published in this <a href=\"https://arxiv.org/abs/1911.00068\" target=\"_blank\">paper</a> | <a href=\"https://l7.curtisnorthcutt.com/confident-learning\" target=\"_blank\">blog</a>.  </p>\n<p>I have made a dataset and demo notebook.  <br>\ndataset: <a href=\"https://www.kaggle.com/tmhrkt/cleanlab\" target=\"_blank\">https://www.kaggle.com/tmhrkt/cleanlab</a>  <br>\nnotebook: <a href=\"https://www.kaggle.com/tmhrkt/cassava-cleanlab-with-efficientnet-b0\" target=\"_blank\">https://www.kaggle.com/tmhrkt/cassava-cleanlab-with-efficientnet-b0</a>  <br>\nI hope it helps you :)</p>",
      "rawMarkdown": "I found this package while searching ways to deal with noisy data.  \ngithub: https://github.com/cgnorthcutt/cleanlab  \ndetail: https://l7.curtisnorthcutt.com/confident-learning  \npaper: https://arxiv.org/abs/1911.00068  \n  \n\n``cleanlab`` is a machine learning python package for **learning with noisy labels** and **finding label errors in datasets**. ``cleanlab`` CLEANs LABels. It is powered by the theory of **confident learning**, published in this [paper](https://arxiv.org/abs/1911.00068) | [blog](https://l7.curtisnorthcutt.com/confident-learning).  \n\n\nI have made a dataset and demo notebook.  \ndataset: https://www.kaggle.com/tmhrkt/cleanlab  \nnotebook: https://www.kaggle.com/tmhrkt/cassava-cleanlab-with-efficientnet-b0  \nI hope it helps you :)",
      "votes": 36
    },
    {
      "id": 1151094,
      "postDate": "2021-01-13T06:17:47.927Z",
      "content": "<p>The simplest way to clean up the false labels is to predict the training set and verification set again. If the confidence level is lower than a certain level, it will be judged as false labels. But this confidence is not easy to grasp, too much cleaning will lose important features, too little cleaning, the final effect of training image is very little. I remember that I have done similar work for this project before, at least 300-400 images have been cleaned up, and at most about 1000 images have been cleaned up, but all of them failed to achieve good results. Since the image noise of training set, verification set and test set is noisy, their distribution is similar. What should we do if the private test set also has error labels similar to the training set? So cleaning up the label may reduce the lb. of course, if the private test set is clean, the noise is small, and there are few wrong labels, then cleaning up the training set label is useful.</p>",
      "rawMarkdown": "The simplest way to clean up the false labels is to predict the training set and verification set again. If the confidence level is lower than a certain level, it will be judged as false labels. But this confidence is not easy to grasp, too much cleaning will lose important features, too little cleaning, the final effect of training image is very little. I remember that I have done similar work for this project before, at least 300-400 images have been cleaned up, and at most about 1000 images have been cleaned up, but all of them failed to achieve good results. Since the image noise of training set, verification set and test set is noisy, their distribution is similar. What should we do if the private test set also has error labels similar to the training set? So cleaning up the label may reduce the lb. of course, if the private test set is clean, the noise is small, and there are few wrong labels, then cleaning up the training set label is useful.",
      "votes": 5
    },
    {
      "id": 1169294,
      "postDate": "2021-01-25T12:21:07.143Z",
      "content": "<p><a href=\"https://www.kaggle.com/zhangeng\" target=\"_blank\">@zhangeng</a>  i agree with your thoughts, i think we should let model fit the noisy label insteading of cleaning up the noisy label</p>",
      "rawMarkdown": "@zhangeng  i agree with your thoughts, i think we should let model fit the noisy label insteading of cleaning up the noisy label"
    },
    {
      "id": 1154427,
      "postDate": "2021-01-15T16:20:46.050Z",
      "content": "<p>Hey! I'm trying to implement it, but if my understanding is correct i'm training a classifier on top a classifier right? How would i go about saving this cleanlab model and then loading it for inference in the test set?</p>",
      "rawMarkdown": "Hey! I'm trying to implement it, but if my understanding is correct i'm training a classifier on top a classifier right? How would i go about saving this cleanlab model and then loading it for inference in the test set?",
      "replies": [
        {
          "id": 1155111,
          "postDate": "2021-01-16T08:04:02.423Z",
          "content": "<p>You can load and predict from pretrained weights by doing this  </p>\n<pre><code>cassava_model = CassavaImgClassifier(params[\"model_name\"], pretrained=False)\ncassava_model.load_state_dict(torch.load(\"path/to/pretrained/weights.pth\"))\nclf = Classifier(cassava_model, params)\npredictions = clf.predict_proba(sample_submission.loc[:, \"image_id\"].values, phase=\"test\")\n</code></pre>",
          "rawMarkdown": "You can load and predict from pretrained weights by doing this  \n```\ncassava_model = CassavaImgClassifier(params[\"model_name\"], pretrained=False)\ncassava_model.load_state_dict(torch.load(\"path/to/pretrained/weights.pth\"))\nclf = Classifier(cassava_model, params)\npredictions = clf.predict_proba(sample_submission.loc[:, \"image_id\"].values, phase=\"test\")\n```",
          "votes": 1
        },
        {
          "id": 1156181,
          "postDate": "2021-01-17T02:00:23.970Z",
          "content": "<p>Thanks for your reply, but i'm still wondering about how you get the weights from the cleanlab classifier to a file so that i can then load it into another notebook using your code.</p>\n<p>Thanks in advance!</p>",
          "rawMarkdown": "Thanks for your reply, but i'm still wondering about how you get the weights from the cleanlab classifier to a file so that i can then load it into another notebook using your code.\n\nThanks in advance!"
        },
        {
          "id": 1156270,
          "postDate": "2021-01-17T03:59:36.587Z",
          "content": "<p><code>LearningWithNoisyLabels</code> returns <code>LearningWithNoisyLabels.clf</code>, which is <code>Classifier</code> class with trained model in my notebook. <code>Classifier</code> class has <code>model</code>, which is pytorch model.  So you can get the trained weights from <code>Classifier.model.state_dict()</code> or <code>LearningWithNoisyLabels.clf.model.state_dict()</code> and save them. Then, you can upload the trained weights to your own dataset and load them in your inference notebook.</p>",
          "rawMarkdown": "`LearningWithNoisyLabels` returns `LearningWithNoisyLabels.clf`, which is `Classifier` class with trained model in my notebook. `Classifier` class has `model`, which is pytorch model.  So you can get the trained weights from `Classifier.model.state_dict()` or `LearningWithNoisyLabels.clf.model.state_dict()` and save them. Then, you can upload the trained weights to your own dataset and load them in your inference notebook.\n\n"
        }
      ]
    },
    {
      "id": 1152138,
      "postDate": "2021-01-13T21:03:01.787Z",
      "content": "<p>Hey thanks for the good work. <br>\nDoes the dataset contain a csv file with cleaned labels?</p>",
      "rawMarkdown": "Hey thanks for the good work. \nDoes the dataset contain a csv file with cleaned labels?",
      "replies": [
        {
          "id": 1152262,
          "postDate": "2021-01-14T01:46:47.717Z",
          "content": "<p>No. The dataset is just a copy of <a href=\"https://github.com/cgnorthcutt/cleanlab\" target=\"_blank\">github</a> repository. </p>",
          "rawMarkdown": "No. The dataset is just a copy of [github](https://github.com/cgnorthcutt/cleanlab) repository. "
        }
      ]
    },
    {
      "id": 1151437,
      "postDate": "2021-01-13T10:23:22.237Z",
      "content": "<p>Really interesting. Thanks for sharing. What about performance and accuracy? </p>",
      "rawMarkdown": "Really interesting. Thanks for sharing. What about performance and accuracy? ",
      "replies": [
        {
          "id": 1151461,
          "postDate": "2021-01-13T10:44:11.313Z",
          "content": "<p>The performance is better than the model without <code>cleanlab</code>. See my comment in this discussion :)</p>",
          "rawMarkdown": "The performance is better than the model without `cleanlab`. See my comment in this discussion :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1148743,
      "postDate": "2021-01-11T11:06:50.403Z",
      "content": "<p>Thanks for sharing! This seems very interesting. Have you tested this on your submissions yet? How were the results?</p>",
      "rawMarkdown": "Thanks for sharing! This seems very interesting. Have you tested this on your submissions yet? How were the results?",
      "replies": [
        {
          "id": 1148945,
          "postDate": "2021-01-11T13:45:36.743Z",
          "content": "<p>I'm now training. I'll update soon.</p>",
          "rawMarkdown": "I'm now training. I'll update soon.",
          "votes": 1
        },
        {
          "id": 1150085,
          "postDate": "2021-01-12T11:04:10.670Z",
          "content": "<p>Training with cleanlab,<br>\nCV: 0.8808<br>\nLB: 0.885</p>\n<p>Training without cleanlab,<br>\nCV: 0.8786<br>\nLB: 0.882  </p>\n<p>They are the same condition except using cleanlab. </p>",
          "rawMarkdown": "Training with cleanlab,\nCV: 0.8808\nLB: 0.885\n\nTraining without cleanlab,\nCV: 0.8786\nLB: 0.882  \n\nThey are the same condition except using cleanlab. ",
          "votes": 4
        },
        {
          "id": 1150355,
          "postDate": "2021-01-12T14:28:46.737Z",
          "content": "<p>Is that single model single fold or an ensemble? Thanks for sharing!</p>",
          "rawMarkdown": "Is that single model single fold or an ensemble? Thanks for sharing!"
        },
        {
          "id": 1150942,
          "postDate": "2021-01-13T01:37:42.487Z",
          "content": "<p>It is 5 fold averaging with one seed. Parameters are the same as the sharing notebook.</p>",
          "rawMarkdown": "It is 5 fold averaging with one seed. Parameters are the same as the sharing notebook.",
          "votes": 1
        },
        {
          "id": 1155144,
          "postDate": "2021-01-16T08:28:07.440Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1152447,
      "postDate": "2021-01-14T06:53:46.857Z",
      "content": "<p>Thanks for sharing !!! <a href=\"https://www.kaggle.com/tmhrkt\" target=\"_blank\">@tmhrkt</a> </p>",
      "rawMarkdown": "Thanks for sharing !!! @tmhrkt ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1151094,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2021-01-13T06:17:47.927000",
      "content": "<p>The simplest way to clean up the false labels is to predict the training set and verification set again. If the confidence level is lower than a certain level, it will be judged as false labels. But this confidence is not easy to grasp, too much cleaning will lose important features, too little cleaning, the final effect of training image is very little. I remember that I have done similar work for this project before, at least 300-400 images have been cleaned up, and at most about 1000 images have been cleaned up, but all of them failed to achieve good results. Since the image noise of training set, verification set and test set is noisy, their distribution is similar. What should we do if the private test set also has error labels similar to the training set? So cleaning up the label may reduce the lb. of course, if the private test set is clean, the noise is small, and there are few wrong labels, then cleaning up the training set label is useful.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1169294,
      "author_name": "Jackie Mai",
      "author_url": "",
      "post_date": "2021-01-25T12:21:07.143000",
      "content": "<p><a href=\"https://www.kaggle.com/zhangeng\" target=\"_blank\">@zhangeng</a>  i agree with your thoughts, i think we should let model fit the noisy label insteading of cleaning up the noisy label</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1154427,
      "author_name": "Gabriel Prado",
      "author_url": "",
      "post_date": "2021-01-15T16:20:46.050000",
      "content": "<p>Hey! I'm trying to implement it, but if my understanding is correct i'm training a classifier on top a classifier right? How would i go about saving this cleanlab model and then loading it for inference in the test set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1155111,
          "author_name": "ktm",
          "author_url": "",
          "post_date": "2021-01-16T08:04:02.423000",
          "content": "<p>You can load and predict from pretrained weights by doing this  </p>\n<pre><code>cassava_model = CassavaImgClassifier(params[\"model_name\"], pretrained=False)\ncassava_model.load_state_dict(torch.load(\"path/to/pretrained/weights.pth\"))\nclf = Classifier(cassava_model, params)\npredictions = clf.predict_proba(sample_submission.loc[:, \"image_id\"].values, phase=\"test\")\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1156181,
          "author_name": "Gabriel Prado",
          "author_url": "",
          "post_date": "2021-01-17T02:00:23.970000",
          "content": "<p>Thanks for your reply, but i'm still wondering about how you get the weights from the cleanlab classifier to a file so that i can then load it into another notebook using your code.</p>\n<p>Thanks in advance!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1156270,
          "author_name": "ktm",
          "author_url": "",
          "post_date": "2021-01-17T03:59:36.587000",
          "content": "<p><code>LearningWithNoisyLabels</code> returns <code>LearningWithNoisyLabels.clf</code>, which is <code>Classifier</code> class with trained model in my notebook. <code>Classifier</code> class has <code>model</code>, which is pytorch model.  So you can get the trained weights from <code>Classifier.model.state_dict()</code> or <code>LearningWithNoisyLabels.clf.model.state_dict()</code> and save them. Then, you can upload the trained weights to your own dataset and load them in your inference notebook.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1152138,
      "author_name": "SohamTamba",
      "author_url": "",
      "post_date": "2021-01-13T21:03:01.787000",
      "content": "<p>Hey thanks for the good work. <br>\nDoes the dataset contain a csv file with cleaned labels?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1152262,
          "author_name": "ktm",
          "author_url": "",
          "post_date": "2021-01-14T01:46:47.717000",
          "content": "<p>No. The dataset is just a copy of <a href=\"https://github.com/cgnorthcutt/cleanlab\" target=\"_blank\">github</a> repository. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1151437,
      "author_name": "Md. Masud Rana",
      "author_url": "",
      "post_date": "2021-01-13T10:23:22.237000",
      "content": "<p>Really interesting. Thanks for sharing. What about performance and accuracy? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1151461,
          "author_name": "ktm",
          "author_url": "",
          "post_date": "2021-01-13T10:44:11.313000",
          "content": "<p>The performance is better than the model without <code>cleanlab</code>. See my comment in this discussion :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1148743,
      "author_name": "Junyi Ng",
      "author_url": "",
      "post_date": "2021-01-11T11:06:50.403000",
      "content": "<p>Thanks for sharing! This seems very interesting. Have you tested this on your submissions yet? How were the results?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1148945,
          "author_name": "ktm",
          "author_url": "",
          "post_date": "2021-01-11T13:45:36.743000",
          "content": "<p>I'm now training. I'll update soon.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1150085,
          "author_name": "ktm",
          "author_url": "",
          "post_date": "2021-01-12T11:04:10.670000",
          "content": "<p>Training with cleanlab,<br>\nCV: 0.8808<br>\nLB: 0.885</p>\n<p>Training without cleanlab,<br>\nCV: 0.8786<br>\nLB: 0.882  </p>\n<p>They are the same condition except using cleanlab. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1150355,
          "author_name": "Gabriel Prado",
          "author_url": "",
          "post_date": "2021-01-12T14:28:46.737000",
          "content": "<p>Is that single model single fold or an ensemble? Thanks for sharing!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1150942,
          "author_name": "ktm",
          "author_url": "",
          "post_date": "2021-01-13T01:37:42.487000",
          "content": "<p>It is 5 fold averaging with one seed. Parameters are the same as the sharing notebook.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1155144,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-16T08:28:07.440000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1152447,
      "author_name": "Satyajit Pattnaik",
      "author_url": "",
      "post_date": "2021-01-14T06:53:46.857000",
      "content": "<p>Thanks for sharing !!! <a href=\"https://www.kaggle.com/tmhrkt\" target=\"_blank\">@tmhrkt</a> </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1148645": "I found this package while searching ways to deal with noisy data.  \ngithub: https://github.com/cgnorthcutt/cleanlab  \ndetail: https://l7.curtisnorthcutt.com/confident-learning  \npaper: https://arxiv.org/abs/1911.00068  \n  \n\n``cleanlab`` is a machine learning python package for **learning with noisy labels** and **finding label errors in datasets**. ``cleanlab`` CLEANs LABels. It is powered by the theory of **confident learning**, published in this [paper](https://arxiv.org/abs/1911.00068) | [blog](https://l7.curtisnorthcutt.com/confident-learning).  \n\n\nI have made a dataset and demo notebook.  \ndataset: https://www.kaggle.com/tmhrkt/cleanlab  \nnotebook: https://www.kaggle.com/tmhrkt/cassava-cleanlab-with-efficientnet-b0  \nI hope it helps you :)",
    "1151094": "The simplest way to clean up the false labels is to predict the training set and verification set again. If the confidence level is lower than a certain level, it will be judged as false labels. But this confidence is not easy to grasp, too much cleaning will lose important features, too little cleaning, the final effect of training image is very little. I remember that I have done similar work for this project before, at least 300-400 images have been cleaned up, and at most about 1000 images have been cleaned up, but all of them failed to achieve good results. Since the image noise of training set, verification set and test set is noisy, their distribution is similar. What should we do if the private test set also has error labels similar to the training set? So cleaning up the label may reduce the lb. of course, if the private test set is clean, the noise is small, and there are few wrong labels, then cleaning up the training set label is useful.",
    "1169294": "@zhangeng  i agree with your thoughts, i think we should let model fit the noisy label insteading of cleaning up the noisy label",
    "1154427": "Hey! I'm trying to implement it, but if my understanding is correct i'm training a classifier on top a classifier right? How would i go about saving this cleanlab model and then loading it for inference in the test set?",
    "1152138": "Hey thanks for the good work. \nDoes the dataset contain a csv file with cleaned labels?",
    "1151437": "Really interesting. Thanks for sharing. What about performance and accuracy? ",
    "1148743": "Thanks for sharing! This seems very interesting. Have you tested this on your submissions yet? How were the results?",
    "1152447": "Thanks for sharing !!! @tmhrkt "
  }
}