{
  "id": 208445,
  "title": "PENCIL:  Probabilistic End-to-end Noise Correction for Learning with Noisy Labels",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/208445",
  "author_name": "",
  "post_date": "2021-01-03T14:08:08.583612400Z",
  "votes": 16,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I found this Paper while exploring ways to deal with the Noisy Data.<br>\nPaper - <a href=\"https://arxiv.org/abs/1903.07788\" target=\"_blank\">https://arxiv.org/abs/1903.07788</a><br>\nThere is even a Pytorch Implementation<br>\nImplementation - <a href=\"https://github.com/yikun2019/PENCIL\" target=\"_blank\">https://github.com/yikun2019/PENCIL</a> </p>\n<p>Abstract:</p>\n<blockquote>\n  <p>Deep learning has achieved excellent performance in various computer vision tasks, but requires a lot of training examples with clean labels. It is easy to collect a dataset with noisy labels, but such noise makes networks overfit seriously and accuracies drop dramatically. To address this problem, we propose an end-to-end framework called PENCIL, which can update both network parameters and label estimations as label distributions. PENCIL is independent of the backbone network structure and does not need an auxiliary clean dataset or prior information about noise, thus it is more general and robust than existing methods and is easy to apply. PENCIL outperforms previous state-of-the-art methods by large margins on both synthetic and real-world datasets with different noise types and noise rates. Experiments show that PENCIL is robust on clean datasets, too.</p>\n</blockquote>\n<p>Hope it helps</p>",
  "messages": [
    {
      "id": "1136883",
      "postDate": "01/03/2021 14:08:08",
      "content": "<p>I found this Paper while exploring ways to deal with the Noisy Data.<br>\nPaper - <a href=\"https://arxiv.org/abs/1903.07788\" target=\"_blank\">https://arxiv.org/abs/1903.07788</a><br>\nThere is even a Pytorch Implementation<br>\nImplementation - <a href=\"https://github.com/yikun2019/PENCIL\" target=\"_blank\">https://github.com/yikun2019/PENCIL</a> </p>\n<p>Abstract:</p>\n<blockquote>\n  <p>Deep learning has achieved excellent performance in various computer vision tasks, but requires a lot of training examples with clean labels. It is easy to collect a dataset with noisy labels, but such noise makes networks overfit seriously and accuracies drop dramatically. To address this problem, we propose an end-to-end framework called PENCIL, which can update both network parameters and label estimations as label distributions. PENCIL is independent of the backbone network structure and does not need an auxiliary clean dataset or prior information about noise, thus it is more general and robust than existing methods and is easy to apply. PENCIL outperforms previous state-of-the-art methods by large margins on both synthetic and real-world datasets with different noise types and noise rates. Experiments show that PENCIL is robust on clean datasets, too.</p>\n</blockquote>\n<p>Hope it helps</p>",
      "rawMarkdown": "I found this Paper while exploring ways to deal with the Noisy Data.\nPaper - https://arxiv.org/abs/1903.07788\nThere is even a Pytorch Implementation\nImplementation - https://github.com/yikun2019/PENCIL \n\nAbstract:\n> Deep learning has achieved excellent performance in various computer vision tasks, but requires a lot of training examples with clean labels. It is easy to collect a dataset with noisy labels, but such noise makes networks overfit seriously and accuracies drop dramatically. To address this problem, we propose an end-to-end framework called PENCIL, which can update both network parameters and label estimations as label distributions. PENCIL is independent of the backbone network structure and does not need an auxiliary clean dataset or prior information about noise, thus it is more general and robust than existing methods and is easy to apply. PENCIL outperforms previous state-of-the-art methods by large margins on both synthetic and real-world datasets with different noise types and noise rates. Experiments show that PENCIL is robust on clean datasets, too.\n\nHope it helps",
      "votes": null
    },
    {
      "id": "1138112",
      "postDate": "01/04/2021 12:02:26",
      "content": "<p>I am wondering if dealing with noisy data is useful since I have tried several simple denoise methods, but both by cv or lb didn't improve. I think we should put efforts on find out the distribution of noisy data instead of training a network with 'clean data'. <br>\n(That's my personal view since i haven't reach very high cv/lb yet😷</p>",
      "rawMarkdown": "I am wondering if dealing with noisy data is useful since I have tried several simple denoise methods, but both by cv or lb didn't improve. I think we should put efforts on find out the distribution of noisy data instead of training a network with 'clean data'. \n(That's my personal view since i haven't reach very high cv/lb yet😷",
      "votes": null
    },
    {
      "id": "1144990",
      "postDate": "01/08/2021 19:49:59",
      "content": "<p>Which methods did you try so far for noise reduction?</p>",
      "rawMarkdown": "Which methods did you try so far for noise reduction?",
      "votes": null
    },
    {
      "id": "1145151",
      "postDate": "01/09/2021 00:03:39",
      "content": "<p><a href=\"https://www.kaggle.com/kewang777\" target=\"_blank\">@kewang777</a> For me all of them worked so im not sure about that haha!</p>",
      "rawMarkdown": "kewang777 For me all of them worked so im not sure about that haha!",
      "votes": null
    },
    {
      "id": "1145282",
      "postDate": "01/09/2021 03:34:03",
      "content": "<p>I've tried co-teaching, co-teaching+ and boostrapping loss but so far none of them work for me. However some agmentation like cutmix did help.</p>",
      "rawMarkdown": "I've tried co-teaching, co-teaching+ and boostrapping loss but so far none of them work for me. However some agmentation like cutmix did help.",
      "votes": null
    },
    {
      "id": "1145285",
      "postDate": "01/09/2021 03:39:07",
      "content": "<p>Do you mind to share the methods you've tried? Maybe my denoise methods weren't good enough. </p>",
      "rawMarkdown": "Do you mind to share the methods you've tried? Maybe my denoise methods weren't good enough.",
      "votes": null
    },
    {
      "id": "1145304",
      "postDate": "01/09/2021 04:01:35",
      "content": "<p>Ive tried label smoothing, cutmix, bitemperedloss, symetric loss, and knowledge distillation (change out of fold labels)</p>",
      "rawMarkdown": "Ive tried label smoothing, cutmix, bitemperedloss, symetric loss, and knowledge distillation (change out of fold labels)",
      "votes": null
    },
    {
      "id": "1145311",
      "postDate": "01/09/2021 04:11:05",
      "content": "<p>Thank you soooooo much. I'll give it a try.</p>",
      "rawMarkdown": "Thank you soooooo much. I'll give it a try.",
      "votes": null
    },
    {
      "id": "1145447",
      "postDate": "01/09/2021 06:34:15",
      "content": "<p><a href=\"https://www.kaggle.com/yannmajewski\" target=\"_blank\">@yannmajewski</a> How exactly would knowledge distillation help? Won't the teacher have memorised the wrong label due to its larger memorisation?</p>",
      "rawMarkdown": "yannmajewski How exactly would knowledge distillation help? Won't the teacher have memorised the wrong label due to its larger memorisation?",
      "votes": null
    },
    {
      "id": "1146001",
      "postDate": "01/09/2021 13:33:35",
      "content": "<p>You have to do a out of fold prediction for each image and change the label if the softmax prediction is higher than a treshhold and if you predicted a different label than the ground truth. Also i dont use a bigger model, i use the same model and retrain it with the corrected labels. </p>\n<p>Im not sure knowledge distillation is the right name for this technique, its kinda like pseudo labeling the train set by only changing the labels that seem wrong</p>",
      "rawMarkdown": "You have to do a out of fold prediction for each image and change the label if the softmax prediction is higher than a treshhold and if you predicted a different label than the ground truth. Also i dont use a bigger model, i use the same model and retrain it with the corrected labels. \n\nIm not sure knowledge distillation is the right name for this technique, its kinda like pseudo labeling the train set by only changing the labels that seem wrong",
      "votes": null
    },
    {
      "id": "1152401",
      "postDate": "01/14/2021 05:45:28",
      "content": "<p>I've tried this method recently. My local cv boosts from 0.893~0.908 to 0.904~0.925, however my lb doesn't change that much…Does it mean self-distillation tend to overfit my cv data? May be there is some leakage since the pseud-labels are predicted from model trained on oof data？</p>",
      "rawMarkdown": "I've tried this method recently. My local cv boosts from 0.893~0.908 to 0.904~0.925, however my lb doesn't change that much...Does it mean self-distillation tend to overfit my cv data? May be there is some leakage since the pseud-labels are predicted from model trained on oof data？",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1138112,
      "author_name": "kewang777",
      "author_url": "",
      "post_date": "01/04/2021 12:02:26",
      "content": "<p>I am wondering if dealing with noisy data is useful since I have tried several simple denoise methods, but both by cv or lb didn't improve. I think we should put efforts on find out the distribution of noisy data instead of training a network with 'clean data'. <br>\n(That's my personal view since i haven't reach very high cv/lb yet😷</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144990,
          "author_name": "mh3891",
          "author_url": "",
          "post_date": "01/08/2021 19:49:59",
          "content": "<p>Which methods did you try so far for noise reduction?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145151,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "01/09/2021 00:03:39",
          "content": "<p><a href=\"https://www.kaggle.com/kewang777\" target=\"_blank\">@kewang777</a> For me all of them worked so im not sure about that haha!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145282,
          "author_name": "kewang777",
          "author_url": "",
          "post_date": "01/09/2021 03:34:03",
          "content": "<p>I've tried co-teaching, co-teaching+ and boostrapping loss but so far none of them work for me. However some agmentation like cutmix did help.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145285,
          "author_name": "kewang777",
          "author_url": "",
          "post_date": "01/09/2021 03:39:07",
          "content": "<p>Do you mind to share the methods you've tried? Maybe my denoise methods weren't good enough. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145304,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "01/09/2021 04:01:35",
          "content": "<p>Ive tried label smoothing, cutmix, bitemperedloss, symetric loss, and knowledge distillation (change out of fold labels)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145311,
          "author_name": "kewang777",
          "author_url": "",
          "post_date": "01/09/2021 04:11:05",
          "content": "<p>Thank you soooooo much. I'll give it a try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145447,
          "author_name": "bluesky314",
          "author_url": "",
          "post_date": "01/09/2021 06:34:15",
          "content": "<p><a href=\"https://www.kaggle.com/yannmajewski\" target=\"_blank\">@yannmajewski</a> How exactly would knowledge distillation help? Won't the teacher have memorised the wrong label due to its larger memorisation?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146001,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "01/09/2021 13:33:35",
          "content": "<p>You have to do a out of fold prediction for each image and change the label if the softmax prediction is higher than a treshhold and if you predicted a different label than the ground truth. Also i dont use a bigger model, i use the same model and retrain it with the corrected labels. </p>\n<p>Im not sure knowledge distillation is the right name for this technique, its kinda like pseudo labeling the train set by only changing the labels that seem wrong</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1152401,
          "author_name": "kewang777",
          "author_url": "",
          "post_date": "01/14/2021 05:45:28",
          "content": "<p>I've tried this method recently. My local cv boosts from 0.893~0.908 to 0.904~0.925, however my lb doesn't change that much…Does it mean self-distillation tend to overfit my cv data? May be there is some leakage since the pseud-labels are predicted from model trained on oof data？</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1136883": "I found this Paper while exploring ways to deal with the Noisy Data.\nPaper - https://arxiv.org/abs/1903.07788\nThere is even a Pytorch Implementation\nImplementation - https://github.com/yikun2019/PENCIL \n\nAbstract:\n> Deep learning has achieved excellent performance in various computer vision tasks, but requires a lot of training examples with clean labels. It is easy to collect a dataset with noisy labels, but such noise makes networks overfit seriously and accuracies drop dramatically. To address this problem, we propose an end-to-end framework called PENCIL, which can update both network parameters and label estimations as label distributions. PENCIL is independent of the backbone network structure and does not need an auxiliary clean dataset or prior information about noise, thus it is more general and robust than existing methods and is easy to apply. PENCIL outperforms previous state-of-the-art methods by large margins on both synthetic and real-world datasets with different noise types and noise rates. Experiments show that PENCIL is robust on clean datasets, too.\n\nHope it helps",
    "1138112": "I am wondering if dealing with noisy data is useful since I have tried several simple denoise methods, but both by cv or lb didn't improve. I think we should put efforts on find out the distribution of noisy data instead of training a network with 'clean data'. \n(That's my personal view since i haven't reach very high cv/lb yet😷",
    "1144990": "Which methods did you try so far for noise reduction?",
    "1145151": "kewang777 For me all of them worked so im not sure about that haha!",
    "1145282": "I've tried co-teaching, co-teaching+ and boostrapping loss but so far none of them work for me. However some agmentation like cutmix did help.",
    "1145285": "Do you mind to share the methods you've tried? Maybe my denoise methods weren't good enough.",
    "1145304": "Ive tried label smoothing, cutmix, bitemperedloss, symetric loss, and knowledge distillation (change out of fold labels)",
    "1145311": "Thank you soooooo much. I'll give it a try.",
    "1145447": "yannmajewski How exactly would knowledge distillation help? Won't the teacher have memorised the wrong label due to its larger memorisation?",
    "1146001": "You have to do a out of fold prediction for each image and change the label if the softmax prediction is higher than a treshhold and if you predicted a different label than the ground truth. Also i dont use a bigger model, i use the same model and retrain it with the corrected labels. \n\nIm not sure knowledge distillation is the right name for this technique, its kinda like pseudo labeling the train set by only changing the labels that seem wrong",
    "1152401": "I've tried this method recently. My local cv boosts from 0.893~0.908 to 0.904~0.925, however my lb doesn't change that much...Does it mean self-distillation tend to overfit my cv data? May be there is some leakage since the pseud-labels are predicted from model trained on oof data？"
  },
  "source": "meta"
}