{
  "id": 212683,
  "title": "Paper: Deep Learning is Robust to Massive Label Noise",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/212683",
  "author_name": "Mark Wijkhuizen",
  "post_date": "2021-01-19T19:37:06.701000",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I came along <a href=\"https://arxiv.org/pdf/1705.10694.pdf\" target=\"_blank\">this</a> paper which discussed the robustness of deep learning on label noise. The most interesting finding in my view was the effect of batch size on model accuracy, as shown in the figure below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Fd350f6df0f3744020c47bfe0e6565873%2Fpaper%20image.png?generation=1611084480888468&amp;alt=media\" alt=\"\"></p>\n<p>This graph shows the accuracy on the MNIST dataset when manually and uniformly adding noise to the labels. The H_a batch size is the full dataset, which practically keeps its original accuracy when adding noise.<br>\nThe results are interesting and, as we have TPU's to our disposal, using an insanely high batch size is possible. As soon as I have some results I will share them with you.</p>\n<p>Feel free to share your results as well ;)</p>\n<p><em>Reference: Rolnick, D., Veit, A., Belongie, S., &amp; Shavit, N. Deep learning is robust to massive label noise (2017). arXiv preprint arXiv:1705.10694.</em></p>",
  "messages": [
    {
      "id": 1160296,
      "postDate": "2021-01-19T19:37:06.703Z",
      "content": "<p>I came along <a href=\"https://arxiv.org/pdf/1705.10694.pdf\" target=\"_blank\">this</a> paper which discussed the robustness of deep learning on label noise. The most interesting finding in my view was the effect of batch size on model accuracy, as shown in the figure below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Fd350f6df0f3744020c47bfe0e6565873%2Fpaper%20image.png?generation=1611084480888468&amp;alt=media\" alt=\"\"></p>\n<p>This graph shows the accuracy on the MNIST dataset when manually and uniformly adding noise to the labels. The H_a batch size is the full dataset, which practically keeps its original accuracy when adding noise.<br>\nThe results are interesting and, as we have TPU's to our disposal, using an insanely high batch size is possible. As soon as I have some results I will share them with you.</p>\n<p>Feel free to share your results as well ;)</p>\n<p><em>Reference: Rolnick, D., Veit, A., Belongie, S., &amp; Shavit, N. Deep learning is robust to massive label noise (2017). arXiv preprint arXiv:1705.10694.</em></p>",
      "rawMarkdown": "I came along [this](https://arxiv.org/pdf/1705.10694.pdf) paper which discussed the robustness of deep learning on label noise. The most interesting finding in my view was the effect of batch size on model accuracy, as shown in the figure below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Fd350f6df0f3744020c47bfe0e6565873%2Fpaper%20image.png?generation=1611084480888468&alt=media)\n\nThis graph shows the accuracy on the MNIST dataset when manually and uniformly adding noise to the labels. The H_a batch size is the full dataset, which practically keeps its original accuracy when adding noise.\nThe results are interesting and, as we have TPU's to our disposal, using an insanely high batch size is possible. As soon as I have some results I will share them with you.\n\nFeel free to share your results as well ;)\n\n*Reference: Rolnick, D., Veit, A., Belongie, S., & Shavit, N. Deep learning is robust to massive label noise (2017). arXiv preprint arXiv:1705.10694.*",
      "votes": 2
    },
    {
      "id": 1162440,
      "postDate": "2021-01-21T06:36:53.827Z",
      "content": "<p>Moreover, it's been known that large batch sizes converge to very sharp minima, easily leading to overfitting. So there's no hard rule as to whether small or large batch sizes generalize better, it's just another hyperparameter you have to tune.</p>",
      "rawMarkdown": "Moreover, it's been known that large batch sizes converge to very sharp minima, easily leading to overfitting. So there's no hard rule as to whether small or large batch sizes generalize better, it's just another hyperparameter you have to tune."
    },
    {
      "id": 1160445,
      "postDate": "2021-01-19T23:01:41.647Z",
      "content": "<blockquote>\n  <p>For example, on MNIST we obtain test accuracy above 90 percent even after each clean training example  has  been  diluted with 100 randomly-labeled examples. Such behavior holds across multiple patterns of label noise, even when erroneous labels are biased towards confusing classes. We show that training in this regime requires a significant but manageable increase in dataset size that is related to the factor by which correct labels have been diluted.</p>\n</blockquote>\n<p>Really striking that the models can still perform well when clean labels are vastly outnumbered by noise. Thanks for sharing the article. I guess the bigger issue in the case of this competition is the noisy test set labels.</p>",
      "rawMarkdown": "> For example, on MNIST we obtain test accuracy above 90 percent even after each clean training example  has  been  diluted with 100 randomly-labeled examples. Such behavior holds across multiple patterns of label noise, even when erroneous labels are biased towards confusing classes. We show that training in this regime requires a significant but manageable increase in dataset size that is related to the factor by which correct labels have been diluted.\n\nReally striking that the models can still perform well when clean labels are vastly outnumbered by noise. Thanks for sharing the article. I guess the bigger issue in the case of this competition is the noisy test set labels.",
      "replies": [
        {
          "id": 1160451,
          "postDate": "2021-01-19T23:12:27.710Z",
          "content": "<p>MNIST is  the worst example for  noisy label benchmark IMHO. I would even say for any kind of benchmark.  </p>\n<p>Clothing1M dataset is way much more appropriate to mimic real world examples. </p>",
          "rawMarkdown": "MNIST is  the worst example for  noisy label benchmark IMHO. I would even say for any kind of benchmark.  \n\nClothing1M dataset is way much more appropriate to mimic real world examples. ",
          "votes": 3
        },
        {
          "id": 1161225,
          "postDate": "2021-01-20T12:26:06.893Z",
          "content": "<p>Could you elaborate your view, why would Clothing1M be a more suitable dataset than MNIST?<br>\nYour rank in this competition is quite impressive, looking forward to your explanation.</p>",
          "rawMarkdown": "Could you elaborate your view, why would Clothing1M be a more suitable dataset than MNIST?\nYour rank in this competition is quite impressive, looking forward to your explanation."
        },
        {
          "id": 1161305,
          "postDate": "2021-01-20T13:32:21.537Z",
          "content": "<p>Because there is no hard example in MNIST.  Even the simplest CNN can achieve +99% Accuracy. </p>\n<p>I guess they increase the batch size to let the model see and learn from the majority of easy examples per minibatch and,  in worst case, just memorize the labels they manually corrupted.  You can still achieve high accuracy in MNIST with that. </p>\n<p>As for Clothing1M, it contains real noise and not  manually flipped noisy labels</p>",
          "rawMarkdown": "Because there is no hard example in MNIST.  Even the simplest CNN can achieve +99% Accuracy. \n\nI guess they increase the batch size to let the model see and learn from the majority of easy examples per minibatch and,  in worst case, just memorize the labels they manually corrupted.  You can still achieve high accuracy in MNIST with that. \n\n\nAs for Clothing1M, it contains real noise and not  manually flipped noisy labels\n"
        },
        {
          "id": 1161617,
          "postDate": "2021-01-20T16:42:27.467Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1161619,
          "postDate": "2021-01-20T16:43:14.230Z",
          "content": "<p>Serigne - Does that mean you think what they found (using huge batch sizes) would not be useful for this specific competition?</p>",
          "rawMarkdown": "Serigne - Does that mean you think what they found (using huge batch sizes) would not be useful for this specific competition?"
        },
        {
          "id": 1161779,
          "postDate": "2021-01-20T18:38:37.297Z",
          "content": "<p>I think there are a couple of significant differences between the situation in the paper and the current competition:</p>\n<ul>\n<li>The paper looks at cases where \"noisy\" labels outnumber clean labels, which does not seem to be the case here.</li>\n<li>I think the paper assumes the test set labels are clean, which is probably not the case here.</li>\n</ul>\n<p>What that means for the potential effect of batch size in the competition, I can't say, but I suppose it would be easy enough to test empirically. Taking the paper at face value, I'd assume that, in this case, larger batch size -&gt; less learning of noisy labels in the training data. But again, this wouldn't address the noisy test set labels issue.</p>",
          "rawMarkdown": "I think there are a couple of significant differences between the situation in the paper and the current competition:\n- The paper looks at cases where \"noisy\" labels outnumber clean labels, which does not seem to be the case here.\n- I think the paper assumes the test set labels are clean, which is probably not the case here.\n\nWhat that means for the potential effect of batch size in the competition, I can't say, but I suppose it would be easy enough to test empirically. Taking the paper at face value, I'd assume that, in this case, larger batch size -> less learning of noisy labels in the training data. But again, this wouldn't address the noisy test set labels issue.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1162440,
      "author_name": "Junyi Ng",
      "author_url": "",
      "post_date": "2021-01-21T06:36:53.827000",
      "content": "<p>Moreover, it's been known that large batch sizes converge to very sharp minima, easily leading to overfitting. So there's no hard rule as to whether small or large batch sizes generalize better, it's just another hyperparameter you have to tune.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1160445,
      "author_name": "DanJL",
      "author_url": "",
      "post_date": "2021-01-19T23:01:41.647000",
      "content": "<blockquote>\n  <p>For example, on MNIST we obtain test accuracy above 90 percent even after each clean training example  has  been  diluted with 100 randomly-labeled examples. Such behavior holds across multiple patterns of label noise, even when erroneous labels are biased towards confusing classes. We show that training in this regime requires a significant but manageable increase in dataset size that is related to the factor by which correct labels have been diluted.</p>\n</blockquote>\n<p>Really striking that the models can still perform well when clean labels are vastly outnumbered by noise. Thanks for sharing the article. I guess the bigger issue in the case of this competition is the noisy test set labels.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1160451,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-01-19T23:12:27.710000",
          "content": "<p>MNIST is  the worst example for  noisy label benchmark IMHO. I would even say for any kind of benchmark.  </p>\n<p>Clothing1M dataset is way much more appropriate to mimic real world examples. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1161225,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2021-01-20T12:26:06.893000",
          "content": "<p>Could you elaborate your view, why would Clothing1M be a more suitable dataset than MNIST?<br>\nYour rank in this competition is quite impressive, looking forward to your explanation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1161305,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-01-20T13:32:21.537000",
          "content": "<p>Because there is no hard example in MNIST.  Even the simplest CNN can achieve +99% Accuracy. </p>\n<p>I guess they increase the batch size to let the model see and learn from the majority of easy examples per minibatch and,  in worst case, just memorize the labels they manually corrupted.  You can still achieve high accuracy in MNIST with that. </p>\n<p>As for Clothing1M, it contains real noise and not  manually flipped noisy labels</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1161617,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-20T16:42:27.467000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1161619,
          "author_name": "impulsecorp",
          "author_url": "",
          "post_date": "2021-01-20T16:43:14.230000",
          "content": "<p>Serigne - Does that mean you think what they found (using huge batch sizes) would not be useful for this specific competition?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1161779,
          "author_name": "DanJL",
          "author_url": "",
          "post_date": "2021-01-20T18:38:37.297000",
          "content": "<p>I think there are a couple of significant differences between the situation in the paper and the current competition:</p>\n<ul>\n<li>The paper looks at cases where \"noisy\" labels outnumber clean labels, which does not seem to be the case here.</li>\n<li>I think the paper assumes the test set labels are clean, which is probably not the case here.</li>\n</ul>\n<p>What that means for the potential effect of batch size in the competition, I can't say, but I suppose it would be easy enough to test empirically. Taking the paper at face value, I'd assume that, in this case, larger batch size -&gt; less learning of noisy labels in the training data. But again, this wouldn't address the noisy test set labels issue.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1160296": "I came along [this](https://arxiv.org/pdf/1705.10694.pdf) paper which discussed the robustness of deep learning on label noise. The most interesting finding in my view was the effect of batch size on model accuracy, as shown in the figure below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Fd350f6df0f3744020c47bfe0e6565873%2Fpaper%20image.png?generation=1611084480888468&alt=media)\n\nThis graph shows the accuracy on the MNIST dataset when manually and uniformly adding noise to the labels. The H_a batch size is the full dataset, which practically keeps its original accuracy when adding noise.\nThe results are interesting and, as we have TPU's to our disposal, using an insanely high batch size is possible. As soon as I have some results I will share them with you.\n\nFeel free to share your results as well ;)\n\n*Reference: Rolnick, D., Veit, A., Belongie, S., & Shavit, N. Deep learning is robust to massive label noise (2017). arXiv preprint arXiv:1705.10694.*",
    "1162440": "Moreover, it's been known that large batch sizes converge to very sharp minima, easily leading to overfitting. So there's no hard rule as to whether small or large batch sizes generalize better, it's just another hyperparameter you have to tune.",
    "1160445": "> For example, on MNIST we obtain test accuracy above 90 percent even after each clean training example  has  been  diluted with 100 randomly-labeled examples. Such behavior holds across multiple patterns of label noise, even when erroneous labels are biased towards confusing classes. We show that training in this regime requires a significant but manageable increase in dataset size that is related to the factor by which correct labels have been diluted.\n\nReally striking that the models can still perform well when clean labels are vastly outnumbered by noise. Thanks for sharing the article. I guess the bigger issue in the case of this competition is the noisy test set labels."
  }
}