{
  "id": 202017,
  "title": "if you are not getting lb 0.901 and above, here is why",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/202017",
  "author_name": "hengck23",
  "post_date": "2020-12-07T21:36:10.486000",
  "votes": 344,
  "comment_count": 102,
  "views": 0,
  "content": "<p>It is now apparent that the dataset is noisy (for both train and public/private test).<br>\nYou also note that with the same model (e.g. efficientnet-b4) the performances of kaggers are very different. It is the training process that matters.</p>\n<p>you are now training in noise and the objective is not to train for the lowest loss but one that is consistent between train and test. This competition is like estimating the noise in the dataset and apply correct training parameters to get to the bayes error.</p>\n<p>there are common techniques like early stopping and label smooth, apply SVM etc.</p>\n<p>but there are also more advantage techniques if you google for \"training in the presence of noisy labels\". but do note that unlike cases in most academic papers where the test is clean, our test is also noisy. you may also want to read about testing in the presence of noisy data</p>\n<p>here I introduce one method that works for me:</p>\n<ul>\n<li>training early stopping (i.e. not choosing the lowest validation loss): lb0.899</li>\n<li>training tempered loss (without early stopping): lb0.901</li>\n<li>training normal loss (without early stopping): lb0.898</li>\n</ul>\n<p><a href=\"https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\" target=\"_blank\">https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html</a></p>\n<ul>\n<li>log loss is modified so that outliers (noise labels) are not penalized that much.</li>\n<li>I think our case is close to large-margin noise</li>\n</ul>\n<p><img src=\"https://1.bp.blogspot.com/-MSsW9QUCyXM/XWQNInb2ztI/AAAAAAAAEjc/_yu48eAnTjMWQtUbjNqhOlqurgapjSIzACLcBGAs/s640/Screenshot%2B2019-08-26%2Bat%2B9.39.48%2BAM.png\" alt=\"\"></p>\n<p><img src=\"https://1.bp.blogspot.com/-wTnN2w3ENVc/XWQOEFA67TI/AAAAAAAAEjo/0W-4vcYefTcwn7tURpmQ6zMUBrUmJ5e-wCLcBGAs/s640/Screenshot%2B2019-08-26%2Bat%2B9.37.41%2BAM.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1105431,
      "postDate": "2020-12-07T21:36:10.487Z",
      "content": "<p>It is now apparent that the dataset is noisy (for both train and public/private test).<br>\nYou also note that with the same model (e.g. efficientnet-b4) the performances of kaggers are very different. It is the training process that matters.</p>\n<p>you are now training in noise and the objective is not to train for the lowest loss but one that is consistent between train and test. This competition is like estimating the noise in the dataset and apply correct training parameters to get to the bayes error.</p>\n<p>there are common techniques like early stopping and label smooth, apply SVM etc.</p>\n<p>but there are also more advantage techniques if you google for \"training in the presence of noisy labels\". but do note that unlike cases in most academic papers where the test is clean, our test is also noisy. you may also want to read about testing in the presence of noisy data</p>\n<p>here I introduce one method that works for me:</p>\n<ul>\n<li>training early stopping (i.e. not choosing the lowest validation loss): lb0.899</li>\n<li>training tempered loss (without early stopping): lb0.901</li>\n<li>training normal loss (without early stopping): lb0.898</li>\n</ul>\n<p><a href=\"https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\" target=\"_blank\">https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html</a></p>\n<ul>\n<li>log loss is modified so that outliers (noise labels) are not penalized that much.</li>\n<li>I think our case is close to large-margin noise</li>\n</ul>\n<p><img src=\"https://1.bp.blogspot.com/-MSsW9QUCyXM/XWQNInb2ztI/AAAAAAAAEjc/_yu48eAnTjMWQtUbjNqhOlqurgapjSIzACLcBGAs/s640/Screenshot%2B2019-08-26%2Bat%2B9.39.48%2BAM.png\" alt=\"\"></p>\n<p><img src=\"https://1.bp.blogspot.com/-wTnN2w3ENVc/XWQOEFA67TI/AAAAAAAAEjo/0W-4vcYefTcwn7tURpmQ6zMUBrUmJ5e-wCLcBGAs/s640/Screenshot%2B2019-08-26%2Bat%2B9.37.41%2BAM.png\" alt=\"\"></p>",
      "rawMarkdown": "It is now apparent that the dataset is noisy (for both train and public/private test).\nYou also note that with the same model (e.g. efficientnet-b4) the performances of kaggers are very different. It is the training process that matters.\n\nyou are now training in noise and the objective is not to train for the lowest loss but one that is consistent between train and test. This competition is like estimating the noise in the dataset and apply correct training parameters to get to the bayes error.\n\nthere are common techniques like early stopping and label smooth, apply SVM etc.\n\nbut there are also more advantage techniques if you google for \"training in the presence of noisy labels\". but do note that unlike cases in most academic papers where the test is clean, our test is also noisy. you may also want to read about testing in the presence of noisy data\n\nhere I introduce one method that works for me:\n-  training early stopping (i.e. not choosing the lowest validation loss): lb0.899\n-  training tempered loss (without early stopping): lb0.901\n-  training normal loss (without early stopping): lb0.898\n\n\nhttps://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\n- log loss is modified so that outliers (noise labels) are not penalized that much.\n- I think our case is close to large-margin noise\n\n![](https://1.bp.blogspot.com/-MSsW9QUCyXM/XWQNInb2ztI/AAAAAAAAEjc/_yu48eAnTjMWQtUbjNqhOlqurgapjSIzACLcBGAs/s640/Screenshot%2B2019-08-26%2Bat%2B9.39.48%2BAM.png)\n\n![](https://1.bp.blogspot.com/-wTnN2w3ENVc/XWQOEFA67TI/AAAAAAAAEjo/0W-4vcYefTcwn7tURpmQ6zMUBrUmJ5e-wCLcBGAs/s640/Screenshot%2B2019-08-26%2Bat%2B9.37.41%2BAM.png)",
      "votes": 344
    },
    {
      "id": 1105542,
      "postDate": "2020-12-08T01:42:34.543Z",
      "content": "<p>Pytorch implementation of Robust Bi-Tempered Logistic Loss Based on Bregman Divergences:<br>\n<a href=\"https://github.com/mlpanda/bi-tempered-loss-pytorch\" target=\"_blank\">https://github.com/mlpanda/bi-tempered-loss-pytorch</a></p>",
      "rawMarkdown": "Pytorch implementation of Robust Bi-Tempered Logistic Loss Based on Bregman Divergences:\nhttps://github.com/mlpanda/bi-tempered-loss-pytorch",
      "votes": 25
    },
    {
      "id": 1105940,
      "postDate": "2020-12-08T10:58:18.780Z",
      "content": "<p>What I done in the past when dealing with noisy labels, was to use the OOF prediction of the model trained with all data after softmax , and eliminate images where the softmax value is too small for the correct label. After eliminating a small quantity of training images, retrain from scratch with the remaining one.<br>\nFor example if after softmax you got the predictions (0.1, 0.3, 0.5, 0.05, 0.05) and if the correct label is 4 that image is eliminated from the training set.<br>\nI did not try it yet to this competition but it's on my todo list.</p>",
      "rawMarkdown": "What I done in the past when dealing with noisy labels, was to use the OOF prediction of the model trained with all data after softmax , and eliminate images where the softmax value is too small for the correct label. After eliminating a small quantity of training images, retrain from scratch with the remaining one.\nFor example if after softmax you got the predictions (0.1, 0.3, 0.5, 0.05, 0.05) and if the correct label is 4 that image is eliminated from the training set.\nI did not try it yet to this competition but it's on my todo list.",
      "votes": 26,
      "replies": [
        {
          "id": 1105953,
          "postDate": "2020-12-08T11:12:06.177Z",
          "content": "<p>cleaning label is also one of the methods I will be trying later. If the test data is clean, I would recommend this type of method. </p>\n<p>but do note that the test data is also dirty. so any elimination of train data does changes some distrubtion</p>",
          "rawMarkdown": "cleaning label is also one of the methods I will be trying later. If the test data is clean, I would recommend this type of method. \n\nbut do note that the test data is also dirty. so any elimination of train data does changes some distrubtion",
          "votes": 5
        },
        {
          "id": 1106122,
          "postDate": "2020-12-08T14:52:37.103Z",
          "content": "<p>I tried that and filtered the low softmax predictions, it didnt seem to have an effect on the lb. I think since the test set is probably noisy we must train with the noise and figure out a way to manage it!</p>",
          "rawMarkdown": "I tried that and filtered the low softmax predictions, it didnt seem to have an effect on the lb. I think since the test set is probably noisy we must train with the noise and figure out a way to manage it!"
        },
        {
          "id": 1106229,
          "postDate": "2020-12-08T16:09:27.777Z",
          "content": "<p>Even if the public/private test set is more or less noisy, we should do our best to remove excessive noisy images. It's all about the ratio of noisy images to all images. If it's under a limit it doesn't hurt (sometimes quite the opposite) but if there are a big part of the training images, it can severely affect the model. </p>",
          "rawMarkdown": "Even if the public/private test set is more or less noisy, we should do our best to remove excessive noisy images. It's all about the ratio of noisy images to all images. If it's under a limit it doesn't hurt (sometimes quite the opposite) but if there are a big part of the training images, it can severely affect the model. ",
          "votes": 4
        },
        {
          "id": 1177229,
          "postDate": "2021-01-30T06:31:21.967Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a> , <br>\nSo, what did you find after removing the noise in the training data? <br>\nDid it hurt the LB? <br>\nIs there less noise in the test set, so removing train noise improves the LB?</p>",
          "rawMarkdown": "Hello @vladvdv , \nSo, what did you find after removing the noise in the training data? \nDid it hurt the LB? \nIs there less noise in the test set, so removing train noise improves the LB?"
        }
      ]
    },
    {
      "id": 1111114,
      "postDate": "2020-12-13T12:29:17.453Z",
      "content": "<p>my svm experiment:</p>\n<pre><code>efficient-net-b4 logistic classifier results:\n\ntrain\nloss : 0.22858\nacc  : 0.93024\n\nvalid\nloss : 0.31352\nacc  : 0.90257\n\n----\nuse efficient-net-b4 features and train a one-vs-rest linear svm classifier\nsvc = svm.LinearSVC(C=1, verbose=1, max_iter=100000, loss='squared_hinge', penalty='l2', dual=True )\n\nC=1\ntrain accuracy 0.9468948998072092\nvalid accuracy 0.8974299065420561\n\nC=0.1\ntrain accuracy 0.9369632529064672\nvalid accuracy 0.8985981308411215\n\nC=0.001\ntrain accuracy 0.9316469007419524\nvalid accuracy 0.9023364485981309\n\nC=0.0001\ntrain accuracy 0.9293100426476603\nvalid accuracy 0.9042056074766355\n\nC=0.00001\ntrain accuracy 0.9268563416486534\nvalid accuracy 0.8985981308411215\n</code></pre>",
      "rawMarkdown": "my svm experiment:\n\n```\nefficient-net-b4 logistic classifier results:\n\ntrain\nloss : 0.22858\nacc  : 0.93024\n\nvalid\nloss : 0.31352\nacc  : 0.90257\n\n----\nuse efficient-net-b4 features and train a one-vs-rest linear svm classifier\nsvc = svm.LinearSVC(C=1, verbose=1, max_iter=100000, loss='squared_hinge', penalty='l2', dual=True )\n\nC=1\ntrain accuracy 0.9468948998072092\nvalid accuracy 0.8974299065420561\n\nC=0.1\ntrain accuracy 0.9369632529064672\nvalid accuracy 0.8985981308411215\n\nC=0.001\ntrain accuracy 0.9316469007419524\nvalid accuracy 0.9023364485981309\n\nC=0.0001\ntrain accuracy 0.9293100426476603\nvalid accuracy 0.9042056074766355\n\nC=0.00001\ntrain accuracy 0.9268563416486534\nvalid accuracy 0.8985981308411215\n```",
      "votes": 17,
      "replies": [
        {
          "id": 1111345,
          "postDate": "2020-12-13T16:34:38.397Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1111486,
          "postDate": "2020-12-13T18:32:41.740Z",
          "content": "<p>I wonder what is the benefit of a svm classifier compared to a normal fc layer trained with the CNN! Good experiments as always <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>!</p>",
          "rawMarkdown": "I wonder what is the benefit of a svm classifier compared to a normal fc layer trained with the CNN! Good experiments as always @hengck23!"
        },
        {
          "id": 1111877,
          "postDate": "2020-12-14T05:23:54.617Z",
          "content": "<p>Probably margin loss is the main difference.</p>",
          "rawMarkdown": "Probably margin loss is the main difference.",
          "votes": 1
        },
        {
          "id": 1111959,
          "postDate": "2020-12-14T06:54:59.727Z",
          "content": "<p>i would think it is the regularisation.<br>\nit shifts the solution of parameter = min(classification loss) a little.</p>\n<p>which is also with you should stop early in training deep network because you are not seeking the min point</p>",
          "rawMarkdown": "i would think it is the regularisation.\nit shifts the solution of parameter = min(classification loss) a little.\n\nwhich is also with you should stop early in training deep network because you are not seeking the min point",
          "votes": 9
        }
      ]
    },
    {
      "id": 1107824,
      "postDate": "2020-12-10T01:40:22.510Z",
      "content": "<p>probing the class distribution for public test set:</p>\n<pre><code>    label\n         0      cbb =    0.048\n         1     cbsd =    0.106\n         2      cgm =   ???\n         3      cmd =   ???\n         4  healthy =    0.139\n</code></pre>",
      "rawMarkdown": "probing the class distribution for public test set:\n\n```\n\n\tlabel\n\t\t 0      cbb =    0.048\n\t\t 1     cbsd =    0.106\n\t\t 2      cgm =   ???\n\t\t 3      cmd =   ???\n\t\t 4  healthy =    0.139\n\n\n```",
      "votes": 13,
      "replies": [
        {
          "id": 1107826,
          "postDate": "2020-12-10T01:43:58.537Z",
          "content": "<p>Class 2: 0.103</p>\n<p>Class 3: 0.602</p>",
          "rawMarkdown": "Class 2: 0.103\n\nClass 3: 0.602",
          "votes": 15
        }
      ]
    },
    {
      "id": 1116899,
      "postDate": "2020-12-17T14:49:47.760Z",
      "content": "<p>so,how to choose the params of 't1' and 't2',Thanks!</p>",
      "rawMarkdown": "so,how to choose the params of 't1' and 't2',Thanks!",
      "votes": 11,
      "replies": [
        {
          "id": 1189729,
          "postDate": "2021-02-07T07:57:45.370Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1106775,
      "postDate": "2020-12-09T05:44:48.823Z",
      "content": "<p><a href=\"https://www.youtube.com/watch?v=8mpBHbjG4E4\" target=\"_blank\">https://www.youtube.com/watch?v=8mpBHbjG4E4</a><br>\nDeep Learning with Label Noise - Kevin McGuinness - UPC TelecomBCN Barcelona 2019</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4405f84043b9dc1dd1ca197fd25f7e8d%2FSelection_155.png?generation=1607492627025997&amp;alt=media\" alt=\"\"></p>\n<p>i think someone needs to prepare some small set of clean labels and share in kaggle …</p>",
      "rawMarkdown": "https://www.youtube.com/watch?v=8mpBHbjG4E4\nDeep Learning with Label Noise - Kevin McGuinness - UPC TelecomBCN Barcelona 2019\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4405f84043b9dc1dd1ca197fd25f7e8d%2FSelection_155.png?generation=1607492627025997&alt=media)\n\ni think someone needs to prepare some small set of clean labels and share in kaggle ...",
      "votes": 7,
      "replies": [
        {
          "id": 1106792,
          "postDate": "2020-12-09T06:08:39.967Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2F4518ec950c16c00491f1afd0d4f736ce%2FScreen%20Shot%202020-12-09%20at%2012.08.14%20AM.png?generation=1607494111001485&amp;alt=media\" alt=\"\"> - started :)</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2F4518ec950c16c00491f1afd0d4f736ce%2FScreen%20Shot%202020-12-09%20at%2012.08.14%20AM.png?generation=1607494111001485&alt=media) - started :)"
        },
        {
          "id": 1107284,
          "postDate": "2020-12-09T15:03:36.063Z",
          "content": "<p>will making a set of clean labels be any useful as the test is also mislabeled?</p>",
          "rawMarkdown": "will making a set of clean labels be any useful as the test is also mislabeled?",
          "votes": 1
        },
        {
          "id": 1107816,
          "postDate": "2020-12-10T01:30:39.087Z",
          "content": "<p>for analysis/experiment and estimate noise label in test, etc</p>",
          "rawMarkdown": "for analysis/experiment and estimate noise label in test, etc"
        }
      ]
    },
    {
      "id": 1105835,
      "postDate": "2020-12-08T08:38:30.490Z",
      "content": "<p>I am afraid the removed part of test data was among the cleanest part of the test set. <br>\nBefore it was removed, my loss and training process supposed to train on noisy labels and doing inference on clean test data. And it was more effective. </p>\n<p>Now things become harder and we have likely more proportion of noise in the test data since that part was removed. </p>",
      "rawMarkdown": "I am afraid the removed part of test data was among the cleanest part of the test set. \nBefore it was removed, my loss and training process supposed to train on noisy labels and doing inference on clean test data. And it was more effective. \n\nNow things become harder and we have likely more proportion of noise in the test data since that part was removed. ",
      "votes": 7
    },
    {
      "id": 1106007,
      "postDate": "2020-12-08T12:02:21.957Z",
      "content": "<p>i find a way to estimate or probe the noise label in public test:</p>\n<p>for example, I am interested in the confusion matrix of the ground truth of class-A.</p>\n<ol>\n<li><p>first a made a submission that set prediction = A for all test samples. I can back compute the number of ground truth that is A. for example we find that there are 500 instances of A</p></li>\n<li><p>use your model to make prediction. make a submission. e.g. we have lb score of s0</p></li>\n<li><p>select the very high confidence prediction of A. this step is very important. we need to make sure that those selected are really A (e.g. select those with probability &gt;0.95). The number selected also needs to be large enough. e.g. i can select about 80 high confident instances of predicted A.</p></li>\n<li><p>I make a new submission such that the selected 80 are set to prediction = random but not A. The lb score should drop. in fact, if they are really A, I can compute the theoretical drop in score. i probably need to do a few times with different random seed. e.g. theoretical drop in score = 80/500</p></li>\n<li><p>what if I set selected 80 are set to prediction = B.  or set prediction = C? the drop in lb score will reflect there are how many misclassifications from A to B, A to C, etc …</p></li>\n</ol>\n<hr>\n<p>but here is the catch:</p>\n<ul>\n<li><p>you only have test public lb score and not all the test lb score.</p></li>\n<li><p>you are going to waste many submissions (and maybe not worth it) if you are teaming up, you need to conserve submission</p></li>\n<li><p>but there is a 2019 server (late submission) for you to play with. here both public and private lb scores are exposed.i suspect 2019 data is less noisy though.</p></li>\n<li><p>you can always play the 2020 server after the competition is over as a post-competition analysis.</p></li>\n</ul>\n<hr>\n<p>you can compare the procedure above on your validation set and public set, and see if you would get the same numbers.  if the numbers are very different, the validation and test set are different from the trained model point of view.</p>\n<hr>\n<p>after doing this type of probing for a few competitions, you can ask yourself. does pseudo label really works?<br>\ni.e. how much can we know the labels with high confidence using our learned model?</p>\n<p>if you are interested, you can google like \"how can we test 2 distribution are the same or not?\", \"how can we know the noise level in dataset without ground truth label\", etc …</p>",
      "rawMarkdown": "i find a way to estimate or probe the noise label in public test:\n\nfor example, I am interested in the confusion matrix of the ground truth of class-A.\n\n1. first a made a submission that set prediction = A for all test samples. I can back compute the number of ground truth that is A. for example we find that there are 500 instances of A\n\n2. use your model to make prediction. make a submission. e.g. we have lb score of s0\n\n3. select the very high confidence prediction of A. this step is very important. we need to make sure that those selected are really A (e.g. select those with probability >0.95). The number selected also needs to be large enough. e.g. i can select about 80 high confident instances of predicted A.\n\n4. I make a new submission such that the selected 80 are set to prediction = random but not A. The lb score should drop. in fact, if they are really A, I can compute the theoretical drop in score. i probably need to do a few times with different random seed. e.g. theoretical drop in score = 80/500\n\n5. what if I set selected 80 are set to prediction = B.  or set prediction = C? the drop in lb score will reflect there are how many misclassifications from A to B, A to C, etc ...\n\n---\n\nbut here is the catch:\n\n- you only have test public lb score and not all the test lb score.\n\n- you are going to waste many submissions (and maybe not worth it) if you are teaming up, you need to conserve submission\n\n- but there is a 2019 server (late submission) for you to play with. here both public and private lb scores are exposed.i suspect 2019 data is less noisy though.\n\n- you can always play the 2020 server after the competition is over as a post-competition analysis.\n\n\n\n---\n\nyou can compare the procedure above on your validation set and public set, and see if you would get the same numbers.  if the numbers are very different, the validation and test set are different from the trained model point of view.\n\n---\n\nafter doing this type of probing for a few competitions, you can ask yourself. does pseudo label really works?\ni.e. how much can we know the labels with high confidence using our learned model?\n\nif you are interested, you can google like \"how can we test 2 distribution are the same or not?\", \"how can we know the noise level in dataset without ground truth label\", etc ...",
      "votes": 8,
      "replies": [
        {
          "id": 1106300,
          "postDate": "2020-12-08T17:32:49.627Z",
          "content": "<p>I think testing on 2019 data was a very good and valid point. The model that scored 0.85 LB in 2020 comp scored just 0.61 LB on 2019. On to my debugging hat now! Thanks!</p>",
          "rawMarkdown": "I think testing on 2019 data was a very good and valid point. The model that scored 0.85 LB in 2020 comp scored just 0.61 LB on 2019. On to my debugging hat now! Thanks!"
        },
        {
          "id": 1106532,
          "postDate": "2020-12-08T22:50:34.347Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a></p>\n<p>I think that would not be necessary. There is a way to bypass it.</p>\n<p>I have given it a rough maths in my head so it should have a pretty good chance to work.</p>\n<p>Using technique to approximate how much noise is in the local validation set and then using it to make a stable cv where you know how much your model is able to get noise out of. It should be a reliable cv to evaluate your models with.</p>\n<p>That is not the limit, there are so many obvious implementations you can do with it. If you want me to give you more examples, just ask for it.</p>",
          "rawMarkdown": "@hengck23\n\nI think that would not be necessary. There is a way to bypass it.\n\nI have given it a rough maths in my head so it should have a pretty good chance to work.\n\nUsing technique to approximate how much noise is in the local validation set and then using it to make a stable cv where you know how much your model is able to get noise out of. It should be a reliable cv to evaluate your models with.\n\nThat is not the limit, there are so many obvious implementations you can do with it. If you want me to give you more examples, just ask for it.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1110019,
      "postDate": "2020-12-12T10:43:16.240Z",
      "content": "<p>Training setting: On our benchmark, all methods are trained on the noisy training sets of two noise types (blue and red) under 10 noise levels (from 0% to 80%), and tested on the same clean validation set.</p>\n<p>you can google for more paper that reference to the dataset \"Controlled Noisy Web Labels\"</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F340143642066172cb62fa8be402c3dfd%2FSelection_025.png?generation=1607769793038638&amp;alt=media\" alt=\"\"></p>\n<p>\"Beyond Synthetic Noise: Deep Learning on Controlled Noisy Labels\" - ICML 2020<br>\n<a href=\"http://proceedings.mlr.press/v119/jiang20c/jiang20c.pdf\" target=\"_blank\">http://proceedings.mlr.press/v119/jiang20c/jiang20c.pdf</a><br>\n<a href=\"https://google.github.io/controlled-noisy-web-labels/index.html\" target=\"_blank\">https://google.github.io/controlled-noisy-web-labels/index.html</a></p>",
      "rawMarkdown": "Training setting: On our benchmark, all methods are trained on the noisy training sets of two noise types (blue and red) under 10 noise levels (from 0% to 80%), and tested on the same clean validation set.\n\n\n\nyou can google for more paper that reference to the dataset \"Controlled Noisy Web Labels\"\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F340143642066172cb62fa8be402c3dfd%2FSelection_025.png?generation=1607769793038638&alt=media)\n\n\n\n\"Beyond Synthetic Noise: Deep Learning on Controlled Noisy Labels\" - ICML 2020\nhttp://proceedings.mlr.press/v119/jiang20c/jiang20c.pdf\nhttps://google.github.io/controlled-noisy-web-labels/index.html\n",
      "votes": 5,
      "replies": [
        {
          "id": 1110020,
          "postDate": "2020-12-12T10:46:44.213Z",
          "content": "<p>why you need early stopping</p>\n<p>quote: \"DNNs may not learn patterns first on red label noise. Arpit et al. (2017) found that DNNs learn patterns first, revealing an interesting property that DNNs are able to automatically learn generalizable “patterns” in the early training stage before memorizing all noisy training labels.\"</p>\n<p>see also: <a href=\"https://github.com/shengliu66/ELR\" target=\"_blank\">https://github.com/shengliu66/ELR</a></p>\n<p>why you need large learning rate:</p>\n<p>quote: \"ImageNet architectures generalize on noisy labels when the networks are fine-tuned. Kornblith et al. (2019) found that fine-tuning better architectures trained on ImageNet tend to perform better on downstream tasks of clean training labels. It is important to verify whether this holds on noisy training labels because if so, one can conveniently transfer better architectures to better overcome the noisy labels\"</p>",
          "rawMarkdown": "why you need early stopping\n\nquote: \"DNNs may not learn patterns first on red label noise. Arpit et al. (2017) found that DNNs learn patterns first, revealing an interesting property that DNNs are able to automatically learn generalizable “patterns” in the early training stage before memorizing all noisy training labels.\"\n\nsee also: https://github.com/shengliu66/ELR\n\nwhy you need large learning rate:\n\nquote: \"ImageNet architectures generalize on noisy labels when the networks are fine-tuned. Kornblith et al. (2019) found that fine-tuning better architectures trained on ImageNet tend to perform better on downstream tasks of clean training labels. It is important to verify whether this holds on noisy training labels because if so, one can conveniently transfer better architectures to better overcome the noisy labels\"\n",
          "votes": 8
        },
        {
          "id": 1110046,
          "postDate": "2020-12-12T11:27:43.060Z",
          "content": "<p>I used ELR_plus few days ago but got no improvement.  </p>\n<p>May be I didn't tune it correctly</p>",
          "rawMarkdown": "I used ELR_plus few days ago but got no improvement.  \n\nMay be I didn't tune it correctly"
        },
        {
          "id": 1112940,
          "postDate": "2020-12-15T03:12:06.863Z",
          "content": "<p>That's one of the best papers I have read recently, so many insights, very well structured, clean and easy to understand. Thanks a lot for sharing it <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>! <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> I used default params from the paper in my case it improves compared to CE baseline. Perhaps, your comparison is not made to a baseline model.</p>",
          "rawMarkdown": "That's one of the best papers I have read recently, so many insights, very well structured, clean and easy to understand. Thanks a lot for sharing it @hengck23! @serigne I used default params from the paper in my case it improves compared to CE baseline. Perhaps, your comparison is not made to a baseline model.",
          "votes": 2
        },
        {
          "id": 1113275,
          "postDate": "2020-12-15T10:46:11.047Z",
          "content": "<p>so you think the main noise is blue noise ,not red noise? otherwise, early stop is not working!</p>",
          "rawMarkdown": "so you think the main noise is blue noise ,not red noise? otherwise, early stop is not working!\n"
        },
        {
          "id": 1119268,
          "postDate": "2020-12-19T21:21:00.100Z",
          "content": "<p>I just need to change the loss function to use this right?</p>",
          "rawMarkdown": "I just need to change the loss function to use this right?"
        },
        {
          "id": 1120665,
          "postDate": "2020-12-21T03:17:59.440Z",
          "content": "<p>I also used elr_loss,my local cv is increase,but the lb score is decrease.May be there have some problem?</p>",
          "rawMarkdown": "I also used elr_loss,my local cv is increase,but the lb score is decrease.May be there have some problem?"
        },
        {
          "id": 1121450,
          "postDate": "2020-12-21T16:41:04.600Z",
          "content": "<p>Hi,did you use the tempered loss to get some improve in cv?Thanks!</p>",
          "rawMarkdown": "Hi,did you use the tempered loss to get some improve in cv?Thanks!"
        }
      ]
    },
    {
      "id": 1105698,
      "postDate": "2020-12-08T05:17:42.827Z",
      "content": "<p>Here is another research work regarding this - <a href=\"https://arxiv.org/pdf/2002.06541v1.pdf\" target=\"_blank\">Learning Not to Learn in the Presence of Noisy Labels</a>. They showed <strong>gambler’s loss</strong> is an effective choice to deal with noisy labels. </p>",
      "rawMarkdown": "Here is another research work regarding this - [Learning Not to Learn in the Presence of Noisy Labels](https://arxiv.org/pdf/2002.06541v1.pdf). They showed **gambler’s loss** is an effective choice to deal with noisy labels. ",
      "votes": 5,
      "replies": [
        {
          "id": 1106313,
          "postDate": "2020-12-08T17:44:43.447Z",
          "content": "<p>thanks, I'll read it :D</p>",
          "rawMarkdown": "thanks, I'll read it :D"
        },
        {
          "id": 1119616,
          "postDate": "2020-12-20T08:35:27.117Z",
          "content": "<p>hey i have implemented GAmblers loss <br>\nPLEASE HAVE Look<br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/205424\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/205424</a></p>",
          "rawMarkdown": "hey i have implemented GAmblers loss \nPLEASE HAVE Look\nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/205424",
          "votes": 1
        }
      ]
    },
    {
      "id": 1110165,
      "postDate": "2020-12-12T13:37:06.870Z",
      "content": "<p>just an idea</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F47cb2ad82014dbba920450f861c3f72d%2FSelection_027.png?generation=1607780223050350&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "just an idea\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F47cb2ad82014dbba920450f861c3f72d%2FSelection_027.png?generation=1607780223050350&alt=media)",
      "votes": 6,
      "replies": [
        {
          "id": 1110200,
          "postDate": "2020-12-12T14:08:36.850Z",
          "content": "<p>I am trying and thinking similarly.<br>\nThanks for sharing your ideas.</p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
          "rawMarkdown": "I am trying and thinking similarly.\nThanks for sharing your ideas.\n\n@hengck23 "
        },
        {
          "id": 1110212,
          "postDate": "2020-12-12T14:19:02.027Z",
          "content": "<p>similar idea includes trying to find subclusters within each class. <br>\nsome subclusters are clean, some are correlated noise (i.e. similarly mislabeled in both train and test), some are  uncorrelated noise.</p>\n<p>e.g. we have K clusters in one class, then we can extract K embeddings for an input test sample. classifier is based on the K embeddings</p>\n<p>some reference paper:<br>\n1) clustering <br>\n\"Unsupervised Visual Representation Learning with SwAV\"<br>\n\"Self-labelling via simultaneous clustering and representation learning\"</p>\n<p>2) sub-class classifier <br>\n\"SoftTriple Loss: Deep Metric Learning Without Triplet Sampling\"</p>",
          "rawMarkdown": "similar idea includes trying to find subclusters within each class. \nsome subclusters are clean, some are correlated noise (i.e. similarly mislabeled in both train and test), some are  uncorrelated noise.\n\ne.g. we have K clusters in one class, then we can extract K embeddings for an input test sample. classifier is based on the K embeddings\n\nsome reference paper:\n1) clustering \n\"Unsupervised Visual Representation Learning with SwAV\"\n\"Self-labelling via simultaneous clustering and representation learning\"\n\n2) sub-class classifier \n\"SoftTriple Loss: Deep Metric Learning Without Triplet Sampling\"",
          "votes": 5
        },
        {
          "id": 1111143,
          "postDate": "2020-12-13T13:06:11.217Z",
          "content": "<p>For SoftTriple there is a great repo: <a href=\"https://github.com/KevinMusgrave/pytorch-metric-learning\" target=\"_blank\">https://github.com/KevinMusgrave/pytorch-metric-learning</a></p>",
          "rawMarkdown": "For SoftTriple there is a great repo: https://github.com/KevinMusgrave/pytorch-metric-learning",
          "votes": 1
        }
      ]
    },
    {
      "id": 1106021,
      "postDate": "2020-12-08T12:10:13.510Z",
      "content": "<p>qoute : \"Our method is the first data-recalibrating method which is guaranteed to converge to a well behaved classifier\" … \"We provide theoretical guarantees showing that for a wide variety of (unknown) noise patterns, a classifier trained with this strategy converges to be consistent with the Bayes classifier\"</p>\n<p>wow, guaranteed ???</p>\n<p><a href=\"https://openreview.net/pdf?id=ZPa2SyGcbwh\" target=\"_blank\">https://openreview.net/pdf?id=ZPa2SyGcbwh</a><br>\nICLR 2021 -LEARNING WITH FEATURE DEPENDENT LABEL NOISE:<br>\nA PROGRESSIVE APPROACH</p>",
      "rawMarkdown": "qoute : \"Our method is the first data-recalibrating method which is guaranteed to converge to a well behaved classifier\" ... \"We provide theoretical guarantees showing that for a wide variety of (unknown) noise patterns, a classifier trained with this strategy converges to be consistent with the Bayes classifier\"\n\nwow, guaranteed ???\n\nhttps://openreview.net/pdf?id=ZPa2SyGcbwh\nICLR 2021 -LEARNING WITH FEATURE DEPENDENT LABEL NOISE:\nA PROGRESSIVE APPROACH",
      "votes": 3
    },
    {
      "id": 1111359,
      "postDate": "2020-12-13T16:51:00.420Z",
      "content": "<p>Can anyone share notebook which use bi tempered loss with pytorch?  Thanks in advance.  </p>",
      "rawMarkdown": "Can anyone share notebook which use bi tempered loss with pytorch?  Thanks in advance.  ",
      "votes": 4
    },
    {
      "id": 1105982,
      "postDate": "2020-12-08T11:42:48.450Z",
      "content": "<p>I took the liberty of uploading the semi-official TF implementation</p>\n<p><a href=\"https://www.kaggle.com/konradb/bitemperedloss\" target=\"_blank\">https://www.kaggle.com/konradb/bitemperedloss</a></p>",
      "rawMarkdown": "I took the liberty of uploading the semi-official TF implementation\n\nhttps://www.kaggle.com/konradb/bitemperedloss\n",
      "votes": 4,
      "replies": [
        {
          "id": 1111941,
          "postDate": "2020-12-14T06:39:31.557Z",
          "content": "<p>Hey thanks for sharing, did you manage to make it work in your pipeline ? if Yes can you share it with us ?</p>",
          "rawMarkdown": "Hey thanks for sharing, did you manage to make it work in your pipeline ? if Yes can you share it with us ?"
        },
        {
          "id": 1112906,
          "postDate": "2020-12-15T02:07:53.283Z",
          "content": "<p>Anyone tried implementing this? Not quite sure what to pass the loss function for the \"activations\" and \"labels\" parameters. Any tips on how to use it with EfficientNet (or any model..) in Keras/TF?</p>",
          "rawMarkdown": "Anyone tried implementing this? Not quite sure what to pass the loss function for the \"activations\" and \"labels\" parameters. Any tips on how to use it with EfficientNet (or any model..) in Keras/TF?",
          "votes": 1
        },
        {
          "id": 1146430,
          "postDate": "2021-01-09T19:10:39.330Z",
          "content": "<p><a href=\"https://www.kaggle.com/aziz69\" target=\"_blank\">@aziz69</a>  Here is the BiTemperedLogisticLoss <a href=\"https://www.kaggle.com/durbin164/tpu-bitempered-logistic-loss-keras-tensorflow\" target=\"_blank\">kernel </a>implementation. </p>\n<p>Another Implementation is <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209773\" target=\"_blank\">here.</a> </p>",
          "rawMarkdown": "@aziz69  Here is the BiTemperedLogisticLoss [kernel ](https://www.kaggle.com/durbin164/tpu-bitempered-logistic-loss-keras-tensorflow)implementation. \n\nAnother Implementation is [here.](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209773) "
        }
      ]
    },
    {
      "id": 1109013,
      "postDate": "2020-12-11T08:44:24.927Z",
      "content": "<p>is average the best choice in ensemble in the presence of noise? … there could be outlier scores!</p>",
      "rawMarkdown": "is average the best choice in ensemble in the presence of noise? ... there could be outlier scores!",
      "votes": 1,
      "replies": [
        {
          "id": 1112582,
          "postDate": "2020-12-14T18:20:48.120Z",
          "content": "<p>What about using confusion matrices and weighting models according to how well they predict each class? So if model A is slightly better than model B at predicting class 1 correctly model A would be weighted higher for probabilities regarding class 1, and so on. It might be that one model is better overall, but that the second model still scores better for at least one of the classes.</p>",
          "rawMarkdown": "What about using confusion matrices and weighting models according to how well they predict each class? So if model A is slightly better than model B at predicting class 1 correctly model A would be weighted higher for probabilities regarding class 1, and so on. It might be that one model is better overall, but that the second model still scores better for at least one of the classes.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1106614,
      "postDate": "2020-12-09T00:43:53.513Z",
      "content": "<p>Great share! I did some sort of label smoothing to get 0.901 and will definitely try tempered loss to see how they combine together, probably label smoothing also allows to make logistic loss less sensitive. I wonder what is the upper bound accuracy for the hidden test set. Wouldn't also manually removing noisy labels, e.g. mislabeled healthy images in dataset, help model to create a better decision boundary?</p>",
      "rawMarkdown": "Great share! I did some sort of label smoothing to get 0.901 and will definitely try tempered loss to see how they combine together, probably label smoothing also allows to make logistic loss less sensitive. I wonder what is the upper bound accuracy for the hidden test set. Wouldn't also manually removing noisy labels, e.g. mislabeled healthy images in dataset, help model to create a better decision boundary?",
      "votes": 1,
      "replies": [
        {
          "id": 1106728,
          "postDate": "2020-12-09T04:24:01.327Z",
          "content": "<p>\". I wonder what is the upper bound accuracy for the hidden test set. \"</p>\n<p>the leaderboard provide you a clue</p>",
          "rawMarkdown": "\". I wonder what is the upper bound accuracy for the hidden test set. \"\n\nthe leaderboard provide you a clue",
          "votes": 1
        },
        {
          "id": 1107538,
          "postDate": "2020-12-09T19:00:05.713Z",
          "content": "<p>Mind sharing the label smoothing factor that worked for you? I'm trying 0.1 at the moment and can report back on the effect after the experiment.</p>",
          "rawMarkdown": "Mind sharing the label smoothing factor that worked for you? I'm trying 0.1 at the moment and can report back on the effect after the experiment."
        },
        {
          "id": 1107785,
          "postDate": "2020-12-10T00:42:48.623Z",
          "content": "<p>Using label smoothing of 0.1 improved my public score by 0.01.</p>",
          "rawMarkdown": "Using label smoothing of 0.1 improved my public score by 0.01.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1105442,
      "postDate": "2020-12-07T21:52:30.650Z",
      "content": "<p>Would you kindly explain why a dataset is <strong>noisy</strong>? I searched on Internet but i still don't get it.</p>",
      "rawMarkdown": "Would you kindly explain why a dataset is **noisy**? I searched on Internet but i still don't get it.",
      "votes": 1,
      "replies": [
        {
          "id": 1105445,
          "postDate": "2020-12-07T22:04:11.403Z",
          "content": "<p>the label is wrong.</p>\n<ul>\n<li><p>if you look at the images by class e.g. healthy leaf, you would find that some images are not healthy at all. a few forum posts show some examples of such images</p></li>\n<li><p>I am not sure who or how the images are labelled. i read in some papers that unless the person is well trained, it is difficult for a person to judge the diseases of the plant.</p></li>\n</ul>\n<hr>\n<p>from machine learning point of view, the tsne plot of the embedding will show that the data points are not separable by class</p>",
          "rawMarkdown": "the label is wrong.\n\n- if you look at the images by class e.g. healthy leaf, you would find that some images are not healthy at all. a few forum posts show some examples of such images\n\n- I am not sure who or how the images are labelled. i read in some papers that unless the person is well trained, it is difficult for a person to judge the diseases of the plant.\n\n---\n\nfrom machine learning point of view, the tsne plot of the embedding will show that the data points are not separable by class",
          "votes": 15
        },
        {
          "id": 1105495,
          "postDate": "2020-12-07T23:47:16.670Z",
          "content": "<p>Then, in a machine learning context, a noisy dataset has wrong labeled items?</p>",
          "rawMarkdown": "Then, in a machine learning context, a noisy dataset has wrong labeled items?"
        },
        {
          "id": 1105496,
          "postDate": "2020-12-07T23:47:32.087Z",
          "content": "<p>The figure below shows the TSNE embeddings for the efficientnet-b3 model without the classifier. I couldn't separate the classes, there is always some overlap between the classes.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F337744%2Ff3471414927b45e6632dab3aecc4ed20%2FScreen%20Shot%202020-12-07%20at%206.41.33%20PM.png?generation=1607384517717402&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "The figure below shows the TSNE embeddings for the efficientnet-b3 model without the classifier. I couldn't separate the classes, there is always some overlap between the classes.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F337744%2Ff3471414927b45e6632dab3aecc4ed20%2FScreen%20Shot%202020-12-07%20at%206.41.33%20PM.png?generation=1607384517717402&alt=media)",
          "votes": 19
        },
        {
          "id": 1105531,
          "postDate": "2020-12-08T01:18:15.567Z",
          "content": "<p>embedding that cannot separate can also methods that the model or loss is not good enough.</p>\n<p>you should physically inspect the images and the labels that overlap to confirm the problem.</p>",
          "rawMarkdown": "embedding that cannot separate can also methods that the model or loss is not good enough.\n\nyou should physically inspect the images and the labels that overlap to confirm the problem.",
          "votes": 4
        },
        {
          "id": 1105537,
          "postDate": "2020-12-08T01:32:36.630Z",
          "content": "<p>Of course, embedding does not necessarily separate the classes. The figures above show the tsne of the output of a pretrained model - so it is model-dependent. However it is also not straightforward to train a tsne, it requires tweaking some hyper-parameters - and this is what I meant. Non-optimal parameters produce non-optimal clusters with noise.</p>",
          "rawMarkdown": "Of course, embedding does not necessarily separate the classes. The figures above show the tsne of the output of a pretrained model - so it is model-dependent. However it is also not straightforward to train a tsne, it requires tweaking some hyper-parameters - and this is what I meant. Non-optimal parameters produce non-optimal clusters with noise.",
          "votes": 3
        },
        {
          "id": 1105543,
          "postDate": "2020-12-08T01:46:21.607Z",
          "content": "<p>i think you can say noise means \"not having the true values\" in machine learning. e.g. wrong labels, wrong measurements in input, etc. more correctly, we should say that the cassava competition has noisy labels. noisy dataset can mean noisy x (input) or noisy y (label) or both</p>\n<p>i want to add, there is always some noise in data (e.g. imagenet). we are interested to know how much noise will affect accuracy results. that is the signal-to-noise problem.</p>\n<p>sometimes we inject noise in training (e.g. flip labels) noise is not necessarily bad if it is in the appropriate amount.</p>",
          "rawMarkdown": "i think you can say noise means \"not having the true values\" in machine learning. e.g. wrong labels, wrong measurements in input, etc. more correctly, we should say that the cassava competition has noisy labels. noisy dataset can mean noisy x (input) or noisy y (label) or both\n\ni want to add, there is always some noise in data (e.g. imagenet). we are interested to know how much noise will affect accuracy results. that is the signal-to-noise problem.\n\nsometimes we inject noise in training (e.g. flip labels) noise is not necessarily bad if it is in the appropriate amount.",
          "votes": 6
        },
        {
          "id": 1105548,
          "postDate": "2020-12-08T01:55:52.700Z",
          "content": "<p>I agree with all of your points.</p>",
          "rawMarkdown": "I agree with all of your points."
        },
        {
          "id": 1111940,
          "postDate": "2020-12-14T06:38:27.193Z",
          "content": "<p><a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a>  thanks for sharing your results, could you post a notebook on how to visualize embeddings using tsne ?</p>",
          "rawMarkdown": "@tolgadincer  thanks for sharing your results, could you post a notebook on how to visualize embeddings using tsne ?"
        },
        {
          "id": 1112528,
          "postDate": "2020-12-14T17:20:44.250Z",
          "content": "<p><a href=\"https://www.kaggle.com/aziz69\" target=\"_blank\">@aziz69</a> I made my TSNE notebook publicly available in <a href=\"https://www.kaggle.com/tolgadincer/cldc-tsne\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "@aziz69 I made my TSNE notebook publicly available in [here](https://www.kaggle.com/tolgadincer/cldc-tsne).",
          "votes": 4
        },
        {
          "id": 1182699,
          "postDate": "2021-02-02T14:58:45.257Z",
          "content": "<p><a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a>  Here is my TSNE, i am using effnetb4 noisy student, with ELR loss, single model 0.896. Does BiTemperedLoss give better results compared to ELR loss? How did you tune the parameters t1 and t2 ? </p>\n<p><a href=\"https://ibb.co/72wXjyx\"><img src=\"https://i.ibb.co/pJtxbX6/tsne.png\" alt=\"tsne\"></a></p>",
          "rawMarkdown": "@tolgadincer  Here is my TSNE, i am using effnetb4 noisy student, with ELR loss, single model 0.896. Does BiTemperedLoss give better results compared to ELR loss? How did you tune the parameters t1 and t2 ? \n\n<a href=\"https://ibb.co/72wXjyx\"><img src=\"https://i.ibb.co/pJtxbX6/tsne.png\" alt=\"tsne\" border=\"0\"></a>\n\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 1133029,
      "postDate": "2020-12-30T21:36:21.147Z",
      "content": "<p>If the test dataset is noisy - it suggests that the leader-board closely ranked top results are somewhat random</p>",
      "rawMarkdown": "If the test dataset is noisy - it suggests that the leader-board closely ranked top results are somewhat random",
      "votes": 2
    },
    {
      "id": 1109544,
      "postDate": "2020-12-11T20:02:54.223Z",
      "content": "<p>This paper might be of interest from google. Sharpness-Aware Minimization for Efficiently Improving Generalization <a href=\"https://arxiv.org/pdf/2010.01412.pdf\" target=\"_blank\">paper</a> and the <a href=\"https://github.com/google-research/sam\" target=\"_blank\">code </a></p>",
      "rawMarkdown": "This paper might be of interest from google. Sharpness-Aware Minimization for Efficiently Improving Generalization [paper](https://arxiv.org/pdf/2010.01412.pdf) and the [code ](https://github.com/google-research/sam)",
      "votes": 2,
      "replies": [
        {
          "id": 1110123,
          "postDate": "2020-12-12T12:54:31.713Z",
          "content": "<p><a href=\"https://www.kaggle.com/sso2000\" target=\"_blank\">@sso2000</a> </p>\n<p>thanks for the  Sharpness-Aware Minimization paper. i briefly read through it and feel that it should work. they are adding a reglarizer to the network parameter. </p>\n<p>it did indeed work on my validation set, an increase of +0.002. At the same time, LB drop by +0.002.</p>\n<p>I think the LB is not only noisy but also contains some data not from the 2020 train set.  I recall there is a rescore of some LB for removal of some inclass dataset  (which i think is that 2019 set). 2019 did indeed exist in the 2020 test set … those with 2019 labels are removed and maybe those without 2019 still remains?</p>\n<p>my next experiment is to try SAM on combined 2019+2020 train dataset</p>",
          "rawMarkdown": "@sso2000 \n\nthanks for the  Sharpness-Aware Minimization paper. i briefly read through it and feel that it should work. they are adding a reglarizer to the network parameter. \n\nit did indeed work on my validation set, an increase of +0.002. At the same time, LB drop by +0.002.\n\nI think the LB is not only noisy but also contains some data not from the 2020 train set.  I recall there is a rescore of some LB for removal of some inclass dataset  (which i think is that 2019 set). 2019 did indeed exist in the 2020 test set ... those with 2019 labels are removed and maybe those without 2019 still remains?\n\nmy next experiment is to try SAM on combined 2019+2020 train dataset",
          "votes": 1
        }
      ]
    },
    {
      "id": 1106755,
      "postDate": "2020-12-09T05:09:59.240Z",
      "content": "<blockquote>\n  <p>training early stopping (i.e. not choosing the lowest validation loss): lb0.899<br>\n  Hey <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> can you elaborate this one? you're saying is to inference with last checkpoint and not lowest val_loss right? I am having hard time getting my EFFNETB4 above 0.891, but i have seen 0.30~ val_loss with it also the val_categorical_accuracy almost matches with LB. Currently, even my B7 couldn't pass 0.893. Lol..<br>\n  My Setup&gt;&gt; Cutmixup, Some_tf_augs, ROT_SHEAR_AUGS, Cosine LR(0.00001&gt;0.0004) and Good old Adam. </p>\n</blockquote>",
      "rawMarkdown": "> training early stopping (i.e. not choosing the lowest validation loss): lb0.899\nHey @hengck23 can you elaborate this one? you're saying is to inference with last checkpoint and not lowest val_loss right? I am having hard time getting my EFFNETB4 above 0.891, but i have seen 0.30~ val_loss with it also the val_categorical_accuracy almost matches with LB. Currently, even my B7 couldn't pass 0.893. Lol..\nMy Setup>> Cutmixup, Some_tf_augs, ROT_SHEAR_AUGS, Cosine LR(0.00001>0.0004) and Good old Adam. ",
      "votes": 2
    },
    {
      "id": 1106628,
      "postDate": "2020-12-09T01:18:32.017Z",
      "content": "<p>Looks like small margin can also have worse decision boundary compared to large, I think manually handling mislabeled training samples should have the same affect of correcting the decision boundary. Unless, there is no systematic error which generalizes to test set why would we like to keep mislabeled samples in training? It would be like gambling if label noise is random…</p>\n<p>Also don't we have risk of overfitting to noise we have in our training set and public LB? e.g. imagine tuning t1 and t2 </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2Ff6615268e1a185767cd88dcd9e0c51f5%2FScreen%20Shot%202020-12-08%20at%207.17.03%20PM.png?generation=1607476710189661&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2F67392bfc4f32e054c30ac0575e3503e5%2FScreen%20Shot%202020-12-08%20at%207.17.28%20PM.png?generation=1607476703235679&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Looks like small margin can also have worse decision boundary compared to large, I think manually handling mislabeled training samples should have the same affect of correcting the decision boundary. Unless, there is no systematic error which generalizes to test set why would we like to keep mislabeled samples in training? It would be like gambling if label noise is random...\n\nAlso don't we have risk of overfitting to noise we have in our training set and public LB? e.g. imagine tuning t1 and t2 \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2Ff6615268e1a185767cd88dcd9e0c51f5%2FScreen%20Shot%202020-12-08%20at%207.17.03%20PM.png?generation=1607476710189661&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2F67392bfc4f32e054c30ac0575e3503e5%2FScreen%20Shot%202020-12-08%20at%207.17.28%20PM.png?generation=1607476703235679&alt=media)",
      "votes": 2
    },
    {
      "id": 1105956,
      "postDate": "2020-12-08T11:14:18.380Z",
      "content": "<p>here are the papers that will clear your doubts. such tricks are common among kagglers without theoretical justification in the past</p>\n<ul>\n<li><p>CAN GRADIENT CLIPPING MITIGATE LABEL NOISE? ICLR 2020<br>\n<a href=\"https://openreview.net/pdf?id=rklB76EKPr\" target=\"_blank\">https://openreview.net/pdf?id=rklB76EKPr</a></p></li>\n<li><p>Does label smoothing mitigate label noise?<br>\n<a href=\"https://arxiv.org/abs/2003.02819\" target=\"_blank\">https://arxiv.org/abs/2003.02819</a></p></li>\n</ul>",
      "rawMarkdown": "here are the papers that will clear your doubts. such tricks are common among kagglers without theoretical justification in the past\n\n\n-  CAN GRADIENT CLIPPING MITIGATE LABEL NOISE? ICLR 2020\nhttps://openreview.net/pdf?id=rklB76EKPr\n\n- Does label smoothing mitigate label noise?\nhttps://arxiv.org/abs/2003.02819\n",
      "votes": 2
    },
    {
      "id": 1106246,
      "postDate": "2020-12-08T16:36:51.803Z",
      "content": "<p>New to machine learning,  learned a lot from this thread.  Thanks.</p>",
      "rawMarkdown": "New to machine learning,  learned a lot from this thread.  Thanks."
    },
    {
      "id": 1230533,
      "postDate": "2021-03-08T08:00:37.210Z",
      "content": "<p>Hey Thanks for pointing out this issue. But I have few questions like : </p>\n<ol>\n<li>How did you figure out if there is noise in the data ? </li>\n<li>What do you call noise in the data ?</li>\n<li>What are various kind of noises in the data ?</li>\n</ol>",
      "rawMarkdown": "Hey Thanks for pointing out this issue. But I have few questions like : \n\n1. How did you figure out if there is noise in the data ? \n2. What do you call noise in the data ?\n3. What are various kind of noises in the data ?"
    },
    {
      "id": 1219308,
      "postDate": "2021-02-26T16:53:18.050Z",
      "content": "<p>Thank you for the information. Very helpful.</p>",
      "rawMarkdown": "Thank you for the information. Very helpful."
    },
    {
      "id": 1190111,
      "postDate": "2021-02-07T13:50:11.510Z",
      "content": "<p>I have a question about bi tempered logistic loss hyperarameter you choice to train model. Thank you</p>",
      "rawMarkdown": "I have a question about bi tempered logistic loss hyperarameter you choice to train model. Thank you"
    },
    {
      "id": 1180502,
      "postDate": "2021-02-01T09:51:47.033Z",
      "content": "<p>May I ask a novice question? If using early stopping but not monitoring the validation loss, which would be the criteria for stopping the training process?</p>",
      "rawMarkdown": "May I ask a novice question? If using early stopping but not monitoring the validation loss, which would be the criteria for stopping the training process?",
      "replies": [
        {
          "id": 1180531,
          "postDate": "2021-02-01T10:08:40.283Z",
          "content": "<p>Validation accuracy</p>",
          "rawMarkdown": "Validation accuracy",
          "votes": 2
        }
      ]
    },
    {
      "id": 1177882,
      "postDate": "2021-01-30T15:05:29.930Z",
      "content": "<p>Nice topic. I found it after an \"initial\" phase of just addressing technical issue, really enlightening.</p>",
      "rawMarkdown": "Nice topic. I found it after an \"initial\" phase of just addressing technical issue, really enlightening."
    },
    {
      "id": 1167162,
      "postDate": "2021-01-24T05:18:52.223Z",
      "content": "<p>Thanks a lot for sharing your experience！</p>",
      "rawMarkdown": "Thanks a lot for sharing your experience！"
    },
    {
      "id": 1131925,
      "postDate": "2020-12-30T04:40:40.473Z",
      "content": "<p>Can someone please share a sample code of how to use <strong>tempered loss</strong> ?</p>",
      "rawMarkdown": "Can someone please share a sample code of how to use **tempered loss** ?",
      "replies": [
        {
          "id": 1132111,
          "postDate": "2020-12-30T07:09:53.597Z",
          "content": "<p><a href=\"https://github.com/fhopfmueller/bi-tempered-loss-pytorch/blob/master/bi_tempered_loss_pytorch.py\" target=\"_blank\">https://github.com/fhopfmueller/bi-tempered-loss-pytorch/blob/master/bi_tempered_loss_pytorch.py</a></p>\n<p>Implementation-<a href=\"https://www.kaggle.com/debarshichanda/cassava-bitempered-logistic-loss\" target=\"_blank\">https://www.kaggle.com/debarshichanda/cassava-bitempered-logistic-loss</a></p>",
          "rawMarkdown": "https://github.com/fhopfmueller/bi-tempered-loss-pytorch/blob/master/bi_tempered_loss_pytorch.py\n\nImplementation-https://www.kaggle.com/debarshichanda/cassava-bitempered-logistic-loss",
          "votes": 1
        },
        {
          "id": 1146433,
          "postDate": "2021-01-09T19:14:32.027Z",
          "content": "<p>If you like Tensorflow you will like <a href=\"https://www.kaggle.com/durbin164/tpu-bitempered-logistic-loss-keras-tensorflow\" target=\"_blank\">this kernel</a> and t<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209773\" target=\"_blank\">his discussion</a> for Tempered loss. </p>",
          "rawMarkdown": "If you like Tensorflow you will like [this kernel](https://www.kaggle.com/durbin164/tpu-bitempered-logistic-loss-keras-tensorflow) and t[his discussion](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209773) for Tempered loss. ",
          "votes": 1
        },
        {
          "id": 1189745,
          "postDate": "2021-02-07T08:06:42.730Z",
          "content": "<p><a href=\"https://github.com/mlpanda/bi-tempered-loss-pytorch\" target=\"_blank\">https://github.com/mlpanda/bi-tempered-loss-pytorch</a></p>",
          "rawMarkdown": "https://github.com/mlpanda/bi-tempered-loss-pytorch"
        }
      ]
    },
    {
      "id": 1130723,
      "postDate": "2020-12-29T09:13:11.437Z",
      "content": "<p>Thank you for your inspiring discussion! I figured out why some labeling was weird. This is my first time confronting noisy data like this, so your various approaches helped me a lot. Thank you!</p>",
      "rawMarkdown": "Thank you for your inspiring discussion! I figured out why some labeling was weird. This is my first time confronting noisy data like this, so your various approaches helped me a lot. Thank you!"
    },
    {
      "id": 1128902,
      "postDate": "2020-12-27T21:40:24.373Z",
      "content": "<p>I'm not sure, do you recommend ignoring early stopping completely and just training for some set number of epochs? <br>\nIf so, then why?</p>",
      "rawMarkdown": "I'm not sure, do you recommend ignoring early stopping completely and just training for some set number of epochs? \nIf so, then why?"
    },
    {
      "id": 1124995,
      "postDate": "2020-12-24T10:13:46.633Z",
      "content": "<p>how you use a pre-train model like efficient nets as we can not use the internet in this competition? I try to submit a model with efficientnetB3 backbone but as code try to download the model, it needs internet and stops me from submitting my result</p>",
      "rawMarkdown": "how you use a pre-train model like efficient nets as we can not use the internet in this competition? I try to submit a model with efficientnetB3 backbone but as code try to download the model, it needs internet and stops me from submitting my result",
      "replies": [
        {
          "id": 1125166,
          "postDate": "2020-12-24T12:46:22.430Z",
          "content": "<p>People usually train with internet on, and then submit with internet off. During submission/inference you can simply load your trained models from a kaggle dataset, without needing to initialize model from the internet. If we really need the imagenet pretrained weights during submission then for that you need to create a dataset and load from that dataset.</p>",
          "rawMarkdown": "People usually train with internet on, and then submit with internet off. During submission/inference you can simply load your trained models from a kaggle dataset, without needing to initialize model from the internet. If we really need the imagenet pretrained weights during submission then for that you need to create a dataset and load from that dataset.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1113742,
      "postDate": "2020-12-15T16:51:11.273Z",
      "content": "<p>I am a beginner this is a very hard post to understand for me:(<br>\nWhat do you suggest me to have understanding of these things? </p>",
      "rawMarkdown": "I am a beginner this is a very hard post to understand for me:(\nWhat do you suggest me to have understanding of these things? "
    },
    {
      "id": 1113513,
      "postDate": "2020-12-15T14:14:07.337Z",
      "content": "<p>By looking at the data and experimenting which type of noise it is? small, large or random?</p>",
      "rawMarkdown": "By looking at the data and experimenting which type of noise it is? small, large or random?"
    },
    {
      "id": 1109900,
      "postDate": "2020-12-12T07:50:33.010Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>. The points were absolutely worth noting! Thanks a lot for sharing the insights. But one thing I would like to know is what if the hidden private data too is labeled in a similar fashion? I mean, it's all human labeling of the leaves, might as well be the case that what we see as noise, could also be available in the private dataset? In that case, if we do handle the claimed \"noise\" in this part of the data, it will surely gonna be a bad predictor of the hidden data. How would you approach this problem??</p>",
      "rawMarkdown": "Hi @hengck23. The points were absolutely worth noting! Thanks a lot for sharing the insights. But one thing I would like to know is what if the hidden private data too is labeled in a similar fashion? I mean, it's all human labeling of the leaves, might as well be the case that what we see as noise, could also be available in the private dataset? In that case, if we do handle the claimed \"noise\" in this part of the data, it will surely gonna be a bad predictor of the hidden data. How would you approach this problem??"
    },
    {
      "id": 1109481,
      "postDate": "2020-12-11T18:24:17.770Z",
      "content": "<p>Please give me the code reference or any sample code for label smoothing for this dataset. Thanks in advance.  </p>",
      "rawMarkdown": "Please give me the code reference or any sample code for label smoothing for this dataset. Thanks in advance.  ",
      "replies": [
        {
          "id": 1109501,
          "postDate": "2020-12-11T18:56:22.080Z",
          "content": "<p>If you use <code>tf.keras</code>, you can use this:</p>\n<p><code>\nloss = tf.keras.losses.CategoricalCrossentropy(\n    from_logits=False, label_smoothing=0.0001,\n    name='categorical_crossentropy'\n    )\nmodel.compile(optimizer=opt,loss=loss,metrics=['categorical_accuracy'])\n</code></p>\n<p><a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy</a></p>\n<p><a href=\"https://www.kaggle.com/aifahim\" target=\"_blank\">@aifahim</a> </p>",
          "rawMarkdown": "If you use `tf.keras`, you can use this:\n\n`\nloss = tf.keras.losses.CategoricalCrossentropy(\n    from_logits=False, label_smoothing=0.0001,\n    name='categorical_crossentropy'\n    )\nmodel.compile(optimizer=opt,loss=loss,metrics=['categorical_accuracy'])\n`\n\nhttps://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy\n\n@aifahim ",
          "votes": 2
        },
        {
          "id": 1109691,
          "postDate": "2020-12-12T01:06:56.310Z",
          "content": "<p>Thank you so much for sharing. If possible please share for the pytorch. </p>",
          "rawMarkdown": "Thank you so much for sharing. If possible please share for the pytorch. "
        },
        {
          "id": 1109747,
          "postDate": "2020-12-12T03:03:06.990Z",
          "content": "<p>Check this discussion.<br>\n<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/166833#930136\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/166833#930136</a></p>\n<p>I used it.</p>",
          "rawMarkdown": "Check this discussion.\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/166833#930136\n\nI used it.",
          "votes": 1
        },
        {
          "id": 1110126,
          "postDate": "2020-12-12T12:59:09.140Z",
          "content": "<p>Thank you so much Kim</p>",
          "rawMarkdown": "Thank you so much Kim"
        },
        {
          "id": 1118468,
          "postDate": "2020-12-19T05:09:02.627Z",
          "content": "<p>Hi. This code asumes 1-hot encoding, that is, each disease gets a binary column. </p>\n<p>The Sparse.CategoricalCrossentropy function does not accept label smoothing. Should we hot-encode the response vector and later convert ir back into a [0-4] vector?</p>",
          "rawMarkdown": "Hi. This code asumes 1-hot encoding, that is, each disease gets a binary column. \n\nThe Sparse.CategoricalCrossentropy function does not accept label smoothing. Should we hot-encode the response vector and later convert ir back into a [0-4] vector?",
          "votes": 1
        },
        {
          "id": 1120057,
          "postDate": "2020-12-20T15:06:40.203Z",
          "content": "<p><a href=\"https://www.kaggle.com/aifahim\" target=\"_blank\">@aifahim</a> Check this discussion<br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203103\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203103</a><br>\nHope this helps</p>",
          "rawMarkdown": "@aifahim Check this discussion\n[https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203103](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203103)\nHope this helps",
          "votes": 1
        },
        {
          "id": 1122733,
          "postDate": "2020-12-22T17:04:20.737Z",
          "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> .</p>",
          "rawMarkdown": "Thank you so much @debarshichanda .",
          "votes": 1
        }
      ]
    },
    {
      "id": 1105804,
      "postDate": "2020-12-08T08:07:22.460Z",
      "content": "<p>Thank you I learned a lot from this thread</p>",
      "rawMarkdown": "Thank you I learned a lot from this thread"
    },
    {
      "id": 1105579,
      "postDate": "2020-12-08T02:54:02.123Z",
      "content": "<p>👍👍👍 Can not agree more!</p>",
      "rawMarkdown": "👍👍👍 Can not agree more!"
    },
    {
      "id": 1113297,
      "postDate": "2020-12-15T11:02:43.243Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1106963,
      "postDate": "2020-12-09T08:57:22.187Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1940986,
      "postDate": "2022-09-15T17:31:18.360Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 5
    },
    {
      "id": 1186197,
      "postDate": "2021-02-04T16:37:15.677Z",
      "content": "<p>Thanks for the insight!</p>",
      "rawMarkdown": "Thanks for the insight!"
    },
    {
      "id": 1162081,
      "postDate": "2021-01-21T01:08:12.063Z",
      "content": "<p>Thank you for your discussion </p>",
      "rawMarkdown": "Thank you for your discussion "
    },
    {
      "id": 1112712,
      "postDate": "2020-12-14T20:35:52.903Z",
      "content": "<p>Thanks, learn too much from this</p>",
      "rawMarkdown": "Thanks, learn too much from this"
    }
  ],
  "comments": [
    {
      "id": 1105542,
      "author_name": "Leon",
      "author_url": "",
      "post_date": "2020-12-08T01:42:34.543000",
      "content": "<p>Pytorch implementation of Robust Bi-Tempered Logistic Loss Based on Bregman Divergences:<br>\n<a href=\"https://github.com/mlpanda/bi-tempered-loss-pytorch\" target=\"_blank\">https://github.com/mlpanda/bi-tempered-loss-pytorch</a></p>",
      "votes": 25,
      "replies": []
    },
    {
      "id": 1105940,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2020-12-08T10:58:18.780000",
      "content": "<p>What I done in the past when dealing with noisy labels, was to use the OOF prediction of the model trained with all data after softmax , and eliminate images where the softmax value is too small for the correct label. After eliminating a small quantity of training images, retrain from scratch with the remaining one.<br>\nFor example if after softmax you got the predictions (0.1, 0.3, 0.5, 0.05, 0.05) and if the correct label is 4 that image is eliminated from the training set.<br>\nI did not try it yet to this competition but it's on my todo list.</p>",
      "votes": 26,
      "replies": [
        {
          "id": 1105953,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-08T11:12:06.177000",
          "content": "<p>cleaning label is also one of the methods I will be trying later. If the test data is clean, I would recommend this type of method. </p>\n<p>but do note that the test data is also dirty. so any elimination of train data does changes some distrubtion</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1106122,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-12-08T14:52:37.103000",
          "content": "<p>I tried that and filtered the low softmax predictions, it didnt seem to have an effect on the lb. I think since the test set is probably noisy we must train with the noise and figure out a way to manage it!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1106229,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-12-08T16:09:27.777000",
          "content": "<p>Even if the public/private test set is more or less noisy, we should do our best to remove excessive noisy images. It's all about the ratio of noisy images to all images. If it's under a limit it doesn't hurt (sometimes quite the opposite) but if there are a big part of the training images, it can severely affect the model. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1177229,
          "author_name": "Kishan Joshi",
          "author_url": "",
          "post_date": "2021-01-30T06:31:21.967000",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/vladvdv\" target=\"_blank\">@vladvdv</a> , <br>\nSo, what did you find after removing the noise in the training data? <br>\nDid it hurt the LB? <br>\nIs there less noise in the test set, so removing train noise improves the LB?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1111114,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-13T12:29:17.453000",
      "content": "<p>my svm experiment:</p>\n<pre><code>efficient-net-b4 logistic classifier results:\n\ntrain\nloss : 0.22858\nacc  : 0.93024\n\nvalid\nloss : 0.31352\nacc  : 0.90257\n\n----\nuse efficient-net-b4 features and train a one-vs-rest linear svm classifier\nsvc = svm.LinearSVC(C=1, verbose=1, max_iter=100000, loss='squared_hinge', penalty='l2', dual=True )\n\nC=1\ntrain accuracy 0.9468948998072092\nvalid accuracy 0.8974299065420561\n\nC=0.1\ntrain accuracy 0.9369632529064672\nvalid accuracy 0.8985981308411215\n\nC=0.001\ntrain accuracy 0.9316469007419524\nvalid accuracy 0.9023364485981309\n\nC=0.0001\ntrain accuracy 0.9293100426476603\nvalid accuracy 0.9042056074766355\n\nC=0.00001\ntrain accuracy 0.9268563416486534\nvalid accuracy 0.8985981308411215\n</code></pre>",
      "votes": 17,
      "replies": [
        {
          "id": 1111345,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-13T16:34:38.397000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1111486,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-12-13T18:32:41.740000",
          "content": "<p>I wonder what is the benefit of a svm classifier compared to a normal fc layer trained with the CNN! Good experiments as always <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1111877,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-12-14T05:23:54.617000",
          "content": "<p>Probably margin loss is the main difference.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1111959,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-14T06:54:59.727000",
          "content": "<p>i would think it is the regularisation.<br>\nit shifts the solution of parameter = min(classification loss) a little.</p>\n<p>which is also with you should stop early in training deep network because you are not seeking the min point</p>",
          "votes": 9,
          "replies": []
        }
      ]
    },
    {
      "id": 1107824,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-10T01:40:22.510000",
      "content": "<p>probing the class distribution for public test set:</p>\n<pre><code>    label\n         0      cbb =    0.048\n         1     cbsd =    0.106\n         2      cgm =   ???\n         3      cmd =   ???\n         4  healthy =    0.139\n</code></pre>",
      "votes": 13,
      "replies": [
        {
          "id": 1107826,
          "author_name": "Harveen Singh Chadha",
          "author_url": "",
          "post_date": "2020-12-10T01:43:58.537000",
          "content": "<p>Class 2: 0.103</p>\n<p>Class 3: 0.602</p>",
          "votes": 15,
          "replies": []
        }
      ]
    },
    {
      "id": 1116899,
      "author_name": "Bcw93",
      "author_url": "",
      "post_date": "2020-12-17T14:49:47.760000",
      "content": "<p>so,how to choose the params of 't1' and 't2',Thanks!</p>",
      "votes": 11,
      "replies": [
        {
          "id": 1189729,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-07T07:57:45.370000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1106775,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-09T05:44:48.823000",
      "content": "<p><a href=\"https://www.youtube.com/watch?v=8mpBHbjG4E4\" target=\"_blank\">https://www.youtube.com/watch?v=8mpBHbjG4E4</a><br>\nDeep Learning with Label Noise - Kevin McGuinness - UPC TelecomBCN Barcelona 2019</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4405f84043b9dc1dd1ca197fd25f7e8d%2FSelection_155.png?generation=1607492627025997&amp;alt=media\" alt=\"\"></p>\n<p>i think someone needs to prepare some small set of clean labels and share in kaggle …</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1106792,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-12-09T06:08:39.967000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2F4518ec950c16c00491f1afd0d4f736ce%2FScreen%20Shot%202020-12-09%20at%2012.08.14%20AM.png?generation=1607494111001485&amp;alt=media\" alt=\"\"> - started :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107284,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2020-12-09T15:03:36.063000",
          "content": "<p>will making a set of clean labels be any useful as the test is also mislabeled?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107816,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-10T01:30:39.087000",
          "content": "<p>for analysis/experiment and estimate noise label in test, etc</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1105835,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-12-08T08:38:30.490000",
      "content": "<p>I am afraid the removed part of test data was among the cleanest part of the test set. <br>\nBefore it was removed, my loss and training process supposed to train on noisy labels and doing inference on clean test data. And it was more effective. </p>\n<p>Now things become harder and we have likely more proportion of noise in the test data since that part was removed. </p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1106007,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-08T12:02:21.957000",
      "content": "<p>i find a way to estimate or probe the noise label in public test:</p>\n<p>for example, I am interested in the confusion matrix of the ground truth of class-A.</p>\n<ol>\n<li><p>first a made a submission that set prediction = A for all test samples. I can back compute the number of ground truth that is A. for example we find that there are 500 instances of A</p></li>\n<li><p>use your model to make prediction. make a submission. e.g. we have lb score of s0</p></li>\n<li><p>select the very high confidence prediction of A. this step is very important. we need to make sure that those selected are really A (e.g. select those with probability &gt;0.95). The number selected also needs to be large enough. e.g. i can select about 80 high confident instances of predicted A.</p></li>\n<li><p>I make a new submission such that the selected 80 are set to prediction = random but not A. The lb score should drop. in fact, if they are really A, I can compute the theoretical drop in score. i probably need to do a few times with different random seed. e.g. theoretical drop in score = 80/500</p></li>\n<li><p>what if I set selected 80 are set to prediction = B.  or set prediction = C? the drop in lb score will reflect there are how many misclassifications from A to B, A to C, etc …</p></li>\n</ol>\n<hr>\n<p>but here is the catch:</p>\n<ul>\n<li><p>you only have test public lb score and not all the test lb score.</p></li>\n<li><p>you are going to waste many submissions (and maybe not worth it) if you are teaming up, you need to conserve submission</p></li>\n<li><p>but there is a 2019 server (late submission) for you to play with. here both public and private lb scores are exposed.i suspect 2019 data is less noisy though.</p></li>\n<li><p>you can always play the 2020 server after the competition is over as a post-competition analysis.</p></li>\n</ul>\n<hr>\n<p>you can compare the procedure above on your validation set and public set, and see if you would get the same numbers.  if the numbers are very different, the validation and test set are different from the trained model point of view.</p>\n<hr>\n<p>after doing this type of probing for a few competitions, you can ask yourself. does pseudo label really works?<br>\ni.e. how much can we know the labels with high confidence using our learned model?</p>\n<p>if you are interested, you can google like \"how can we test 2 distribution are the same or not?\", \"how can we know the noise level in dataset without ground truth label\", etc …</p>",
      "votes": 8,
      "replies": [
        {
          "id": 1106300,
          "author_name": "Harveen Singh Chadha",
          "author_url": "",
          "post_date": "2020-12-08T17:32:49.627000",
          "content": "<p>I think testing on 2019 data was a very good and valid point. The model that scored 0.85 LB in 2020 comp scored just 0.61 LB on 2019. On to my debugging hat now! Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1106532,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-12-08T22:50:34.347000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a></p>\n<p>I think that would not be necessary. There is a way to bypass it.</p>\n<p>I have given it a rough maths in my head so it should have a pretty good chance to work.</p>\n<p>Using technique to approximate how much noise is in the local validation set and then using it to make a stable cv where you know how much your model is able to get noise out of. It should be a reliable cv to evaluate your models with.</p>\n<p>That is not the limit, there are so many obvious implementations you can do with it. If you want me to give you more examples, just ask for it.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1110019,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-12T10:43:16.240000",
      "content": "<p>Training setting: On our benchmark, all methods are trained on the noisy training sets of two noise types (blue and red) under 10 noise levels (from 0% to 80%), and tested on the same clean validation set.</p>\n<p>you can google for more paper that reference to the dataset \"Controlled Noisy Web Labels\"</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F340143642066172cb62fa8be402c3dfd%2FSelection_025.png?generation=1607769793038638&amp;alt=media\" alt=\"\"></p>\n<p>\"Beyond Synthetic Noise: Deep Learning on Controlled Noisy Labels\" - ICML 2020<br>\n<a href=\"http://proceedings.mlr.press/v119/jiang20c/jiang20c.pdf\" target=\"_blank\">http://proceedings.mlr.press/v119/jiang20c/jiang20c.pdf</a><br>\n<a href=\"https://google.github.io/controlled-noisy-web-labels/index.html\" target=\"_blank\">https://google.github.io/controlled-noisy-web-labels/index.html</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 1110020,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-12T10:46:44.213000",
          "content": "<p>why you need early stopping</p>\n<p>quote: \"DNNs may not learn patterns first on red label noise. Arpit et al. (2017) found that DNNs learn patterns first, revealing an interesting property that DNNs are able to automatically learn generalizable “patterns” in the early training stage before memorizing all noisy training labels.\"</p>\n<p>see also: <a href=\"https://github.com/shengliu66/ELR\" target=\"_blank\">https://github.com/shengliu66/ELR</a></p>\n<p>why you need large learning rate:</p>\n<p>quote: \"ImageNet architectures generalize on noisy labels when the networks are fine-tuned. Kornblith et al. (2019) found that fine-tuning better architectures trained on ImageNet tend to perform better on downstream tasks of clean training labels. It is important to verify whether this holds on noisy training labels because if so, one can conveniently transfer better architectures to better overcome the noisy labels\"</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1110046,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-12-12T11:27:43.060000",
          "content": "<p>I used ELR_plus few days ago but got no improvement.  </p>\n<p>May be I didn't tune it correctly</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1112940,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-12-15T03:12:06.863000",
          "content": "<p>That's one of the best papers I have read recently, so many insights, very well structured, clean and easy to understand. Thanks a lot for sharing it <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>! <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> I used default params from the paper in my case it improves compared to CE baseline. Perhaps, your comparison is not made to a baseline model.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1113275,
          "author_name": "bbbarbel",
          "author_url": "",
          "post_date": "2020-12-15T10:46:11.047000",
          "content": "<p>so you think the main noise is blue noise ,not red noise? otherwise, early stop is not working!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119268,
          "author_name": "Debarshi Chanda",
          "author_url": "",
          "post_date": "2020-12-19T21:21:00.100000",
          "content": "<p>I just need to change the loss function to use this right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1120665,
          "author_name": "Bcw93",
          "author_url": "",
          "post_date": "2020-12-21T03:17:59.440000",
          "content": "<p>I also used elr_loss,my local cv is increase,but the lb score is decrease.May be there have some problem?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1121450,
          "author_name": "Bcw93",
          "author_url": "",
          "post_date": "2020-12-21T16:41:04.600000",
          "content": "<p>Hi,did you use the tempered loss to get some improve in cv?Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1105698,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-12-08T05:17:42.827000",
      "content": "<p>Here is another research work regarding this - <a href=\"https://arxiv.org/pdf/2002.06541v1.pdf\" target=\"_blank\">Learning Not to Learn in the Presence of Noisy Labels</a>. They showed <strong>gambler’s loss</strong> is an effective choice to deal with noisy labels. </p>",
      "votes": 5,
      "replies": [
        {
          "id": 1106313,
          "author_name": "Rakka Alhazimi",
          "author_url": "",
          "post_date": "2020-12-08T17:44:43.447000",
          "content": "<p>thanks, I'll read it :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119616,
          "author_name": "rajan",
          "author_url": "",
          "post_date": "2020-12-20T08:35:27.117000",
          "content": "<p>hey i have implemented GAmblers loss <br>\nPLEASE HAVE Look<br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/205424\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/205424</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1110165,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-12T13:37:06.870000",
      "content": "<p>just an idea</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F47cb2ad82014dbba920450f861c3f72d%2FSelection_027.png?generation=1607780223050350&amp;alt=media\" alt=\"\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 1110200,
          "author_name": "Heroseo",
          "author_url": "",
          "post_date": "2020-12-12T14:08:36.850000",
          "content": "<p>I am trying and thinking similarly.<br>\nThanks for sharing your ideas.</p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1110212,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-12T14:19:02.027000",
          "content": "<p>similar idea includes trying to find subclusters within each class. <br>\nsome subclusters are clean, some are correlated noise (i.e. similarly mislabeled in both train and test), some are  uncorrelated noise.</p>\n<p>e.g. we have K clusters in one class, then we can extract K embeddings for an input test sample. classifier is based on the K embeddings</p>\n<p>some reference paper:<br>\n1) clustering <br>\n\"Unsupervised Visual Representation Learning with SwAV\"<br>\n\"Self-labelling via simultaneous clustering and representation learning\"</p>\n<p>2) sub-class classifier <br>\n\"SoftTriple Loss: Deep Metric Learning Without Triplet Sampling\"</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1111143,
          "author_name": "mknzfr",
          "author_url": "",
          "post_date": "2020-12-13T13:06:11.217000",
          "content": "<p>For SoftTriple there is a great repo: <a href=\"https://github.com/KevinMusgrave/pytorch-metric-learning\" target=\"_blank\">https://github.com/KevinMusgrave/pytorch-metric-learning</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1106021,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-08T12:10:13.510000",
      "content": "<p>qoute : \"Our method is the first data-recalibrating method which is guaranteed to converge to a well behaved classifier\" … \"We provide theoretical guarantees showing that for a wide variety of (unknown) noise patterns, a classifier trained with this strategy converges to be consistent with the Bayes classifier\"</p>\n<p>wow, guaranteed ???</p>\n<p><a href=\"https://openreview.net/pdf?id=ZPa2SyGcbwh\" target=\"_blank\">https://openreview.net/pdf?id=ZPa2SyGcbwh</a><br>\nICLR 2021 -LEARNING WITH FEATURE DEPENDENT LABEL NOISE:<br>\nA PROGRESSIVE APPROACH</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1111359,
      "author_name": "AIFahim",
      "author_url": "",
      "post_date": "2020-12-13T16:51:00.420000",
      "content": "<p>Can anyone share notebook which use bi tempered loss with pytorch?  Thanks in advance.  </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1105982,
      "author_name": "Konrad Banachewicz",
      "author_url": "",
      "post_date": "2020-12-08T11:42:48.450000",
      "content": "<p>I took the liberty of uploading the semi-official TF implementation</p>\n<p><a href=\"https://www.kaggle.com/konradb/bitemperedloss\" target=\"_blank\">https://www.kaggle.com/konradb/bitemperedloss</a></p>",
      "votes": 4,
      "replies": [
        {
          "id": 1111941,
          "author_name": "Aziz_Belaweid",
          "author_url": "",
          "post_date": "2020-12-14T06:39:31.557000",
          "content": "<p>Hey thanks for sharing, did you manage to make it work in your pipeline ? if Yes can you share it with us ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1112906,
          "author_name": "Thomas Brekk Unnvik",
          "author_url": "",
          "post_date": "2020-12-15T02:07:53.283000",
          "content": "<p>Anyone tried implementing this? Not quite sure what to pass the loss function for the \"activations\" and \"labels\" parameters. Any tips on how to use it with EfficientNet (or any model..) in Keras/TF?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1146430,
          "author_name": "Md. Masud Rana",
          "author_url": "",
          "post_date": "2021-01-09T19:10:39.330000",
          "content": "<p><a href=\"https://www.kaggle.com/aziz69\" target=\"_blank\">@aziz69</a>  Here is the BiTemperedLogisticLoss <a href=\"https://www.kaggle.com/durbin164/tpu-bitempered-logistic-loss-keras-tensorflow\" target=\"_blank\">kernel </a>implementation. </p>\n<p>Another Implementation is <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209773\" target=\"_blank\">here.</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1109013,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-11T08:44:24.927000",
      "content": "<p>is average the best choice in ensemble in the presence of noise? … there could be outlier scores!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1112582,
          "author_name": "Thomas Brekk Unnvik",
          "author_url": "",
          "post_date": "2020-12-14T18:20:48.120000",
          "content": "<p>What about using confusion matrices and weighting models according to how well they predict each class? So if model A is slightly better than model B at predicting class 1 correctly model A would be weighted higher for probabilities regarding class 1, and so on. It might be that one model is better overall, but that the second model still scores better for at least one of the classes.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1106614,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2020-12-09T00:43:53.513000",
      "content": "<p>Great share! I did some sort of label smoothing to get 0.901 and will definitely try tempered loss to see how they combine together, probably label smoothing also allows to make logistic loss less sensitive. I wonder what is the upper bound accuracy for the hidden test set. Wouldn't also manually removing noisy labels, e.g. mislabeled healthy images in dataset, help model to create a better decision boundary?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1106728,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-09T04:24:01.327000",
          "content": "<p>\". I wonder what is the upper bound accuracy for the hidden test set. \"</p>\n<p>the leaderboard provide you a clue</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107538,
          "author_name": "Matthew Thomas",
          "author_url": "",
          "post_date": "2020-12-09T19:00:05.713000",
          "content": "<p>Mind sharing the label smoothing factor that worked for you? I'm trying 0.1 at the moment and can report back on the effect after the experiment.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107785,
          "author_name": "Matthew Thomas",
          "author_url": "",
          "post_date": "2020-12-10T00:42:48.623000",
          "content": "<p>Using label smoothing of 0.1 improved my public score by 0.01.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1105442,
      "author_name": "Hiram Coria 🧬",
      "author_url": "",
      "post_date": "2020-12-07T21:52:30.650000",
      "content": "<p>Would you kindly explain why a dataset is <strong>noisy</strong>? I searched on Internet but i still don't get it.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1105445,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-07T22:04:11.403000",
          "content": "<p>the label is wrong.</p>\n<ul>\n<li><p>if you look at the images by class e.g. healthy leaf, you would find that some images are not healthy at all. a few forum posts show some examples of such images</p></li>\n<li><p>I am not sure who or how the images are labelled. i read in some papers that unless the person is well trained, it is difficult for a person to judge the diseases of the plant.</p></li>\n</ul>\n<hr>\n<p>from machine learning point of view, the tsne plot of the embedding will show that the data points are not separable by class</p>",
          "votes": 15,
          "replies": []
        },
        {
          "id": 1105495,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-12-07T23:47:16.670000",
          "content": "<p>Then, in a machine learning context, a noisy dataset has wrong labeled items?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1105496,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-12-07T23:47:32.087000",
          "content": "<p>The figure below shows the TSNE embeddings for the efficientnet-b3 model without the classifier. I couldn't separate the classes, there is always some overlap between the classes.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F337744%2Ff3471414927b45e6632dab3aecc4ed20%2FScreen%20Shot%202020-12-07%20at%206.41.33%20PM.png?generation=1607384517717402&amp;alt=media\" alt=\"\"></p>",
          "votes": 19,
          "replies": []
        },
        {
          "id": 1105531,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-08T01:18:15.567000",
          "content": "<p>embedding that cannot separate can also methods that the model or loss is not good enough.</p>\n<p>you should physically inspect the images and the labels that overlap to confirm the problem.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1105537,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-12-08T01:32:36.630000",
          "content": "<p>Of course, embedding does not necessarily separate the classes. The figures above show the tsne of the output of a pretrained model - so it is model-dependent. However it is also not straightforward to train a tsne, it requires tweaking some hyper-parameters - and this is what I meant. Non-optimal parameters produce non-optimal clusters with noise.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1105543,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-08T01:46:21.607000",
          "content": "<p>i think you can say noise means \"not having the true values\" in machine learning. e.g. wrong labels, wrong measurements in input, etc. more correctly, we should say that the cassava competition has noisy labels. noisy dataset can mean noisy x (input) or noisy y (label) or both</p>\n<p>i want to add, there is always some noise in data (e.g. imagenet). we are interested to know how much noise will affect accuracy results. that is the signal-to-noise problem.</p>\n<p>sometimes we inject noise in training (e.g. flip labels) noise is not necessarily bad if it is in the appropriate amount.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1105548,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-12-08T01:55:52.700000",
          "content": "<p>I agree with all of your points.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1111940,
          "author_name": "Aziz_Belaweid",
          "author_url": "",
          "post_date": "2020-12-14T06:38:27.193000",
          "content": "<p><a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a>  thanks for sharing your results, could you post a notebook on how to visualize embeddings using tsne ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1112528,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-12-14T17:20:44.250000",
          "content": "<p><a href=\"https://www.kaggle.com/aziz69\" target=\"_blank\">@aziz69</a> I made my TSNE notebook publicly available in <a href=\"https://www.kaggle.com/tolgadincer/cldc-tsne\" target=\"_blank\">here</a>.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1182699,
          "author_name": "mane.stoimchev",
          "author_url": "",
          "post_date": "2021-02-02T14:58:45.257000",
          "content": "<p><a href=\"https://www.kaggle.com/tolgadincer\" target=\"_blank\">@tolgadincer</a>  Here is my TSNE, i am using effnetb4 noisy student, with ELR loss, single model 0.896. Does BiTemperedLoss give better results compared to ELR loss? How did you tune the parameters t1 and t2 ? </p>\n<p><a href=\"https://ibb.co/72wXjyx\"><img src=\"https://i.ibb.co/pJtxbX6/tsne.png\" alt=\"tsne\"></a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1133029,
      "author_name": "Nir",
      "author_url": "",
      "post_date": "2020-12-30T21:36:21.147000",
      "content": "<p>If the test dataset is noisy - it suggests that the leader-board closely ranked top results are somewhat random</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1109544,
      "author_name": "s.s.o",
      "author_url": "",
      "post_date": "2020-12-11T20:02:54.223000",
      "content": "<p>This paper might be of interest from google. Sharpness-Aware Minimization for Efficiently Improving Generalization <a href=\"https://arxiv.org/pdf/2010.01412.pdf\" target=\"_blank\">paper</a> and the <a href=\"https://github.com/google-research/sam\" target=\"_blank\">code </a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1110123,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-12T12:54:31.713000",
          "content": "<p><a href=\"https://www.kaggle.com/sso2000\" target=\"_blank\">@sso2000</a> </p>\n<p>thanks for the  Sharpness-Aware Minimization paper. i briefly read through it and feel that it should work. they are adding a reglarizer to the network parameter. </p>\n<p>it did indeed work on my validation set, an increase of +0.002. At the same time, LB drop by +0.002.</p>\n<p>I think the LB is not only noisy but also contains some data not from the 2020 train set.  I recall there is a rescore of some LB for removal of some inclass dataset  (which i think is that 2019 set). 2019 did indeed exist in the 2020 test set … those with 2019 labels are removed and maybe those without 2019 still remains?</p>\n<p>my next experiment is to try SAM on combined 2019+2020 train dataset</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1106755,
      "author_name": "Ankush kuwar",
      "author_url": "",
      "post_date": "2020-12-09T05:09:59.240000",
      "content": "<blockquote>\n  <p>training early stopping (i.e. not choosing the lowest validation loss): lb0.899<br>\n  Hey <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> can you elaborate this one? you're saying is to inference with last checkpoint and not lowest val_loss right? I am having hard time getting my EFFNETB4 above 0.891, but i have seen 0.30~ val_loss with it also the val_categorical_accuracy almost matches with LB. Currently, even my B7 couldn't pass 0.893. Lol..<br>\n  My Setup&gt;&gt; Cutmixup, Some_tf_augs, ROT_SHEAR_AUGS, Cosine LR(0.00001&gt;0.0004) and Good old Adam. </p>\n</blockquote>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1106628,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2020-12-09T01:18:32.017000",
      "content": "<p>Looks like small margin can also have worse decision boundary compared to large, I think manually handling mislabeled training samples should have the same affect of correcting the decision boundary. Unless, there is no systematic error which generalizes to test set why would we like to keep mislabeled samples in training? It would be like gambling if label noise is random…</p>\n<p>Also don't we have risk of overfitting to noise we have in our training set and public LB? e.g. imagine tuning t1 and t2 </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2Ff6615268e1a185767cd88dcd9e0c51f5%2FScreen%20Shot%202020-12-08%20at%207.17.03%20PM.png?generation=1607476710189661&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2F67392bfc4f32e054c30ac0575e3503e5%2FScreen%20Shot%202020-12-08%20at%207.17.28%20PM.png?generation=1607476703235679&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1105956,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-08T11:14:18.380000",
      "content": "<p>here are the papers that will clear your doubts. such tricks are common among kagglers without theoretical justification in the past</p>\n<ul>\n<li><p>CAN GRADIENT CLIPPING MITIGATE LABEL NOISE? ICLR 2020<br>\n<a href=\"https://openreview.net/pdf?id=rklB76EKPr\" target=\"_blank\">https://openreview.net/pdf?id=rklB76EKPr</a></p></li>\n<li><p>Does label smoothing mitigate label noise?<br>\n<a href=\"https://arxiv.org/abs/2003.02819\" target=\"_blank\">https://arxiv.org/abs/2003.02819</a></p></li>\n</ul>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1106246,
      "author_name": "luqing Wang",
      "author_url": "",
      "post_date": "2020-12-08T16:36:51.803000",
      "content": "<p>New to machine learning,  learned a lot from this thread.  Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1230533,
      "author_name": "AbhishekAgrawal",
      "author_url": "",
      "post_date": "2021-03-08T08:00:37.210000",
      "content": "<p>Hey Thanks for pointing out this issue. But I have few questions like : </p>\n<ol>\n<li>How did you figure out if there is noise in the data ? </li>\n<li>What do you call noise in the data ?</li>\n<li>What are various kind of noises in the data ?</li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1219308,
      "author_name": "Vishnu Praneeth",
      "author_url": "",
      "post_date": "2021-02-26T16:53:18.050000",
      "content": "<p>Thank you for the information. Very helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1190111,
      "author_name": "HungNT",
      "author_url": "",
      "post_date": "2021-02-07T13:50:11.510000",
      "content": "<p>I have a question about bi tempered logistic loss hyperarameter you choice to train model. Thank you</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1180502,
      "author_name": "jcesquivel",
      "author_url": "",
      "post_date": "2021-02-01T09:51:47.033000",
      "content": "<p>May I ask a novice question? If using early stopping but not monitoring the validation loss, which would be the criteria for stopping the training process?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1180531,
          "author_name": "Debarshi Chanda",
          "author_url": "",
          "post_date": "2021-02-01T10:08:40.283000",
          "content": "<p>Validation accuracy</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1177882,
      "author_name": "Etienne R",
      "author_url": "",
      "post_date": "2021-01-30T15:05:29.930000",
      "content": "<p>Nice topic. I found it after an \"initial\" phase of just addressing technical issue, really enlightening.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1167162,
      "author_name": "Lino",
      "author_url": "",
      "post_date": "2021-01-24T05:18:52.223000",
      "content": "<p>Thanks a lot for sharing your experience！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1131925,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-30T04:40:40.473000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1132111,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-30T07:09:53.597000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1146433,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-09T19:14:32.027000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1189745,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-07T08:06:42.730000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1130723,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-29T09:13:11.437000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1128902,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-27T21:40:24.373000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1124995,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-24T10:13:46.633000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1125166,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-24T12:46:22.430000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1113742,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-15T16:51:11.273000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1113513,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-15T14:14:07.337000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1109900,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-12T07:50:33.010000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1109481,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-11T18:24:17.770000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1109501,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-11T18:56:22.080000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1109691,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-12T01:06:56.310000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1109747,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-12T03:03:06.990000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1110126,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-12T12:59:09.140000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1118468,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-19T05:09:02.627000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1120057,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-20T15:06:40.203000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1122733,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-22T17:04:20.737000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1105804,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-08T08:07:22.460000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1105579,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-08T02:54:02.123000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1113297,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-15T11:02:43.243000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1106963,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-09T08:57:22.187000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1940986,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-15T17:31:18.360000",
      "content": "",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1186197,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-04T16:37:15.677000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1162081,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-21T01:08:12.063000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1112712,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-14T20:35:52.903000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1105431": "It is now apparent that the dataset is noisy (for both train and public/private test).\nYou also note that with the same model (e.g. efficientnet-b4) the performances of kaggers are very different. It is the training process that matters.\n\nyou are now training in noise and the objective is not to train for the lowest loss but one that is consistent between train and test. This competition is like estimating the noise in the dataset and apply correct training parameters to get to the bayes error.\n\nthere are common techniques like early stopping and label smooth, apply SVM etc.\n\nbut there are also more advantage techniques if you google for \"training in the presence of noisy labels\". but do note that unlike cases in most academic papers where the test is clean, our test is also noisy. you may also want to read about testing in the presence of noisy data\n\nhere I introduce one method that works for me:\n-  training early stopping (i.e. not choosing the lowest validation loss): lb0.899\n-  training tempered loss (without early stopping): lb0.901\n-  training normal loss (without early stopping): lb0.898\n\n\nhttps://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\n- log loss is modified so that outliers (noise labels) are not penalized that much.\n- I think our case is close to large-margin noise\n\n![](https://1.bp.blogspot.com/-MSsW9QUCyXM/XWQNInb2ztI/AAAAAAAAEjc/_yu48eAnTjMWQtUbjNqhOlqurgapjSIzACLcBGAs/s640/Screenshot%2B2019-08-26%2Bat%2B9.39.48%2BAM.png)\n\n![](https://1.bp.blogspot.com/-wTnN2w3ENVc/XWQOEFA67TI/AAAAAAAAEjo/0W-4vcYefTcwn7tURpmQ6zMUBrUmJ5e-wCLcBGAs/s640/Screenshot%2B2019-08-26%2Bat%2B9.37.41%2BAM.png)",
    "1105542": "Pytorch implementation of Robust Bi-Tempered Logistic Loss Based on Bregman Divergences:\nhttps://github.com/mlpanda/bi-tempered-loss-pytorch",
    "1105940": "What I done in the past when dealing with noisy labels, was to use the OOF prediction of the model trained with all data after softmax , and eliminate images where the softmax value is too small for the correct label. After eliminating a small quantity of training images, retrain from scratch with the remaining one.\nFor example if after softmax you got the predictions (0.1, 0.3, 0.5, 0.05, 0.05) and if the correct label is 4 that image is eliminated from the training set.\nI did not try it yet to this competition but it's on my todo list.",
    "1111114": "my svm experiment:\n\n```\nefficient-net-b4 logistic classifier results:\n\ntrain\nloss : 0.22858\nacc  : 0.93024\n\nvalid\nloss : 0.31352\nacc  : 0.90257\n\n----\nuse efficient-net-b4 features and train a one-vs-rest linear svm classifier\nsvc = svm.LinearSVC(C=1, verbose=1, max_iter=100000, loss='squared_hinge', penalty='l2', dual=True )\n\nC=1\ntrain accuracy 0.9468948998072092\nvalid accuracy 0.8974299065420561\n\nC=0.1\ntrain accuracy 0.9369632529064672\nvalid accuracy 0.8985981308411215\n\nC=0.001\ntrain accuracy 0.9316469007419524\nvalid accuracy 0.9023364485981309\n\nC=0.0001\ntrain accuracy 0.9293100426476603\nvalid accuracy 0.9042056074766355\n\nC=0.00001\ntrain accuracy 0.9268563416486534\nvalid accuracy 0.8985981308411215\n```",
    "1107824": "probing the class distribution for public test set:\n\n```\n\n\tlabel\n\t\t 0      cbb =    0.048\n\t\t 1     cbsd =    0.106\n\t\t 2      cgm =   ???\n\t\t 3      cmd =   ???\n\t\t 4  healthy =    0.139\n\n\n```",
    "1116899": "so,how to choose the params of 't1' and 't2',Thanks!",
    "1106775": "https://www.youtube.com/watch?v=8mpBHbjG4E4\nDeep Learning with Label Noise - Kevin McGuinness - UPC TelecomBCN Barcelona 2019\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4405f84043b9dc1dd1ca197fd25f7e8d%2FSelection_155.png?generation=1607492627025997&alt=media)\n\ni think someone needs to prepare some small set of clean labels and share in kaggle ...",
    "1105835": "I am afraid the removed part of test data was among the cleanest part of the test set. \nBefore it was removed, my loss and training process supposed to train on noisy labels and doing inference on clean test data. And it was more effective. \n\nNow things become harder and we have likely more proportion of noise in the test data since that part was removed. ",
    "1106007": "i find a way to estimate or probe the noise label in public test:\n\nfor example, I am interested in the confusion matrix of the ground truth of class-A.\n\n1. first a made a submission that set prediction = A for all test samples. I can back compute the number of ground truth that is A. for example we find that there are 500 instances of A\n\n2. use your model to make prediction. make a submission. e.g. we have lb score of s0\n\n3. select the very high confidence prediction of A. this step is very important. we need to make sure that those selected are really A (e.g. select those with probability >0.95). The number selected also needs to be large enough. e.g. i can select about 80 high confident instances of predicted A.\n\n4. I make a new submission such that the selected 80 are set to prediction = random but not A. The lb score should drop. in fact, if they are really A, I can compute the theoretical drop in score. i probably need to do a few times with different random seed. e.g. theoretical drop in score = 80/500\n\n5. what if I set selected 80 are set to prediction = B.  or set prediction = C? the drop in lb score will reflect there are how many misclassifications from A to B, A to C, etc ...\n\n---\n\nbut here is the catch:\n\n- you only have test public lb score and not all the test lb score.\n\n- you are going to waste many submissions (and maybe not worth it) if you are teaming up, you need to conserve submission\n\n- but there is a 2019 server (late submission) for you to play with. here both public and private lb scores are exposed.i suspect 2019 data is less noisy though.\n\n- you can always play the 2020 server after the competition is over as a post-competition analysis.\n\n\n\n---\n\nyou can compare the procedure above on your validation set and public set, and see if you would get the same numbers.  if the numbers are very different, the validation and test set are different from the trained model point of view.\n\n---\n\nafter doing this type of probing for a few competitions, you can ask yourself. does pseudo label really works?\ni.e. how much can we know the labels with high confidence using our learned model?\n\nif you are interested, you can google like \"how can we test 2 distribution are the same or not?\", \"how can we know the noise level in dataset without ground truth label\", etc ...",
    "1110019": "Training setting: On our benchmark, all methods are trained on the noisy training sets of two noise types (blue and red) under 10 noise levels (from 0% to 80%), and tested on the same clean validation set.\n\n\n\nyou can google for more paper that reference to the dataset \"Controlled Noisy Web Labels\"\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F340143642066172cb62fa8be402c3dfd%2FSelection_025.png?generation=1607769793038638&alt=media)\n\n\n\n\"Beyond Synthetic Noise: Deep Learning on Controlled Noisy Labels\" - ICML 2020\nhttp://proceedings.mlr.press/v119/jiang20c/jiang20c.pdf\nhttps://google.github.io/controlled-noisy-web-labels/index.html\n",
    "1105698": "Here is another research work regarding this - [Learning Not to Learn in the Presence of Noisy Labels](https://arxiv.org/pdf/2002.06541v1.pdf). They showed **gambler’s loss** is an effective choice to deal with noisy labels. ",
    "1110165": "just an idea\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F47cb2ad82014dbba920450f861c3f72d%2FSelection_027.png?generation=1607780223050350&alt=media)",
    "1106021": "qoute : \"Our method is the first data-recalibrating method which is guaranteed to converge to a well behaved classifier\" ... \"We provide theoretical guarantees showing that for a wide variety of (unknown) noise patterns, a classifier trained with this strategy converges to be consistent with the Bayes classifier\"\n\nwow, guaranteed ???\n\nhttps://openreview.net/pdf?id=ZPa2SyGcbwh\nICLR 2021 -LEARNING WITH FEATURE DEPENDENT LABEL NOISE:\nA PROGRESSIVE APPROACH",
    "1111359": "Can anyone share notebook which use bi tempered loss with pytorch?  Thanks in advance.  ",
    "1105982": "I took the liberty of uploading the semi-official TF implementation\n\nhttps://www.kaggle.com/konradb/bitemperedloss\n",
    "1109013": "is average the best choice in ensemble in the presence of noise? ... there could be outlier scores!",
    "1106614": "Great share! I did some sort of label smoothing to get 0.901 and will definitely try tempered loss to see how they combine together, probably label smoothing also allows to make logistic loss less sensitive. I wonder what is the upper bound accuracy for the hidden test set. Wouldn't also manually removing noisy labels, e.g. mislabeled healthy images in dataset, help model to create a better decision boundary?",
    "1105442": "Would you kindly explain why a dataset is **noisy**? I searched on Internet but i still don't get it.",
    "1133029": "If the test dataset is noisy - it suggests that the leader-board closely ranked top results are somewhat random",
    "1109544": "This paper might be of interest from google. Sharpness-Aware Minimization for Efficiently Improving Generalization [paper](https://arxiv.org/pdf/2010.01412.pdf) and the [code ](https://github.com/google-research/sam)",
    "1106755": "> training early stopping (i.e. not choosing the lowest validation loss): lb0.899\nHey @hengck23 can you elaborate this one? you're saying is to inference with last checkpoint and not lowest val_loss right? I am having hard time getting my EFFNETB4 above 0.891, but i have seen 0.30~ val_loss with it also the val_categorical_accuracy almost matches with LB. Currently, even my B7 couldn't pass 0.893. Lol..\nMy Setup>> Cutmixup, Some_tf_augs, ROT_SHEAR_AUGS, Cosine LR(0.00001>0.0004) and Good old Adam. ",
    "1106628": "Looks like small margin can also have worse decision boundary compared to large, I think manually handling mislabeled training samples should have the same affect of correcting the decision boundary. Unless, there is no systematic error which generalizes to test set why would we like to keep mislabeled samples in training? It would be like gambling if label noise is random...\n\nAlso don't we have risk of overfitting to noise we have in our training set and public LB? e.g. imagine tuning t1 and t2 \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2Ff6615268e1a185767cd88dcd9e0c51f5%2FScreen%20Shot%202020-12-08%20at%207.17.03%20PM.png?generation=1607476710189661&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F558069%2F67392bfc4f32e054c30ac0575e3503e5%2FScreen%20Shot%202020-12-08%20at%207.17.28%20PM.png?generation=1607476703235679&alt=media)",
    "1105956": "here are the papers that will clear your doubts. such tricks are common among kagglers without theoretical justification in the past\n\n\n-  CAN GRADIENT CLIPPING MITIGATE LABEL NOISE? ICLR 2020\nhttps://openreview.net/pdf?id=rklB76EKPr\n\n- Does label smoothing mitigate label noise?\nhttps://arxiv.org/abs/2003.02819\n",
    "1106246": "New to machine learning,  learned a lot from this thread.  Thanks.",
    "1230533": "Hey Thanks for pointing out this issue. But I have few questions like : \n\n1. How did you figure out if there is noise in the data ? \n2. What do you call noise in the data ?\n3. What are various kind of noises in the data ?",
    "1219308": "Thank you for the information. Very helpful.",
    "1190111": "I have a question about bi tempered logistic loss hyperarameter you choice to train model. Thank you",
    "1180502": "May I ask a novice question? If using early stopping but not monitoring the validation loss, which would be the criteria for stopping the training process?",
    "1177882": "Nice topic. I found it after an \"initial\" phase of just addressing technical issue, really enlightening.",
    "1167162": "Thanks a lot for sharing your experience！",
    "1131925": "Can someone please share a sample code of how to use **tempered loss** ?",
    "1130723": "Thank you for your inspiring discussion! I figured out why some labeling was weird. This is my first time confronting noisy data like this, so your various approaches helped me a lot. Thank you!",
    "1128902": "I'm not sure, do you recommend ignoring early stopping completely and just training for some set number of epochs? \nIf so, then why?",
    "1124995": "how you use a pre-train model like efficient nets as we can not use the internet in this competition? I try to submit a model with efficientnetB3 backbone but as code try to download the model, it needs internet and stops me from submitting my result",
    "1113742": "I am a beginner this is a very hard post to understand for me:(\nWhat do you suggest me to have understanding of these things? ",
    "1113513": "By looking at the data and experimenting which type of noise it is? small, large or random?",
    "1109900": "Hi @hengck23. The points were absolutely worth noting! Thanks a lot for sharing the insights. But one thing I would like to know is what if the hidden private data too is labeled in a similar fashion? I mean, it's all human labeling of the leaves, might as well be the case that what we see as noise, could also be available in the private dataset? In that case, if we do handle the claimed \"noise\" in this part of the data, it will surely gonna be a bad predictor of the hidden data. How would you approach this problem??",
    "1109481": "Please give me the code reference or any sample code for label smoothing for this dataset. Thanks in advance.  ",
    "1105804": "Thank you I learned a lot from this thread",
    "1105579": "👍👍👍 Can not agree more!",
    "1113297": "",
    "1106963": "",
    "1940986": "Thanks for sharing",
    "1186197": "Thanks for the insight!",
    "1162081": "Thank you for your discussion ",
    "1112712": "Thanks, learn too much from this"
  }
}