{
  "id": 218205,
  "title": "Try Online Uncertainty Sample Mining - OUSMLoss (pytorch Implementation)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/218205",
  "author_name": "",
  "post_date": "2021-02-09T17:09:09.048554600Z",
  "votes": 25,
  "comment_count": 29,
  "views": 0,
  "content": "<p>if you read this paper <a href=\"https://arxiv.org/pdf/1901.07759.pdf\" target=\"_blank\">ROBUST LEARNING AT NOISY LABELED MEDICAL IMAGES:\nAPPLIED TO SKIN LESION CLASSIFICATION</a> , you will understand that OUSMLoss is a good choice for tackling label noise,<br>\nfrom this link,you can try OUSMLoss implementation : <a href=\"https://github.com/analokmaus/kaggle-panda-challenge-public/blob/208caf4c83a5ab9d181e66eee447cd2e475d58dc/models/noisy_loss.py\" target=\"_blank\">https://github.com/analokmaus/kaggle-panda-challenge-public/blob/208caf4c83a5ab9d181e66eee447cd2e475d58dc/models/noisy_loss.py</a><br>\nfor handling noisy labels issue, we tried but couldn't get better result than tempered loss, if you have tried it, can you please tell us if it works for your model or not? </p>\n<p>which loss function works best for you(you can keep it secret,if you don't want to share though) <strong>(asking for a friend)</strong></p>\n<p>if you simply want to use OUSMLoss  without  nn module implementation and with simple function implementation, then you can try the code below : </p>\n<pre><code>def _check_input_type(x, y, loss):\n    if loss in ['MSELoss', 'MAELoss', 'huber']:\n        if x.shape[-1] == 1:\n            return x.squeeze(), y.float()\n        else:\n            return x, y.float()\n    elif loss in ['CrossEntropy']:\n        return x, y.long()\n    else:\n        return x, y\n</code></pre>\n<pre><code>def soft_cross_entropy_loss(logits, targets, weights=1, reduction='none'):\n    if len(targets.shape) == 1 or targets.shape[1] == 1:\n        onehot_targets = torch.eye(logits.shape[1])[targets].to(logits.device)\n    else:\n        onehot_targets = targets\n    loss = -torch.sum(onehot_targets * F.log_softmax(logits, 1), 1)\n    if reduction == 'none':\n        return loss\n    elif reduction == 'sum':\n        return loss.sum()\n    elif reduction == 'mean':\n        return loss.mean()\n</code></pre>\n<pre><code>def ousm(logits, targets, indices=None):\n    logits, targets = _check_input_type(logits, targets, 'CrossEntropy')\n    bs = logits.shape[0]\n    k = 5\n    if bs - k &gt; 0:\n      losses = soft_cross_entropy_loss(logits, targets)\n      if len(losses.shape) == 2:\n          losses = losses.mean(1)\n      _, idxs = losses.topk(bs-k, largest=False)\n      losses = losses.index_select(0, idxs)\n\n      return losses.mean()\n</code></pre>\n<p>thank you.</p>",
  "messages": [
    {
      "id": "1193509",
      "postDate": "02/09/2021 17:09:09",
      "content": "<p>if you read this paper <a href=\"https://arxiv.org/pdf/1901.07759.pdf\" target=\"_blank\">ROBUST LEARNING AT NOISY LABELED MEDICAL IMAGES:\nAPPLIED TO SKIN LESION CLASSIFICATION</a> , you will understand that OUSMLoss is a good choice for tackling label noise,<br>\nfrom this link,you can try OUSMLoss implementation : <a href=\"https://github.com/analokmaus/kaggle-panda-challenge-public/blob/208caf4c83a5ab9d181e66eee447cd2e475d58dc/models/noisy_loss.py\" target=\"_blank\">https://github.com/analokmaus/kaggle-panda-challenge-public/blob/208caf4c83a5ab9d181e66eee447cd2e475d58dc/models/noisy_loss.py</a><br>\nfor handling noisy labels issue, we tried but couldn't get better result than tempered loss, if you have tried it, can you please tell us if it works for your model or not? </p>\n<p>which loss function works best for you(you can keep it secret,if you don't want to share though) <strong>(asking for a friend)</strong></p>\n<p>if you simply want to use OUSMLoss  without  nn module implementation and with simple function implementation, then you can try the code below : </p>\n<pre><code>def _check_input_type(x, y, loss):\n    if loss in ['MSELoss', 'MAELoss', 'huber']:\n        if x.shape[-1] == 1:\n            return x.squeeze(), y.float()\n        else:\n            return x, y.float()\n    elif loss in ['CrossEntropy']:\n        return x, y.long()\n    else:\n        return x, y\n</code></pre>\n<pre><code>def soft_cross_entropy_loss(logits, targets, weights=1, reduction='none'):\n    if len(targets.shape) == 1 or targets.shape[1] == 1:\n        onehot_targets = torch.eye(logits.shape[1])[targets].to(logits.device)\n    else:\n        onehot_targets = targets\n    loss = -torch.sum(onehot_targets * F.log_softmax(logits, 1), 1)\n    if reduction == 'none':\n        return loss\n    elif reduction == 'sum':\n        return loss.sum()\n    elif reduction == 'mean':\n        return loss.mean()\n</code></pre>\n<pre><code>def ousm(logits, targets, indices=None):\n    logits, targets = _check_input_type(logits, targets, 'CrossEntropy')\n    bs = logits.shape[0]\n    k = 5\n    if bs - k &gt; 0:\n      losses = soft_cross_entropy_loss(logits, targets)\n      if len(losses.shape) == 2:\n          losses = losses.mean(1)\n      _, idxs = losses.topk(bs-k, largest=False)\n      losses = losses.index_select(0, idxs)\n\n      return losses.mean()\n</code></pre>\n<p>thank you.</p>",
      "rawMarkdown": "if you read this paper [ROBUST LEARNING AT NOISY LABELED MEDICAL IMAGES:\nAPPLIED TO SKIN LESION CLASSIFICATION](https://arxiv.org/pdf/1901.07759.pdf) , you will understand that OUSMLoss is a good choice for tackling label noise,\nfrom this link,you can try OUSMLoss implementation : https://github.com/analokmaus/kaggle-panda-challenge-public/blob/208caf4c83a5ab9d181e66eee447cd2e475d58dc/models/noisy_loss.py\nfor handling noisy labels issue, we tried but couldn't get better result than tempered loss, if you have tried it, can you please tell us if it works for your model or not? \n\nwhich loss function works best for you(you can keep it secret,if you don't want to share though) **(asking for a friend)**\n\nif you simply want to use OUSMLoss  without  nn module implementation and with simple function implementation, then you can try the code below : \n\n\n```\ndef _check_input_type(x, y, loss):\n    if loss in ['MSELoss', 'MAELoss', 'huber']:\n        if x.shape[-1] == 1:\n            return x.squeeze(), y.float()\n        else:\n            return x, y.float()\n    elif loss in ['CrossEntropy']:\n        return x, y.long()\n    else:\n        return x, y\n```\n\n```\ndef soft_cross_entropy_loss(logits, targets, weights=1, reduction='none'):\n    if len(targets.shape) == 1 or targets.shape[1] == 1:\n        onehot_targets = torch.eye(logits.shape[1])[targets].to(logits.device)\n    else:\n        onehot_targets = targets\n    loss = -torch.sum(onehot_targets * F.log_softmax(logits, 1), 1)\n    if reduction == 'none':\n        return loss\n    elif reduction == 'sum':\n        return loss.sum()\n    elif reduction == 'mean':\n        return loss.mean()\n```\n\n\n```\ndef ousm(logits, targets, indices=None):\n    logits, targets = _check_input_type(logits, targets, 'CrossEntropy')\n    bs = logits.shape[0]\n    k = 5\n    if bs - k > 0:\n      losses = soft_cross_entropy_loss(logits, targets)\n      if len(losses.shape) == 2:\n          losses = losses.mean(1)\n      _, idxs = losses.topk(bs-k, largest=False)\n      losses = losses.index_select(0, idxs)\n      \n      return losses.mean()\n```\n\nthank you.",
      "votes": null
    },
    {
      "id": "1194401",
      "postDate": "02/10/2021 07:22:06",
      "content": "<p>thx - interesting idea! …and no comments… I think people are tired because its so hard to achieve better results in this competition - like I am 😃</p>",
      "rawMarkdown": "thx - interesting idea! ...and no comments... I think people are tired because its so hard to achieve better results in this competition - like I am 😃",
      "votes": null
    },
    {
      "id": "1194410",
      "postDate": "02/10/2021 07:26:46",
      "content": "<p>IMHO i think it's  not so hard to get a decent public  lb score close to 0. 905 but it will definitely be very difficult to get a good private lb score</p>",
      "rawMarkdown": "IMHO i think it's  not so hard to get a decent public  lb score close to 0. 905 but it will definitely be very difficult to get a good private lb score",
      "votes": null
    },
    {
      "id": "1194464",
      "postDate": "02/10/2021 07:55:34",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> - Question --&gt; Why do you have largest = True in topk loss. The implementation says choose low uncertainty samples, so shouldn't it be largest=False. Am I misunderstanding something here?</p>",
      "rawMarkdown": "mobassir - Question --> Why do you have largest = True in topk loss. The implementation says choose low uncertainty samples, so shouldn't it be largest=False. Am I misunderstanding something here?",
      "votes": null
    },
    {
      "id": "1194467",
      "postDate": "02/10/2021 07:58:25",
      "content": "<p>for you my friend.. I am rather a beginner, constantly out of computing resources and this is my first pytorch project…</p>",
      "rawMarkdown": "for you my friend.. I am rather a beginner, constantly out of computing resources and this is my first pytorch project...",
      "votes": null
    },
    {
      "id": "1194508",
      "postDate": "02/10/2021 08:08:33",
      "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>  sorry, it will be largest = false,i have modified the code,thanks</p>",
      "rawMarkdown": "pheadrus  sorry, it will be largest = false,i have modified the code,thanks",
      "votes": null
    },
    {
      "id": "1194519",
      "postDate": "02/10/2021 08:11:45",
      "content": "<p><a href=\"https://www.kaggle.com/romanweilguny\" target=\"_blank\">@romanweilguny</a> you have kaggle gpu,tpu and colab as well,you have enough computing resource for this competition i believe,you need to use them efficiently,for example instead of tweaking hyperparameter for 5 fold trainining,try to add new things in your model and train that for few epochs and for single fold,if you see it goes well with cv only then go for 5 fold long training,,i see a lot of kaggler tweak simple hyperparameter and keep training over and over again,this is not the best way of utilizing time and limited resources imho,sorry if i am wrong</p>",
      "rawMarkdown": "romanweilguny you have kaggle gpu,tpu and colab as well,you have enough computing resource for this competition i believe,you need to use them efficiently,for example instead of tweaking hyperparameter for 5 fold trainining,try to add new things in your model and train that for few epochs and for single fold,if you see it goes well with cv only then go for 5 fold long training,,i see a lot of kaggler tweak simple hyperparameter and keep training over and over again,this is not the best way of utilizing time and limited resources imho,sorry if i am wrong",
      "votes": null
    },
    {
      "id": "1194768",
      "postDate": "02/10/2021 11:04:47",
      "content": "<p>The best loss function that worked best for me in this competition is <code>hair_loss</code>. The pytorch implementation is shown below:</p>\n<pre><code>def hair_loss(logits, targets, reduction='all'):\n    if reduction == 'all':\n        return 'You lost all your hair'\n</code></pre>",
      "rawMarkdown": "The best loss function that worked best for me in this competition is `hair_loss`. The pytorch implementation is shown below:\n```\ndef hair_loss(logits, targets, reduction='all'):\n    if reduction == 'all':\n        return 'You lost all your hair'\n```",
      "votes": null
    },
    {
      "id": "1194810",
      "postDate": "02/10/2021 11:35:11",
      "content": "<p>This is part  of my general loss function  . While the idea is theoretically sound and could help a bit to online denoising, I didn't notice any significant improvement using it (in comparison to other parts of my loss function)</p>\n<p>Anyway I think <code>k=5</code> is a bit too high, particularly if you use small batch size </p>",
      "rawMarkdown": "This is part  of my general loss function  . While the idea is theoretically sound and could help a bit to online denoising, I didn't notice any significant improvement using it (in comparison to other parts of my loss function)\n\nAnyway I think `k=5` is a bit too high, particularly if you use small batch size",
      "votes": null
    },
    {
      "id": "1194812",
      "postDate": "02/10/2021 11:36:22",
      "content": "<p>There are even some who tweak random seed  ^^</p>\n<p>FWIW I use colab only (and Kaggle just for inference)</p>",
      "rawMarkdown": "There are even some who tweak random seed  ^^\n\nFWIW I use colab only (and Kaggle just for inference)",
      "votes": null
    },
    {
      "id": "1194869",
      "postDate": "02/10/2021 12:19:35",
      "content": "<p>How much is your CV, LB increase when use your custom loss func?. Thank you!</p>",
      "rawMarkdown": "How much is your CV, LB increase when use your custom loss func?. Thank you!",
      "votes": null
    },
    {
      "id": "1194881",
      "postDate": "02/10/2021 12:36:19",
      "content": "<p>😂😂😂😂😂 new paper is coming 😆😆😆</p>",
      "rawMarkdown": "😂😂😂😂😂 new paper is coming 😆😆😆",
      "votes": null
    },
    {
      "id": "1194889",
      "postDate": "02/10/2021 12:47:40",
      "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> i do agree with you, we tried k =  1,2 and bigger batch size but nothing works for us, in pandas competition <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169230\" target=\"_blank\">6th place solution</a> team said that this loos function worked well for them,not sure why it is not working in this competition, is it only label noise or something else too that is causing trouble? during EDA i checked that for each class sample can contain many different looking images but they all belong to same class, it doesn't look like 5 class classification problem but &gt;5 class classification problem..sorry if i am wrong,,it can be another reason why models struggle to learn,,but i can't understand why large models are not working well in this competition</p>",
      "rawMarkdown": "serigne i do agree with you, we tried k =  1,2 and bigger batch size but nothing works for us, in pandas competition [6th place solution](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169230) team said that this loos function worked well for them,not sure why it is not working in this competition, is it only label noise or something else too that is causing trouble? during EDA i checked that for each class sample can contain many different looking images but they all belong to same class, it doesn't look like 5 class classification problem but >5 class classification problem..sorry if i am wrong,,it can be another reason why models struggle to learn,,but i can't understand why large models are not working well in this competition",
      "votes": null
    },
    {
      "id": "1194971",
      "postDate": "02/10/2021 13:49:08",
      "content": "<p>maybe there's one class here named \"unknown\". If you check the CropNet trained model in tensorflow, the model use six classes</p>",
      "rawMarkdown": "maybe there's one class here named \"unknown\". If you check the CropNet trained model in tensorflow, the model use six classes",
      "votes": null
    },
    {
      "id": "1194985",
      "postDate": "02/10/2021 13:52:00",
      "content": "<p><a href=\"https://www.kaggle.com/projdev\" target=\"_blank\">@projdev</a> which notebook you are talking about? can you share the link of that notebook?</p>",
      "rawMarkdown": "projdev which notebook you are talking about? can you share the link of that notebook?",
      "votes": null
    },
    {
      "id": "1195510",
      "postDate": "02/10/2021 21:13:42",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> thanks for sharing, this is the first time that I've ever read that paper, so, if I'm not wrong, you are weighting the samples just right here.</p>\n<pre><code> loss = -torch.sum(onehot_targets * F.log_softmax(logits, 1), 1)\n</code></pre>\n<p>And then dropout the noisiest samples (k samples) here.</p>\n<pre><code>_, idxs = losses.topk(bs-k, largest=False)\n</code></pre>\n<p>The method looks good, I'm going to try with then  right away. Thanks</p>",
      "rawMarkdown": "Hi @mobassir thanks for sharing, this is the first time that I've ever read that paper, so, if I'm not wrong, you are weighting the samples just right here.\n```\n loss = -torch.sum(onehot_targets * F.log_softmax(logits, 1), 1)\n```\nAnd then dropout the noisiest samples (k samples) here.\n```\n_, idxs = losses.topk(bs-k, largest=False)\n```\nThe method looks good, I'm going to try with then  right away. Thanks",
      "votes": null
    },
    {
      "id": "1195619",
      "postDate": "02/11/2021 00:36:09",
      "content": "<p>You may check CropNet trained model in tensorflow hub</p>",
      "rawMarkdown": "You may check CropNet trained model in tensorflow hub",
      "votes": null
    },
    {
      "id": "1196780",
      "postDate": "02/11/2021 16:22:36",
      "content": "<p>Pardon me for this stupid question, how do you manage long training sessions in colab? Wish there was a commit feature in colab.</p>",
      "rawMarkdown": "Pardon me for this stupid question, how do you manage long training sessions in colab? Wish there was a commit feature in colab.",
      "votes": null
    },
    {
      "id": "1196786",
      "postDate": "02/11/2021 16:28:21",
      "content": "<p>get colab pro and save models directly to your google drive</p>",
      "rawMarkdown": "get colab pro and save models directly to your google drive",
      "votes": null
    },
    {
      "id": "1196826",
      "postDate": "02/11/2021 17:24:30",
      "content": "<p>I use colab Pro </p>\n<p>V100 +AMP is about 3 or 4x times faster than kaggle P100<br>\nSessions can still run up to 24hours. </p>",
      "rawMarkdown": "I use colab Pro \n\nV100 +AMP is about 3 or 4x times faster than kaggle P100\nSessions can still run up to 24hours.",
      "votes": null
    },
    {
      "id": "1196901",
      "postDate": "02/11/2021 18:39:57",
      "content": "<p>Awesome 👍🔥👀😄</p>",
      "rawMarkdown": "Awesome 👍🔥👀😄",
      "votes": null
    },
    {
      "id": "1196910",
      "postDate": "02/11/2021 18:48:10",
      "content": "<p>It works so great! Finally my training and validation loss are perfectly correlated…</p>",
      "rawMarkdown": "It works so great! Finally my training and validation loss are perfectly correlated...",
      "votes": null
    },
    {
      "id": "1197037",
      "postDate": "02/11/2021 21:44:38",
      "content": "<p>colab pro is still only available in the us and canada</p>",
      "rawMarkdown": "colab pro is still only available in the us and canada",
      "votes": null
    },
    {
      "id": "1197044",
      "postDate": "02/11/2021 21:57:14",
      "content": "<p>VPN for registration is enough to get the account outside us..</p>",
      "rawMarkdown": "VPN for registration is enough to get the account outside us..",
      "votes": null
    },
    {
      "id": "1197049",
      "postDate": "02/11/2021 22:03:02",
      "content": "<p>just be creative, use gdrive synchronization and save all your model automatically on your gdrive, doesn't matter if this on gpu or TPU, is all the same. If the VM stop just restore your training, load the last weights and you are in business. If accelerator are restricted just use another account and sync with your gdrive from your main account. </p>",
      "rawMarkdown": "just be creative, use gdrive synchronization and save all your model automatically on your gdrive, doesn't matter if this on gpu or TPU, is all the same. If the VM stop just restore your training, load the last weights and you are in business. If accelerator are restricted just use another account and sync with your gdrive from your main account.",
      "votes": null
    },
    {
      "id": "1197070",
      "postDate": "02/11/2021 22:54:12",
      "content": "<p>I didn't use VPN to register at Colab Pro from outside US.  </p>\n<p>But it was at the first days. May be something had changed since then. </p>",
      "rawMarkdown": "I didn't use VPN to register at Colab Pro from outside US.  \n\nBut it was at the first days. May be something had changed since then.",
      "votes": null
    },
    {
      "id": "1198060",
      "postDate": "02/12/2021 17:13:20",
      "content": "<p>I've implemented on tensorflow models but the results did not improved as i concerned.</p>\n<pre><code>class OUSM(Loss): #inherit parent class\n    \"\"\"\n    If you want to read about it please check the link https://arxiv.org/pdf/1901.07759.pdf\n    \"\"\"\n    #initialize instance attributes\n    def __init__(self, label_smoothing = .4,):\n        super(OUSM, self).__init__( )\n        self.cce_Loss = tf.keras.losses.CategoricalCrossentropy(label_smoothing = label_smoothing,reduction=tf.keras.losses.Reduction.NONE)\n    #compute loss\n    def call(self, y_true, y_pred):\n        loss =  self.cce_Loss(y_true, y_pred)\n        _ , indxs = tf.raw_ops.TopKV2( input = loss,k=1)\n        value      = tf.subtract(tf.math.reduce_sum(loss), \n                                 tf.math.reduce_sum(tf.gather(loss, indxs)) \n                                 )\n        return  tf.reshape(value ,())   \n</code></pre>",
      "rawMarkdown": "I've implemented on tensorflow models but the results did not improved as i concerned.\n```\nclass OUSM(Loss): #inherit parent class\n    \"\"\"\n    If you want to read about it please check the link https://arxiv.org/pdf/1901.07759.pdf\n    \"\"\"\n    #initialize instance attributes\n    def __init__(self, label_smoothing = .4,):\n        super(OUSM, self).__init__( )\n        self.cce_Loss = tf.keras.losses.CategoricalCrossentropy(label_smoothing = label_smoothing,reduction=tf.keras.losses.Reduction.NONE)\n    #compute loss\n    def call(self, y_true, y_pred):\n        loss =  self.cce_Loss(y_true, y_pred)\n        _ , indxs = tf.raw_ops.TopKV2( input = loss,k=1)\n        value      = tf.subtract(tf.math.reduce_sum(loss), \n                                 tf.math.reduce_sum(tf.gather(loss, indxs)) \n                                 )\n        return  tf.reshape(value ,())   \n\n```",
      "votes": null
    },
    {
      "id": "1198070",
      "postDate": "02/12/2021 17:26:21",
      "content": "<p>Good work, hope you will be able to use it in another competition and get good result with it,thanks </p>",
      "rawMarkdown": "Good work, hope you will be able to use it in another competition and get good result with it,thanks",
      "votes": null
    },
    {
      "id": "1198097",
      "postDate": "02/12/2021 17:55:36",
      "content": "<p>Maybe this is not a good context of use but the idea looks interesting, more tools to face problems 👍👍. Anyway, Thanks for sharing this idea . </p>",
      "rawMarkdown": "Maybe this is not a good context of use but the idea looks interesting, more tools to face problems 👍👍. Anyway, Thanks for sharing this idea .",
      "votes": null
    },
    {
      "id": "1199369",
      "postDate": "02/13/2021 19:17:58",
      "content": "<p>Tried OUSMLoss on resnext50, got almost similar score (-.001) I guess it can be used to add variation… If time permits.<br>\nThanks for posting</p>",
      "rawMarkdown": "Tried OUSMLoss on resnext50, got almost similar score (-.001) I guess it can be used to add variation... If time permits.\nThanks for posting",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1194401,
      "author_name": "romanweilguny",
      "author_url": "",
      "post_date": "02/10/2021 07:22:06",
      "content": "<p>thx - interesting idea! …and no comments… I think people are tired because its so hard to achieve better results in this competition - like I am 😃</p>",
      "votes": null,
      "replies": [
        {
          "id": 1194410,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "02/10/2021 07:26:46",
          "content": "<p>IMHO i think it's  not so hard to get a decent public  lb score close to 0. 905 but it will definitely be very difficult to get a good private lb score</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1194467,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "02/10/2021 07:58:25",
          "content": "<p>for you my friend.. I am rather a beginner, constantly out of computing resources and this is my first pytorch project…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1194519,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "02/10/2021 08:11:45",
          "content": "<p><a href=\"https://www.kaggle.com/romanweilguny\" target=\"_blank\">@romanweilguny</a> you have kaggle gpu,tpu and colab as well,you have enough computing resource for this competition i believe,you need to use them efficiently,for example instead of tweaking hyperparameter for 5 fold trainining,try to add new things in your model and train that for few epochs and for single fold,if you see it goes well with cv only then go for 5 fold long training,,i see a lot of kaggler tweak simple hyperparameter and keep training over and over again,this is not the best way of utilizing time and limited resources imho,sorry if i am wrong</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1194812,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "02/10/2021 11:36:22",
          "content": "<p>There are even some who tweak random seed  ^^</p>\n<p>FWIW I use colab only (and Kaggle just for inference)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1196780,
          "author_name": "abhinand05",
          "author_url": "",
          "post_date": "02/11/2021 16:22:36",
          "content": "<p>Pardon me for this stupid question, how do you manage long training sessions in colab? Wish there was a commit feature in colab.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1196786,
              "author_name": "alexanderriedel",
              "author_url": "",
              "post_date": "02/11/2021 16:28:21",
              "content": "<p>get colab pro and save models directly to your google drive</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1196826,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "02/11/2021 17:24:30",
          "content": "<p>I use colab Pro </p>\n<p>V100 +AMP is about 3 or 4x times faster than kaggle P100<br>\nSessions can still run up to 24hours. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1197037,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "02/11/2021 21:44:38",
          "content": "<p>colab pro is still only available in the us and canada</p>",
          "votes": null,
          "replies": [
            {
              "id": 1197044,
              "author_name": "alexanderriedel",
              "author_url": "",
              "post_date": "02/11/2021 21:57:14",
              "content": "<p>VPN for registration is enough to get the account outside us..</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1197049,
          "author_name": "scarecrow2020",
          "author_url": "",
          "post_date": "02/11/2021 22:03:02",
          "content": "<p>just be creative, use gdrive synchronization and save all your model automatically on your gdrive, doesn't matter if this on gpu or TPU, is all the same. If the VM stop just restore your training, load the last weights and you are in business. If accelerator are restricted just use another account and sync with your gdrive from your main account. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1197070,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "02/11/2021 22:54:12",
          "content": "<p>I didn't use VPN to register at Colab Pro from outside US.  </p>\n<p>But it was at the first days. May be something had changed since then. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1194464,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "02/10/2021 07:55:34",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> - Question --&gt; Why do you have largest = True in topk loss. The implementation says choose low uncertainty samples, so shouldn't it be largest=False. Am I misunderstanding something here?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1194508,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "02/10/2021 08:08:33",
          "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a>  sorry, it will be largest = false,i have modified the code,thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1194768,
      "author_name": "underwearfitting",
      "author_url": "",
      "post_date": "02/10/2021 11:04:47",
      "content": "<p>The best loss function that worked best for me in this competition is <code>hair_loss</code>. The pytorch implementation is shown below:</p>\n<pre><code>def hair_loss(logits, targets, reduction='all'):\n    if reduction == 'all':\n        return 'You lost all your hair'\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1194881,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "02/10/2021 12:36:19",
          "content": "<p>😂😂😂😂😂 new paper is coming 😆😆😆</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1196901,
          "author_name": "saurabhbagchi",
          "author_url": "",
          "post_date": "02/11/2021 18:39:57",
          "content": "<p>Awesome 👍🔥👀😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1196910,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "02/11/2021 18:48:10",
          "content": "<p>It works so great! Finally my training and validation loss are perfectly correlated…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1194810,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "02/10/2021 11:35:11",
      "content": "<p>This is part  of my general loss function  . While the idea is theoretically sound and could help a bit to online denoising, I didn't notice any significant improvement using it (in comparison to other parts of my loss function)</p>\n<p>Anyway I think <code>k=5</code> is a bit too high, particularly if you use small batch size </p>",
      "votes": null,
      "replies": [
        {
          "id": 1194869,
          "author_name": "hungkhoi",
          "author_url": "",
          "post_date": "02/10/2021 12:19:35",
          "content": "<p>How much is your CV, LB increase when use your custom loss func?. Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1194889,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "02/10/2021 12:47:40",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> i do agree with you, we tried k =  1,2 and bigger batch size but nothing works for us, in pandas competition <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169230\" target=\"_blank\">6th place solution</a> team said that this loos function worked well for them,not sure why it is not working in this competition, is it only label noise or something else too that is causing trouble? during EDA i checked that for each class sample can contain many different looking images but they all belong to same class, it doesn't look like 5 class classification problem but &gt;5 class classification problem..sorry if i am wrong,,it can be another reason why models struggle to learn,,but i can't understand why large models are not working well in this competition</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1194971,
          "author_name": "projdev",
          "author_url": "",
          "post_date": "02/10/2021 13:49:08",
          "content": "<p>maybe there's one class here named \"unknown\". If you check the CropNet trained model in tensorflow, the model use six classes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1194985,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "02/10/2021 13:52:00",
          "content": "<p><a href=\"https://www.kaggle.com/projdev\" target=\"_blank\">@projdev</a> which notebook you are talking about? can you share the link of that notebook?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1195619,
          "author_name": "projdev",
          "author_url": "",
          "post_date": "02/11/2021 00:36:09",
          "content": "<p>You may check CropNet trained model in tensorflow hub</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1195510,
      "author_name": "scarecrow2020",
      "author_url": "",
      "post_date": "02/10/2021 21:13:42",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> thanks for sharing, this is the first time that I've ever read that paper, so, if I'm not wrong, you are weighting the samples just right here.</p>\n<pre><code> loss = -torch.sum(onehot_targets * F.log_softmax(logits, 1), 1)\n</code></pre>\n<p>And then dropout the noisiest samples (k samples) here.</p>\n<pre><code>_, idxs = losses.topk(bs-k, largest=False)\n</code></pre>\n<p>The method looks good, I'm going to try with then  right away. Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1198060,
      "author_name": "scarecrow2020",
      "author_url": "",
      "post_date": "02/12/2021 17:13:20",
      "content": "<p>I've implemented on tensorflow models but the results did not improved as i concerned.</p>\n<pre><code>class OUSM(Loss): #inherit parent class\n    \"\"\"\n    If you want to read about it please check the link https://arxiv.org/pdf/1901.07759.pdf\n    \"\"\"\n    #initialize instance attributes\n    def __init__(self, label_smoothing = .4,):\n        super(OUSM, self).__init__( )\n        self.cce_Loss = tf.keras.losses.CategoricalCrossentropy(label_smoothing = label_smoothing,reduction=tf.keras.losses.Reduction.NONE)\n    #compute loss\n    def call(self, y_true, y_pred):\n        loss =  self.cce_Loss(y_true, y_pred)\n        _ , indxs = tf.raw_ops.TopKV2( input = loss,k=1)\n        value      = tf.subtract(tf.math.reduce_sum(loss), \n                                 tf.math.reduce_sum(tf.gather(loss, indxs)) \n                                 )\n        return  tf.reshape(value ,())   \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1198070,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "02/12/2021 17:26:21",
          "content": "<p>Good work, hope you will be able to use it in another competition and get good result with it,thanks </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1198097,
          "author_name": "scarecrow2020",
          "author_url": "",
          "post_date": "02/12/2021 17:55:36",
          "content": "<p>Maybe this is not a good context of use but the idea looks interesting, more tools to face problems 👍👍. Anyway, Thanks for sharing this idea . </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1199369,
      "author_name": "albernard",
      "author_url": "",
      "post_date": "02/13/2021 19:17:58",
      "content": "<p>Tried OUSMLoss on resnext50, got almost similar score (-.001) I guess it can be used to add variation… If time permits.<br>\nThanks for posting</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1193509": "if you read this paper [ROBUST LEARNING AT NOISY LABELED MEDICAL IMAGES:\nAPPLIED TO SKIN LESION CLASSIFICATION](https://arxiv.org/pdf/1901.07759.pdf) , you will understand that OUSMLoss is a good choice for tackling label noise,\nfrom this link,you can try OUSMLoss implementation : https://github.com/analokmaus/kaggle-panda-challenge-public/blob/208caf4c83a5ab9d181e66eee447cd2e475d58dc/models/noisy_loss.py\nfor handling noisy labels issue, we tried but couldn't get better result than tempered loss, if you have tried it, can you please tell us if it works for your model or not? \n\nwhich loss function works best for you(you can keep it secret,if you don't want to share though) **(asking for a friend)**\n\nif you simply want to use OUSMLoss  without  nn module implementation and with simple function implementation, then you can try the code below : \n\n\n```\ndef _check_input_type(x, y, loss):\n    if loss in ['MSELoss', 'MAELoss', 'huber']:\n        if x.shape[-1] == 1:\n            return x.squeeze(), y.float()\n        else:\n            return x, y.float()\n    elif loss in ['CrossEntropy']:\n        return x, y.long()\n    else:\n        return x, y\n```\n\n```\ndef soft_cross_entropy_loss(logits, targets, weights=1, reduction='none'):\n    if len(targets.shape) == 1 or targets.shape[1] == 1:\n        onehot_targets = torch.eye(logits.shape[1])[targets].to(logits.device)\n    else:\n        onehot_targets = targets\n    loss = -torch.sum(onehot_targets * F.log_softmax(logits, 1), 1)\n    if reduction == 'none':\n        return loss\n    elif reduction == 'sum':\n        return loss.sum()\n    elif reduction == 'mean':\n        return loss.mean()\n```\n\n\n```\ndef ousm(logits, targets, indices=None):\n    logits, targets = _check_input_type(logits, targets, 'CrossEntropy')\n    bs = logits.shape[0]\n    k = 5\n    if bs - k > 0:\n      losses = soft_cross_entropy_loss(logits, targets)\n      if len(losses.shape) == 2:\n          losses = losses.mean(1)\n      _, idxs = losses.topk(bs-k, largest=False)\n      losses = losses.index_select(0, idxs)\n      \n      return losses.mean()\n```\n\nthank you.",
    "1194401": "thx - interesting idea! ...and no comments... I think people are tired because its so hard to achieve better results in this competition - like I am 😃",
    "1194410": "IMHO i think it's  not so hard to get a decent public  lb score close to 0. 905 but it will definitely be very difficult to get a good private lb score",
    "1194464": "mobassir - Question --> Why do you have largest = True in topk loss. The implementation says choose low uncertainty samples, so shouldn't it be largest=False. Am I misunderstanding something here?",
    "1194467": "for you my friend.. I am rather a beginner, constantly out of computing resources and this is my first pytorch project...",
    "1194508": "pheadrus  sorry, it will be largest = false,i have modified the code,thanks",
    "1194519": "romanweilguny you have kaggle gpu,tpu and colab as well,you have enough computing resource for this competition i believe,you need to use them efficiently,for example instead of tweaking hyperparameter for 5 fold trainining,try to add new things in your model and train that for few epochs and for single fold,if you see it goes well with cv only then go for 5 fold long training,,i see a lot of kaggler tweak simple hyperparameter and keep training over and over again,this is not the best way of utilizing time and limited resources imho,sorry if i am wrong",
    "1194768": "The best loss function that worked best for me in this competition is `hair_loss`. The pytorch implementation is shown below:\n```\ndef hair_loss(logits, targets, reduction='all'):\n    if reduction == 'all':\n        return 'You lost all your hair'\n```",
    "1194810": "This is part  of my general loss function  . While the idea is theoretically sound and could help a bit to online denoising, I didn't notice any significant improvement using it (in comparison to other parts of my loss function)\n\nAnyway I think `k=5` is a bit too high, particularly if you use small batch size",
    "1194812": "There are even some who tweak random seed  ^^\n\nFWIW I use colab only (and Kaggle just for inference)",
    "1194869": "How much is your CV, LB increase when use your custom loss func?. Thank you!",
    "1194881": "😂😂😂😂😂 new paper is coming 😆😆😆",
    "1194889": "serigne i do agree with you, we tried k =  1,2 and bigger batch size but nothing works for us, in pandas competition [6th place solution](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169230) team said that this loos function worked well for them,not sure why it is not working in this competition, is it only label noise or something else too that is causing trouble? during EDA i checked that for each class sample can contain many different looking images but they all belong to same class, it doesn't look like 5 class classification problem but >5 class classification problem..sorry if i am wrong,,it can be another reason why models struggle to learn,,but i can't understand why large models are not working well in this competition",
    "1194971": "maybe there's one class here named \"unknown\". If you check the CropNet trained model in tensorflow, the model use six classes",
    "1194985": "projdev which notebook you are talking about? can you share the link of that notebook?",
    "1195510": "Hi @mobassir thanks for sharing, this is the first time that I've ever read that paper, so, if I'm not wrong, you are weighting the samples just right here.\n```\n loss = -torch.sum(onehot_targets * F.log_softmax(logits, 1), 1)\n```\nAnd then dropout the noisiest samples (k samples) here.\n```\n_, idxs = losses.topk(bs-k, largest=False)\n```\nThe method looks good, I'm going to try with then  right away. Thanks",
    "1195619": "You may check CropNet trained model in tensorflow hub",
    "1196780": "Pardon me for this stupid question, how do you manage long training sessions in colab? Wish there was a commit feature in colab.",
    "1196786": "get colab pro and save models directly to your google drive",
    "1196826": "I use colab Pro \n\nV100 +AMP is about 3 or 4x times faster than kaggle P100\nSessions can still run up to 24hours.",
    "1196901": "Awesome 👍🔥👀😄",
    "1196910": "It works so great! Finally my training and validation loss are perfectly correlated...",
    "1197037": "colab pro is still only available in the us and canada",
    "1197044": "VPN for registration is enough to get the account outside us..",
    "1197049": "just be creative, use gdrive synchronization and save all your model automatically on your gdrive, doesn't matter if this on gpu or TPU, is all the same. If the VM stop just restore your training, load the last weights and you are in business. If accelerator are restricted just use another account and sync with your gdrive from your main account.",
    "1197070": "I didn't use VPN to register at Colab Pro from outside US.  \n\nBut it was at the first days. May be something had changed since then.",
    "1198060": "I've implemented on tensorflow models but the results did not improved as i concerned.\n```\nclass OUSM(Loss): #inherit parent class\n    \"\"\"\n    If you want to read about it please check the link https://arxiv.org/pdf/1901.07759.pdf\n    \"\"\"\n    #initialize instance attributes\n    def __init__(self, label_smoothing = .4,):\n        super(OUSM, self).__init__( )\n        self.cce_Loss = tf.keras.losses.CategoricalCrossentropy(label_smoothing = label_smoothing,reduction=tf.keras.losses.Reduction.NONE)\n    #compute loss\n    def call(self, y_true, y_pred):\n        loss =  self.cce_Loss(y_true, y_pred)\n        _ , indxs = tf.raw_ops.TopKV2( input = loss,k=1)\n        value      = tf.subtract(tf.math.reduce_sum(loss), \n                                 tf.math.reduce_sum(tf.gather(loss, indxs)) \n                                 )\n        return  tf.reshape(value ,())   \n\n```",
    "1198070": "Good work, hope you will be able to use it in another competition and get good result with it,thanks",
    "1198097": "Maybe this is not a good context of use but the idea looks interesting, more tools to face problems 👍👍. Anyway, Thanks for sharing this idea .",
    "1199369": "Tried OUSMLoss on resnext50, got almost similar score (-.001) I guess it can be used to add variation... If time permits.\nThanks for posting"
  },
  "source": "meta"
}