{
  "id": 135317,
  "title": "[pytorch] how to do OHEM?  (loss = loss_focal_ohem*a + loss_softmax*b + loss_triplet*c ？)",
  "url": "/competitions/bengaliai-cv19/discussion/135317",
  "author_name": "",
  "post_date": "2020-03-13T06:49:48.752295100Z",
  "votes": 13,
  "comment_count": 8,
  "views": 0,
  "content": "<p>[pytorch] ohem loss implementation:  <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a></p>\n\n<p>Here's what I thought:</p>\n\n<p>```\n            if epoch &lt; 30 :\n                rate_ohem = 1.0\n            elif epoch &lt; 60:\n                rate_ohem = 0.8\n            elif epoch &lt; 80:\n                rate_ohem = 0.75\n            elif epoch &lt; 100:\n                rate_ohem = 0.7\n             else:\n                rate_ohem = 0.6</p>\n\n<pre><code>         loss = cutmix_criterion(logit[0],logit[1],logit[2], targets, rate_ohem) \n</code></pre>\n\n<p>```</p>\n\n<p>```\ndef ohem_loss( rate, cls_pred, cls_target ):</p>\n\n<pre><code>batch_size = cls_pred.size(0) \nohem_cls_loss = F.cross_entropy(cls_pred, cls_target, reduction='none', ignore_index=-1)\n\nsorted_ohem_loss, idx = torch.sort(ohem_cls_loss, descending=True)\nkeep_num = min(sorted_ohem_loss.size()[0], int(batch_size*rate) )\nif keep_num &amp;lt; sorted_ohem_loss.size()[0]:\n    keep_idx_cuda = idx[:keep_num]\n    ohem_cls_loss = ohem_cls_loss[keep_idx_cuda]\ncls_loss = ohem_cls_loss.sum() / keep_num\nreturn cls_loss\n</code></pre>\n\n<p>def rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)</p>\n\n<pre><code># uniform\ncx = np.random.randint(W)\ncy = np.random.randint(H)\n\nbbx1 = np.clip(cx - cut_w // 2, 0, W)\nbby1 = np.clip(cy - cut_h // 2, 0, H)\nbbx2 = np.clip(cx + cut_w // 2, 0, W)\nbby2 = np.clip(cy + cut_h // 2, 0, H)\n\nreturn bbx1, bby1, bbx2, bby2\n</code></pre>\n\n<p>def cutmix(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]</p>\n\n<pre><code>lam = np.random.beta(alpha, alpha)\nbbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\ndata[:, :, bbx1:bbx2, bby1:bby2] = data[indices, :, bbx1:bbx2, bby1:bby2]\n# adjust lambda to exactly match pixel ratio\nlam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n\ntargets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\nreturn data, targets\n</code></pre>\n\n<h1>loss</h1>\n\n<p>def cutmix_criterion(preds1,preds2,preds3, targets, rate=0.7):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    # criterion = nn.CrossEntropyLoss(reduction='mean')\n    criterion = ohem_loss\n    return [ lam * criterion(rate, preds1, targets1) + (1 - lam) * criterion(rate, preds1, targets2), lam * criterion(rate, preds2, targets3) + (1 - lam) * criterion(rate, preds2, targets4), lam * criterion(rate, preds3, targets5) + (1 - lam) * criterion(rate, preds3, targets6) ]\n```</p>",
  "messages": [
    {
      "id": "770613",
      "postDate": "03/13/2020 06:49:48",
      "content": "<p>[pytorch] ohem loss implementation:  <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a></p>\n\n<p>Here's what I thought:</p>\n\n<p>```\n            if epoch &lt; 30 :\n                rate_ohem = 1.0\n            elif epoch &lt; 60:\n                rate_ohem = 0.8\n            elif epoch &lt; 80:\n                rate_ohem = 0.75\n            elif epoch &lt; 100:\n                rate_ohem = 0.7\n             else:\n                rate_ohem = 0.6</p>\n\n<pre><code>         loss = cutmix_criterion(logit[0],logit[1],logit[2], targets, rate_ohem) \n</code></pre>\n\n<p>```</p>\n\n<p>```\ndef ohem_loss( rate, cls_pred, cls_target ):</p>\n\n<pre><code>batch_size = cls_pred.size(0) \nohem_cls_loss = F.cross_entropy(cls_pred, cls_target, reduction='none', ignore_index=-1)\n\nsorted_ohem_loss, idx = torch.sort(ohem_cls_loss, descending=True)\nkeep_num = min(sorted_ohem_loss.size()[0], int(batch_size*rate) )\nif keep_num &amp;lt; sorted_ohem_loss.size()[0]:\n    keep_idx_cuda = idx[:keep_num]\n    ohem_cls_loss = ohem_cls_loss[keep_idx_cuda]\ncls_loss = ohem_cls_loss.sum() / keep_num\nreturn cls_loss\n</code></pre>\n\n<p>def rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)</p>\n\n<pre><code># uniform\ncx = np.random.randint(W)\ncy = np.random.randint(H)\n\nbbx1 = np.clip(cx - cut_w // 2, 0, W)\nbby1 = np.clip(cy - cut_h // 2, 0, H)\nbbx2 = np.clip(cx + cut_w // 2, 0, W)\nbby2 = np.clip(cy + cut_h // 2, 0, H)\n\nreturn bbx1, bby1, bbx2, bby2\n</code></pre>\n\n<p>def cutmix(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]</p>\n\n<pre><code>lam = np.random.beta(alpha, alpha)\nbbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\ndata[:, :, bbx1:bbx2, bby1:bby2] = data[indices, :, bbx1:bbx2, bby1:bby2]\n# adjust lambda to exactly match pixel ratio\nlam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n\ntargets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\nreturn data, targets\n</code></pre>\n\n<h1>loss</h1>\n\n<p>def cutmix_criterion(preds1,preds2,preds3, targets, rate=0.7):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    # criterion = nn.CrossEntropyLoss(reduction='mean')\n    criterion = ohem_loss\n    return [ lam * criterion(rate, preds1, targets1) + (1 - lam) * criterion(rate, preds1, targets2), lam * criterion(rate, preds2, targets3) + (1 - lam) * criterion(rate, preds2, targets4), lam * criterion(rate, preds3, targets5) + (1 - lam) * criterion(rate, preds3, targets6) ]\n```</p>",
      "rawMarkdown": "[pytorch] ohem loss implementation:  https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\n\nHere's what I thought:\n\n```\n            if epoch &lt; 30 :\n                rate_ohem = 1.0\n            elif epoch &lt; 60:\n                rate_ohem = 0.8\n            elif epoch &lt; 80:\n                rate_ohem = 0.75\n            elif epoch &lt; 100:\n                rate_ohem = 0.7\n             else:\n                rate_ohem = 0.6\n\n             loss = cutmix_criterion(logit[0],logit[1],logit[2], targets, rate_ohem) \n\n```\n\n\n```\ndef ohem_loss( rate, cls_pred, cls_target ):\n\n    batch_size = cls_pred.size(0) \n    ohem_cls_loss = F.cross_entropy(cls_pred, cls_target, reduction='none', ignore_index=-1)\n\n    sorted_ohem_loss, idx = torch.sort(ohem_cls_loss, descending=True)\n    keep_num = min(sorted_ohem_loss.size()[0], int(batch_size*rate) )\n    if keep_num &lt; sorted_ohem_loss.size()[0]:\n        keep_idx_cuda = idx[:keep_num]\n        ohem_cls_loss = ohem_cls_loss[keep_idx_cuda]\n    cls_loss = ohem_cls_loss.sum() / keep_num\n    return cls_loss\n\ndef rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)\n\n    # uniform\n    cx = np.random.randint(W)\n    cy = np.random.randint(H)\n\n    bbx1 = np.clip(cx - cut_w // 2, 0, W)\n    bby1 = np.clip(cy - cut_h // 2, 0, H)\n    bbx2 = np.clip(cx + cut_w // 2, 0, W)\n    bby2 = np.clip(cy + cut_h // 2, 0, H)\n\n    return bbx1, bby1, bbx2, bby2\ndef cutmix(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]\n\n    lam = np.random.beta(alpha, alpha)\n    bbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\n    data[:, :, bbx1:bbx2, bby1:bby2] = data[indices, :, bbx1:bbx2, bby1:bby2]\n    # adjust lambda to exactly match pixel ratio\n    lam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n\n    targets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n    return data, targets\n# loss \ndef cutmix_criterion(preds1,preds2,preds3, targets, rate=0.7):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    # criterion = nn.CrossEntropyLoss(reduction='mean')\n    criterion = ohem_loss\n    return [ lam * criterion(rate, preds1, targets1) + (1 - lam) * criterion(rate, preds1, targets2), lam * criterion(rate, preds2, targets3) + (1 - lam) * criterion(rate, preds2, targets4), lam * criterion(rate, preds3, targets5) + (1 - lam) * criterion(rate, preds3, targets6) ]\n```",
      "votes": null
    },
    {
      "id": "770624",
      "postDate": "03/13/2020 07:00:58",
      "content": "<p>or：\n<code>\nloss = loss_focal_ohem*a + loss_softmax*b + loss_triplet*c ？\n</code>\n[pytorch] focal loss + ohem implementation: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128665\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128665</a></p>",
      "rawMarkdown": "or：\n```\nloss = loss_focal_ohem*a + loss_softmax*b + loss_triplet*c ？\n```\n[pytorch] focal loss + ohem implementation: https://www.kaggle.com/c/bengaliai-cv19/discussion/128665",
      "votes": null
    },
    {
      "id": "770626",
      "postDate": "03/13/2020 07:01:02",
      "content": "<p>a question, since I used this loss in training phase, it is also necessary to apply it in the validation phase, am I correct? If yes, I'm encountering a \"nan\" loss value during validation. Did you also encounter this issue when using ohem loss?</p>",
      "rawMarkdown": "a question, since I used this loss in training phase, it is also necessary to apply it in the validation phase, am I correct? If yes, I'm encountering a \"nan\" loss value during validation. Did you also encounter this issue when using ohem loss?",
      "votes": null
    },
    {
      "id": "770628",
      "postDate": "03/13/2020 07:07:15",
      "content": "<p>sorry，I haven't encounter  this issue. you may refer to following tips：\n（1）check your code.\n（2）see here: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a> .</p>",
      "rawMarkdown": "sorry，I haven't encounter  this issue. you may refer to following tips：\n（1）check your code.\n（2）see here: https://www.kaggle.com/c/bengaliai-cv19/discussion/128637 .",
      "votes": null
    },
    {
      "id": "771188",
      "postDate": "03/13/2020 21:15:09",
      "content": "<p>I did something similar, except that I was starting with <code>rate = 0.8</code> and trained for <code>140</code> epochs. I decreased the rate by <code>0.1</code> every <code>20-25</code> epochs reaching <code>0.3</code> rate at epoch <code>120</code>. </p>\n\n<p>The results were worse than the exact same configuration but with cross-entropy. I guess my mistake was that I needed to start with <code>1.0</code> and train with it until I reach <code>0.97</code>. From that point I think OHEM can be beneficial.</p>",
      "rawMarkdown": "I did something similar, except that I was starting with `rate = 0.8` and trained for `140` epochs. I decreased the rate by `0.1` every `20-25` epochs reaching `0.3` rate at epoch `120`. \n\nThe results were worse than the exact same configuration but with cross-entropy. I guess my mistake was that I needed to start with `1.0` and train with it until I reach `0.97`. From that point I think OHEM can be beneficial.",
      "votes": null
    },
    {
      "id": "771379",
      "postDate": "03/14/2020 05:18:48",
      "content": "<p>Hmm! <a href=\"/machinelp\">@machinelp</a> Thanks for this post. I did similar experiments but with some difference. What i did was  <code>ohem_rate =pow(.2,current epoch number/total number of epochs)</code> Still model is training but it looks promising for me.</p>",
      "rawMarkdown": "Hmm! @machinelp Thanks for this post. I did similar experiments but with some difference. What i did was  `ohem_rate =pow(.2,current epoch number/total number of epochs)` Still model is training but it looks promising for me.",
      "votes": null
    },
    {
      "id": "771381",
      "postDate": "03/14/2020 05:20:55",
      "content": "<p>👍 </p>",
      "rawMarkdown": "👍",
      "votes": null
    },
    {
      "id": "774906",
      "postDate": "03/16/2020 03:42:54",
      "content": "<p>I went with <code>curr_ohem_rate*decay_rate every x epochs</code> until a set minimum, it produced worse loss (train and val) and cv but better lb. Though all metrics were all highly unstable once the rate dropped low.</p>\n\n<p>Given more time I could tweak it more but I've given up on it since I only started experimenting with mining a few days ago.</p>",
      "rawMarkdown": "I went with `curr_ohem_rate*decay_rate every x epochs` until a set minimum, it produced worse loss (train and val) and cv but better lb. Though all metrics were all highly unstable once the rate dropped low.\n\nGiven more time I could tweak it more but I've given up on it since I only started experimenting with mining a few days ago.",
      "votes": null
    },
    {
      "id": "775806",
      "postDate": "03/17/2020 01:13:56",
      "content": "<p>very good !</p>",
      "rawMarkdown": "very good !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 770624,
      "author_name": "machinelp",
      "author_url": "",
      "post_date": "03/13/2020 07:00:58",
      "content": "<p>or：\n<code>\nloss = loss_focal_ohem*a + loss_softmax*b + loss_triplet*c ？\n</code>\n[pytorch] focal loss + ohem implementation: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128665\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128665</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 770626,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "03/13/2020 07:01:02",
      "content": "<p>a question, since I used this loss in training phase, it is also necessary to apply it in the validation phase, am I correct? If yes, I'm encountering a \"nan\" loss value during validation. Did you also encounter this issue when using ohem loss?</p>",
      "votes": null,
      "replies": [
        {
          "id": 770628,
          "author_name": "machinelp",
          "author_url": "",
          "post_date": "03/13/2020 07:07:15",
          "content": "<p>sorry，I haven't encounter  this issue. you may refer to following tips：\n（1）check your code.\n（2）see here: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a> .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 771188,
      "author_name": "lightnezzofbeing",
      "author_url": "",
      "post_date": "03/13/2020 21:15:09",
      "content": "<p>I did something similar, except that I was starting with <code>rate = 0.8</code> and trained for <code>140</code> epochs. I decreased the rate by <code>0.1</code> every <code>20-25</code> epochs reaching <code>0.3</code> rate at epoch <code>120</code>. </p>\n\n<p>The results were worse than the exact same configuration but with cross-entropy. I guess my mistake was that I needed to start with <code>1.0</code> and train with it until I reach <code>0.97</code>. From that point I think OHEM can be beneficial.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 771379,
      "author_name": "ratan123",
      "author_url": "",
      "post_date": "03/14/2020 05:18:48",
      "content": "<p>Hmm! <a href=\"/machinelp\">@machinelp</a> Thanks for this post. I did similar experiments but with some difference. What i did was  <code>ohem_rate =pow(.2,current epoch number/total number of epochs)</code> Still model is training but it looks promising for me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 771381,
          "author_name": "machinelp",
          "author_url": "",
          "post_date": "03/14/2020 05:20:55",
          "content": "<p>👍 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 774906,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "03/16/2020 03:42:54",
          "content": "<p>I went with <code>curr_ohem_rate*decay_rate every x epochs</code> until a set minimum, it produced worse loss (train and val) and cv but better lb. Though all metrics were all highly unstable once the rate dropped low.</p>\n\n<p>Given more time I could tweak it more but I've given up on it since I only started experimenting with mining a few days ago.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 775806,
      "author_name": "likaiheng",
      "author_url": "",
      "post_date": "03/17/2020 01:13:56",
      "content": "<p>very good !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "770613": "[pytorch] ohem loss implementation:  https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\n\nHere's what I thought:\n\n```\n            if epoch &lt; 30 :\n                rate_ohem = 1.0\n            elif epoch &lt; 60:\n                rate_ohem = 0.8\n            elif epoch &lt; 80:\n                rate_ohem = 0.75\n            elif epoch &lt; 100:\n                rate_ohem = 0.7\n             else:\n                rate_ohem = 0.6\n\n             loss = cutmix_criterion(logit[0],logit[1],logit[2], targets, rate_ohem) \n\n```\n\n\n```\ndef ohem_loss( rate, cls_pred, cls_target ):\n\n    batch_size = cls_pred.size(0) \n    ohem_cls_loss = F.cross_entropy(cls_pred, cls_target, reduction='none', ignore_index=-1)\n\n    sorted_ohem_loss, idx = torch.sort(ohem_cls_loss, descending=True)\n    keep_num = min(sorted_ohem_loss.size()[0], int(batch_size*rate) )\n    if keep_num &lt; sorted_ohem_loss.size()[0]:\n        keep_idx_cuda = idx[:keep_num]\n        ohem_cls_loss = ohem_cls_loss[keep_idx_cuda]\n    cls_loss = ohem_cls_loss.sum() / keep_num\n    return cls_loss\n\ndef rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)\n\n    # uniform\n    cx = np.random.randint(W)\n    cy = np.random.randint(H)\n\n    bbx1 = np.clip(cx - cut_w // 2, 0, W)\n    bby1 = np.clip(cy - cut_h // 2, 0, H)\n    bbx2 = np.clip(cx + cut_w // 2, 0, W)\n    bby2 = np.clip(cy + cut_h // 2, 0, H)\n\n    return bbx1, bby1, bbx2, bby2\ndef cutmix(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]\n\n    lam = np.random.beta(alpha, alpha)\n    bbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\n    data[:, :, bbx1:bbx2, bby1:bby2] = data[indices, :, bbx1:bbx2, bby1:bby2]\n    # adjust lambda to exactly match pixel ratio\n    lam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n\n    targets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n    return data, targets\n# loss \ndef cutmix_criterion(preds1,preds2,preds3, targets, rate=0.7):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    # criterion = nn.CrossEntropyLoss(reduction='mean')\n    criterion = ohem_loss\n    return [ lam * criterion(rate, preds1, targets1) + (1 - lam) * criterion(rate, preds1, targets2), lam * criterion(rate, preds2, targets3) + (1 - lam) * criterion(rate, preds2, targets4), lam * criterion(rate, preds3, targets5) + (1 - lam) * criterion(rate, preds3, targets6) ]\n```",
    "770624": "or：\n```\nloss = loss_focal_ohem*a + loss_softmax*b + loss_triplet*c ？\n```\n[pytorch] focal loss + ohem implementation: https://www.kaggle.com/c/bengaliai-cv19/discussion/128665",
    "770626": "a question, since I used this loss in training phase, it is also necessary to apply it in the validation phase, am I correct? If yes, I'm encountering a \"nan\" loss value during validation. Did you also encounter this issue when using ohem loss?",
    "770628": "sorry，I haven't encounter  this issue. you may refer to following tips：\n（1）check your code.\n（2）see here: https://www.kaggle.com/c/bengaliai-cv19/discussion/128637 .",
    "771188": "I did something similar, except that I was starting with `rate = 0.8` and trained for `140` epochs. I decreased the rate by `0.1` every `20-25` epochs reaching `0.3` rate at epoch `120`. \n\nThe results were worse than the exact same configuration but with cross-entropy. I guess my mistake was that I needed to start with `1.0` and train with it until I reach `0.97`. From that point I think OHEM can be beneficial.",
    "771379": "Hmm! @machinelp Thanks for this post. I did similar experiments but with some difference. What i did was  `ohem_rate =pow(.2,current epoch number/total number of epochs)` Still model is training but it looks promising for me.",
    "771381": "👍",
    "774906": "I went with `curr_ohem_rate*decay_rate every x epochs` until a set minimum, it produced worse loss (train and val) and cv but better lb. Though all metrics were all highly unstable once the rate dropped low.\n\nGiven more time I could tweak it more but I've given up on it since I only started experimenting with mining a few days ago.",
    "775806": "very good !"
  },
  "source": "meta"
}