{
  "id": 198620,
  "title": "Fmix v.s. Cutmix Simple Explanation and Visualizations",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/198620",
  "author_name": "",
  "post_date": "2020-11-22T05:01:10.398416900Z",
  "votes": 40,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi Kagglers,</p>\n<p>In Bengali competition, a winning solution was to use FMix to do image augmentations.<br>\nIn this competition, I also applied Cutmix and Fmix to see how it actually looks like in our dataset.</p>\n<h3>Short Explanation</h3>\n<p><strong>Cutmix</strong> is an augmentation strategy that takes two images (say images with names A and B) and uses a sampled <em>rectangular</em> part from image B to replace the part of image A.</p>\n<p><strong>FMix</strong> is a variant of Cutmix, which uses a sampled <em>irregular</em> part from image B to replace the part of image A. What makes it irregularly-shaped is that it sampled the images from Fourier space to mix training examples.</p>\n<h3>More Details</h3>\n<p>For more visualizations and implementations details, you could check my kernel here: <a href=\"https://www.kaggle.com/khyeh0719/cutmix-v-s-fmix-with-visualization\" target=\"_blank\">https://www.kaggle.com/khyeh0719/cutmix-v-s-fmix-with-visualization</a></p>\n<p>Since I haven't found any existing fmix library in Kaggle dataset, and for those who would like to do training in submission kernel, I've uploaded the fmix library as a dataset here: <a href=\"https://www.kaggle.com/khyeh0719/image-fmix\" target=\"_blank\">https://www.kaggle.com/khyeh0719/image-fmix</a>. </p>\n<p>To use the library in Kaggle kernel, add this dataset to your kernel, and use the following simple codes to do the trick without an actual installation. You could also check my visualization kernel as well to see how to import the Fmix library.</p>\n<pre><code>package_path = '../input/image-fmix/FMix-master'\nimport sys; sys.path.append(package_path)\nfrom fmix import sample_mask\n\n# fmix API\ndef fmix(data, targets, alpha, decay_power, shape, max_soft=0.0, reformulate=False):\n    lam, mask = sample_mask(alpha, decay_power, shape, max_soft, reformulate)\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets = targets[indices]\n    x1 = torch.from_numpy(mask)*data\n    x2 = torch.from_numpy(1-mask)*shuffled_data\n    targets=(targets, shuffled_targets, lam)\n\n    return (x1+x2), targets\n</code></pre>",
  "messages": [
    {
      "id": "1086827",
      "postDate": "11/22/2020 05:01:10",
      "content": "<p>Hi Kagglers,</p>\n<p>In Bengali competition, a winning solution was to use FMix to do image augmentations.<br>\nIn this competition, I also applied Cutmix and Fmix to see how it actually looks like in our dataset.</p>\n<h3>Short Explanation</h3>\n<p><strong>Cutmix</strong> is an augmentation strategy that takes two images (say images with names A and B) and uses a sampled <em>rectangular</em> part from image B to replace the part of image A.</p>\n<p><strong>FMix</strong> is a variant of Cutmix, which uses a sampled <em>irregular</em> part from image B to replace the part of image A. What makes it irregularly-shaped is that it sampled the images from Fourier space to mix training examples.</p>\n<h3>More Details</h3>\n<p>For more visualizations and implementations details, you could check my kernel here: <a href=\"https://www.kaggle.com/khyeh0719/cutmix-v-s-fmix-with-visualization\" target=\"_blank\">https://www.kaggle.com/khyeh0719/cutmix-v-s-fmix-with-visualization</a></p>\n<p>Since I haven't found any existing fmix library in Kaggle dataset, and for those who would like to do training in submission kernel, I've uploaded the fmix library as a dataset here: <a href=\"https://www.kaggle.com/khyeh0719/image-fmix\" target=\"_blank\">https://www.kaggle.com/khyeh0719/image-fmix</a>. </p>\n<p>To use the library in Kaggle kernel, add this dataset to your kernel, and use the following simple codes to do the trick without an actual installation. You could also check my visualization kernel as well to see how to import the Fmix library.</p>\n<pre><code>package_path = '../input/image-fmix/FMix-master'\nimport sys; sys.path.append(package_path)\nfrom fmix import sample_mask\n\n# fmix API\ndef fmix(data, targets, alpha, decay_power, shape, max_soft=0.0, reformulate=False):\n    lam, mask = sample_mask(alpha, decay_power, shape, max_soft, reformulate)\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets = targets[indices]\n    x1 = torch.from_numpy(mask)*data\n    x2 = torch.from_numpy(1-mask)*shuffled_data\n    targets=(targets, shuffled_targets, lam)\n\n    return (x1+x2), targets\n</code></pre>",
      "rawMarkdown": "Hi Kagglers,\n\nIn Bengali competition, a winning solution was to use FMix to do image augmentations.\nIn this competition, I also applied Cutmix and Fmix to see how it actually looks like in our dataset.\n\n### Short Explanation \n**Cutmix** is an augmentation strategy that takes two images (say images with names A and B) and uses a sampled *rectangular* part from image B to replace the part of image A.\n\n**FMix** is a variant of Cutmix, which uses a sampled *irregular* part from image B to replace the part of image A. What makes it irregularly-shaped is that it sampled the images from Fourier space to mix training examples.\n\n### More Details\nFor more visualizations and implementations details, you could check my kernel here: https://www.kaggle.com/khyeh0719/cutmix-v-s-fmix-with-visualization\n\nSince I haven't found any existing fmix library in Kaggle dataset, and for those who would like to do training in submission kernel, I've uploaded the fmix library as a dataset here: https://www.kaggle.com/khyeh0719/image-fmix. \n\nTo use the library in Kaggle kernel, add this dataset to your kernel, and use the following simple codes to do the trick without an actual installation. You could also check my visualization kernel as well to see how to import the Fmix library.\n\n```\npackage_path = '../input/image-fmix/FMix-master'\nimport sys; sys.path.append(package_path)\nfrom fmix import sample_mask\n\n# fmix API\ndef fmix(data, targets, alpha, decay_power, shape, max_soft=0.0, reformulate=False):\n    lam, mask = sample_mask(alpha, decay_power, shape, max_soft, reformulate)\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets = targets[indices]\n    x1 = torch.from_numpy(mask)*data\n    x2 = torch.from_numpy(1-mask)*shuffled_data\n    targets=(targets, shuffled_targets, lam)\n    \n    return (x1+x2), targets\n```",
      "votes": null
    },
    {
      "id": "1088156",
      "postDate": "11/23/2020 11:37:42",
      "content": "<p>Hello! There is a question to explore… <br>\nI found that fmix seems to be a bit slow, about 2 times slower than cutmix. Is it a problem with the algorithm itself, the efficiency of my program, or a framework? Researching…Questioning…</p>",
      "rawMarkdown": "Hello! There is a question to explore... \nI found that fmix seems to be a bit slow, about 2 times slower than cutmix. Is it a problem with the algorithm itself, the efficiency of my program, or a framework? Researching...Questioning...",
      "votes": null
    },
    {
      "id": "1088167",
      "postDate": "11/23/2020 11:50:11",
      "content": "<p>Hi, in cutmix, we only need to select a random box and do image replacement. In fmix, we need to generate low freq image to determine the irregularly shaped masks, and then do image replacement. You could check here: <a href=\"https://github.com/ecs-vlc/FMix/blob/615f77c414311fe4b2f788419b1047730afa1142/fmix.py#L56\" target=\"_blank\">https://github.com/ecs-vlc/FMix/blob/615f77c414311fe4b2f788419b1047730afa1142/fmix.py#L56</a>. I guess most of time is due to the algorithms itself.</p>",
      "rawMarkdown": "Hi, in cutmix, we only need to select a random box and do image replacement. In fmix, we need to generate low freq image to determine the irregularly shaped masks, and then do image replacement. You could check here: https://github.com/ecs-vlc/FMix/blob/615f77c414311fe4b2f788419b1047730afa1142/fmix.py#L56. I guess most of time is due to the algorithms itself.",
      "votes": null
    },
    {
      "id": "1089229",
      "postDate": "11/24/2020 10:07:35",
      "content": "<p>hello.After using fmix enhancement, it seems that it is not much different from cutmix. Of course, I only did two experiments. Cutmix = &gt; 0.8, fmix = &gt; 0.793, guy！ How much has it improved?</p>",
      "rawMarkdown": "hello.After using fmix enhancement, it seems that it is not much different from cutmix. Of course, I only did two experiments. Cutmix = > 0.8, fmix = > 0.793, guy！ How much has it improved?",
      "votes": null
    },
    {
      "id": "1097062",
      "postDate": "12/01/2020 00:16:19",
      "content": "<p><a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> Question about the label once we apply cutmix . From the paper we have </p>\n<pre><code>1: for each iteration do\n2: input, target = get minibatch(dataset) . input is N×C×W×H size tensor, target is N×K size tensor.\n3: if mode == training then\n4: input s, target s = shuffle minibatch(input, target) . CutMix starts here.\n5: lambda = Unif(0,1)\n6: r x = Unif(0,W)\n7: r y = Unif(0,H)\n8: r w = Sqrt(1 - lambda)\n9: r h = Sqrt(1 - lambda)\n10: x1 = Round(Clip(r x - r w / 2, min=0))\n11: x2 = Round(Clip(r x + r w / 2, max=W))\n12: y1 = Round(Clip(r y - r h / 2, min=0))\n13: y2 = Round(Clip(r y + r h / 2, min=H))\n14: input[:, :, x1:x2, y1:y2] = input s[:, :, x1:x2, y1:y2]\n15: lambda = 1 - (x2-x1)*(y2-y1)/(W*H) . Adjust lambda to the exact area ratio.\n16: target = lambda * target + (1 - lambda) * target_s . CutMix ends.\n17: end if\n18: output = model forward(input)\n19: loss = compute loss(output, target)\n20: model update()\n21: end for\n</code></pre>\n<p>In step 16 , target might end up being a float . How are you handling this ?<br>\ne.g. target = [ 0 0 0 0 1 ] target_s = [0 1 0 0 0] , if lamda = 0.33 , your final target will be<br>\n[0 0.67 0 0 0.33]</p>\n<p>In this case do you just take the majority class and normalize it as  [0 1 0 0 0]</p>",
      "rawMarkdown": "khyeh0719 Question about the label once we apply cutmix . From the paper we have \n\n```\n1: for each iteration do\n2: input, target = get minibatch(dataset) . input is N×C×W×H size tensor, target is N×K size tensor.\n3: if mode == training then\n4: input s, target s = shuffle minibatch(input, target) . CutMix starts here.\n5: lambda = Unif(0,1)\n6: r x = Unif(0,W)\n7: r y = Unif(0,H)\n8: r w = Sqrt(1 - lambda)\n9: r h = Sqrt(1 - lambda)\n10: x1 = Round(Clip(r x - r w / 2, min=0))\n11: x2 = Round(Clip(r x + r w / 2, max=W))\n12: y1 = Round(Clip(r y - r h / 2, min=0))\n13: y2 = Round(Clip(r y + r h / 2, min=H))\n14: input[:, :, x1:x2, y1:y2] = input s[:, :, x1:x2, y1:y2]\n15: lambda = 1 - (x2-x1)*(y2-y1)/(W*H) . Adjust lambda to the exact area ratio.\n16: target = lambda * target + (1 - lambda) * target_s . CutMix ends.\n17: end if\n18: output = model forward(input)\n19: loss = compute loss(output, target)\n20: model update()\n21: end for\n```\n\nIn step 16 , target might end up being a float . How are you handling this ?\ne.g. target = [ 0 0 0 0 1 ] target_s = [0 1 0 0 0] , if lamda = 0.33 , your final target will be\n[0 0.67 0 0 0.33]\n\nIn this case do you just take the majority class and normalize it as  [0 1 0 0 0]",
      "votes": null
    },
    {
      "id": "1097661",
      "postDate": "12/01/2020 07:03:44",
      "content": "<p>yes, use soft labels for multiclass problems.<br>\nYou could modify the loss function to support this in multiclassification problem, as in my training kernel here: <a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\" target=\"_blank\">https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug</a></p>\n<p>reference: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173733\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173733</a></p>\n<pre><code>class MyCrossEntropyLoss(_WeightedLoss):\n    def __init__(self, weight=None, reduction='mean'):\n        super().__init__(weight=weight, reduction=reduction)\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n        lsm = F.log_softmax(inputs, -1)\n\n        if self.weight is not None:\n            lsm = lsm * self.weight.unsqueeze(0)\n\n        loss = -(targets * lsm).sum(-1)\n\n        if  self.reduction == 'sum':\n            loss = loss.sum()\n        elif  self.reduction == 'mean':\n            loss = loss.mean()\n\n        return loss\n</code></pre>",
      "rawMarkdown": "yes, use soft labels for multiclass problems.\nYou could modify the loss function to support this in multiclassification problem, as in my training kernel here: https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\n\nreference: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173733\n\n```\nclass MyCrossEntropyLoss(_WeightedLoss):\n    def __init__(self, weight=None, reduction='mean'):\n        super().__init__(weight=weight, reduction=reduction)\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n        lsm = F.log_softmax(inputs, -1)\n\n        if self.weight is not None:\n            lsm = lsm * self.weight.unsqueeze(0)\n\n        loss = -(targets * lsm).sum(-1)\n\n        if  self.reduction == 'sum':\n            loss = loss.sum()\n        elif  self.reduction == 'mean':\n            loss = loss.mean()\n\n        return loss\n```",
      "votes": null
    },
    {
      "id": "1097695",
      "postDate": "12/01/2020 07:30:51",
      "content": "<p>Thanks for the reply. Let me give it a try </p>",
      "rawMarkdown": "Thanks for the reply. Let me give it a try",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1088156,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "11/23/2020 11:37:42",
      "content": "<p>Hello! There is a question to explore… <br>\nI found that fmix seems to be a bit slow, about 2 times slower than cutmix. Is it a problem with the algorithm itself, the efficiency of my program, or a framework? Researching…Questioning…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1088167,
          "author_name": "khyeh0719",
          "author_url": "",
          "post_date": "11/23/2020 11:50:11",
          "content": "<p>Hi, in cutmix, we only need to select a random box and do image replacement. In fmix, we need to generate low freq image to determine the irregularly shaped masks, and then do image replacement. You could check here: <a href=\"https://github.com/ecs-vlc/FMix/blob/615f77c414311fe4b2f788419b1047730afa1142/fmix.py#L56\" target=\"_blank\">https://github.com/ecs-vlc/FMix/blob/615f77c414311fe4b2f788419b1047730afa1142/fmix.py#L56</a>. I guess most of time is due to the algorithms itself.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1089229,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "11/24/2020 10:07:35",
      "content": "<p>hello.After using fmix enhancement, it seems that it is not much different from cutmix. Of course, I only did two experiments. Cutmix = &gt; 0.8, fmix = &gt; 0.793, guy！ How much has it improved?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1097062,
      "author_name": "venkat555",
      "author_url": "",
      "post_date": "12/01/2020 00:16:19",
      "content": "<p><a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> Question about the label once we apply cutmix . From the paper we have </p>\n<pre><code>1: for each iteration do\n2: input, target = get minibatch(dataset) . input is N×C×W×H size tensor, target is N×K size tensor.\n3: if mode == training then\n4: input s, target s = shuffle minibatch(input, target) . CutMix starts here.\n5: lambda = Unif(0,1)\n6: r x = Unif(0,W)\n7: r y = Unif(0,H)\n8: r w = Sqrt(1 - lambda)\n9: r h = Sqrt(1 - lambda)\n10: x1 = Round(Clip(r x - r w / 2, min=0))\n11: x2 = Round(Clip(r x + r w / 2, max=W))\n12: y1 = Round(Clip(r y - r h / 2, min=0))\n13: y2 = Round(Clip(r y + r h / 2, min=H))\n14: input[:, :, x1:x2, y1:y2] = input s[:, :, x1:x2, y1:y2]\n15: lambda = 1 - (x2-x1)*(y2-y1)/(W*H) . Adjust lambda to the exact area ratio.\n16: target = lambda * target + (1 - lambda) * target_s . CutMix ends.\n17: end if\n18: output = model forward(input)\n19: loss = compute loss(output, target)\n20: model update()\n21: end for\n</code></pre>\n<p>In step 16 , target might end up being a float . How are you handling this ?<br>\ne.g. target = [ 0 0 0 0 1 ] target_s = [0 1 0 0 0] , if lamda = 0.33 , your final target will be<br>\n[0 0.67 0 0 0.33]</p>\n<p>In this case do you just take the majority class and normalize it as  [0 1 0 0 0]</p>",
      "votes": null,
      "replies": [
        {
          "id": 1097661,
          "author_name": "khyeh0719",
          "author_url": "",
          "post_date": "12/01/2020 07:03:44",
          "content": "<p>yes, use soft labels for multiclass problems.<br>\nYou could modify the loss function to support this in multiclassification problem, as in my training kernel here: <a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\" target=\"_blank\">https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug</a></p>\n<p>reference: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173733\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173733</a></p>\n<pre><code>class MyCrossEntropyLoss(_WeightedLoss):\n    def __init__(self, weight=None, reduction='mean'):\n        super().__init__(weight=weight, reduction=reduction)\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n        lsm = F.log_softmax(inputs, -1)\n\n        if self.weight is not None:\n            lsm = lsm * self.weight.unsqueeze(0)\n\n        loss = -(targets * lsm).sum(-1)\n\n        if  self.reduction == 'sum':\n            loss = loss.sum()\n        elif  self.reduction == 'mean':\n            loss = loss.mean()\n\n        return loss\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1097695,
          "author_name": "venkat555",
          "author_url": "",
          "post_date": "12/01/2020 07:30:51",
          "content": "<p>Thanks for the reply. Let me give it a try </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1086827": "Hi Kagglers,\n\nIn Bengali competition, a winning solution was to use FMix to do image augmentations.\nIn this competition, I also applied Cutmix and Fmix to see how it actually looks like in our dataset.\n\n### Short Explanation \n**Cutmix** is an augmentation strategy that takes two images (say images with names A and B) and uses a sampled *rectangular* part from image B to replace the part of image A.\n\n**FMix** is a variant of Cutmix, which uses a sampled *irregular* part from image B to replace the part of image A. What makes it irregularly-shaped is that it sampled the images from Fourier space to mix training examples.\n\n### More Details\nFor more visualizations and implementations details, you could check my kernel here: https://www.kaggle.com/khyeh0719/cutmix-v-s-fmix-with-visualization\n\nSince I haven't found any existing fmix library in Kaggle dataset, and for those who would like to do training in submission kernel, I've uploaded the fmix library as a dataset here: https://www.kaggle.com/khyeh0719/image-fmix. \n\nTo use the library in Kaggle kernel, add this dataset to your kernel, and use the following simple codes to do the trick without an actual installation. You could also check my visualization kernel as well to see how to import the Fmix library.\n\n```\npackage_path = '../input/image-fmix/FMix-master'\nimport sys; sys.path.append(package_path)\nfrom fmix import sample_mask\n\n# fmix API\ndef fmix(data, targets, alpha, decay_power, shape, max_soft=0.0, reformulate=False):\n    lam, mask = sample_mask(alpha, decay_power, shape, max_soft, reformulate)\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets = targets[indices]\n    x1 = torch.from_numpy(mask)*data\n    x2 = torch.from_numpy(1-mask)*shuffled_data\n    targets=(targets, shuffled_targets, lam)\n    \n    return (x1+x2), targets\n```",
    "1088156": "Hello! There is a question to explore... \nI found that fmix seems to be a bit slow, about 2 times slower than cutmix. Is it a problem with the algorithm itself, the efficiency of my program, or a framework? Researching...Questioning...",
    "1088167": "Hi, in cutmix, we only need to select a random box and do image replacement. In fmix, we need to generate low freq image to determine the irregularly shaped masks, and then do image replacement. You could check here: https://github.com/ecs-vlc/FMix/blob/615f77c414311fe4b2f788419b1047730afa1142/fmix.py#L56. I guess most of time is due to the algorithms itself.",
    "1089229": "hello.After using fmix enhancement, it seems that it is not much different from cutmix. Of course, I only did two experiments. Cutmix = > 0.8, fmix = > 0.793, guy！ How much has it improved?",
    "1097062": "khyeh0719 Question about the label once we apply cutmix . From the paper we have \n\n```\n1: for each iteration do\n2: input, target = get minibatch(dataset) . input is N×C×W×H size tensor, target is N×K size tensor.\n3: if mode == training then\n4: input s, target s = shuffle minibatch(input, target) . CutMix starts here.\n5: lambda = Unif(0,1)\n6: r x = Unif(0,W)\n7: r y = Unif(0,H)\n8: r w = Sqrt(1 - lambda)\n9: r h = Sqrt(1 - lambda)\n10: x1 = Round(Clip(r x - r w / 2, min=0))\n11: x2 = Round(Clip(r x + r w / 2, max=W))\n12: y1 = Round(Clip(r y - r h / 2, min=0))\n13: y2 = Round(Clip(r y + r h / 2, min=H))\n14: input[:, :, x1:x2, y1:y2] = input s[:, :, x1:x2, y1:y2]\n15: lambda = 1 - (x2-x1)*(y2-y1)/(W*H) . Adjust lambda to the exact area ratio.\n16: target = lambda * target + (1 - lambda) * target_s . CutMix ends.\n17: end if\n18: output = model forward(input)\n19: loss = compute loss(output, target)\n20: model update()\n21: end for\n```\n\nIn step 16 , target might end up being a float . How are you handling this ?\ne.g. target = [ 0 0 0 0 1 ] target_s = [0 1 0 0 0] , if lamda = 0.33 , your final target will be\n[0 0.67 0 0 0.33]\n\nIn this case do you just take the majority class and normalize it as  [0 1 0 0 0]",
    "1097661": "yes, use soft labels for multiclass problems.\nYou could modify the loss function to support this in multiclassification problem, as in my training kernel here: https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-train-amp-aug\n\nreference: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173733\n\n```\nclass MyCrossEntropyLoss(_WeightedLoss):\n    def __init__(self, weight=None, reduction='mean'):\n        super().__init__(weight=weight, reduction=reduction)\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets):\n        lsm = F.log_softmax(inputs, -1)\n\n        if self.weight is not None:\n            lsm = lsm * self.weight.unsqueeze(0)\n\n        loss = -(targets * lsm).sum(-1)\n\n        if  self.reduction == 'sum':\n            loss = loss.sum()\n        elif  self.reduction == 'mean':\n            loss = loss.mean()\n\n        return loss\n```",
    "1097695": "Thanks for the reply. Let me give it a try"
  },
  "source": "meta"
}