{
  "id": 132006,
  "title": "Cutmix-Mixup Callback in fast.ai",
  "url": "/competitions/bengaliai-cv19/discussion/132006",
  "author_name": "Nicholas Lyu",
  "post_date": "2020-02-23T11:23:07.455000",
  "votes": 16,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Finally through the bubble.\nHi! <a href=\"/iafoss\">@iafoss</a>  Iafoss's fast.ai kernel has been of great help to me. I noticed that there is no cutmix callback implemented. Here is my implementation of cutmix+mixup callback, also borrowing some good-quality code from <a href=\"/machinelp\">@machinelp</a> ; I arbitrarily called the callback cutmixup.</p>\n\n<p>```\nclass CutMixUpCallback(LearnerCallback):\n    \"Callback that creates the mixed-up input and target.\"\n    def <strong>init</strong>(self, learn:Learner, alpha:float=args['mixup_alpha'], cutmix_prob:float=0.5, cutmix_alpha=args['cutmix_alpha']):\n        super().<strong>init</strong>(learn)\n        self.alpha = alpha\n        self.cutmix_prob = cutmix_prob\n        self.cutmix_alpha = cutmix_alpha</p>\n\n<pre><code>def on_train_begin(self, **kwargs):\n    self.learn.loss_func = MixLoss(self.learn.loss_func)\n\ndef on_batch_begin(self, last_input, last_target, train, **kwargs):\n    if not train: \n        return\n\n    # Implement cutmix\n    if np.random.rand() &amp;lt; self.cutmix_prob:\n        B, W, H = last_input.size(0), last_input.size(1), last_input.size(2)\n        indices = torch.randperm(B)\n        Lam = []\n        last_input_ = last_input.clone()\n        for i in range(B):\n            lam = np.random.beta(self.cutmix_alpha, self.cutmix_alpha)\n            cut_rat = np.sqrt(1. - lam)\n            cut_w = np.int(W * cut_rat)\n            cut_h = np.int(H * cut_rat)\n\n            # uniform\n            cx = np.random.randint(W)\n            cy = np.random.randint(H)\n\n            bbx1 = np.clip(cx - cut_w // 2, 0, W)\n            bby1 = np.clip(cy - cut_h // 2, 0, H)\n            bbx2 = np.clip(cx + cut_w // 2, 0, W)\n            bby2 = np.clip(cy + cut_h // 2, 0, H)\n            last_input[i, bbx1:bbx2, bby1:bby2] = last_input_[indices[i], bbx1:bbx2, bby1:bby2]\n            # adjust lambda to exactly match pixel ratio\n            Lam.append(1 - (bbx2 - bbx1) * (bby2 - bby1) / (W * H))\n        Lam = torch.tensor(Lam).cuda()\n        new_input = last_input\n        new_target = torch.cat([last_target.float(), last_target[indices].float(), Lam[:,None].float()], 1)\n    else:\n        lambd = np.random.beta(self.alpha, self.alpha, last_target.size(0))\n        lambd = np.concatenate([lambd[:,None], 1-lambd[:,None]], 1).max(1)\n        lambd = last_input.new(lambd)\n        shuffle = torch.randperm(last_target.size(0)).to(last_input.device)\n        # No stack x, compute x\n        out_shape = [lambd.size(0)] + [1 for _ in range(len(last_input.shape) - 1)]\n        new_input = (last_input * lambd.view(out_shape) + last_input[shuffle] * (1-lambd).view(out_shape))\n        # Stack y\n        new_target = torch.cat([last_target.float(), last_target[shuffle].float(), lambd[:,None].float()], 1)\n    return {'last_input': new_input, 'last_target': new_target}  \n\ndef on_train_end(self, **kwargs):\n    self.learn.loss_func = self.learn.loss_func.get_old()\n</code></pre>\n\n<p>```</p>\n\n<p>And here is the corresponding MixLoss criterion</p>\n\n<p><code>``\nclass MixLoss(Module):\n    \"Adapt the loss function</code>crit` to go with mixup &amp; cutmix.\"\n    def <strong>init</strong>(self, crit, reduction='mean'):\n        super().<strong>init</strong>()\n        if hasattr(crit, 'reduction'): \n            self.crit = crit\n            self.old_red = crit.reduction\n            setattr(self.crit, 'reduction', 'none')\n        else: \n            self.crit = partial(crit, reduction='none')\n            self.old_crit = crit\n        self.reduction = reduction</p>\n\n<pre><code>def forward(self, output, target):\n    if len(target.shape) == 2 and target.shape[1] == 7:\n        loss1, loss2 = self.crit(output,target[:,0:3]), self.crit(output,target[:,3:6])\n        d = loss1 * target[:,-1] + loss2 * (1-target[:,-1])\n    else:  \n        d = self.crit(output, target)\n\n    if self.reduction == 'mean':    \n        return d.mean()\n    elif self.reduction == 'sum':   \n        return d.sum()\n    return d\n\ndef get_old(self):\n    if hasattr(self, 'old_crit'):  return self.old_crit\n    elif hasattr(self, 'old_red'): \n        setattr(self.crit, 'reduction', self.old_red)\n        return self.crit\n</code></pre>\n\n<p>```</p>\n\n<p>Using it in a training loop:\n<code>\nlr = args['lr']\nlearn.fit_one_cycle(args['total_epochs'], max_lr=lr, \n    pct_start=0.0, div_factor=100, callbacks = [logger, SaveModelCallback(learn,monitor='metric_tot',\n    mode='max',experiment=args['experiment_name']),CutMixUpCallback(learn)])\n</code></p>\n\n<p>My first time using fast.ai. Hope you find this useful :P</p>",
  "messages": [
    {
      "id": 754305,
      "postDate": "2020-02-23T11:23:07.457Z",
      "content": "<p>Finally through the bubble.\nHi! <a href=\"/iafoss\">@iafoss</a>  Iafoss's fast.ai kernel has been of great help to me. I noticed that there is no cutmix callback implemented. Here is my implementation of cutmix+mixup callback, also borrowing some good-quality code from <a href=\"/machinelp\">@machinelp</a> ; I arbitrarily called the callback cutmixup.</p>\n\n<p>```\nclass CutMixUpCallback(LearnerCallback):\n    \"Callback that creates the mixed-up input and target.\"\n    def <strong>init</strong>(self, learn:Learner, alpha:float=args['mixup_alpha'], cutmix_prob:float=0.5, cutmix_alpha=args['cutmix_alpha']):\n        super().<strong>init</strong>(learn)\n        self.alpha = alpha\n        self.cutmix_prob = cutmix_prob\n        self.cutmix_alpha = cutmix_alpha</p>\n\n<pre><code>def on_train_begin(self, **kwargs):\n    self.learn.loss_func = MixLoss(self.learn.loss_func)\n\ndef on_batch_begin(self, last_input, last_target, train, **kwargs):\n    if not train: \n        return\n\n    # Implement cutmix\n    if np.random.rand() &amp;lt; self.cutmix_prob:\n        B, W, H = last_input.size(0), last_input.size(1), last_input.size(2)\n        indices = torch.randperm(B)\n        Lam = []\n        last_input_ = last_input.clone()\n        for i in range(B):\n            lam = np.random.beta(self.cutmix_alpha, self.cutmix_alpha)\n            cut_rat = np.sqrt(1. - lam)\n            cut_w = np.int(W * cut_rat)\n            cut_h = np.int(H * cut_rat)\n\n            # uniform\n            cx = np.random.randint(W)\n            cy = np.random.randint(H)\n\n            bbx1 = np.clip(cx - cut_w // 2, 0, W)\n            bby1 = np.clip(cy - cut_h // 2, 0, H)\n            bbx2 = np.clip(cx + cut_w // 2, 0, W)\n            bby2 = np.clip(cy + cut_h // 2, 0, H)\n            last_input[i, bbx1:bbx2, bby1:bby2] = last_input_[indices[i], bbx1:bbx2, bby1:bby2]\n            # adjust lambda to exactly match pixel ratio\n            Lam.append(1 - (bbx2 - bbx1) * (bby2 - bby1) / (W * H))\n        Lam = torch.tensor(Lam).cuda()\n        new_input = last_input\n        new_target = torch.cat([last_target.float(), last_target[indices].float(), Lam[:,None].float()], 1)\n    else:\n        lambd = np.random.beta(self.alpha, self.alpha, last_target.size(0))\n        lambd = np.concatenate([lambd[:,None], 1-lambd[:,None]], 1).max(1)\n        lambd = last_input.new(lambd)\n        shuffle = torch.randperm(last_target.size(0)).to(last_input.device)\n        # No stack x, compute x\n        out_shape = [lambd.size(0)] + [1 for _ in range(len(last_input.shape) - 1)]\n        new_input = (last_input * lambd.view(out_shape) + last_input[shuffle] * (1-lambd).view(out_shape))\n        # Stack y\n        new_target = torch.cat([last_target.float(), last_target[shuffle].float(), lambd[:,None].float()], 1)\n    return {'last_input': new_input, 'last_target': new_target}  \n\ndef on_train_end(self, **kwargs):\n    self.learn.loss_func = self.learn.loss_func.get_old()\n</code></pre>\n\n<p>```</p>\n\n<p>And here is the corresponding MixLoss criterion</p>\n\n<p><code>``\nclass MixLoss(Module):\n    \"Adapt the loss function</code>crit` to go with mixup &amp; cutmix.\"\n    def <strong>init</strong>(self, crit, reduction='mean'):\n        super().<strong>init</strong>()\n        if hasattr(crit, 'reduction'): \n            self.crit = crit\n            self.old_red = crit.reduction\n            setattr(self.crit, 'reduction', 'none')\n        else: \n            self.crit = partial(crit, reduction='none')\n            self.old_crit = crit\n        self.reduction = reduction</p>\n\n<pre><code>def forward(self, output, target):\n    if len(target.shape) == 2 and target.shape[1] == 7:\n        loss1, loss2 = self.crit(output,target[:,0:3]), self.crit(output,target[:,3:6])\n        d = loss1 * target[:,-1] + loss2 * (1-target[:,-1])\n    else:  \n        d = self.crit(output, target)\n\n    if self.reduction == 'mean':    \n        return d.mean()\n    elif self.reduction == 'sum':   \n        return d.sum()\n    return d\n\ndef get_old(self):\n    if hasattr(self, 'old_crit'):  return self.old_crit\n    elif hasattr(self, 'old_red'): \n        setattr(self.crit, 'reduction', self.old_red)\n        return self.crit\n</code></pre>\n\n<p>```</p>\n\n<p>Using it in a training loop:\n<code>\nlr = args['lr']\nlearn.fit_one_cycle(args['total_epochs'], max_lr=lr, \n    pct_start=0.0, div_factor=100, callbacks = [logger, SaveModelCallback(learn,monitor='metric_tot',\n    mode='max',experiment=args['experiment_name']),CutMixUpCallback(learn)])\n</code></p>\n\n<p>My first time using fast.ai. Hope you find this useful :P</p>",
      "rawMarkdown": "Finally through the bubble.\nHi! @iafoss  Iafoss's fast.ai kernel has been of great help to me. I noticed that there is no cutmix callback implemented. Here is my implementation of cutmix+mixup callback, also borrowing some good-quality code from @machinelp ; I arbitrarily called the callback cutmixup.\n\n```\nclass CutMixUpCallback(LearnerCallback):\n    \"Callback that creates the mixed-up input and target.\"\n    def __init__(self, learn:Learner, alpha:float=args['mixup_alpha'], cutmix_prob:float=0.5, cutmix_alpha=args['cutmix_alpha']):\n        super().__init__(learn)\n        self.alpha = alpha\n        self.cutmix_prob = cutmix_prob\n        self.cutmix_alpha = cutmix_alpha\n    \n    def on_train_begin(self, **kwargs):\n        self.learn.loss_func = MixLoss(self.learn.loss_func)\n        \n    def on_batch_begin(self, last_input, last_target, train, **kwargs):\n        if not train: \n            return\n        \n        # Implement cutmix\n        if np.random.rand() &lt; self.cutmix_prob:\n            B, W, H = last_input.size(0), last_input.size(1), last_input.size(2)\n            indices = torch.randperm(B)\n            Lam = []\n            last_input_ = last_input.clone()\n            for i in range(B):\n                lam = np.random.beta(self.cutmix_alpha, self.cutmix_alpha)\n                cut_rat = np.sqrt(1. - lam)\n                cut_w = np.int(W * cut_rat)\n                cut_h = np.int(H * cut_rat)\n\n                # uniform\n                cx = np.random.randint(W)\n                cy = np.random.randint(H)\n\n                bbx1 = np.clip(cx - cut_w // 2, 0, W)\n                bby1 = np.clip(cy - cut_h // 2, 0, H)\n                bbx2 = np.clip(cx + cut_w // 2, 0, W)\n                bby2 = np.clip(cy + cut_h // 2, 0, H)\n                last_input[i, bbx1:bbx2, bby1:bby2] = last_input_[indices[i], bbx1:bbx2, bby1:bby2]\n                # adjust lambda to exactly match pixel ratio\n                Lam.append(1 - (bbx2 - bbx1) * (bby2 - bby1) / (W * H))\n            Lam = torch.tensor(Lam).cuda()\n            new_input = last_input\n            new_target = torch.cat([last_target.float(), last_target[indices].float(), Lam[:,None].float()], 1)\n        else:\n            lambd = np.random.beta(self.alpha, self.alpha, last_target.size(0))\n            lambd = np.concatenate([lambd[:,None], 1-lambd[:,None]], 1).max(1)\n            lambd = last_input.new(lambd)\n            shuffle = torch.randperm(last_target.size(0)).to(last_input.device)\n            # No stack x, compute x\n            out_shape = [lambd.size(0)] + [1 for _ in range(len(last_input.shape) - 1)]\n            new_input = (last_input * lambd.view(out_shape) + last_input[shuffle] * (1-lambd).view(out_shape))\n            # Stack y\n            new_target = torch.cat([last_target.float(), last_target[shuffle].float(), lambd[:,None].float()], 1)\n        return {'last_input': new_input, 'last_target': new_target}  \n    \n    def on_train_end(self, **kwargs):\n        self.learn.loss_func = self.learn.loss_func.get_old()\n```\n\nAnd here is the corresponding MixLoss criterion\n\n```\nclass MixLoss(Module):\n    \"Adapt the loss function `crit` to go with mixup &amp; cutmix.\"\n    def __init__(self, crit, reduction='mean'):\n        super().__init__()\n        if hasattr(crit, 'reduction'): \n            self.crit = crit\n            self.old_red = crit.reduction\n            setattr(self.crit, 'reduction', 'none')\n        else: \n            self.crit = partial(crit, reduction='none')\n            self.old_crit = crit\n        self.reduction = reduction\n        \n    def forward(self, output, target):\n        if len(target.shape) == 2 and target.shape[1] == 7:\n            loss1, loss2 = self.crit(output,target[:,0:3]), self.crit(output,target[:,3:6])\n            d = loss1 * target[:,-1] + loss2 * (1-target[:,-1])\n        else:  \n            d = self.crit(output, target)\n            \n        if self.reduction == 'mean':    \n            return d.mean()\n        elif self.reduction == 'sum':   \n            return d.sum()\n        return d\n    \n    def get_old(self):\n        if hasattr(self, 'old_crit'):  return self.old_crit\n        elif hasattr(self, 'old_red'): \n            setattr(self.crit, 'reduction', self.old_red)\n            return self.crit\n```\n\nUsing it in a training loop:\n```\nlr = args['lr']\nlearn.fit_one_cycle(args['total_epochs'], max_lr=lr, \n    pct_start=0.0, div_factor=100, callbacks = [logger, SaveModelCallback(learn,monitor='metric_tot',\n    mode='max',experiment=args['experiment_name']),CutMixUpCallback(learn)])\n```\n\nMy first time using fast.ai. Hope you find this useful :P",
      "votes": 16
    },
    {
      "id": 754374,
      "postDate": "2020-02-23T13:25:44.460Z",
      "content": "<p>👍 </p>",
      "rawMarkdown": "👍 ",
      "votes": 2
    },
    {
      "id": 754956,
      "postDate": "2020-02-24T09:15:56.337Z",
      "content": "<p>Thx 🙏 I found mistakes in my implementation. Gonna try my new version and then yours </p>",
      "rawMarkdown": "Thx 🙏 I found mistakes in my implementation. Gonna try my new version and then yours "
    },
    {
      "id": 754911,
      "postDate": "2020-02-24T07:55:31.763Z",
      "content": "<p>I actually tried something similar but was plagued with \"tensor of dimensions while concatenating\" error. I was not able to resolve it. Does this implementation work?</p>",
      "rawMarkdown": "I actually tried something similar but was plagued with \"tensor of dimensions while concatenating\" error. I was not able to resolve it. Does this implementation work?",
      "replies": [
        {
          "id": 755076,
          "postDate": "2020-02-24T12:03:30.820Z",
          "content": "<p><a href=\"/venky2506\">@venky2506</a> I am not sure I have encountered this error...For your information I am using one-channel images to feed into my networks, which corresponds to tensor size of BxHxW. Maybe this is the problem?</p>",
          "rawMarkdown": "@venky2506 I am not sure I have encountered this error...For your information I am using one-channel images to feed into my networks, which corresponds to tensor size of BxHxW. Maybe this is the problem?"
        },
        {
          "id": 755165,
          "postDate": "2020-02-24T14:08:06.750Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 755167,
          "postDate": "2020-02-24T14:10:33.170Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 755291,
          "postDate": "2020-02-24T16:27:26.060Z",
          "content": "<p>I am getting this error \"invalid argument 0: Tensors must have same number of dimensions: got 3 and 2 at /opt/conda/conda-bld/pytorch_1570711300255/work/aten/src/TH/generic/THTensor.cpp:680\" when I run </p>\n\n<p><code>learn.fit_one_cycle(25,max_lr=slice(1e-4,1e-7), wd=[1e-3,0.1e-1], pct_start=0.0, \n    div_factor=100, callbacks = [logger, SaveModelCallback(learn,every='improvement',name='best_weights_epoch126',monitor='metric_tot'),MixUpCallback(learn),EarlyStoppingCallback(learn,monitor='metric_tot',patience=10,min_delta=1e-4,),\n                                 ReduceLROnPlateauCallback(learn,monitor='metric_tot',min_delta=1e-4,min_lr=1e-12,patience=4)\n                                ,CutMixUpCallback(learn)])</code></p>\n\n<p>How do you resolve it?</p>",
          "rawMarkdown": "I am getting this error \"invalid argument 0: Tensors must have same number of dimensions: got 3 and 2 at /opt/conda/conda-bld/pytorch_1570711300255/work/aten/src/TH/generic/THTensor.cpp:680\" when I run \n\n`learn.fit_one_cycle(25,max_lr=slice(1e-4,1e-7), wd=[1e-3,0.1e-1], pct_start=0.0, \n    div_factor=100, callbacks = [logger, SaveModelCallback(learn,every='improvement',name='best_weights_epoch126',monitor='metric_tot'),MixUpCallback(learn),EarlyStoppingCallback(learn,monitor='metric_tot',patience=10,min_delta=1e-4,),\n                                 ReduceLROnPlateauCallback(learn,monitor='metric_tot',min_delta=1e-4,min_lr=1e-12,patience=4)\n                                ,CutMixUpCallback(learn)])`\n\nHow do you resolve it?"
        },
        {
          "id": 756283,
          "postDate": "2020-02-25T15:24:20.087Z",
          "content": "<p>You try to run learn with MixUpCallback and CutMixUpCallback. That's the first thing worth fixing.</p>",
          "rawMarkdown": "You try to run learn with MixUpCallback and CutMixUpCallback. That's the first thing worth fixing."
        },
        {
          "id": 756386,
          "postDate": "2020-02-25T17:16:06.803Z",
          "content": "<p><a href=\"/venky2506\">@venky2506</a>, CutMixUpCallback is all you need because that combine cutmix+mixup as stated in the main post by <a href=\"/roguekk007\">@roguekk007</a>. Remove MixUpCallback and it should work.</p>",
          "rawMarkdown": "@venky2506, CutMixUpCallback is all you need because that combine cutmix+mixup as stated in the main post by @roguekk007. Remove MixUpCallback and it should work."
        },
        {
          "id": 756397,
          "postDate": "2020-02-25T17:29:25.597Z",
          "content": "<p>😅 . Did not notice that the first time. Worked perfectly</p>",
          "rawMarkdown": "😅 . Did not notice that the first time. Worked perfectly"
        },
        {
          "id": 756399,
          "postDate": "2020-02-25T17:30:58.700Z",
          "content": "<p>Why does it throw this error btw?</p>",
          "rawMarkdown": "Why does it throw this error btw?"
        },
        {
          "id": 756707,
          "postDate": "2020-02-26T01:32:54.510Z",
          "content": "<p><a href=\"/venky2506\">@venky2506</a> My guess is that you are feeding 3-channel images during dataloading. The input for my model was BxWxH, not Bx3xWxH, you can check this. Otherwise, I honestly have no idea</p>",
          "rawMarkdown": "@venky2506 My guess is that you are feeding 3-channel images during dataloading. The input for my model was BxWxH, not Bx3xWxH, you can check this. Otherwise, I honestly have no idea"
        }
      ]
    },
    {
      "id": 754758,
      "postDate": "2020-02-24T02:47:29.910Z",
      "content": "<p>It's impressive! What's more, is fast.ai belong to automated ml tools(AMLT)?</p>",
      "rawMarkdown": "It's impressive! What's more, is fast.ai belong to automated ml tools(AMLT)?",
      "replies": [
        {
          "id": 754891,
          "postDate": "2020-02-24T07:19:02.793Z",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> Unfortunately, I don't think so. Trying to squeeze in gold-zone first before using automatedML :P</p>",
          "rawMarkdown": "@cnzengshiyuan Unfortunately, I don't think so. Trying to squeeze in gold-zone first before using automatedML :P",
          "votes": 1
        },
        {
          "id": 754899,
          "postDate": "2020-02-24T07:30:57.183Z",
          "content": "<p>Thank you for your reply! I didn't try fast.ai as thinking it might be amlt 😂 <br>\nBut now I think I can set about it(fast.ai) and try something new :)😄  Thanks lol</p>",
          "rawMarkdown": "Thank you for your reply! I didn't try fast.ai as thinking it might be amlt 😂   \nBut now I think I can set about it(fast.ai) and try something new :)😄  Thanks lol"
        }
      ]
    },
    {
      "id": 754460,
      "postDate": "2020-02-23T15:19:42.183Z",
      "content": "<p>very nice =) I have slightly different implementation =) Gonna try yours and see if it improves =) </p>",
      "rawMarkdown": "very nice =) I have slightly different implementation =) Gonna try yours and see if it improves =) "
    }
  ],
  "comments": [
    {
      "id": 754374,
      "author_name": "MachineLP",
      "author_url": "",
      "post_date": "2020-02-23T13:25:44.460000",
      "content": "<p>👍 </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 754956,
      "author_name": "Mishunyayev Nikita",
      "author_url": "",
      "post_date": "2020-02-24T09:15:56.337000",
      "content": "<p>Thx 🙏 I found mistakes in my implementation. Gonna try my new version and then yours </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 754911,
      "author_name": "Son of Anton v3.0",
      "author_url": "",
      "post_date": "2020-02-24T07:55:31.763000",
      "content": "<p>I actually tried something similar but was plagued with \"tensor of dimensions while concatenating\" error. I was not able to resolve it. Does this implementation work?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 755076,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-02-24T12:03:30.820000",
          "content": "<p><a href=\"/venky2506\">@venky2506</a> I am not sure I have encountered this error...For your information I am using one-channel images to feed into my networks, which corresponds to tensor size of BxHxW. Maybe this is the problem?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755165,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-24T14:08:06.750000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755167,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-24T14:10:33.170000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755291,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-02-24T16:27:26.060000",
          "content": "<p>I am getting this error \"invalid argument 0: Tensors must have same number of dimensions: got 3 and 2 at /opt/conda/conda-bld/pytorch_1570711300255/work/aten/src/TH/generic/THTensor.cpp:680\" when I run </p>\n\n<p><code>learn.fit_one_cycle(25,max_lr=slice(1e-4,1e-7), wd=[1e-3,0.1e-1], pct_start=0.0, \n    div_factor=100, callbacks = [logger, SaveModelCallback(learn,every='improvement',name='best_weights_epoch126',monitor='metric_tot'),MixUpCallback(learn),EarlyStoppingCallback(learn,monitor='metric_tot',patience=10,min_delta=1e-4,),\n                                 ReduceLROnPlateauCallback(learn,monitor='metric_tot',min_delta=1e-4,min_lr=1e-12,patience=4)\n                                ,CutMixUpCallback(learn)])</code></p>\n\n<p>How do you resolve it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756283,
          "author_name": "Mishunyayev Nikita",
          "author_url": "",
          "post_date": "2020-02-25T15:24:20.087000",
          "content": "<p>You try to run learn with MixUpCallback and CutMixUpCallback. That's the first thing worth fixing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756386,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-02-25T17:16:06.803000",
          "content": "<p><a href=\"/venky2506\">@venky2506</a>, CutMixUpCallback is all you need because that combine cutmix+mixup as stated in the main post by <a href=\"/roguekk007\">@roguekk007</a>. Remove MixUpCallback and it should work.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756397,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-02-25T17:29:25.597000",
          "content": "<p>😅 . Did not notice that the first time. Worked perfectly</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756399,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-02-25T17:30:58.700000",
          "content": "<p>Why does it throw this error btw?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756707,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-02-26T01:32:54.510000",
          "content": "<p><a href=\"/venky2506\">@venky2506</a> My guess is that you are feeding 3-channel images during dataloading. The input for my model was BxWxH, not Bx3xWxH, you can check this. Otherwise, I honestly have no idea</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 754758,
      "author_name": "Shiyuan Zeng",
      "author_url": "",
      "post_date": "2020-02-24T02:47:29.910000",
      "content": "<p>It's impressive! What's more, is fast.ai belong to automated ml tools(AMLT)?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 754891,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-02-24T07:19:02.793000",
          "content": "<p><a href=\"/cnzengshiyuan\">@cnzengshiyuan</a> Unfortunately, I don't think so. Trying to squeeze in gold-zone first before using automatedML :P</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 754899,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-02-24T07:30:57.183000",
          "content": "<p>Thank you for your reply! I didn't try fast.ai as thinking it might be amlt 😂 <br>\nBut now I think I can set about it(fast.ai) and try something new :)😄  Thanks lol</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 754460,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2020-02-23T15:19:42.183000",
      "content": "<p>very nice =) I have slightly different implementation =) Gonna try yours and see if it improves =) </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "754305": "Finally through the bubble.\nHi! @iafoss  Iafoss's fast.ai kernel has been of great help to me. I noticed that there is no cutmix callback implemented. Here is my implementation of cutmix+mixup callback, also borrowing some good-quality code from @machinelp ; I arbitrarily called the callback cutmixup.\n\n```\nclass CutMixUpCallback(LearnerCallback):\n    \"Callback that creates the mixed-up input and target.\"\n    def __init__(self, learn:Learner, alpha:float=args['mixup_alpha'], cutmix_prob:float=0.5, cutmix_alpha=args['cutmix_alpha']):\n        super().__init__(learn)\n        self.alpha = alpha\n        self.cutmix_prob = cutmix_prob\n        self.cutmix_alpha = cutmix_alpha\n    \n    def on_train_begin(self, **kwargs):\n        self.learn.loss_func = MixLoss(self.learn.loss_func)\n        \n    def on_batch_begin(self, last_input, last_target, train, **kwargs):\n        if not train: \n            return\n        \n        # Implement cutmix\n        if np.random.rand() &lt; self.cutmix_prob:\n            B, W, H = last_input.size(0), last_input.size(1), last_input.size(2)\n            indices = torch.randperm(B)\n            Lam = []\n            last_input_ = last_input.clone()\n            for i in range(B):\n                lam = np.random.beta(self.cutmix_alpha, self.cutmix_alpha)\n                cut_rat = np.sqrt(1. - lam)\n                cut_w = np.int(W * cut_rat)\n                cut_h = np.int(H * cut_rat)\n\n                # uniform\n                cx = np.random.randint(W)\n                cy = np.random.randint(H)\n\n                bbx1 = np.clip(cx - cut_w // 2, 0, W)\n                bby1 = np.clip(cy - cut_h // 2, 0, H)\n                bbx2 = np.clip(cx + cut_w // 2, 0, W)\n                bby2 = np.clip(cy + cut_h // 2, 0, H)\n                last_input[i, bbx1:bbx2, bby1:bby2] = last_input_[indices[i], bbx1:bbx2, bby1:bby2]\n                # adjust lambda to exactly match pixel ratio\n                Lam.append(1 - (bbx2 - bbx1) * (bby2 - bby1) / (W * H))\n            Lam = torch.tensor(Lam).cuda()\n            new_input = last_input\n            new_target = torch.cat([last_target.float(), last_target[indices].float(), Lam[:,None].float()], 1)\n        else:\n            lambd = np.random.beta(self.alpha, self.alpha, last_target.size(0))\n            lambd = np.concatenate([lambd[:,None], 1-lambd[:,None]], 1).max(1)\n            lambd = last_input.new(lambd)\n            shuffle = torch.randperm(last_target.size(0)).to(last_input.device)\n            # No stack x, compute x\n            out_shape = [lambd.size(0)] + [1 for _ in range(len(last_input.shape) - 1)]\n            new_input = (last_input * lambd.view(out_shape) + last_input[shuffle] * (1-lambd).view(out_shape))\n            # Stack y\n            new_target = torch.cat([last_target.float(), last_target[shuffle].float(), lambd[:,None].float()], 1)\n        return {'last_input': new_input, 'last_target': new_target}  \n    \n    def on_train_end(self, **kwargs):\n        self.learn.loss_func = self.learn.loss_func.get_old()\n```\n\nAnd here is the corresponding MixLoss criterion\n\n```\nclass MixLoss(Module):\n    \"Adapt the loss function `crit` to go with mixup &amp; cutmix.\"\n    def __init__(self, crit, reduction='mean'):\n        super().__init__()\n        if hasattr(crit, 'reduction'): \n            self.crit = crit\n            self.old_red = crit.reduction\n            setattr(self.crit, 'reduction', 'none')\n        else: \n            self.crit = partial(crit, reduction='none')\n            self.old_crit = crit\n        self.reduction = reduction\n        \n    def forward(self, output, target):\n        if len(target.shape) == 2 and target.shape[1] == 7:\n            loss1, loss2 = self.crit(output,target[:,0:3]), self.crit(output,target[:,3:6])\n            d = loss1 * target[:,-1] + loss2 * (1-target[:,-1])\n        else:  \n            d = self.crit(output, target)\n            \n        if self.reduction == 'mean':    \n            return d.mean()\n        elif self.reduction == 'sum':   \n            return d.sum()\n        return d\n    \n    def get_old(self):\n        if hasattr(self, 'old_crit'):  return self.old_crit\n        elif hasattr(self, 'old_red'): \n            setattr(self.crit, 'reduction', self.old_red)\n            return self.crit\n```\n\nUsing it in a training loop:\n```\nlr = args['lr']\nlearn.fit_one_cycle(args['total_epochs'], max_lr=lr, \n    pct_start=0.0, div_factor=100, callbacks = [logger, SaveModelCallback(learn,monitor='metric_tot',\n    mode='max',experiment=args['experiment_name']),CutMixUpCallback(learn)])\n```\n\nMy first time using fast.ai. Hope you find this useful :P",
    "754374": "👍 ",
    "754956": "Thx 🙏 I found mistakes in my implementation. Gonna try my new version and then yours ",
    "754911": "I actually tried something similar but was plagued with \"tensor of dimensions while concatenating\" error. I was not able to resolve it. Does this implementation work?",
    "754758": "It's impressive! What's more, is fast.ai belong to automated ml tools(AMLT)?",
    "754460": "very nice =) I have slightly different implementation =) Gonna try yours and see if it improves =) "
  }
}