{
  "id": 37738,
  "title": "Small minibatches? Consider batch renormalization.",
  "url": "/competitions/carvana-image-masking-challenge/discussion/37738",
  "author_name": "",
  "post_date": "2017-08-08T12:09:20.481599900Z",
  "votes": 7,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Batch normalization works great on moderately sized minibatches, but its performance suffers on small minibatches. In the <a href=\"https://arxiv.org/abs/1702.03275\">batch renormalization paper</a>, they showed that batch <strong>normalization</strong> achieved 78.3% on Inception-v3 with batch size 32, but only 74.2% with batch size 4. In contrast, batch <strong>renormalization</strong> achieved 76.5% with batch size 4.</p>\n\n<p>This is relevant to this competition because the high image resolution (1918x1280) means you may only be able to fit a few full size images in your video card's memory, thus limiting the max batch size.</p>\n\n<p><img src=\"https://i.imgur.com/dH4mPCO.png\" alt=\"figures 1 and 2 from the paper\" title=\"\"></p>",
  "messages": [
    {
      "id": "211240",
      "postDate": "08/08/2017 12:09:20",
      "content": "<p>Batch normalization works great on moderately sized minibatches, but its performance suffers on small minibatches. In the <a href=\"https://arxiv.org/abs/1702.03275\">batch renormalization paper</a>, they showed that batch <strong>normalization</strong> achieved 78.3% on Inception-v3 with batch size 32, but only 74.2% with batch size 4. In contrast, batch <strong>renormalization</strong> achieved 76.5% with batch size 4.</p>\n\n<p>This is relevant to this competition because the high image resolution (1918x1280) means you may only be able to fit a few full size images in your video card's memory, thus limiting the max batch size.</p>\n\n<p><img src=\"https://i.imgur.com/dH4mPCO.png\" alt=\"figures 1 and 2 from the paper\" title=\"\"></p>",
      "rawMarkdown": "Batch normalization works great on moderately sized minibatches, but its performance suffers on small minibatches. In the [batch renormalization paper](https://arxiv.org/abs/1702.03275), they showed that batch **normalization** achieved 78.3% on Inception-v3 with batch size 32, but only 74.2% with batch size 4. In contrast, batch **renormalization** achieved 76.5% with batch size 4.\n\nThis is relevant to this competition because the high image resolution (1918x1280) means you may only be able to fit a few full size images in your video card's memory, thus limiting the max batch size.\n\n![figures 1 and 2 from the paper](https://i.imgur.com/dH4mPCO.png)",
      "votes": null
    },
    {
      "id": "211274",
      "postDate": "08/08/2017 14:08:26",
      "content": "<p>another possible alternative:\nWeight Normalization from <a href=\"https://arxiv.org/abs/1602.07868\">https://arxiv.org/abs/1602.07868</a></p>\n\n<p><a href=\"https://github.com/pytorch/pytorch/blob/master/torch/nn/utils/weight_norm.py\">https://github.com/pytorch/pytorch/blob/master/torch/nn/utils/weight_norm.py</a></p>\n\n<p><a href=\"http://forums.fast.ai/t/pytorch-weight-norm/4619\">http://forums.fast.ai/t/pytorch-weight-norm/4619</a></p>",
      "rawMarkdown": "another possible alternative:\nWeight Normalization from https://arxiv.org/abs/1602.07868\n\n\nhttps://github.com/pytorch/pytorch/blob/master/torch/nn/utils/weight_norm.py\n\nhttp://forums.fast.ai/t/pytorch-weight-norm/4619",
      "votes": null
    },
    {
      "id": "211324",
      "postDate": "08/08/2017 15:56:34",
      "content": "<p>this paper also removes BN</p>\n\n<p>An Effective Training Method For Deep Convolutional Neural Network</p>\n\n<p><a href=\"https://arxiv.org/abs/1708.01666\">https://arxiv.org/abs/1708.01666</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/211324/7022/Selection_038.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "this paper also removes BN\n\nAn Effective Training Method For Deep Convolutional Neural Network\n\nhttps://arxiv.org/abs/1708.01666\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/211324/7022/Selection_038.png",
      "votes": null
    },
    {
      "id": "212829",
      "postDate": "08/12/2017 19:39:23",
      "content": "<p>I tried to reimplement this in PyTorch. First, I started with this module:</p>\n\n<pre><code>class NG(nn.Module):\n    # https://arxiv.org/pdf/1708.01666.pdf\n\n    def __init__(self, num_parameters, init=-1):\n        super().__init__()\n        self.num_parameters = num_parameters\n        self.weight = nn.Parameter(torch.Tensor(num_parameters).fill_(init))\n\n    def forward(self, x):\n        shape = list(x.size())\n        shape[0] = 1\n        shape[2] = 1   # 1D\n        shape[3] = 1   # 2D\n        #shape[4] = 1  # 3D\n\n        lt_mask = x.lt(self.weight.view(shape)).float()\n        return self.weight.view(shape) * lt_mask + (1-lt_mask) * x\n</code></pre>\n\n<p>Then I simply deleted every instance of BNorm, and replaced all PReLU(previous_layers_out_channel) with NG(previous_layers_out_channel). The result? The network still converges, albeit a lot slower and less accurately, i.e. worse loss. But here is the confusing part: my network's memory footprint jumped from 4823MiB (BNorm/PReLU) to 5751MiB (NG). How does that make sense? If anything, I'd assume NG would introduce the exact same number of parameters as PReLU with the added benefit of no BN.</p>\n\n<p>Is my implementation flawed? I would never use the non-linearity generator if it converges slower, has a worse loss, and uses more memory.</p>\n\n<p>EDIT:\nFor completeness, to use NG in one's models, one is supposed to:</p>\n\n<blockquote>\n  <p>Given a normalized input image vector (all pixels are rescaled to [-1, 1]) and some proper weight\n  initializations (Xavier initialization [15] or “MSRA” initialization [1])</p>\n</blockquote>\n\n<p>Which I had done in my test.</p>",
      "rawMarkdown": "I tried to reimplement this in PyTorch. First, I started with this module:\n\n    class NG(nn.Module):\n        # https://arxiv.org/pdf/1708.01666.pdf\n        \n        def __init__(self, num_parameters, init=-1):\n            super().__init__()\n            self.num_parameters = num_parameters\n            self.weight = nn.Parameter(torch.Tensor(num_parameters).fill_(init))\n    \n        def forward(self, x):\n            shape = list(x.size())\n            shape[0] = 1\n            shape[2] = 1   # 1D\n            shape[3] = 1   # 2D\n            #shape[4] = 1  # 3D\n            \n            lt_mask = x.lt(self.weight.view(shape)).float()\n            return self.weight.view(shape) * lt_mask + (1-lt_mask) * x\n\nThen I simply deleted every instance of BNorm, and replaced all PReLU(previous_layers_out_channel) with NG(previous_layers_out_channel). The result? The network still converges, albeit a lot slower and less accurately, i.e. worse loss. But here is the confusing part: my network's memory footprint jumped from 4823MiB (BNorm/PReLU) to 5751MiB (NG). How does that make sense? If anything, I'd assume NG would introduce the exact same number of parameters as PReLU with the added benefit of no BN.\n\nIs my implementation flawed? I would never use the non-linearity generator if it converges slower, has a worse loss, and uses more memory.\n\nEDIT:\nFor completeness, to use NG in one's models, one is supposed to:\n\n&gt; Given a normalized input image vector (all pixels are rescaled to [-1, 1]) and some proper weight\ninitializations (Xavier initialization [15] or “MSRA” initialization [1])\n\nWhich I had done in my test.",
      "votes": null
    },
    {
      "id": "212832",
      "postDate": "08/12/2017 19:45:04",
      "content": "<p>I also tried this on my model. Again, every where a conv is followed by a BN, I removed the BN and wrapped the conv layer in a weight_norm(). I don't even remember the final results anymore except that they were underwhelming.</p>\n\n<p>I tried a bunch of other experiments as well. For example, for the first time ever, I was able to get SELU converging in models that had both skip connections and residual connections. But again, convergence was a LOT slower, and I wasn't able to get as accurate scores as my baseline BN/PReLU model. Also with SELU since BN is removed, the network is <strong>very</strong> particular about the LR. I was using something like 0.000001. Even 0.0001 would cause the network to output all zeros. This made it difficult to use all the fancy optimizers like yellowfin and snapshot ensembling.</p>",
      "rawMarkdown": "I also tried this on my model. Again, every where a conv is followed by a BN, I removed the BN and wrapped the conv layer in a weight_norm(). I don't even remember the final results anymore except that they were underwhelming.\n\nI tried a bunch of other experiments as well. For example, for the first time ever, I was able to get SELU converging in models that had both skip connections and residual connections. But again, convergence was a LOT slower, and I wasn't able to get as accurate scores as my baseline BN/PReLU model. Also with SELU since BN is removed, the network is **very** particular about the LR. I was using something like 0.000001. Even 0.0001 would cause the network to output all zeros. This made it difficult to use all the fancy optimizers like yellowfin and snapshot ensembling.",
      "votes": null
    },
    {
      "id": "212893",
      "postDate": "08/13/2017 01:29:46",
      "content": "<p>how about x=F.relu(x-t,inplace=True)+t. this hould gives about same memory?</p>",
      "rawMarkdown": "how about x=F.relu(x-t,inplace=True)+t. this hould gives about same memory?",
      "votes": null
    },
    {
      "id": "212898",
      "postDate": "08/13/2017 01:46:37",
      "content": "<p>i my experiments, it seems that BN don't consume much memory in pytorch. i compare memory at run time using nvidia-smi for conv-bn and conv. for larger effective batch size, i use gradient accumulation.</p>\n\n<p>it seems that pytorch don't have good renormalisation implementation yet.</p>",
      "rawMarkdown": "i my experiments, it seems that BN don't consume much memory in pytorch. i compare memory at run time using nvidia-smi for conv-bn and conv. for larger effective batch size, i use gradient accumulation.\n\nit seems that pytorch don't have good renormalisation implementation yet.",
      "votes": null
    },
    {
      "id": "213090",
      "postDate": "08/13/2017 17:31:07",
      "content": "<p>I can confirm that the much cleaner <code>F.relu</code>, with <code>self.t</code> wrapped in a parameter + appropriately broadcasted does indeed produce the exact same param count as <code>nn.PReLU</code>. Nice!</p>\n\n<p>BN+NG converges at around the same speed as BN+PReLU, but ever so slightly less accurately. Using the same large LRs, NG <em>without</em> BN initially trained very slowly, but then collapsed to all background after about 4-5 epochs. This was sorely disappointing, since in the paper they supposedly used rates as large as 1.5. Anyhow, using a modest 0.0001 rate allowed convergence... and then lowering to 0.000001 allowed the net to get to <em>almost</em> the accuracy of the BN+PReLU model while using 556 MiB less run-time memory according to nvidia-smi. I'm guessing that 556 MiB came from the BN. BN+PReLU was at 4823 and NG-only model was at 4267 MiB.</p>\n\n<p>Take home for my set of tests here is that NG will have a slower convergence due to being sensitive to LR (counter to the paper); and it will result in a network with a slightly smaller footprint along with slightly less accuracy. 556 MiB isn't enough to make or break my hand-me-down 8GB card, so I'll be sticking with BN+PReLU for the time being.</p>\n\n<p>This is the only attempted BatchRenorm2d I've been able to locate for PyTorch, but looks like it's still buggy. There are a few working 1d implementations. Would love to throw a 2d layer in to test once one is stable.</p>\n\n<p><a href=\"https://discuss.pytorch.org/t/computing-the-gradients-for-batch-renormalization/828/3\">https://discuss.pytorch.org/t/computing-the-gradients-for-batch-renormalization/828/3</a></p>",
      "rawMarkdown": "I can confirm that the much cleaner `F.relu`, with `self.t` wrapped in a parameter + appropriately broadcasted does indeed produce the exact same param count as `nn.PReLU`. Nice!\n\nBN+NG converges at around the same speed as BN+PReLU, but ever so slightly less accurately. Using the same large LRs, NG *without* BN initially trained very slowly, but then collapsed to all background after about 4-5 epochs. This was sorely disappointing, since in the paper they supposedly used rates as large as 1.5. Anyhow, using a modest 0.0001 rate allowed convergence... and then lowering to 0.000001 allowed the net to get to *almost* the accuracy of the BN+PReLU model while using 556 MiB less run-time memory according to nvidia-smi. I'm guessing that 556 MiB came from the BN. BN+PReLU was at 4823 and NG-only model was at 4267 MiB.\n\nTake home for my set of tests here is that NG will have a slower convergence due to being sensitive to LR (counter to the paper); and it will result in a network with a slightly smaller footprint along with slightly less accuracy. 556 MiB isn't enough to make or break my hand-me-down 8GB card, so I'll be sticking with BN+PReLU for the time being.\n\nThis is the only attempted BatchRenorm2d I've been able to locate for PyTorch, but looks like it's still buggy. There are a few working 1d implementations. Would love to throw a 2d layer in to test once one is stable.\n\nhttps://discuss.pytorch.org/t/computing-the-gradients-for-batch-renormalization/828/3",
      "votes": null
    },
    {
      "id": "213096",
      "postDate": "08/13/2017 17:52:26",
      "content": "<p>thanks for the report!</p>\n\n<p>Another possibility is to start with BN+Conv first. Then absorb the BN parameters into a pure Conv layer since BN is essentially a linear operation. (<a href=\"https://github.com/sanghoon/pva-faster-rcnn/issues/5\">https://github.com/sanghoon/pva-faster-rcnn/issues/5</a>). </p>\n\n<p>BN+Conv is to trained an initial, CNN to initialize the weights. Absorbed conv is for fine-tunning with lower rates.  (My guess is that if the rate is small, changes are small and maybe BN is not required)</p>\n\n<p>However, it is rather troublesome to do this 2-stage approach.</p>",
      "rawMarkdown": "thanks for the report!\n\nAnother possibility is to start with BN+Conv first. Then absorb the BN parameters into a pure Conv layer since BN is essentially a linear operation. (https://github.com/sanghoon/pva-faster-rcnn/issues/5). \n\nBN+Conv is to trained an initial, CNN to initialize the weights. Absorbed conv is for fine-tunning with lower rates.  (My guess is that if the rate is small, changes are small and maybe BN is not required)\n\nHowever, it is rather troublesome to do this 2-stage approach.",
      "votes": null
    },
    {
      "id": "215918",
      "postDate": "08/23/2017 16:45:59",
      "content": "<p>I tried using Batch Renormalization from keras implementation, but the score didn't improve.(batchsize = 4)\nHow about using other Normalization/Regularization e.g. dropout or L2 penalize?</p>",
      "rawMarkdown": "I tried using Batch Renormalization from keras implementation, but the score didn't improve.(batchsize = 4)\nHow about using other Normalization/Regularization e.g. dropout or L2 penalize?",
      "votes": null
    },
    {
      "id": "221120",
      "postDate": "09/14/2017 07:00:36",
      "content": "<p>How did try it? Pre-training? I also tried using it from keras implementation, and the score improved a little.(batch size=2, input size=1024x1024)</p>",
      "rawMarkdown": "How did try it? Pre-training? I also tried using it from keras implementation, and the score improved a little.(batch size=2, input size=1024x1024)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 211274,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/08/2017 14:08:26",
      "content": "<p>another possible alternative:\nWeight Normalization from <a href=\"https://arxiv.org/abs/1602.07868\">https://arxiv.org/abs/1602.07868</a></p>\n\n<p><a href=\"https://github.com/pytorch/pytorch/blob/master/torch/nn/utils/weight_norm.py\">https://github.com/pytorch/pytorch/blob/master/torch/nn/utils/weight_norm.py</a></p>\n\n<p><a href=\"http://forums.fast.ai/t/pytorch-weight-norm/4619\">http://forums.fast.ai/t/pytorch-weight-norm/4619</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 212832,
          "author_name": "authman",
          "author_url": "",
          "post_date": "08/12/2017 19:45:04",
          "content": "<p>I also tried this on my model. Again, every where a conv is followed by a BN, I removed the BN and wrapped the conv layer in a weight_norm(). I don't even remember the final results anymore except that they were underwhelming.</p>\n\n<p>I tried a bunch of other experiments as well. For example, for the first time ever, I was able to get SELU converging in models that had both skip connections and residual connections. But again, convergence was a LOT slower, and I wasn't able to get as accurate scores as my baseline BN/PReLU model. Also with SELU since BN is removed, the network is <strong>very</strong> particular about the LR. I was using something like 0.000001. Even 0.0001 would cause the network to output all zeros. This made it difficult to use all the fancy optimizers like yellowfin and snapshot ensembling.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 211324,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/08/2017 15:56:34",
      "content": "<p>this paper also removes BN</p>\n\n<p>An Effective Training Method For Deep Convolutional Neural Network</p>\n\n<p><a href=\"https://arxiv.org/abs/1708.01666\">https://arxiv.org/abs/1708.01666</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/211324/7022/Selection_038.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 212829,
          "author_name": "authman",
          "author_url": "",
          "post_date": "08/12/2017 19:39:23",
          "content": "<p>I tried to reimplement this in PyTorch. First, I started with this module:</p>\n\n<pre><code>class NG(nn.Module):\n    # https://arxiv.org/pdf/1708.01666.pdf\n\n    def __init__(self, num_parameters, init=-1):\n        super().__init__()\n        self.num_parameters = num_parameters\n        self.weight = nn.Parameter(torch.Tensor(num_parameters).fill_(init))\n\n    def forward(self, x):\n        shape = list(x.size())\n        shape[0] = 1\n        shape[2] = 1   # 1D\n        shape[3] = 1   # 2D\n        #shape[4] = 1  # 3D\n\n        lt_mask = x.lt(self.weight.view(shape)).float()\n        return self.weight.view(shape) * lt_mask + (1-lt_mask) * x\n</code></pre>\n\n<p>Then I simply deleted every instance of BNorm, and replaced all PReLU(previous_layers_out_channel) with NG(previous_layers_out_channel). The result? The network still converges, albeit a lot slower and less accurately, i.e. worse loss. But here is the confusing part: my network's memory footprint jumped from 4823MiB (BNorm/PReLU) to 5751MiB (NG). How does that make sense? If anything, I'd assume NG would introduce the exact same number of parameters as PReLU with the added benefit of no BN.</p>\n\n<p>Is my implementation flawed? I would never use the non-linearity generator if it converges slower, has a worse loss, and uses more memory.</p>\n\n<p>EDIT:\nFor completeness, to use NG in one's models, one is supposed to:</p>\n\n<blockquote>\n  <p>Given a normalized input image vector (all pixels are rescaled to [-1, 1]) and some proper weight\n  initializations (Xavier initialization [15] or “MSRA” initialization [1])</p>\n</blockquote>\n\n<p>Which I had done in my test.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 212893,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/13/2017 01:29:46",
          "content": "<p>how about x=F.relu(x-t,inplace=True)+t. this hould gives about same memory?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 212898,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/13/2017 01:46:37",
          "content": "<p>i my experiments, it seems that BN don't consume much memory in pytorch. i compare memory at run time using nvidia-smi for conv-bn and conv. for larger effective batch size, i use gradient accumulation.</p>\n\n<p>it seems that pytorch don't have good renormalisation implementation yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213090,
          "author_name": "authman",
          "author_url": "",
          "post_date": "08/13/2017 17:31:07",
          "content": "<p>I can confirm that the much cleaner <code>F.relu</code>, with <code>self.t</code> wrapped in a parameter + appropriately broadcasted does indeed produce the exact same param count as <code>nn.PReLU</code>. Nice!</p>\n\n<p>BN+NG converges at around the same speed as BN+PReLU, but ever so slightly less accurately. Using the same large LRs, NG <em>without</em> BN initially trained very slowly, but then collapsed to all background after about 4-5 epochs. This was sorely disappointing, since in the paper they supposedly used rates as large as 1.5. Anyhow, using a modest 0.0001 rate allowed convergence... and then lowering to 0.000001 allowed the net to get to <em>almost</em> the accuracy of the BN+PReLU model while using 556 MiB less run-time memory according to nvidia-smi. I'm guessing that 556 MiB came from the BN. BN+PReLU was at 4823 and NG-only model was at 4267 MiB.</p>\n\n<p>Take home for my set of tests here is that NG will have a slower convergence due to being sensitive to LR (counter to the paper); and it will result in a network with a slightly smaller footprint along with slightly less accuracy. 556 MiB isn't enough to make or break my hand-me-down 8GB card, so I'll be sticking with BN+PReLU for the time being.</p>\n\n<p>This is the only attempted BatchRenorm2d I've been able to locate for PyTorch, but looks like it's still buggy. There are a few working 1d implementations. Would love to throw a 2d layer in to test once one is stable.</p>\n\n<p><a href=\"https://discuss.pytorch.org/t/computing-the-gradients-for-batch-renormalization/828/3\">https://discuss.pytorch.org/t/computing-the-gradients-for-batch-renormalization/828/3</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 213096,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/13/2017 17:52:26",
          "content": "<p>thanks for the report!</p>\n\n<p>Another possibility is to start with BN+Conv first. Then absorb the BN parameters into a pure Conv layer since BN is essentially a linear operation. (<a href=\"https://github.com/sanghoon/pva-faster-rcnn/issues/5\">https://github.com/sanghoon/pva-faster-rcnn/issues/5</a>). </p>\n\n<p>BN+Conv is to trained an initial, CNN to initialize the weights. Absorbed conv is for fine-tunning with lower rates.  (My guess is that if the rate is small, changes are small and maybe BN is not required)</p>\n\n<p>However, it is rather troublesome to do this 2-stage approach.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 215918,
      "author_name": "lyakaap",
      "author_url": "",
      "post_date": "08/23/2017 16:45:59",
      "content": "<p>I tried using Batch Renormalization from keras implementation, but the score didn't improve.(batchsize = 4)\nHow about using other Normalization/Regularization e.g. dropout or L2 penalize?</p>",
      "votes": null,
      "replies": [
        {
          "id": 221120,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "09/14/2017 07:00:36",
          "content": "<p>How did try it? Pre-training? I also tried using it from keras implementation, and the score improved a little.(batch size=2, input size=1024x1024)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "211240": "Batch normalization works great on moderately sized minibatches, but its performance suffers on small minibatches. In the [batch renormalization paper](https://arxiv.org/abs/1702.03275), they showed that batch **normalization** achieved 78.3% on Inception-v3 with batch size 32, but only 74.2% with batch size 4. In contrast, batch **renormalization** achieved 76.5% with batch size 4.\n\nThis is relevant to this competition because the high image resolution (1918x1280) means you may only be able to fit a few full size images in your video card's memory, thus limiting the max batch size.\n\n![figures 1 and 2 from the paper](https://i.imgur.com/dH4mPCO.png)",
    "211274": "another possible alternative:\nWeight Normalization from https://arxiv.org/abs/1602.07868\n\n\nhttps://github.com/pytorch/pytorch/blob/master/torch/nn/utils/weight_norm.py\n\nhttp://forums.fast.ai/t/pytorch-weight-norm/4619",
    "211324": "this paper also removes BN\n\nAn Effective Training Method For Deep Convolutional Neural Network\n\nhttps://arxiv.org/abs/1708.01666\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/211324/7022/Selection_038.png",
    "212829": "I tried to reimplement this in PyTorch. First, I started with this module:\n\n    class NG(nn.Module):\n        # https://arxiv.org/pdf/1708.01666.pdf\n        \n        def __init__(self, num_parameters, init=-1):\n            super().__init__()\n            self.num_parameters = num_parameters\n            self.weight = nn.Parameter(torch.Tensor(num_parameters).fill_(init))\n    \n        def forward(self, x):\n            shape = list(x.size())\n            shape[0] = 1\n            shape[2] = 1   # 1D\n            shape[3] = 1   # 2D\n            #shape[4] = 1  # 3D\n            \n            lt_mask = x.lt(self.weight.view(shape)).float()\n            return self.weight.view(shape) * lt_mask + (1-lt_mask) * x\n\nThen I simply deleted every instance of BNorm, and replaced all PReLU(previous_layers_out_channel) with NG(previous_layers_out_channel). The result? The network still converges, albeit a lot slower and less accurately, i.e. worse loss. But here is the confusing part: my network's memory footprint jumped from 4823MiB (BNorm/PReLU) to 5751MiB (NG). How does that make sense? If anything, I'd assume NG would introduce the exact same number of parameters as PReLU with the added benefit of no BN.\n\nIs my implementation flawed? I would never use the non-linearity generator if it converges slower, has a worse loss, and uses more memory.\n\nEDIT:\nFor completeness, to use NG in one's models, one is supposed to:\n\n&gt; Given a normalized input image vector (all pixels are rescaled to [-1, 1]) and some proper weight\ninitializations (Xavier initialization [15] or “MSRA” initialization [1])\n\nWhich I had done in my test.",
    "212832": "I also tried this on my model. Again, every where a conv is followed by a BN, I removed the BN and wrapped the conv layer in a weight_norm(). I don't even remember the final results anymore except that they were underwhelming.\n\nI tried a bunch of other experiments as well. For example, for the first time ever, I was able to get SELU converging in models that had both skip connections and residual connections. But again, convergence was a LOT slower, and I wasn't able to get as accurate scores as my baseline BN/PReLU model. Also with SELU since BN is removed, the network is **very** particular about the LR. I was using something like 0.000001. Even 0.0001 would cause the network to output all zeros. This made it difficult to use all the fancy optimizers like yellowfin and snapshot ensembling.",
    "212893": "how about x=F.relu(x-t,inplace=True)+t. this hould gives about same memory?",
    "212898": "i my experiments, it seems that BN don't consume much memory in pytorch. i compare memory at run time using nvidia-smi for conv-bn and conv. for larger effective batch size, i use gradient accumulation.\n\nit seems that pytorch don't have good renormalisation implementation yet.",
    "213090": "I can confirm that the much cleaner `F.relu`, with `self.t` wrapped in a parameter + appropriately broadcasted does indeed produce the exact same param count as `nn.PReLU`. Nice!\n\nBN+NG converges at around the same speed as BN+PReLU, but ever so slightly less accurately. Using the same large LRs, NG *without* BN initially trained very slowly, but then collapsed to all background after about 4-5 epochs. This was sorely disappointing, since in the paper they supposedly used rates as large as 1.5. Anyhow, using a modest 0.0001 rate allowed convergence... and then lowering to 0.000001 allowed the net to get to *almost* the accuracy of the BN+PReLU model while using 556 MiB less run-time memory according to nvidia-smi. I'm guessing that 556 MiB came from the BN. BN+PReLU was at 4823 and NG-only model was at 4267 MiB.\n\nTake home for my set of tests here is that NG will have a slower convergence due to being sensitive to LR (counter to the paper); and it will result in a network with a slightly smaller footprint along with slightly less accuracy. 556 MiB isn't enough to make or break my hand-me-down 8GB card, so I'll be sticking with BN+PReLU for the time being.\n\nThis is the only attempted BatchRenorm2d I've been able to locate for PyTorch, but looks like it's still buggy. There are a few working 1d implementations. Would love to throw a 2d layer in to test once one is stable.\n\nhttps://discuss.pytorch.org/t/computing-the-gradients-for-batch-renormalization/828/3",
    "213096": "thanks for the report!\n\nAnother possibility is to start with BN+Conv first. Then absorb the BN parameters into a pure Conv layer since BN is essentially a linear operation. (https://github.com/sanghoon/pva-faster-rcnn/issues/5). \n\nBN+Conv is to trained an initial, CNN to initialize the weights. Absorbed conv is for fine-tunning with lower rates.  (My guess is that if the rate is small, changes are small and maybe BN is not required)\n\nHowever, it is rather troublesome to do this 2-stage approach.",
    "215918": "I tried using Batch Renormalization from keras implementation, but the score didn't improve.(batchsize = 4)\nHow about using other Normalization/Regularization e.g. dropout or L2 penalize?",
    "221120": "How did try it? Pre-training? I also tried using it from keras implementation, and the score improved a little.(batch size=2, input size=1024x1024)"
  },
  "source": "meta"
}