{
  "id": 79911,
  "title": "Common pitfalls of public kernels",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79911",
  "author_name": "Dmytro Danevskyi",
  "post_date": "2019-02-08T14:53:42.327000",
  "votes": 71,
  "comment_count": 39,
  "views": 0,
  "content": "<p>Now, when this crazy competition is (almost) over, it's a perfect time to rest and do some analysis :)</p>\n\n<p>In this topic, I want to spot some common mistakes made in public kernels created during this competition. Some of these mistakes were repeated so many times that they almost stopped to look queerly :)</p>\n\n<p>Note that I created this topic for educational purposes only, not to make kernels authors look weird or stupid. We all make occasional mistakes from time to time, especially when we don't have much time to do the proper code check.</p>\n\n<p><strong>1. Position-dependent bias</strong></p>\n\n<p>Many pytorch kernels use <code>Attention</code> class that has the following line:</p>\n\n<p><code>self.b = nn.Parameter(torch.zeros(step_dim))</code></p>\n\n<p><code>step_dim</code> is chosen to be 70, which is the length to which all questions are either padded or truncated.</p>\n\n<p>Thus, the bias term has the individual component for each sequence step. This is weird, at least because different questions have a different number of tokens. What if you need to process a question which is longer than 70 tokens? It's clear that a model should not be constrained to work only for questions of some particular length. Another argument is that each bias component is \"tied\" to a specific token position in a sentence, but an absolute position rarely has something to do with the token meaning or any other linguistic property.</p>\n\n<p>A more reasonable way to use a bias term is to have a single scalar bias that is shared among all sequence elements. </p>\n\n<p><strong>2. \"Spatial\" dropout</strong></p>\n\n<p>The same pytorch code has another suspicious component: the use of 2d dropout. It's used since pytorch doesn't provide a proper implementation of 1d dropout (that zeros out whole channels of a sequence of 1d vectors). So one option is to use a 2d version on unsqueezed embeddings. However, the implementation used in public kernels is completely wrong :)</p>\n\n<p>2d dropout layer is initialized as follows:</p>\n\n<p><code>self.embedding_dropout = nn.Dropout2d(0.1)</code></p>\n\n<p>and then used like this:</p>\n\n<p><code>h = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 0)))</code></p>\n\n<p>Let's see what is going on here. First, you have a batch of embeddings, say it has a shape of <code>(N, T, K)</code>. <code>N</code> is the number of elements in a batch, <code>T</code> is the sequence length and <code>K</code> is the feature (embedding) dimension. If we apply <code>unsqueeze()</code> operation on the <code>0</code> axis now, the tensor now has a shape of <code>(1, N, T, K)</code>. 2D dropout zeros random channels (which is the second axis in pytorch). Oh, wait... The second axis is the batch dimension! In this case, 2D dropout will zero several batch elements, not channels. </p>\n\n<p>A proper way to apply 1d dropout, in this case, is to reshape the tensor into <code>(N, K, 1, T)</code> tensor, apply the 2d dropout and then make it have <code>(N, T, K)</code> shape again.</p>\n\n<p>Please contribute if you also noticed any wrong use of common operations!</p>\n\n<p>Good luck and happy kaggling! :0</p>",
  "messages": [
    {
      "id": 468241,
      "postDate": "2019-02-08T14:53:42.327Z",
      "content": "<p>Now, when this crazy competition is (almost) over, it's a perfect time to rest and do some analysis :)</p>\n\n<p>In this topic, I want to spot some common mistakes made in public kernels created during this competition. Some of these mistakes were repeated so many times that they almost stopped to look queerly :)</p>\n\n<p>Note that I created this topic for educational purposes only, not to make kernels authors look weird or stupid. We all make occasional mistakes from time to time, especially when we don't have much time to do the proper code check.</p>\n\n<p><strong>1. Position-dependent bias</strong></p>\n\n<p>Many pytorch kernels use <code>Attention</code> class that has the following line:</p>\n\n<p><code>self.b = nn.Parameter(torch.zeros(step_dim))</code></p>\n\n<p><code>step_dim</code> is chosen to be 70, which is the length to which all questions are either padded or truncated.</p>\n\n<p>Thus, the bias term has the individual component for each sequence step. This is weird, at least because different questions have a different number of tokens. What if you need to process a question which is longer than 70 tokens? It's clear that a model should not be constrained to work only for questions of some particular length. Another argument is that each bias component is \"tied\" to a specific token position in a sentence, but an absolute position rarely has something to do with the token meaning or any other linguistic property.</p>\n\n<p>A more reasonable way to use a bias term is to have a single scalar bias that is shared among all sequence elements. </p>\n\n<p><strong>2. \"Spatial\" dropout</strong></p>\n\n<p>The same pytorch code has another suspicious component: the use of 2d dropout. It's used since pytorch doesn't provide a proper implementation of 1d dropout (that zeros out whole channels of a sequence of 1d vectors). So one option is to use a 2d version on unsqueezed embeddings. However, the implementation used in public kernels is completely wrong :)</p>\n\n<p>2d dropout layer is initialized as follows:</p>\n\n<p><code>self.embedding_dropout = nn.Dropout2d(0.1)</code></p>\n\n<p>and then used like this:</p>\n\n<p><code>h = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 0)))</code></p>\n\n<p>Let's see what is going on here. First, you have a batch of embeddings, say it has a shape of <code>(N, T, K)</code>. <code>N</code> is the number of elements in a batch, <code>T</code> is the sequence length and <code>K</code> is the feature (embedding) dimension. If we apply <code>unsqueeze()</code> operation on the <code>0</code> axis now, the tensor now has a shape of <code>(1, N, T, K)</code>. 2D dropout zeros random channels (which is the second axis in pytorch). Oh, wait... The second axis is the batch dimension! In this case, 2D dropout will zero several batch elements, not channels. </p>\n\n<p>A proper way to apply 1d dropout, in this case, is to reshape the tensor into <code>(N, K, 1, T)</code> tensor, apply the 2d dropout and then make it have <code>(N, T, K)</code> shape again.</p>\n\n<p>Please contribute if you also noticed any wrong use of common operations!</p>\n\n<p>Good luck and happy kaggling! :0</p>",
      "rawMarkdown": "Now, when this crazy competition is (almost) over, it's a perfect time to rest and do some analysis :)\n\nIn this topic, I want to spot some common mistakes made in public kernels created during this competition. Some of these mistakes were repeated so many times that they almost stopped to look queerly :)\n\nNote that I created this topic for educational purposes only, not to make kernels authors look weird or stupid. We all make occasional mistakes from time to time, especially when we don't have much time to do the proper code check.\n\n**1. Position-dependent bias**\n\nMany pytorch kernels use `Attention` class that has the following line:\n\n`self.b = nn.Parameter(torch.zeros(step_dim))`\n\n`step_dim` is chosen to be 70, which is the length to which all questions are either padded or truncated.\n\nThus, the bias term has the individual component for each sequence step. This is weird, at least because different questions have a different number of tokens. What if you need to process a question which is longer than 70 tokens? It's clear that a model should not be constrained to work only for questions of some particular length. Another argument is that each bias component is \"tied\" to a specific token position in a sentence, but an absolute position rarely has something to do with the token meaning or any other linguistic property.\n\nA more reasonable way to use a bias term is to have a single scalar bias that is shared among all sequence elements. \n\n**2. \"Spatial\" dropout**\n\nThe same pytorch code has another suspicious component: the use of 2d dropout. It's used since pytorch doesn't provide a proper implementation of 1d dropout (that zeros out whole channels of a sequence of 1d vectors). So one option is to use a 2d version on unsqueezed embeddings. However, the implementation used in public kernels is completely wrong :)\n\n2d dropout layer is initialized as follows:\n\n`self.embedding_dropout = nn.Dropout2d(0.1)`\n\nand then used like this:\n\n`h = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 0)))`\n\nLet's see what is going on here. First, you have a batch of embeddings, say it has a shape of `(N, T, K)`. `N` is the number of elements in a batch, `T` is the sequence length and `K` is the feature (embedding) dimension. If we apply `unsqueeze()` operation on the `0` axis now, the tensor now has a shape of `(1, N, T, K)`. 2D dropout zeros random channels (which is the second axis in pytorch). Oh, wait... The second axis is the batch dimension! In this case, 2D dropout will zero several batch elements, not channels. \n\nA proper way to apply 1d dropout, in this case, is to reshape the tensor into `(N, K, 1, T)` tensor, apply the 2d dropout and then make it have `(N, T, K)` shape again.\n\nPlease contribute if you also noticed any wrong use of common operations!\n\nGood luck and happy kaggling! :0",
      "votes": 71
    },
    {
      "id": 468329,
      "postDate": "2019-02-08T17:22:46.143Z",
      "content": "<p>Both of these issues originated in <a href=\"https://www.kaggle.com/bminixhofer/deterministic-neural-networks-using-pytorch\">my kernel</a>. </p>\n\n<p>It is really questionable whether position-dependent bias makes sense. In the popular Keras attention script it is implemented the same way though:</p>\n\n<pre><code>  if self.bias:\n          self.b = self.add_weight((input_shape[1],),\n                                   initializer='zero',\n                                   name='{}_b'.format(self.name),\n                                   regularizer=self.b_regularizer,\n                                   constraint=self.b_constraint)\n</code></pre>\n\n<p>With a quick google search I was not able to find out whether bias is made position dependent in the original paper. Do you know about that?</p>\n\n<p>Now the second issue is certainly a mistake I made when writing the kernel. Sorry to everyone whose kernel this code spread to. I fixed this in a new version for people who might use my kernel as a reference later by doing:</p>\n\n<pre><code>h_embedding = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 2)), 2)\n</code></pre>\n\n<p>Is that correct?</p>\n\n<p>Thanks for pointing out the mistakes!</p>",
      "rawMarkdown": "Both of these issues originated in [my kernel](https://www.kaggle.com/bminixhofer/deterministic-neural-networks-using-pytorch). \n\nIt is really questionable whether position-dependent bias makes sense. In the popular Keras attention script it is implemented the same way though:\n\n      if self.bias:\n              self.b = self.add_weight((input_shape[1],),\n                                       initializer='zero',\n                                       name='{}_b'.format(self.name),\n                                       regularizer=self.b_regularizer,\n                                       constraint=self.b_constraint)\n\nWith a quick google search I was not able to find out whether bias is made position dependent in the original paper. Do you know about that?\n\nNow the second issue is certainly a mistake I made when writing the kernel. Sorry to everyone whose kernel this code spread to. I fixed this in a new version for people who might use my kernel as a reference later by doing:\n\n    h_embedding = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 2)), 2)\n\nIs that correct?\n\nThanks for pointing out the mistakes!",
      "votes": 11,
      "replies": [
        {
          "id": 468365,
          "postDate": "2019-02-08T19:05:02.573Z",
          "content": "<p><a href=\"/bminixhofer\">@bminixhofer</a> you can refer to this implementation: <a href=\"https://github.com/TobiasLee/Text-Classification/blob/master/models/modules/attention.py\">https://github.com/TobiasLee/Text-Classification/blob/master/models/modules/attention.py</a></p>\n\n<p>It uses a bit more complicated version of attention, where each RNN output is first squashed using <code>tanh(W @ h + b)</code> transformation, but the bias is clearly shared among sequence elements and is not dependent on the sequence length.</p>",
          "rawMarkdown": "@bminixhofer you can refer to this implementation: https://github.com/TobiasLee/Text-Classification/blob/master/models/modules/attention.py\n\nIt uses a bit more complicated version of attention, where each RNN output is first squashed using `tanh(W @ h + b)` transformation, but the bias is clearly shared among sequence elements and is not dependent on the sequence length.",
          "votes": 3
        },
        {
          "id": 468371,
          "postDate": "2019-02-08T19:16:59.143Z",
          "content": "<p><a href=\"/bminixhofer\">@bminixhofer</a> as for your dropout implementation, to make it correct, you need additionally to permute the dimensions after doing the <code>unsqueeze()</code> operation. If you apply unsqueeze on 2nd axis, tensor will have a shape of (N, T, 1, K) and if you then apply 2d dropout it will zero entire embeddings! So you need to add <code>t.permute(0, 3, 2, 1)</code> to make it correct. After doing a dropout, you need to permute the dimensions to original order using <code>t.permute(0, 3, 2, 1)</code> and do the <code>squeeze()</code>.</p>\n\n<p>So the entire operation will look something like this:</p>\n\n<p><code>\nembeddings = ...\nembeddings = embeddings.unsqueeze(2)    # (N, T, 1, K)\nembeddings = embeddings.permute(0, 3, 2, 1)  # (N, K, 1, T)\nembeddings = self.dropout2d(embeddings)  # (N, K, 1, T), some features are masked\nembeddings = embeddings.permute(0, 3, 2, 1)  # (N, T, 1, K)\nembeddings = embeddings.squeeze(2)  # (N, T, K)\n</code></p>\n\n<p>Looks a bit cumbersome, but you can put it into a separate module to make your main <code>forward()</code> function to be more clear.</p>",
          "rawMarkdown": "@bminixhofer as for your dropout implementation, to make it correct, you need additionally to permute the dimensions after doing the `unsqueeze()` operation. If you apply unsqueeze on 2nd axis, tensor will have a shape of (N, T, 1, K) and if you then apply 2d dropout it will zero entire embeddings! So you need to add `t.permute(0, 3, 2, 1)` to make it correct. After doing a dropout, you need to permute the dimensions to original order using `t.permute(0, 3, 2, 1)` and do the `squeeze()`.\n\nSo the entire operation will look something like this:\n\n```\nembeddings = ...\nembeddings = embeddings.unsqueeze(2)    # (N, T, 1, K)\nembeddings = embeddings.permute(0, 3, 2, 1)  # (N, K, 1, T)\nembeddings = self.dropout2d(embeddings)  # (N, K, 1, T), some features are masked\nembeddings = embeddings.permute(0, 3, 2, 1)  # (N, T, 1, K)\nembeddings = embeddings.squeeze(2)  # (N, T, K)\n```\n\nLooks a bit cumbersome, but you can put it into a separate module to make your main `forward()` function to be more clear.\n",
          "votes": 4
        },
        {
          "id": 468392,
          "postDate": "2019-02-08T20:05:57.667Z",
          "content": "<p>suppose if we don't use permute to change the shape, then we are effectively dropping some words in the total sentence.  </p>\n\n<p>Am I right Dmitriy ?</p>",
          "rawMarkdown": "suppose if we don't use permute to change the shape, then we are effectively dropping some words in the total sentence.  \n\nAm I right Dmitriy ?",
          "votes": 2
        },
        {
          "id": 468437,
          "postDate": "2019-02-08T21:25:40.077Z",
          "content": "<p>Ok, that makes sense <a href=\"/ddanevskyi\">@ddanevskyi</a>. I hope spatial dropout gets added to the standard PyTorch library. And yes, without permutation all features of some positions would be dropped <a href=\"/suchith0312\">@suchith0312</a>.</p>",
          "rawMarkdown": "Ok, that makes sense @ddanevskyi. I hope spatial dropout gets added to the standard PyTorch library. And yes, without permutation all features of some positions would be dropped @suchith0312.",
          "votes": 2
        },
        {
          "id": 469365,
          "postDate": "2019-02-11T03:55:13.397Z",
          "content": "<p>Hi, here is a shorter code with the same concept</p>\n\n<pre><code>emb = ...\nemb = self.dropout2d(emb.transpose(1,2).unsqueeze(-1)).squeeze().transpose(1,2)\n</code></pre>",
          "rawMarkdown": "Hi, here is a shorter code with the same concept\n\n    emb = ...\n    emb = self.dropout2d(emb.transpose(1,2).unsqueeze(-1)).squeeze().transpose(1,2)",
          "votes": 3
        },
        {
          "id": 469483,
          "postDate": "2019-02-11T09:57:29.543Z",
          "content": "<p><a href=\"/radream\">@radream</a> right!  I used a more explicit version mainly to be able to write comments</p>",
          "rawMarkdown": "@radream right!  I used a more explicit version mainly to be able to write comments"
        },
        {
          "id": 470043,
          "postDate": "2019-02-12T09:35:46.977Z",
          "content": "<p>Here is the kernel with same implementation. Drastically improved CV 0.6792 to .6842\n<a href=\"https://www.kaggle.com/mlwhiz/third-place-model-for-toxic-spatial-dropout\">https://www.kaggle.com/mlwhiz/third-place-model-for-toxic-spatial-dropout</a></p>",
          "rawMarkdown": "Here is the kernel with same implementation. Drastically improved CV 0.6792 to .6842\nhttps://www.kaggle.com/mlwhiz/third-place-model-for-toxic-spatial-dropout",
          "votes": 1
        },
        {
          "id": 470052,
          "postDate": "2019-02-12T10:15:18.753Z",
          "content": "<p><a href=\"/mlwhiz\">@mlwhiz</a> cool!</p>",
          "rawMarkdown": "@mlwhiz cool!"
        },
        {
          "id": 471372,
          "postDate": "2019-02-14T11:10:25.147Z",
          "content": "<p><a href=\"/bminixhofer\">@bminixhofer</a>, thanks for your contributions in this competition. Your pytorch kernel and the pytorch starter by <a href=\"/hung96ad\">@hung96ad</a> (thanks Hung) helped me start using pytorch in this contest. My pytorch model did not score as well as my keras model but I used it as a backup. So I am sad to not see both of you on the stage-2 LB, you easily could have jumped several places up. Anyway, see you in the next one.</p>",
          "rawMarkdown": "@bminixhofer, thanks for your contributions in this competition. Your pytorch kernel and the pytorch starter by @hung96ad (thanks Hung) helped me start using pytorch in this contest. My pytorch model did not score as well as my keras model but I used it as a backup. So I am sad to not see both of you on the stage-2 LB, you easily could have jumped several places up. Anyway, see you in the next one.",
          "votes": 3
        },
        {
          "id": 471441,
          "postDate": "2019-02-14T13:19:02.657Z",
          "content": "<p>You're welcome! And congratulations on your placement. I grew frustrated with this competition because I was not able to improve my score much further than the public kernels. It's amazing how many unique approaches people found for a competition with only text input and no external data.</p>",
          "rawMarkdown": "You're welcome! And congratulations on your placement. I grew frustrated with this competition because I was not able to improve my score much further than the public kernels. It's amazing how many unique approaches people found for a competition with only text input and no external data.",
          "votes": 3
        },
        {
          "id": 471472,
          "postDate": "2019-02-14T14:11:01.267Z",
          "content": "<p>Thanks <a href=\"/bminixhofer\">@bminixhofer</a>,  I know what you mean. I felt just as frustrated but learning pytorch kept me going for a bit. In fact my best model that placed me at my current position was a submission from 2 months ago. So it is really interesting to see other's creatvity and I would surely use some in the future.</p>",
          "rawMarkdown": "Thanks @bminixhofer,  I know what you mean. I felt just as frustrated but learning pytorch kept me going for a bit. In fact my best model that placed me at my current position was a submission from 2 months ago. So it is really interesting to see other's creatvity and I would surely use some in the future."
        },
        {
          "id": 471930,
          "postDate": "2019-02-15T05:09:56.617Z",
          "content": "<p>thank <a href=\"/sheriytm\">@sheriytm</a> and congratulations on your placement. I tried with many Kernel but my score isn't able to improve. Howerver, I know many approaches from this competition.</p>",
          "rawMarkdown": "thank @sheriytm and congratulations on your placement. I tried with many Kernel but my score isn't able to improve. Howerver, I know many approaches from this competition.",
          "votes": 1
        },
        {
          "id": 471936,
          "postDate": "2019-02-15T05:38:35Z",
          "content": "<p>Thanks <a href=\"/hung96ad\">@hung96ad</a> . I feel the same way about approaches learnt here will be useful for future work.  I am sure I will be revisiting some of these interesting submissions for inspiration just like I have referred to the highly optimized kernels in the Mercari Price competition a few times.</p>",
          "rawMarkdown": "Thanks @hung96ad . I feel the same way about approaches learnt here will be useful for future work.  I am sure I will be revisiting some of these interesting submissions for inspiration just like I have referred to the highly optimized kernels in the Mercari Price competition a few times."
        }
      ]
    },
    {
      "id": 468391,
      "postDate": "2019-02-08T20:04:41.100Z",
      "content": "<p>I used this alternative version of attention from <a href=\"https://discuss.pytorch.org/t/self-attention-on-words-and-masking/5671/4\">https://discuss.pytorch.org/t/self-attention-on-words-and-masking/5671/4</a>. I changed the code a bit to work with pytorch 1.0. This version doesn't use a bias term, but in my opinion, setting a maximum token length of 70 shouldn't decrease the performance of the model by too much. This version however, does have the problem of not masking before running softmax over the the attention weights. Allennlp has a masked_log_softmax function <a href=\"https://github.com/allenai/allennlp/blob/b6cc9d39651273e8ec2a7e334908ffa9de5c2026/allennlp/nn/util.py#L272-L303\">https://github.com/allenai/allennlp/blob/b6cc9d39651273e8ec2a7e334908ffa9de5c2026/allennlp/nn/util.py#L272-L303</a> that correctly masks the input sequence before applying softmax</p>\n\n<pre><code>class SelfAttention(nn.Module):\ndef __init__(self, hidden_size, batch_first=False):\n    super(SelfAttention, self).__init__()\n\n    self.hidden_size = hidden_size\n    self.batch_first = batch_first\n\n    self.att_weights = nn.Parameter(torch.Tensor(1, hidden_size), requires_grad=True)\n    nn.init.xavier_uniform_(self.att_weights.data)\n\ndef get_mask(self):\n    pass\n\ndef forward(self, inputs, lengths):\n    if self.batch_first:\n        batch_size, max_len = inputs.size()[:2]\n    else:\n        max_len, batch_size = inputs.size()[:2]\n\n    # apply attention layer\n    weights = torch.bmm(inputs,\n                        self.att_weights  # (1, hidden_size)\n                        .permute(1, 0)  # (hidden_size, 1)\n                        .unsqueeze(0)  # (1, hidden_size, 1)\n                        .repeat(batch_size, 1, 1) # (batch_size, hidden_size, 1)\n                        )\n\n    attentions = torch.softmax(F.relu(weights.squeeze()), dim=-1)\n\n    # create mask based on the sentence lengths\n    mask = torch.ones(attentions.size(), requires_grad=True).cuda()\n    for i, l in enumerate(lengths):  # skip the first sentence\n        if l &lt; max_len:\n            mask[i, l:] = 0\n\n    # apply mask and renormalize attention scores (weights)\n    masked = attentions * mask\n    _sums = masked.sum(-1).unsqueeze(-1)  # sums per row\n\n    attentions = masked.div(_sums)\n\n    # apply attention weights\n    weighted = torch.mul(inputs, attentions.unsqueeze(-1).expand_as(inputs))\n\n    # get the final fixed vector representations of the sentences\n    representations = weighted.sum(1).squeeze()\n\n    return representations, attentions\n</code></pre>",
      "rawMarkdown": "I used this alternative version of attention from https://discuss.pytorch.org/t/self-attention-on-words-and-masking/5671/4. I changed the code a bit to work with pytorch 1.0. This version doesn't use a bias term, but in my opinion, setting a maximum token length of 70 shouldn't decrease the performance of the model by too much. This version however, does have the problem of not masking before running softmax over the the attention weights. Allennlp has a masked_log_softmax function https://github.com/allenai/allennlp/blob/b6cc9d39651273e8ec2a7e334908ffa9de5c2026/allennlp/nn/util.py#L272-L303 that correctly masks the input sequence before applying softmax\n\n    class SelfAttention(nn.Module):\n    def __init__(self, hidden_size, batch_first=False):\n        super(SelfAttention, self).__init__()\n\n        self.hidden_size = hidden_size\n        self.batch_first = batch_first\n\n        self.att_weights = nn.Parameter(torch.Tensor(1, hidden_size), requires_grad=True)\n        nn.init.xavier_uniform_(self.att_weights.data)\n\n    def get_mask(self):\n        pass\n\n    def forward(self, inputs, lengths):\n        if self.batch_first:\n            batch_size, max_len = inputs.size()[:2]\n        else:\n            max_len, batch_size = inputs.size()[:2]\n            \n        # apply attention layer\n        weights = torch.bmm(inputs,\n                            self.att_weights  # (1, hidden_size)\n                            .permute(1, 0)  # (hidden_size, 1)\n                            .unsqueeze(0)  # (1, hidden_size, 1)\n                            .repeat(batch_size, 1, 1) # (batch_size, hidden_size, 1)\n                            )\n    \n        attentions = torch.softmax(F.relu(weights.squeeze()), dim=-1)\n        \n        # create mask based on the sentence lengths\n        mask = torch.ones(attentions.size(), requires_grad=True).cuda()\n        for i, l in enumerate(lengths):  # skip the first sentence\n            if l &lt; max_len:\n                mask[i, l:] = 0\n\n        # apply mask and renormalize attention scores (weights)\n        masked = attentions * mask\n        _sums = masked.sum(-1).unsqueeze(-1)  # sums per row\n        \n        attentions = masked.div(_sums)\n\n        # apply attention weights\n        weighted = torch.mul(inputs, attentions.unsqueeze(-1).expand_as(inputs))\n\n        # get the final fixed vector representations of the sentences\n        representations = weighted.sum(1).squeeze()\n\n        return representations, attentions",
      "votes": 3
    },
    {
      "id": 468544,
      "postDate": "2019-02-09T05:28:15.340Z",
      "content": "<p>I rerun my 2 submissions, with that dropout correction &amp; guess what, my CV improved by 0.004. Wish I correct it earlier..</p>",
      "rawMarkdown": "I rerun my 2 submissions, with that dropout correction &amp; guess what, my CV improved by 0.004. Wish I correct it earlier..",
      "votes": 4,
      "replies": [
        {
          "id": 468792,
          "postDate": "2019-02-09T17:12:37.027Z",
          "content": "<p>For me also it was 0.0015 on CV score. </p>",
          "rawMarkdown": "For me also it was 0.0015 on CV score. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 471993,
      "postDate": "2019-02-15T07:30:44.263Z",
      "content": "<p><a href=\"/ddanevskyi\">@ddanevskyi</a>, as an experiment I used late submission to check the result of a modified copy of <a href=\"/bminixhofer\">@bminixhofer</a>'s kernel and one with his current update in the dropout mentioned. See the result below:-</p>\n\n<ul>\n<li>My modified copy of the kernel scored </li>\n</ul>\n\n<p>private LB = 0.70299</p>\n\n<p>public LB = 0.69834</p>\n\n<ul>\n<li>My same kernel with the latest corrections for the dropout.</li>\n</ul>\n\n<p>private LB = 0.70220</p>\n\n<p>public LB = 0.69395</p>",
      "rawMarkdown": "@ddanevskyi, as an experiment I used late submission to check the result of a modified copy of @bminixhofer's kernel and one with his current update in the dropout mentioned. See the result below:-\n\n- My modified copy of the kernel scored \n\nprivate LB = 0.70299\n\npublic LB = 0.69834\n\n- My same kernel with the latest corrections for the dropout.\n\nprivate LB = 0.70220\n\npublic LB = 0.69395",
      "votes": 1,
      "replies": [
        {
          "id": 472135,
          "postDate": "2019-02-15T12:08:15.457Z",
          "content": "<p>Well, I just pointed out that the way dropout is used in public kernels is incorrect. Albeit some people reported major improvements from switching to the corrected version, it might not work for you out of the box for several reasons. First, \"correct\" dropout increases the batch size (since it doesn't drop batch elements), so you might need to tune your learning rate schedule or use a lower number of training epochs. Second, \"correct\" dropout <em>does</em> use dropout (haha), so you might need to do some additional hyperparameter tuning to make it work better.</p>",
          "rawMarkdown": "Well, I just pointed out that the way dropout is used in public kernels is incorrect. Albeit some people reported major improvements from switching to the corrected version, it might not work for you out of the box for several reasons. First, \"correct\" dropout increases the batch size (since it doesn't drop batch elements), so you might need to tune your learning rate schedule or use a lower number of training epochs. Second, \"correct\" dropout *does* use dropout (haha), so you might need to do some additional hyperparameter tuning to make it work better.",
          "votes": 2
        },
        {
          "id": 472155,
          "postDate": "2019-02-15T12:31:47.587Z",
          "content": "<p>I agree with your earlier explanation about it being wrong. I was just curious as to how much did it mess with the model. Hence surprised to see the results. Thought others may want to know since <a href=\"/bminixhofer\">@bminixhofer</a>'s kernel does not show scores because he missed stage-2. I forgot to add above that the corrections improved the local CV as well. </p>",
          "rawMarkdown": "I agree with your earlier explanation about it being wrong. I was just curious as to how much did it mess with the model. Hence surprised to see the results. Thought others may want to know since @bminixhofer's kernel does not show scores because he missed stage-2. I forgot to add above that the corrections improved the local CV as well. ",
          "votes": 1
        },
        {
          "id": 472316,
          "postDate": "2019-02-15T17:24:49.217Z",
          "content": "<p>I think you need to retune the model when you changed the spatialdropout. It make the more harder to overfit.</p>",
          "rawMarkdown": "I think you need to retune the model when you changed the spatialdropout. It make the more harder to overfit.",
          "votes": 1
        }
      ]
    },
    {
      "id": 469983,
      "postDate": "2019-02-12T07:03:33.993Z",
      "content": "<p>very good observation :)</p>",
      "rawMarkdown": "very good observation :)",
      "votes": 1
    },
    {
      "id": 468778,
      "postDate": "2019-02-09T16:30:57.037Z",
      "content": "<p>Noooooooooo 😱</p>",
      "rawMarkdown": "Noooooooooo 😱",
      "votes": 1
    },
    {
      "id": 468330,
      "postDate": "2019-02-08T17:23:23.777Z",
      "content": "<p>Thanks for the post. I did copy the code of dropout and my cv score decreased, so didn't use it.</p>\n\n<p>Edit: I started to read more about spatial dropout. I didn't understand somethings quite well. (I may be wrong but this is my view)</p>\n\n<p>So in convolution layer spatial dropout zeros out the entire channel. So if input as 64 channels we randomly zero out some channels out of 64. </p>\n\n<p>Now in case of text, we push the embeddings to in each time step to RNN/LSTM. So here we randomly zero out some of the time steps out of total T steps(here maybe I am wrong). </p>\n\n<p>In pytorch, the nn.dropout2d has channels in 1st axis(if we start counting axis from 0), then we have to shape the tensor into (N, T, 1, K)(because time steps will be channels here) but you wrote (N, K, 1, T). </p>",
      "rawMarkdown": "Thanks for the post. I did copy the code of dropout and my cv score decreased, so didn't use it.\n\nEdit: I started to read more about spatial dropout. I didn't understand somethings quite well. (I may be wrong but this is my view)\n\nSo in convolution layer spatial dropout zeros out the entire channel. So if input as 64 channels we randomly zero out some channels out of 64. \n\nNow in case of text, we push the embeddings to in each time step to RNN/LSTM. So here we randomly zero out some of the time steps out of total T steps(here maybe I am wrong). \n\nIn pytorch, the nn.dropout2d has channels in 1st axis(if we start counting axis from 0), then we have to shape the tensor into (N, T, 1, K)(because time steps will be channels here) but you wrote (N, K, 1, T). ",
      "votes": 1,
      "replies": [
        {
          "id": 468366,
          "postDate": "2019-02-08T19:08:36.250Z",
          "content": "<p><a href=\"/suchith0312\">@suchith0312</a> It has to be (N, K, 1, T). We clearly want to zero <em>features</em>, not timesteps. You can think of this tensor as a batch of images, having K channels, the width of T and the height of 1.</p>",
          "rawMarkdown": "@suchith0312 It has to be (N, K, 1, T). We clearly want to zero *features*, not timesteps. You can think of this tensor as a batch of images, having K channels, the width of T and the height of 1.",
          "votes": 6
        },
        {
          "id": 468389,
          "postDate": "2019-02-08T20:02:57.883Z",
          "content": "<p>Thanks for the reply.  I think you are right as I have been checking keras spatial dropout implementation and they use a noise shape of <code>(batch_size,1, features)</code>. So eventually they are dropping the one dimension of the embedding for all time steps.</p>\n\n<p>But I am not getting the feel of why spatial dropout works in some cases?</p>\n\n<p>Suppose we have a sequence length of 3 and embedding dimension of 6.</p>\n\n<p>Now if the shape is <code>(N,6,1,3)</code> and dropout2d on that will make to drop some out of 6 channels present.\nLet's suppose we have input as <code>[1,2,3,4,5,6] , [7,8,9,10,11,12] , [-1,-2,-3,-4,-5,-6]</code> (embedding of some word), then after applying dropout we can get <code>[1,2,0,0,5,6] , [7,8,0,0,11,12] , [-1,-2,0,0,-5,-6]</code>. This can be embedding of some other word. So by dropping the features, we are changing the word. </p>\n\n<p>So it makes sense to drop some of the words as in a sentence, as these words can be noisy and the meaning of sentence really doesn't depend on these words.</p>",
          "rawMarkdown": "Thanks for the reply.  I think you are right as I have been checking keras spatial dropout implementation and they use a noise shape of `(batch_size,1, features)`. So eventually they are dropping the one dimension of the embedding for all time steps.\n \nBut I am not getting the feel of why spatial dropout works in some cases?\n \nSuppose we have a sequence length of 3 and embedding dimension of 6.\n\nNow if the shape is `(N,6,1,3)` and dropout2d on that will make to drop some out of 6 channels present.\nLet's suppose we have input as `[1,2,3,4,5,6] , [7,8,9,10,11,12] , [-1,-2,-3,-4,-5,-6]` (embedding of some word), then after applying dropout we can get `[1,2,0,0,5,6] , [7,8,0,0,11,12] , [-1,-2,0,0,-5,-6]`. This can be embedding of some other word. So by dropping the features, we are changing the word. \n\nSo it makes sense to drop some of the words as in a sentence, as these words can be noisy and the meaning of sentence really doesn't depend on these words.",
          "votes": 2
        }
      ]
    },
    {
      "id": 468317,
      "postDate": "2019-02-08T17:03:14.830Z",
      "content": "<p>thanks for pointing out the correct usage of dropout! I was so curious why my pytorch and keras giving me very different convergence speed - my guess was the spatial dropout part but didn't pay too much attention and ignored the error = =|||</p>",
      "rawMarkdown": "thanks for pointing out the correct usage of dropout! I was so curious why my pytorch and keras giving me very different convergence speed - my guess was the spatial dropout part but didn't pay too much attention and ignored the error = =|||",
      "votes": 1
    },
    {
      "id": 468294,
      "postDate": "2019-02-08T16:26:52.623Z",
      "content": "<p>Thanks for pointing out the correct usage of the dropout. I saw that in many kernels and was pretty confident what they were doing wasnt correct but didnt know enough about pytorch to say for sure. Glad we stuck with keras throughout</p>",
      "rawMarkdown": "Thanks for pointing out the correct usage of the dropout. I saw that in many kernels and was pretty confident what they were doing wasnt correct but didnt know enough about pytorch to say for sure. Glad we stuck with keras throughout",
      "votes": 1
    },
    {
      "id": 469302,
      "postDate": "2019-02-10T23:02:52.523Z",
      "content": "<p>Haha, that's fascinating! I used hyperopt to find the best possible hyperparameters and it has suggested that I remove embedding dropout completely. </p>\n\n<p>However, this \"dropout\" essentially means sampling of random 90% of samples on every epoch, so it's (sort of) regularization.</p>",
      "rawMarkdown": "Haha, that's fascinating! I used hyperopt to find the best possible hyperparameters and it has suggested that I remove embedding dropout completely. \n\nHowever, this \"dropout\" essentially means sampling of random 90% of samples on every epoch, so it's (sort of) regularization.",
      "votes": 2,
      "replies": [
        {
          "id": 469321,
          "postDate": "2019-02-11T00:52:54.957Z",
          "content": "<p><a href=\"/artyomp\">@artyomp</a>, yeah but the point is that it was not working as intended :-)</p>",
          "rawMarkdown": "@artyomp, yeah but the point is that it was not working as intended :-)",
          "votes": 2
        },
        {
          "id": 469530,
          "postDate": "2019-02-11T11:30:17.243Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 469009,
      "postDate": "2019-02-10T09:44:51.100Z",
      "content": "<p>Thanks <a href=\"/ddanevskyi\">@ddanevskyi</a> for pointing these out. I also noticed  one of the issues but was not quite sure if it was really wrong. My pytorch skills is not as defined as my keras hence I hesitated, I wish I asked anyway.</p>",
      "rawMarkdown": "Thanks @ddanevskyi for pointing these out. I also noticed  one of the issues but was not quite sure if it was really wrong. My pytorch skills is not as defined as my keras hence I hesitated, I wish I asked anyway.",
      "votes": 2
    },
    {
      "id": 471983,
      "postDate": "2019-02-15T07:16:39.673Z",
      "content": "<p>Dmitriy, very good observations and thank you for sharing! But what would be the difference of using regular dropout <code>self.embedding_dropout = nn.Dropout(0.1)</code> with \n<code>h = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 0)))</code> afterwards?</p>",
      "rawMarkdown": "Dmitriy, very good observations and thank you for sharing! But what would be the difference of using regular dropout `self.embedding_dropout = nn.Dropout(0.1)` with \n`h = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 0)))` afterwards?",
      "replies": [
        {
          "id": 472127,
          "postDate": "2019-02-15T11:57:40.743Z",
          "content": "<p>Regular dropout drops elements independently, while \"spatial\" or \"locked\" dropout applies the same mask for all sequence steps. It was demonstrated that the latter performs better for RNN models. See <a href=\"https://arxiv.org/abs/1512.05287\">https://arxiv.org/abs/1512.05287</a> for details.</p>",
          "rawMarkdown": "Regular dropout drops elements independently, while \"spatial\" or \"locked\" dropout applies the same mask for all sequence steps. It was demonstrated that the latter performs better for RNN models. See https://arxiv.org/abs/1512.05287 for details.",
          "votes": 4
        },
        {
          "id": 472165,
          "postDate": "2019-02-15T12:43:25.690Z",
          "content": "<p>Thanks again! Well, regular dropout worked better than the one with the mistake you've pointed out. </p>",
          "rawMarkdown": "Thanks again! Well, regular dropout worked better than the one with the mistake you've pointed out. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 469181,
      "postDate": "2019-02-10T16:55:21.730Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 468718,
      "postDate": "2019-02-09T14:32:45.413Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 470811,
      "postDate": "2019-02-13T16:05:08.313Z",
      "content": "<p>Thanks for sharing Dmitriy!</p>",
      "rawMarkdown": "Thanks for sharing Dmitriy!",
      "votes": 1
    },
    {
      "id": 469712,
      "postDate": "2019-02-11T17:58:58.650Z",
      "content": "<p>Thanks~</p>",
      "rawMarkdown": "Thanks~",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 468329,
      "author_name": "Benjamin Minixhofer",
      "author_url": "",
      "post_date": "2019-02-08T17:22:46.143000",
      "content": "<p>Both of these issues originated in <a href=\"https://www.kaggle.com/bminixhofer/deterministic-neural-networks-using-pytorch\">my kernel</a>. </p>\n\n<p>It is really questionable whether position-dependent bias makes sense. In the popular Keras attention script it is implemented the same way though:</p>\n\n<pre><code>  if self.bias:\n          self.b = self.add_weight((input_shape[1],),\n                                   initializer='zero',\n                                   name='{}_b'.format(self.name),\n                                   regularizer=self.b_regularizer,\n                                   constraint=self.b_constraint)\n</code></pre>\n\n<p>With a quick google search I was not able to find out whether bias is made position dependent in the original paper. Do you know about that?</p>\n\n<p>Now the second issue is certainly a mistake I made when writing the kernel. Sorry to everyone whose kernel this code spread to. I fixed this in a new version for people who might use my kernel as a reference later by doing:</p>\n\n<pre><code>h_embedding = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 2)), 2)\n</code></pre>\n\n<p>Is that correct?</p>\n\n<p>Thanks for pointing out the mistakes!</p>",
      "votes": 11,
      "replies": [
        {
          "id": 468365,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-02-08T19:05:02.573000",
          "content": "<p><a href=\"/bminixhofer\">@bminixhofer</a> you can refer to this implementation: <a href=\"https://github.com/TobiasLee/Text-Classification/blob/master/models/modules/attention.py\">https://github.com/TobiasLee/Text-Classification/blob/master/models/modules/attention.py</a></p>\n\n<p>It uses a bit more complicated version of attention, where each RNN output is first squashed using <code>tanh(W @ h + b)</code> transformation, but the bias is clearly shared among sequence elements and is not dependent on the sequence length.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 468371,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-02-08T19:16:59.143000",
          "content": "<p><a href=\"/bminixhofer\">@bminixhofer</a> as for your dropout implementation, to make it correct, you need additionally to permute the dimensions after doing the <code>unsqueeze()</code> operation. If you apply unsqueeze on 2nd axis, tensor will have a shape of (N, T, 1, K) and if you then apply 2d dropout it will zero entire embeddings! So you need to add <code>t.permute(0, 3, 2, 1)</code> to make it correct. After doing a dropout, you need to permute the dimensions to original order using <code>t.permute(0, 3, 2, 1)</code> and do the <code>squeeze()</code>.</p>\n\n<p>So the entire operation will look something like this:</p>\n\n<p><code>\nembeddings = ...\nembeddings = embeddings.unsqueeze(2)    # (N, T, 1, K)\nembeddings = embeddings.permute(0, 3, 2, 1)  # (N, K, 1, T)\nembeddings = self.dropout2d(embeddings)  # (N, K, 1, T), some features are masked\nembeddings = embeddings.permute(0, 3, 2, 1)  # (N, T, 1, K)\nembeddings = embeddings.squeeze(2)  # (N, T, K)\n</code></p>\n\n<p>Looks a bit cumbersome, but you can put it into a separate module to make your main <code>forward()</code> function to be more clear.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 468392,
          "author_name": "redshellspy",
          "author_url": "",
          "post_date": "2019-02-08T20:05:57.667000",
          "content": "<p>suppose if we don't use permute to change the shape, then we are effectively dropping some words in the total sentence.  </p>\n\n<p>Am I right Dmitriy ?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 468437,
          "author_name": "Benjamin Minixhofer",
          "author_url": "",
          "post_date": "2019-02-08T21:25:40.077000",
          "content": "<p>Ok, that makes sense <a href=\"/ddanevskyi\">@ddanevskyi</a>. I hope spatial dropout gets added to the standard PyTorch library. And yes, without permutation all features of some positions would be dropped <a href=\"/suchith0312\">@suchith0312</a>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 469365,
          "author_name": "Yu-ray Li",
          "author_url": "",
          "post_date": "2019-02-11T03:55:13.397000",
          "content": "<p>Hi, here is a shorter code with the same concept</p>\n\n<pre><code>emb = ...\nemb = self.dropout2d(emb.transpose(1,2).unsqueeze(-1)).squeeze().transpose(1,2)\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 469483,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-02-11T09:57:29.543000",
          "content": "<p><a href=\"/radream\">@radream</a> right!  I used a more explicit version mainly to be able to write comments</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 470043,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2019-02-12T09:35:46.977000",
          "content": "<p>Here is the kernel with same implementation. Drastically improved CV 0.6792 to .6842\n<a href=\"https://www.kaggle.com/mlwhiz/third-place-model-for-toxic-spatial-dropout\">https://www.kaggle.com/mlwhiz/third-place-model-for-toxic-spatial-dropout</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 470052,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-02-12T10:15:18.753000",
          "content": "<p><a href=\"/mlwhiz\">@mlwhiz</a> cool!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 471372,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-02-14T11:10:25.147000",
          "content": "<p><a href=\"/bminixhofer\">@bminixhofer</a>, thanks for your contributions in this competition. Your pytorch kernel and the pytorch starter by <a href=\"/hung96ad\">@hung96ad</a> (thanks Hung) helped me start using pytorch in this contest. My pytorch model did not score as well as my keras model but I used it as a backup. So I am sad to not see both of you on the stage-2 LB, you easily could have jumped several places up. Anyway, see you in the next one.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 471441,
          "author_name": "Benjamin Minixhofer",
          "author_url": "",
          "post_date": "2019-02-14T13:19:02.657000",
          "content": "<p>You're welcome! And congratulations on your placement. I grew frustrated with this competition because I was not able to improve my score much further than the public kernels. It's amazing how many unique approaches people found for a competition with only text input and no external data.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 471472,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-02-14T14:11:01.267000",
          "content": "<p>Thanks <a href=\"/bminixhofer\">@bminixhofer</a>,  I know what you mean. I felt just as frustrated but learning pytorch kept me going for a bit. In fact my best model that placed me at my current position was a submission from 2 months ago. So it is really interesting to see other's creatvity and I would surely use some in the future.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 471930,
          "author_name": "Hung The Nguyen",
          "author_url": "",
          "post_date": "2019-02-15T05:09:56.617000",
          "content": "<p>thank <a href=\"/sheriytm\">@sheriytm</a> and congratulations on your placement. I tried with many Kernel but my score isn't able to improve. Howerver, I know many approaches from this competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 471936,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-02-15T05:38:35",
          "content": "<p>Thanks <a href=\"/hung96ad\">@hung96ad</a> . I feel the same way about approaches learnt here will be useful for future work.  I am sure I will be revisiting some of these interesting submissions for inspiration just like I have referred to the highly optimized kernels in the Mercari Price competition a few times.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 468391,
      "author_name": "bilal2vec",
      "author_url": "",
      "post_date": "2019-02-08T20:04:41.100000",
      "content": "<p>I used this alternative version of attention from <a href=\"https://discuss.pytorch.org/t/self-attention-on-words-and-masking/5671/4\">https://discuss.pytorch.org/t/self-attention-on-words-and-masking/5671/4</a>. I changed the code a bit to work with pytorch 1.0. This version doesn't use a bias term, but in my opinion, setting a maximum token length of 70 shouldn't decrease the performance of the model by too much. This version however, does have the problem of not masking before running softmax over the the attention weights. Allennlp has a masked_log_softmax function <a href=\"https://github.com/allenai/allennlp/blob/b6cc9d39651273e8ec2a7e334908ffa9de5c2026/allennlp/nn/util.py#L272-L303\">https://github.com/allenai/allennlp/blob/b6cc9d39651273e8ec2a7e334908ffa9de5c2026/allennlp/nn/util.py#L272-L303</a> that correctly masks the input sequence before applying softmax</p>\n\n<pre><code>class SelfAttention(nn.Module):\ndef __init__(self, hidden_size, batch_first=False):\n    super(SelfAttention, self).__init__()\n\n    self.hidden_size = hidden_size\n    self.batch_first = batch_first\n\n    self.att_weights = nn.Parameter(torch.Tensor(1, hidden_size), requires_grad=True)\n    nn.init.xavier_uniform_(self.att_weights.data)\n\ndef get_mask(self):\n    pass\n\ndef forward(self, inputs, lengths):\n    if self.batch_first:\n        batch_size, max_len = inputs.size()[:2]\n    else:\n        max_len, batch_size = inputs.size()[:2]\n\n    # apply attention layer\n    weights = torch.bmm(inputs,\n                        self.att_weights  # (1, hidden_size)\n                        .permute(1, 0)  # (hidden_size, 1)\n                        .unsqueeze(0)  # (1, hidden_size, 1)\n                        .repeat(batch_size, 1, 1) # (batch_size, hidden_size, 1)\n                        )\n\n    attentions = torch.softmax(F.relu(weights.squeeze()), dim=-1)\n\n    # create mask based on the sentence lengths\n    mask = torch.ones(attentions.size(), requires_grad=True).cuda()\n    for i, l in enumerate(lengths):  # skip the first sentence\n        if l &lt; max_len:\n            mask[i, l:] = 0\n\n    # apply mask and renormalize attention scores (weights)\n    masked = attentions * mask\n    _sums = masked.sum(-1).unsqueeze(-1)  # sums per row\n\n    attentions = masked.div(_sums)\n\n    # apply attention weights\n    weighted = torch.mul(inputs, attentions.unsqueeze(-1).expand_as(inputs))\n\n    # get the final fixed vector representations of the sentences\n    representations = weighted.sum(1).squeeze()\n\n    return representations, attentions\n</code></pre>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 468544,
      "author_name": "Prashant Kikani",
      "author_url": "",
      "post_date": "2019-02-09T05:28:15.340000",
      "content": "<p>I rerun my 2 submissions, with that dropout correction &amp; guess what, my CV improved by 0.004. Wish I correct it earlier..</p>",
      "votes": 4,
      "replies": [
        {
          "id": 468792,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2019-02-09T17:12:37.027000",
          "content": "<p>For me also it was 0.0015 on CV score. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 471993,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-02-15T07:30:44.263000",
      "content": "<p><a href=\"/ddanevskyi\">@ddanevskyi</a>, as an experiment I used late submission to check the result of a modified copy of <a href=\"/bminixhofer\">@bminixhofer</a>'s kernel and one with his current update in the dropout mentioned. See the result below:-</p>\n\n<ul>\n<li>My modified copy of the kernel scored </li>\n</ul>\n\n<p>private LB = 0.70299</p>\n\n<p>public LB = 0.69834</p>\n\n<ul>\n<li>My same kernel with the latest corrections for the dropout.</li>\n</ul>\n\n<p>private LB = 0.70220</p>\n\n<p>public LB = 0.69395</p>",
      "votes": 1,
      "replies": [
        {
          "id": 472135,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-02-15T12:08:15.457000",
          "content": "<p>Well, I just pointed out that the way dropout is used in public kernels is incorrect. Albeit some people reported major improvements from switching to the corrected version, it might not work for you out of the box for several reasons. First, \"correct\" dropout increases the batch size (since it doesn't drop batch elements), so you might need to tune your learning rate schedule or use a lower number of training epochs. Second, \"correct\" dropout <em>does</em> use dropout (haha), so you might need to do some additional hyperparameter tuning to make it work better.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 472155,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-02-15T12:31:47.587000",
          "content": "<p>I agree with your earlier explanation about it being wrong. I was just curious as to how much did it mess with the model. Hence surprised to see the results. Thought others may want to know since <a href=\"/bminixhofer\">@bminixhofer</a>'s kernel does not show scores because he missed stage-2. I forgot to add above that the corrections improved the local CV as well. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 472316,
          "author_name": "Smilence",
          "author_url": "",
          "post_date": "2019-02-15T17:24:49.217000",
          "content": "<p>I think you need to retune the model when you changed the spatialdropout. It make the more harder to overfit.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 469983,
      "author_name": "Chintankumar Upadhyay",
      "author_url": "",
      "post_date": "2019-02-12T07:03:33.993000",
      "content": "<p>very good observation :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 468778,
      "author_name": "Max Schumacher",
      "author_url": "",
      "post_date": "2019-02-09T16:30:57.037000",
      "content": "<p>Noooooooooo 😱</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 468330,
      "author_name": "redshellspy",
      "author_url": "",
      "post_date": "2019-02-08T17:23:23.777000",
      "content": "<p>Thanks for the post. I did copy the code of dropout and my cv score decreased, so didn't use it.</p>\n\n<p>Edit: I started to read more about spatial dropout. I didn't understand somethings quite well. (I may be wrong but this is my view)</p>\n\n<p>So in convolution layer spatial dropout zeros out the entire channel. So if input as 64 channels we randomly zero out some channels out of 64. </p>\n\n<p>Now in case of text, we push the embeddings to in each time step to RNN/LSTM. So here we randomly zero out some of the time steps out of total T steps(here maybe I am wrong). </p>\n\n<p>In pytorch, the nn.dropout2d has channels in 1st axis(if we start counting axis from 0), then we have to shape the tensor into (N, T, 1, K)(because time steps will be channels here) but you wrote (N, K, 1, T). </p>",
      "votes": 1,
      "replies": [
        {
          "id": 468366,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-02-08T19:08:36.250000",
          "content": "<p><a href=\"/suchith0312\">@suchith0312</a> It has to be (N, K, 1, T). We clearly want to zero <em>features</em>, not timesteps. You can think of this tensor as a batch of images, having K channels, the width of T and the height of 1.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 468389,
          "author_name": "redshellspy",
          "author_url": "",
          "post_date": "2019-02-08T20:02:57.883000",
          "content": "<p>Thanks for the reply.  I think you are right as I have been checking keras spatial dropout implementation and they use a noise shape of <code>(batch_size,1, features)</code>. So eventually they are dropping the one dimension of the embedding for all time steps.</p>\n\n<p>But I am not getting the feel of why spatial dropout works in some cases?</p>\n\n<p>Suppose we have a sequence length of 3 and embedding dimension of 6.</p>\n\n<p>Now if the shape is <code>(N,6,1,3)</code> and dropout2d on that will make to drop some out of 6 channels present.\nLet's suppose we have input as <code>[1,2,3,4,5,6] , [7,8,9,10,11,12] , [-1,-2,-3,-4,-5,-6]</code> (embedding of some word), then after applying dropout we can get <code>[1,2,0,0,5,6] , [7,8,0,0,11,12] , [-1,-2,0,0,-5,-6]</code>. This can be embedding of some other word. So by dropping the features, we are changing the word. </p>\n\n<p>So it makes sense to drop some of the words as in a sentence, as these words can be noisy and the meaning of sentence really doesn't depend on these words.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 468317,
      "author_name": "Smilence",
      "author_url": "",
      "post_date": "2019-02-08T17:03:14.830000",
      "content": "<p>thanks for pointing out the correct usage of dropout! I was so curious why my pytorch and keras giving me very different convergence speed - my guess was the spatial dropout part but didn't pay too much attention and ignored the error = =|||</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 468294,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2019-02-08T16:26:52.623000",
      "content": "<p>Thanks for pointing out the correct usage of the dropout. I saw that in many kernels and was pretty confident what they were doing wasnt correct but didnt know enough about pytorch to say for sure. Glad we stuck with keras throughout</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 469302,
      "author_name": "Artyom Palvelev",
      "author_url": "",
      "post_date": "2019-02-10T23:02:52.523000",
      "content": "<p>Haha, that's fascinating! I used hyperopt to find the best possible hyperparameters and it has suggested that I remove embedding dropout completely. </p>\n\n<p>However, this \"dropout\" essentially means sampling of random 90% of samples on every epoch, so it's (sort of) regularization.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 469321,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-02-11T00:52:54.957000",
          "content": "<p><a href=\"/artyomp\">@artyomp</a>, yeah but the point is that it was not working as intended :-)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 469530,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-11T11:30:17.243000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 469009,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-02-10T09:44:51.100000",
      "content": "<p>Thanks <a href=\"/ddanevskyi\">@ddanevskyi</a> for pointing these out. I also noticed  one of the issues but was not quite sure if it was really wrong. My pytorch skills is not as defined as my keras hence I hesitated, I wish I asked anyway.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 471983,
      "author_name": "Kostya Kravchenko",
      "author_url": "",
      "post_date": "2019-02-15T07:16:39.673000",
      "content": "<p>Dmitriy, very good observations and thank you for sharing! But what would be the difference of using regular dropout <code>self.embedding_dropout = nn.Dropout(0.1)</code> with \n<code>h = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 0)))</code> afterwards?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 472127,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-02-15T11:57:40.743000",
          "content": "<p>Regular dropout drops elements independently, while \"spatial\" or \"locked\" dropout applies the same mask for all sequence steps. It was demonstrated that the latter performs better for RNN models. See <a href=\"https://arxiv.org/abs/1512.05287\">https://arxiv.org/abs/1512.05287</a> for details.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 472165,
          "author_name": "Kostya Kravchenko",
          "author_url": "",
          "post_date": "2019-02-15T12:43:25.690000",
          "content": "<p>Thanks again! Well, regular dropout worked better than the one with the mistake you've pointed out. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 469181,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-10T16:55:21.730000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 468718,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-09T14:32:45.413000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 470811,
      "author_name": "mtodisco10",
      "author_url": "",
      "post_date": "2019-02-13T16:05:08.313000",
      "content": "<p>Thanks for sharing Dmitriy!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 469712,
      "author_name": "Jalen Fu",
      "author_url": "",
      "post_date": "2019-02-11T17:58:58.650000",
      "content": "<p>Thanks~</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "468241": "Now, when this crazy competition is (almost) over, it's a perfect time to rest and do some analysis :)\n\nIn this topic, I want to spot some common mistakes made in public kernels created during this competition. Some of these mistakes were repeated so many times that they almost stopped to look queerly :)\n\nNote that I created this topic for educational purposes only, not to make kernels authors look weird or stupid. We all make occasional mistakes from time to time, especially when we don't have much time to do the proper code check.\n\n**1. Position-dependent bias**\n\nMany pytorch kernels use `Attention` class that has the following line:\n\n`self.b = nn.Parameter(torch.zeros(step_dim))`\n\n`step_dim` is chosen to be 70, which is the length to which all questions are either padded or truncated.\n\nThus, the bias term has the individual component for each sequence step. This is weird, at least because different questions have a different number of tokens. What if you need to process a question which is longer than 70 tokens? It's clear that a model should not be constrained to work only for questions of some particular length. Another argument is that each bias component is \"tied\" to a specific token position in a sentence, but an absolute position rarely has something to do with the token meaning or any other linguistic property.\n\nA more reasonable way to use a bias term is to have a single scalar bias that is shared among all sequence elements. \n\n**2. \"Spatial\" dropout**\n\nThe same pytorch code has another suspicious component: the use of 2d dropout. It's used since pytorch doesn't provide a proper implementation of 1d dropout (that zeros out whole channels of a sequence of 1d vectors). So one option is to use a 2d version on unsqueezed embeddings. However, the implementation used in public kernels is completely wrong :)\n\n2d dropout layer is initialized as follows:\n\n`self.embedding_dropout = nn.Dropout2d(0.1)`\n\nand then used like this:\n\n`h = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 0)))`\n\nLet's see what is going on here. First, you have a batch of embeddings, say it has a shape of `(N, T, K)`. `N` is the number of elements in a batch, `T` is the sequence length and `K` is the feature (embedding) dimension. If we apply `unsqueeze()` operation on the `0` axis now, the tensor now has a shape of `(1, N, T, K)`. 2D dropout zeros random channels (which is the second axis in pytorch). Oh, wait... The second axis is the batch dimension! In this case, 2D dropout will zero several batch elements, not channels. \n\nA proper way to apply 1d dropout, in this case, is to reshape the tensor into `(N, K, 1, T)` tensor, apply the 2d dropout and then make it have `(N, T, K)` shape again.\n\nPlease contribute if you also noticed any wrong use of common operations!\n\nGood luck and happy kaggling! :0",
    "468329": "Both of these issues originated in [my kernel](https://www.kaggle.com/bminixhofer/deterministic-neural-networks-using-pytorch). \n\nIt is really questionable whether position-dependent bias makes sense. In the popular Keras attention script it is implemented the same way though:\n\n      if self.bias:\n              self.b = self.add_weight((input_shape[1],),\n                                       initializer='zero',\n                                       name='{}_b'.format(self.name),\n                                       regularizer=self.b_regularizer,\n                                       constraint=self.b_constraint)\n\nWith a quick google search I was not able to find out whether bias is made position dependent in the original paper. Do you know about that?\n\nNow the second issue is certainly a mistake I made when writing the kernel. Sorry to everyone whose kernel this code spread to. I fixed this in a new version for people who might use my kernel as a reference later by doing:\n\n    h_embedding = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 2)), 2)\n\nIs that correct?\n\nThanks for pointing out the mistakes!",
    "468391": "I used this alternative version of attention from https://discuss.pytorch.org/t/self-attention-on-words-and-masking/5671/4. I changed the code a bit to work with pytorch 1.0. This version doesn't use a bias term, but in my opinion, setting a maximum token length of 70 shouldn't decrease the performance of the model by too much. This version however, does have the problem of not masking before running softmax over the the attention weights. Allennlp has a masked_log_softmax function https://github.com/allenai/allennlp/blob/b6cc9d39651273e8ec2a7e334908ffa9de5c2026/allennlp/nn/util.py#L272-L303 that correctly masks the input sequence before applying softmax\n\n    class SelfAttention(nn.Module):\n    def __init__(self, hidden_size, batch_first=False):\n        super(SelfAttention, self).__init__()\n\n        self.hidden_size = hidden_size\n        self.batch_first = batch_first\n\n        self.att_weights = nn.Parameter(torch.Tensor(1, hidden_size), requires_grad=True)\n        nn.init.xavier_uniform_(self.att_weights.data)\n\n    def get_mask(self):\n        pass\n\n    def forward(self, inputs, lengths):\n        if self.batch_first:\n            batch_size, max_len = inputs.size()[:2]\n        else:\n            max_len, batch_size = inputs.size()[:2]\n            \n        # apply attention layer\n        weights = torch.bmm(inputs,\n                            self.att_weights  # (1, hidden_size)\n                            .permute(1, 0)  # (hidden_size, 1)\n                            .unsqueeze(0)  # (1, hidden_size, 1)\n                            .repeat(batch_size, 1, 1) # (batch_size, hidden_size, 1)\n                            )\n    \n        attentions = torch.softmax(F.relu(weights.squeeze()), dim=-1)\n        \n        # create mask based on the sentence lengths\n        mask = torch.ones(attentions.size(), requires_grad=True).cuda()\n        for i, l in enumerate(lengths):  # skip the first sentence\n            if l &lt; max_len:\n                mask[i, l:] = 0\n\n        # apply mask and renormalize attention scores (weights)\n        masked = attentions * mask\n        _sums = masked.sum(-1).unsqueeze(-1)  # sums per row\n        \n        attentions = masked.div(_sums)\n\n        # apply attention weights\n        weighted = torch.mul(inputs, attentions.unsqueeze(-1).expand_as(inputs))\n\n        # get the final fixed vector representations of the sentences\n        representations = weighted.sum(1).squeeze()\n\n        return representations, attentions",
    "468544": "I rerun my 2 submissions, with that dropout correction &amp; guess what, my CV improved by 0.004. Wish I correct it earlier..",
    "471993": "@ddanevskyi, as an experiment I used late submission to check the result of a modified copy of @bminixhofer's kernel and one with his current update in the dropout mentioned. See the result below:-\n\n- My modified copy of the kernel scored \n\nprivate LB = 0.70299\n\npublic LB = 0.69834\n\n- My same kernel with the latest corrections for the dropout.\n\nprivate LB = 0.70220\n\npublic LB = 0.69395",
    "469983": "very good observation :)",
    "468778": "Noooooooooo 😱",
    "468330": "Thanks for the post. I did copy the code of dropout and my cv score decreased, so didn't use it.\n\nEdit: I started to read more about spatial dropout. I didn't understand somethings quite well. (I may be wrong but this is my view)\n\nSo in convolution layer spatial dropout zeros out the entire channel. So if input as 64 channels we randomly zero out some channels out of 64. \n\nNow in case of text, we push the embeddings to in each time step to RNN/LSTM. So here we randomly zero out some of the time steps out of total T steps(here maybe I am wrong). \n\nIn pytorch, the nn.dropout2d has channels in 1st axis(if we start counting axis from 0), then we have to shape the tensor into (N, T, 1, K)(because time steps will be channels here) but you wrote (N, K, 1, T). ",
    "468317": "thanks for pointing out the correct usage of dropout! I was so curious why my pytorch and keras giving me very different convergence speed - my guess was the spatial dropout part but didn't pay too much attention and ignored the error = =|||",
    "468294": "Thanks for pointing out the correct usage of the dropout. I saw that in many kernels and was pretty confident what they were doing wasnt correct but didnt know enough about pytorch to say for sure. Glad we stuck with keras throughout",
    "469302": "Haha, that's fascinating! I used hyperopt to find the best possible hyperparameters and it has suggested that I remove embedding dropout completely. \n\nHowever, this \"dropout\" essentially means sampling of random 90% of samples on every epoch, so it's (sort of) regularization.",
    "469009": "Thanks @ddanevskyi for pointing these out. I also noticed  one of the issues but was not quite sure if it was really wrong. My pytorch skills is not as defined as my keras hence I hesitated, I wish I asked anyway.",
    "471983": "Dmitriy, very good observations and thank you for sharing! But what would be the difference of using regular dropout `self.embedding_dropout = nn.Dropout(0.1)` with \n`h = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 0)))` afterwards?",
    "469181": "",
    "468718": "",
    "470811": "Thanks for sharing Dmitriy!",
    "469712": "Thanks~"
  }
}