{
  "id": 206620,
  "title": "Handling tasks with multiple questions",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206620",
  "author_name": "Rodolphe Lampe",
  "post_date": "2020-12-25T14:53:38.519000",
  "votes": 12,
  "comment_count": 23,
  "views": 0,
  "content": "<p>How did you handled tasks with many questions ?<br>\nI changed the attention masks so that it doesn't look at questions in the same task_container_id.</p>\n<p>The code is </p>\n<pre><code>from einops import repeat, rearrange\na = repeat(tasks, 'batch length -&gt; batch length repeat', repeat=length)\n\n# We then create a tensor where all rows are T, T1, ..., T_{L-1} where T is\n# just a \"token\" equal to -1 here (it needs to be different of all others T_j)\nb = torch.roll(tasks, 1, dims=1)\nassert (tasks[:, :-1] == b[:, 1:]).all()\nb[:, 0] = -1\nb = repeat(b, 'batch length -&gt; batch repeat length', repeat=length)\n\nd = repeat((a!=b), 'batch length1 length2 -&gt; repeat batch length1 length2', repeat=self.hparams.nhead)\nd = rearrange(d, 'nhead batch l1 l2 -&gt; (nhead batch) l1 l2')\n\n# Now, the tensor a != b gives all the logic about a position being compatible\n# relative to the tasks but we need to add the \"future is not lookable\" logic\nc = (torch.triu(torch.ones(length, length)) == 1).transpose(0, 1).to(self.device)\nc = repeat(c, 'length1 length2 -&gt; repeat length1 length2', repeat=batch_size*self.hparams.nhead)\n\nmask = d &amp; c\nreturn mask.float().masked_fill(mask == 0, float('-inf')).masked_fill(mask == 1, float(0.0))\n</code></pre>",
  "messages": [
    {
      "id": 1126370,
      "postDate": "2020-12-25T14:53:38.520Z",
      "content": "<p>How did you handled tasks with many questions ?<br>\nI changed the attention masks so that it doesn't look at questions in the same task_container_id.</p>\n<p>The code is </p>\n<pre><code>from einops import repeat, rearrange\na = repeat(tasks, 'batch length -&gt; batch length repeat', repeat=length)\n\n# We then create a tensor where all rows are T, T1, ..., T_{L-1} where T is\n# just a \"token\" equal to -1 here (it needs to be different of all others T_j)\nb = torch.roll(tasks, 1, dims=1)\nassert (tasks[:, :-1] == b[:, 1:]).all()\nb[:, 0] = -1\nb = repeat(b, 'batch length -&gt; batch repeat length', repeat=length)\n\nd = repeat((a!=b), 'batch length1 length2 -&gt; repeat batch length1 length2', repeat=self.hparams.nhead)\nd = rearrange(d, 'nhead batch l1 l2 -&gt; (nhead batch) l1 l2')\n\n# Now, the tensor a != b gives all the logic about a position being compatible\n# relative to the tasks but we need to add the \"future is not lookable\" logic\nc = (torch.triu(torch.ones(length, length)) == 1).transpose(0, 1).to(self.device)\nc = repeat(c, 'length1 length2 -&gt; repeat length1 length2', repeat=batch_size*self.hparams.nhead)\n\nmask = d &amp; c\nreturn mask.float().masked_fill(mask == 0, float('-inf')).masked_fill(mask == 1, float(0.0))\n</code></pre>",
      "rawMarkdown": "How did you handled tasks with many questions ?\nI changed the attention masks so that it doesn't look at questions in the same task_container_id.\n\nThe code is \n\n```\nfrom einops import repeat, rearrange\na = repeat(tasks, 'batch length -> batch length repeat', repeat=length)\n\n# We then create a tensor where all rows are T, T1, ..., T_{L-1} where T is\n# just a \"token\" equal to -1 here (it needs to be different of all others T_j)\nb = torch.roll(tasks, 1, dims=1)\nassert (tasks[:, :-1] == b[:, 1:]).all()\nb[:, 0] = -1\nb = repeat(b, 'batch length -> batch repeat length', repeat=length)\n\nd = repeat((a!=b), 'batch length1 length2 -> repeat batch length1 length2', repeat=self.hparams.nhead)\nd = rearrange(d, 'nhead batch l1 l2 -> (nhead batch) l1 l2')\n\n# Now, the tensor a != b gives all the logic about a position being compatible\n# relative to the tasks but we need to add the \"future is not lookable\" logic\nc = (torch.triu(torch.ones(length, length)) == 1).transpose(0, 1).to(self.device)\nc = repeat(c, 'length1 length2 -> repeat length1 length2', repeat=batch_size*self.hparams.nhead)\n\nmask = d & c\nreturn mask.float().masked_fill(mask == 0, float('-inf')).masked_fill(mask == 1, float(0.0))\n```",
      "votes": 12
    },
    {
      "id": 1132108,
      "postDate": "2020-12-30T07:08:02.543Z",
      "content": "<p>Great post, I'm doing similar thing, because data leakage happens without such mask.</p>",
      "rawMarkdown": "Great post, I'm doing similar thing, because data leakage happens without such mask.",
      "votes": 1
    },
    {
      "id": 1126852,
      "postDate": "2020-12-26T03:16:19.183Z",
      "content": "<p>Very cool, thanks for sharing. I had resolved to just use a triangle mask knowing that in prod the performance would degrade a bit. Dunno why I hadn't thought of this 👍🏾. Thank you.</p>",
      "rawMarkdown": "Very cool, thanks for sharing. I had resolved to just use a triangle mask knowing that in prod the performance would degrade a bit. Dunno why I hadn't thought of this 👍🏾. Thank you.",
      "votes": 1,
      "replies": [
        {
          "id": 1131997,
          "postDate": "2020-12-30T05:28:11.040Z",
          "content": "<p>Getting around to merging this in. So <code>tasks</code> is supposed to be [B, SeqLen] holding the container_task_ids?</p>",
          "rawMarkdown": "Getting around to merging this in. So `tasks` is supposed to be [B, SeqLen] holding the container_task_ids?"
        },
        {
          "id": 1133005,
          "postDate": "2020-12-30T21:10:16.520Z",
          "content": "<p>I got this working in my build by making the following changes:</p>\n<pre><code>c = torch.ones([seq_len, seq_len], device=task_container_id.device, dtype=torch.uint8)\nc = c.triu_(1).view(seq_len, seq_len).bool()\nfuture_mask = (c|~d).bool()\n</code></pre>\n<p>This returns a BoolTensor that can directly be fed into MultiHeadAttention layer's <code>forward</code> call, per its <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html?highlight=multiheadattention#torch.nn.MultiheadAttention\" target=\"_blank\">documentation</a>.</p>",
          "rawMarkdown": "I got this working in my build by making the following changes:\n\n```\nc = torch.ones([seq_len, seq_len], device=task_container_id.device, dtype=torch.uint8)\nc = c.triu_(1).view(seq_len, seq_len).bool()\nfuture_mask = (c|~d).bool()\n```\n\nThis returns a BoolTensor that can directly be fed into MultiHeadAttention layer's `forward` call, per its [documentation](https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html?highlight=multiheadattention#torch.nn.MultiheadAttention)."
        },
        {
          "id": 1133302,
          "postDate": "2020-12-31T05:18:10.273Z",
          "content": "<p>it returns nan values as o/p instead of a 0-1 probability<br>\nWhen we have both padding mask as well as peak ahead mask, then all weights are -inf and softmax returns nan. <br>\nUsing the popular SAKT kernel shared has this issue. There is 0 padding everywhere - on encoder as well as decoder. Combine this with peek ahead mask and we could see that for certain case, everything is masked out, resulting in above error. This also messes gradients big time.<br>\nOne solution is to use custom attention layer and catch these instead of depending on the above Torch function. <br>\nInterestingly the public kernel performs well in spite of all the 0 paddings and non-masking of those. Probably the model just learns to ignore these zeros. </p>",
          "rawMarkdown": "it returns nan values as o/p instead of a 0-1 probability\nWhen we have both padding mask as well as peak ahead mask, then all weights are -inf and softmax returns nan. \nUsing the popular SAKT kernel shared has this issue. There is 0 padding everywhere - on encoder as well as decoder. Combine this with peek ahead mask and we could see that for certain case, everything is masked out, resulting in above error. This also messes gradients big time.\nOne solution is to use custom attention layer and catch these instead of depending on the above Torch function. \nInterestingly the public kernel performs well in spite of all the 0 paddings and non-masking of those. Probably the model just learns to ignore these zeros. "
        },
        {
          "id": 1133555,
          "postDate": "2020-12-31T09:57:43.353Z",
          "content": "<p>Why do you use (c|~d) ? The contribution of ~d means you're adding questions to look at instead of removing questions. I don't understand.</p>\n<p>Indeed, there is a problem when the first task contains multiple questions because we're then deactivating all available questions … I don't see how this case should be handled</p>",
          "rawMarkdown": "Why do you use (c|~d) ? The contribution of ~d means you're adding questions to look at instead of removing questions. I don't understand.\n\nIndeed, there is a problem when the first task contains multiple questions because we're then deactivating all available questions ... I don't see how this case should be handled"
        },
        {
          "id": 1134485,
          "postDate": "2021-01-01T10:38:11.137Z",
          "content": "<p>What would be the final mask if we have task_container_id (5, 7) : (batch size=5, seq_len=7) ?</p>\n<pre><code># Input (task_container_id) 5x7\ntensor([[ 108, 109, 110, 110, 110, 110, 111],\n        [ 20,  20,  20,  21, 22, 22, 22],\n        [ 13,  14,  15,  16,  17,  18,  19],\n        [ 10,  11,  11,  12,  14,  13,  14], # Should not happen\n        [  0,   1,   1,  2,  2,  3,  3]])\n\n# Output (True = Ignore) 5x7\ntensor([[False, False, False,  True,  True,  True, False],\n        [False,  True,  True, False, False, False,  True],\n        [False, False, False, False, False, False, False],\n        [False, False,  True, False, False, False, False],\n        [False, False,  True, False,  True, False,  True]])\n</code></pre>",
          "rawMarkdown": "What would be the final mask if we have task_container_id (5, 7) : (batch size=5, seq_len=7) ?\n\n```\n# Input (task_container_id) 5x7\ntensor([[ 108, 109, 110, 110, 110, 110, 111],\n        [ 20,  20,  20,  21, 22, 22, 22],\n        [ 13,  14,  15,  16,  17,  18,  19],\n        [ 10,  11,  11,  12,  14,  13,  14], # Should not happen\n        [  0,   1,   1,  2,  2,  3,  3]])\n\n# Output (True = Ignore) 5x7\ntensor([[False, False, False,  True,  True,  True, False],\n        [False,  True,  True, False, False, False,  True],\n        [False, False, False, False, False, False, False],\n        [False, False,  True, False, False, False, False],\n        [False, False,  True, False,  True, False,  True]])\n```\n\n"
        },
        {
          "id": 1135332,
          "postDate": "2021-01-02T06:20:19.437Z",
          "content": "<p>I think another way to avoid nan when using both padding mask and attention mask is to pad on the right side instead of the left side.</p>",
          "rawMarkdown": "I think another way to avoid nan when using both padding mask and attention mask is to pad on the right side instead of the left side."
        },
        {
          "id": 1135373,
          "postDate": "2021-01-02T07:49:05.727Z",
          "content": "<p>I would like to share my implementation here</p>\n<pre><code>def task_mask(tasks):\n    seq_length=len(tasks)\n    future_mask = np.triu(np.ones((seq_length, seq_length)), k=1).astype('bool')\n    container_mask= np.ones((seq_length, seq_length))\n    container_mask=(container_mask*tasks.reshape(1,-1))==(container_mask*tasks.reshape(-1,1))\n    future_mask=future_mask+container_mask\n    np.fill_diagonal(future_mask,0)\n    return future_mask\n</code></pre>\n<p>I think it's correct and it decreases CV compared to using triu mask. Havent checked effect on lb tho</p>",
          "rawMarkdown": "I would like to share my implementation here\n```\ndef task_mask(tasks):\n    seq_length=len(tasks)\n    future_mask = np.triu(np.ones((seq_length, seq_length)), k=1).astype('bool')\n    container_mask= np.ones((seq_length, seq_length))\n    container_mask=(container_mask*tasks.reshape(1,-1))==(container_mask*tasks.reshape(-1,1))\n    future_mask=future_mask+container_mask\n    np.fill_diagonal(future_mask,0)\n    return future_mask\n\n```\n\nI think it's correct and it decreases CV compared to using triu mask. Havent checked effect on lb tho",
          "votes": 4
        },
        {
          "id": 1135384,
          "postDate": "2021-01-02T07:58:17.487Z",
          "content": "<p><a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">@rodolphelampe</a> cant that issue be resolved if you simply have the diagonal set to False? The first questions in the same container wont be able to see any other questions but that is similar to the situation with the first question when u use the normal future mask and cannot be helped</p>",
          "rawMarkdown": "@rodolphelampe cant that issue be resolved if you simply have the diagonal set to False? The first questions in the same container wont be able to see any other questions but that is similar to the situation with the first question when u use the normal future mask and cannot be helped"
        },
        {
          "id": 1135399,
          "postDate": "2021-01-02T08:13:38.587Z",
          "content": "<p><a href=\"https://www.kaggle.com/Shujun\" target=\"_blank\">@Shujun</a> Thanks the following line works if <code>tasks</code> is (1, seq_len):</p>\n<p><code>container_mask = (container_mask*tasks.reshape(1,-1)) == (container_mask*tasks.reshape(-1,1))</code></p>",
          "rawMarkdown": "@Shujun Thanks the following line works if `tasks` is (1, seq_len):\n\n`container_mask = (container_mask*tasks.reshape(1,-1)) == (container_mask*tasks.reshape(-1,1))`"
        },
        {
          "id": 1135418,
          "postDate": "2021-01-02T08:30:34.220Z",
          "content": "<p>It doesnt have to be actually. The reshape(1,-1) is like an unsqueeze operation </p>",
          "rawMarkdown": "It doesnt have to be actually. The reshape(1,-1) is like an unsqueeze operation "
        },
        {
          "id": 1135442,
          "postDate": "2021-01-02T08:55:25.223Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> did you have to make that change for it to work? for me i always input task as a vector</p>",
          "rawMarkdown": "@mpware did you have to make that change for it to work? for me i always input task as a vector"
        },
        {
          "id": 1135471,
          "postDate": "2021-01-02T09:24:15.917Z",
          "content": "<p>The problem I see is that if we set the diagonal to True (ie authorize to look at it) then, in the decoder, we have access to the previous answer which is not accessible in inference (for the questions in the bundle except for the very first one)</p>",
          "rawMarkdown": "The problem I see is that if we set the diagonal to True (ie authorize to look at it) then, in the decoder, we have access to the previous answer which is not accessible in inference (for the questions in the bundle except for the very first one)"
        },
        {
          "id": 1135472,
          "postDate": "2021-01-02T09:25:33.483Z",
          "content": "<p>For Pytorch <code>nn.MultiHeadAttention</code> I think it could be used as below. However, it fails when used with padding masks enabled, whatever left/right padding, with NaN loss as focused by <a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a>. Not sure how to solve it except with a custom <code>MultiHeadAttention</code>.</p>\n<pre><code>def generate_mask(self, size, diagonal=1):        \n    return torch.triu(torch.ones(size, size)==1, diagonal=diagonal)\n\ndef tasks_mask(self, tasks, seq_length, diagonal=1):\n    future_mask = self.generate_mask(seq_length, diagonal=diagonal).to(tasks.device)\n    container_mask= torch.ones((seq_length, seq_length)).to(tasks.device)\n    container_mask=(container_mask*tasks.reshape(1,-1))==(container_mask*tasks.reshape(-1,1))\n    future_mask=future_mask+container_mask\n    future_mask = future_mask.fill_diagonal_(False)\n    return future_mask\n\ndef tasks_3d_mask(self, tasks, seq_length, diagonal=1):\n    mask_3d = [self.tasks_mask(t, seq_length, diagonal=diagonal) for t in tasks]\n    mask_3d = torch.stack(mask_3d, dim=0)\n    # Need BS*num_heads shape\n    repeat_3d = [mask_3d for t in range(self.nhead)]\n    repeat_3d = torch.cat(repeat_3d)\n    return repeat_3d\n</code></pre>",
          "rawMarkdown": "For Pytorch `nn.MultiHeadAttention` I think it could be used as below. However, it fails when used with padding masks enabled, whatever left/right padding, with NaN loss as focused by @allohvk. Not sure how to solve it except with a custom `MultiHeadAttention`.\n\n```\ndef generate_mask(self, size, diagonal=1):        \n\treturn torch.triu(torch.ones(size, size)==1, diagonal=diagonal)\n\ndef tasks_mask(self, tasks, seq_length, diagonal=1):\n\tfuture_mask = self.generate_mask(seq_length, diagonal=diagonal).to(tasks.device)\n\tcontainer_mask= torch.ones((seq_length, seq_length)).to(tasks.device)\n\tcontainer_mask=(container_mask*tasks.reshape(1,-1))==(container_mask*tasks.reshape(-1,1))\n\tfuture_mask=future_mask+container_mask\n\tfuture_mask = future_mask.fill_diagonal_(False)\n\treturn future_mask\n\ndef tasks_3d_mask(self, tasks, seq_length, diagonal=1):\n\tmask_3d = [self.tasks_mask(t, seq_length, diagonal=diagonal) for t in tasks]\n\tmask_3d = torch.stack(mask_3d, dim=0)\n\t# Need BS*num_heads shape\n\trepeat_3d = [mask_3d for t in range(self.nhead)]\n\trepeat_3d = torch.cat(repeat_3d)\n\treturn repeat_3d\n```\n",
          "votes": 2
        },
        {
          "id": 1135477,
          "postDate": "2021-01-02T09:29:26.863Z",
          "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> My input is (BS, seq_len)</p>",
          "rawMarkdown": "@shujun717 My input is (BS, seq_len)"
        },
        {
          "id": 1135488,
          "postDate": "2021-01-02T09:43:28.687Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> sorry this is supposed to be used with a single sample at a time (I use it in a pytorch Dataset). As for NAN issue, it's just one of the many technical difficulties this competition presents</p>",
          "rawMarkdown": "@mpware sorry this is supposed to be used with a single sample at a time (I use it in a pytorch Dataset). As for NAN issue, it's just one of the many technical difficulties this competition presents",
          "votes": 1
        }
      ]
    },
    {
      "id": 1132777,
      "postDate": "2020-12-30T17:11:42.910Z",
      "content": "<p>Thanks. I was running into issues with Pytorch pad mask feature: <a href=\"https://github.com/pytorch/pytorch/pull/24888\" target=\"_blank\">https://github.com/pytorch/pytorch/pull/24888</a> <br>\nAbove should help I guess..</p>",
      "rawMarkdown": "Thanks. I was running into issues with Pytorch pad mask feature: https://github.com/pytorch/pytorch/pull/24888 \nAbove should help I guess..",
      "replies": [
        {
          "id": 1134491,
          "postDate": "2021-01-01T10:50:24.487Z",
          "content": "<p>Is there an actual patch? That git issue makes it seem like no patch / expected behavior. What's the workaround?</p>",
          "rawMarkdown": "Is there an actual patch? That git issue makes it seem like no patch / expected behavior. What's the workaround?"
        },
        {
          "id": 1134561,
          "postDate": "2021-01-01T11:47:19.017Z",
          "content": "<p>unfortunately, no…<br>\nBasically we are talking of some sort of moving mask here…across each row..For certain questions, the other questions in the same bundle need to be masked. For other questions these are not masked. So as we move along the sequence the pad mask keeps changing dynamically. it is a difficult situation and existing libraries may not support.<br>\ncustom code may work</p>",
          "rawMarkdown": "unfortunately, no...\nBasically we are talking of some sort of moving mask here...across each row..For certain questions, the other questions in the same bundle need to be masked. For other questions these are not masked. So as we move along the sequence the pad mask keeps changing dynamically. it is a difficult situation and existing libraries may not support.\ncustom code may work"
        },
        {
          "id": 1135313,
          "postDate": "2021-01-02T05:55:40.650Z",
          "content": "<p>hmm..there is a workaround..there is going to be some pain coding, but in seq models this feature should make a big diff..so it is worth it<br>\nMy ID started working since yesterday and as of now I am still fine-tuning basic stuff. Once I get to a decent score, I will try n let u know</p>",
          "rawMarkdown": "hmm..there is a workaround..there is going to be some pain coding, but in seq models this feature should make a big diff..so it is worth it\nMy ID started working since yesterday and as of now I am still fine-tuning basic stuff. Once I get to a decent score, I will try n let u know"
        }
      ]
    },
    {
      "id": 1132048,
      "postDate": "2020-12-30T06:18:12.290Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1130445,
      "postDate": "2020-12-29T04:08:13.130Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 1132108,
      "author_name": "mamas",
      "author_url": "",
      "post_date": "2020-12-30T07:08:02.543000",
      "content": "<p>Great post, I'm doing similar thing, because data leakage happens without such mask.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1126852,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-12-26T03:16:19.183000",
      "content": "<p>Very cool, thanks for sharing. I had resolved to just use a triangle mask knowing that in prod the performance would degrade a bit. Dunno why I hadn't thought of this 👍🏾. Thank you.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1131997,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-30T05:28:11.040000",
          "content": "<p>Getting around to merging this in. So <code>tasks</code> is supposed to be [B, SeqLen] holding the container_task_ids?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133005,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-30T21:10:16.520000",
          "content": "<p>I got this working in my build by making the following changes:</p>\n<pre><code>c = torch.ones([seq_len, seq_len], device=task_container_id.device, dtype=torch.uint8)\nc = c.triu_(1).view(seq_len, seq_len).bool()\nfuture_mask = (c|~d).bool()\n</code></pre>\n<p>This returns a BoolTensor that can directly be fed into MultiHeadAttention layer's <code>forward</code> call, per its <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html?highlight=multiheadattention#torch.nn.MultiheadAttention\" target=\"_blank\">documentation</a>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133302,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-31T05:18:10.273000",
          "content": "<p>it returns nan values as o/p instead of a 0-1 probability<br>\nWhen we have both padding mask as well as peak ahead mask, then all weights are -inf and softmax returns nan. <br>\nUsing the popular SAKT kernel shared has this issue. There is 0 padding everywhere - on encoder as well as decoder. Combine this with peek ahead mask and we could see that for certain case, everything is masked out, resulting in above error. This also messes gradients big time.<br>\nOne solution is to use custom attention layer and catch these instead of depending on the above Torch function. <br>\nInterestingly the public kernel performs well in spite of all the 0 paddings and non-masking of those. Probably the model just learns to ignore these zeros. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133555,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2020-12-31T09:57:43.353000",
          "content": "<p>Why do you use (c|~d) ? The contribution of ~d means you're adding questions to look at instead of removing questions. I don't understand.</p>\n<p>Indeed, there is a problem when the first task contains multiple questions because we're then deactivating all available questions … I don't see how this case should be handled</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1134485,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-01T10:38:11.137000",
          "content": "<p>What would be the final mask if we have task_container_id (5, 7) : (batch size=5, seq_len=7) ?</p>\n<pre><code># Input (task_container_id) 5x7\ntensor([[ 108, 109, 110, 110, 110, 110, 111],\n        [ 20,  20,  20,  21, 22, 22, 22],\n        [ 13,  14,  15,  16,  17,  18,  19],\n        [ 10,  11,  11,  12,  14,  13,  14], # Should not happen\n        [  0,   1,   1,  2,  2,  3,  3]])\n\n# Output (True = Ignore) 5x7\ntensor([[False, False, False,  True,  True,  True, False],\n        [False,  True,  True, False, False, False,  True],\n        [False, False, False, False, False, False, False],\n        [False, False,  True, False, False, False, False],\n        [False, False,  True, False,  True, False,  True]])\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135332,
          "author_name": "Moeen Bagheri",
          "author_url": "",
          "post_date": "2021-01-02T06:20:19.437000",
          "content": "<p>I think another way to avoid nan when using both padding mask and attention mask is to pad on the right side instead of the left side.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135373,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2021-01-02T07:49:05.727000",
          "content": "<p>I would like to share my implementation here</p>\n<pre><code>def task_mask(tasks):\n    seq_length=len(tasks)\n    future_mask = np.triu(np.ones((seq_length, seq_length)), k=1).astype('bool')\n    container_mask= np.ones((seq_length, seq_length))\n    container_mask=(container_mask*tasks.reshape(1,-1))==(container_mask*tasks.reshape(-1,1))\n    future_mask=future_mask+container_mask\n    np.fill_diagonal(future_mask,0)\n    return future_mask\n</code></pre>\n<p>I think it's correct and it decreases CV compared to using triu mask. Havent checked effect on lb tho</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1135384,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2021-01-02T07:58:17.487000",
          "content": "<p><a href=\"https://www.kaggle.com/rodolphelampe\" target=\"_blank\">@rodolphelampe</a> cant that issue be resolved if you simply have the diagonal set to False? The first questions in the same container wont be able to see any other questions but that is similar to the situation with the first question when u use the normal future mask and cannot be helped</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135399,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-02T08:13:38.587000",
          "content": "<p><a href=\"https://www.kaggle.com/Shujun\" target=\"_blank\">@Shujun</a> Thanks the following line works if <code>tasks</code> is (1, seq_len):</p>\n<p><code>container_mask = (container_mask*tasks.reshape(1,-1)) == (container_mask*tasks.reshape(-1,1))</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135418,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2021-01-02T08:30:34.220000",
          "content": "<p>It doesnt have to be actually. The reshape(1,-1) is like an unsqueeze operation </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135442,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2021-01-02T08:55:25.223000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> did you have to make that change for it to work? for me i always input task as a vector</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135471,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2021-01-02T09:24:15.917000",
          "content": "<p>The problem I see is that if we set the diagonal to True (ie authorize to look at it) then, in the decoder, we have access to the previous answer which is not accessible in inference (for the questions in the bundle except for the very first one)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135472,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-02T09:25:33.483000",
          "content": "<p>For Pytorch <code>nn.MultiHeadAttention</code> I think it could be used as below. However, it fails when used with padding masks enabled, whatever left/right padding, with NaN loss as focused by <a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a>. Not sure how to solve it except with a custom <code>MultiHeadAttention</code>.</p>\n<pre><code>def generate_mask(self, size, diagonal=1):        \n    return torch.triu(torch.ones(size, size)==1, diagonal=diagonal)\n\ndef tasks_mask(self, tasks, seq_length, diagonal=1):\n    future_mask = self.generate_mask(seq_length, diagonal=diagonal).to(tasks.device)\n    container_mask= torch.ones((seq_length, seq_length)).to(tasks.device)\n    container_mask=(container_mask*tasks.reshape(1,-1))==(container_mask*tasks.reshape(-1,1))\n    future_mask=future_mask+container_mask\n    future_mask = future_mask.fill_diagonal_(False)\n    return future_mask\n\ndef tasks_3d_mask(self, tasks, seq_length, diagonal=1):\n    mask_3d = [self.tasks_mask(t, seq_length, diagonal=diagonal) for t in tasks]\n    mask_3d = torch.stack(mask_3d, dim=0)\n    # Need BS*num_heads shape\n    repeat_3d = [mask_3d for t in range(self.nhead)]\n    repeat_3d = torch.cat(repeat_3d)\n    return repeat_3d\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1135477,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-02T09:29:26.863000",
          "content": "<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> My input is (BS, seq_len)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135488,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2021-01-02T09:43:28.687000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> sorry this is supposed to be used with a single sample at a time (I use it in a pytorch Dataset). As for NAN issue, it's just one of the many technical difficulties this competition presents</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1132777,
      "author_name": "Allohvk",
      "author_url": "",
      "post_date": "2020-12-30T17:11:42.910000",
      "content": "<p>Thanks. I was running into issues with Pytorch pad mask feature: <a href=\"https://github.com/pytorch/pytorch/pull/24888\" target=\"_blank\">https://github.com/pytorch/pytorch/pull/24888</a> <br>\nAbove should help I guess..</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1134491,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2021-01-01T10:50:24.487000",
          "content": "<p>Is there an actual patch? That git issue makes it seem like no patch / expected behavior. What's the workaround?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1134561,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2021-01-01T11:47:19.017000",
          "content": "<p>unfortunately, no…<br>\nBasically we are talking of some sort of moving mask here…across each row..For certain questions, the other questions in the same bundle need to be masked. For other questions these are not masked. So as we move along the sequence the pad mask keeps changing dynamically. it is a difficult situation and existing libraries may not support.<br>\ncustom code may work</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135313,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2021-01-02T05:55:40.650000",
          "content": "<p>hmm..there is a workaround..there is going to be some pain coding, but in seq models this feature should make a big diff..so it is worth it<br>\nMy ID started working since yesterday and as of now I am still fine-tuning basic stuff. Once I get to a decent score, I will try n let u know</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1132048,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-30T06:18:12.290000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1130445,
      "author_name": "Lokesh",
      "author_url": "",
      "post_date": "2020-12-29T04:08:13.130000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1126370": "How did you handled tasks with many questions ?\nI changed the attention masks so that it doesn't look at questions in the same task_container_id.\n\nThe code is \n\n```\nfrom einops import repeat, rearrange\na = repeat(tasks, 'batch length -> batch length repeat', repeat=length)\n\n# We then create a tensor where all rows are T, T1, ..., T_{L-1} where T is\n# just a \"token\" equal to -1 here (it needs to be different of all others T_j)\nb = torch.roll(tasks, 1, dims=1)\nassert (tasks[:, :-1] == b[:, 1:]).all()\nb[:, 0] = -1\nb = repeat(b, 'batch length -> batch repeat length', repeat=length)\n\nd = repeat((a!=b), 'batch length1 length2 -> repeat batch length1 length2', repeat=self.hparams.nhead)\nd = rearrange(d, 'nhead batch l1 l2 -> (nhead batch) l1 l2')\n\n# Now, the tensor a != b gives all the logic about a position being compatible\n# relative to the tasks but we need to add the \"future is not lookable\" logic\nc = (torch.triu(torch.ones(length, length)) == 1).transpose(0, 1).to(self.device)\nc = repeat(c, 'length1 length2 -> repeat length1 length2', repeat=batch_size*self.hparams.nhead)\n\nmask = d & c\nreturn mask.float().masked_fill(mask == 0, float('-inf')).masked_fill(mask == 1, float(0.0))\n```",
    "1132108": "Great post, I'm doing similar thing, because data leakage happens without such mask.",
    "1126852": "Very cool, thanks for sharing. I had resolved to just use a triangle mask knowing that in prod the performance would degrade a bit. Dunno why I hadn't thought of this 👍🏾. Thank you.",
    "1132777": "Thanks. I was running into issues with Pytorch pad mask feature: https://github.com/pytorch/pytorch/pull/24888 \nAbove should help I guess..",
    "1132048": "",
    "1130445": "Thanks for sharing."
  }
}