{
  "id": 210113,
  "title": "2nd Place Solution (LSTM-Encoded SAKT-like TransformerEncoder)",
  "url": "/competitions/riiid-test-answer-prediction/discussion/210113",
  "author_name": "mamas",
  "post_date": "2021-01-09T18:13:42.939000",
  "votes": 207,
  "comment_count": 50,
  "views": 0,
  "content": "<p>Thank you all teams who competed with me, all the people who participated in this competition, and the organizer that hosts such a great competition with the well-designed API!<br>\nCongrats <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a>, who defeats me and becomes the winner in this competition.</p>\n<p>I'm happy because it's my first time I get solo prize!</p>\n<p>It's my 4th kaggle competition and it was fun to compete with my past teammates ( <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a>, <a href=\"https://www.kaggle.com/pocketsuteado\" target=\"_blank\">@pocketsuteado</a>) and people who I competed with past competitions (e.g. <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>, <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>). <br>\nHere, I will explain the summary of my model and my features. </p>\n<p>I uploaded 2 kaggle notebooks for the explanation. As the ensemble is not so important in my solution, I will only explain my single model. </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution\" target=\"_blank\">6 similar models weighted average</a> : 0.817 public/0.818 private</li>\n<li><a href=\"https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold\" target=\"_blank\">single model</a> : 0.814 public/0.816 private</li>\n</ol>\n<h1>Models</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F762938%2F39ec58b701e6d37701cc794763d3a473%2F2nd_place.png?generation=1610213712807228&amp;alt=media\" alt=\"\"></p>\n<h2>Overview</h2>\n<p>my model is similar to SAKT, with 400 sequence length, 512 dimension, 4 nheads. <br>\nI don't use lecture information for the input of transformer model, which I guess is why I lost in this competition. For the query and key/value of SAKT-like model, I used LSTM-encoded features, whose input is as follows.</p>\n<ul>\n<li><p><strong>\"Query\" features</strong><br>\ncontent_id<br>\npart<br>\ntags<br>\nnormalized timedelta<br>\nnormalized log timestamp <br>\ncorrect answer<br>\ntask_container_id delta<br>\ncontent_type_id delta<br>\nnormalized absolute position </p></li>\n<li><p><strong>\"Memory\" features</strong><br>\nexplanation<br>\ncorrectness<br>\nnormalized elapsed time<br>\nuser_answer</p></li>\n</ul>\n<h2>Detailed Explanation of Training/Inference Process</h2>\n<p>I tried a very precise indexing/masking technique to avoid data leakage in training process, I guess which is partly why I became 2nd place. OK, suppose a very simple model of the sequence length = 5, and a task_container_id history of a specific user (without lecture) is like this.<br>\n<code>[0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 5, 8, 7, 7, 6]</code><br>\nAs there is 15 measurements in this user, I made 15/5 = 3 training samples and loss mask.</p>\n<p>(1) input task_container_id: <code>[pad, pad, pad, pad, pad, 0, 0, 0, 1, 1]</code> <br>\n(1) loss_mask: <code>[False, False, False, False, False, True, True, True, True, True]</code><br>\n(2) input task_container_id: <code>[0, 0, 0, 1, 1, 2, 2, 3, 3, 4]</code> <br>\n(2) loss_mask: <code>[False, False, False, False, False, True, True, True, True, True]</code><br>\n(3) input task_container_id: <code>[2, 2, 3, 3, 4, 5, 8, 7, 7, 6]</code> <br>\n(3) loss_mask: <code>[False, False, False, False, False, True, True, True, True, True]</code></p>\n<p>In this competition, the handling of the task_container_id is very important, as <strong>it is not allowed to use \"memory\" features of the same task container id</strong> to avoid leakage.<br>\nSo, after applying LSTM to the features, such fancy indexing is required to avoid the leakage for 3 training samples, where -1 means this position can't attend any position.</p>\n<p>(1) indices: <code>[-1, -1, -1, -1, -1, 4, 4, 4, 7, 7]</code><br>\n(2) indices: <code>[-1, -1, -1, 2, 2, 4, 4, 6, 6, 8]</code><br>\n(3) indices: <code>[-1, -1, 1, 1, 3, 4, 5, 6, 6, 8]</code></p>\n<p>To get this indices very fast in the training/inference phase, I wrote such cython (#1) function (vectorized implementation):</p>\n<pre><code>%%cython\nimport numpy as np\ncimport numpy as np\ncpdef np.ndarray[int] cget_memory_indices(np.ndarray task):\n    cdef Py_ssize_t n = task.shape[1]\n    cdef np.ndarray[int, ndim = 2] res = np.zeros_like(task, dtype = np.int32)\n    cdef np.ndarray[int] tmp_counter = np.full(task.shape[0], -1, dtype = np.int32)\n    cdef np.ndarray[int] u_counter = np.full(task.shape[0], task.shape[1] - 1, dtype = np.int32)\n    for i in range(n):\n        res[:, i] = u_counter\n        tmp_counter += 1\n        if i != n - 1:\n            mask = (task[:, i] != task[:, i + 1])\n            u_counter[mask] = tmp_counter[mask]\n    return res\n</code></pre>\n<p>After applying such fancy indexing to the output of LSTM features, I concatenated them with \"Query\" features and apply MLP, then we can get query for SAKT model. Then, I obtained key/value for SAKT model by applying MLP to query concatenated with \"Memory\" features. </p>\n<p>To train SAKT-like model, precise memory masking is also required to avoid leakage. I used <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html\" target=\"_blank\">3D attention mask</a> of torch.nn.MultiheadAttention. It should be noted <a href=\"https://pytorch.org/docs/stable/generated/torch.repeat_interleave.html\" target=\"_blank\">torch.repeat_interleave</a> must be leveraged to make 3D mask (batchsize * nhead, sequence length, sequence length). <br>\nThe memory mask for these 3 samples are like this. Here, 1 means True and 0 means False. Please remember the <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html\" target=\"_blank\">documentation</a> says </p>\n<pre><code>attn_mask ensure that position i is allowed to attend the unmasked positions. If a BoolTensor is provided, positions with True is not allowed to attend while False values will be unchanged.\n</code></pre>\n<p>(1) memory_mask:</p>\n<pre><code>array([[1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [1, 1, 1, 0, 0, 0, 0, 0, 1, 1],\n       [1, 1, 1, 0, 0, 0, 0, 0, 1, 1]])\n</code></pre>\n<p>(2) memory_mask:</p>\n<pre><code>array([[1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [0, 0, 0, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 1, 1, 0, 0, 0, 0, 0, 1]])\n</code></pre>\n<p>(3)  memory_mask:</p>\n<pre><code>array([[1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [0, 0, 1, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 1, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [1, 0, 0, 0, 0, 0, 1, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 1, 1, 0, 0, 0, 0, 0, 1]])\n</code></pre>\n<p>To make this memory mask very fast in the training phase, I wrote such cython (#2) function (not vectorized implementation):</p>\n<pre><code>%%cython\nimport numpy as np\ncimport numpy as np\ncpdef np.ndarray[int] cget_memory_mask(np.ndarray task, int n_length):\n    cdef Py_ssize_t n = task.shape[0]\n    cdef np.ndarray[int, ndim = 2] res = np.full((n, n), 1, dtype = np.int32)\n    cdef int tmp_counter = 0\n    cdef int u_counter = 0\n    for i in range(n):\n        tmp_counter += 1\n        if i == n - 1 or task[i] != task[i + 1]:\n            res[i - tmp_counter + 1 : i + 1, :u_counter] = 0\n            if u_counter == 0:\n                res[i - tmp_counter + 1 : i + 1, n - 1] = 0\n            if u_counter &gt; n_length:\n                res[i - tmp_counter + 1: i + 1, :(u_counter - n_length)] = 1\n            u_counter += tmp_counter\n            tmp_counter = 0\n    return res\n</code></pre>\n<p>In the inference phase, memory mask is not required. For more details, please look at my <a href=\"https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold\" target=\"_blank\">single model notebook</a>.</p>\n<h1>My Features</h1>\n<p>Transformer is great, but it suffers from a problem that the sequence length cannot be infinite. I mean, it cannot consider the information of very old samples. To tackle this problem and to leverage the lecture information, I made simple features and concatenated them with the output features of SAKT-like model, and applied MLP and sigmoid.<br>\nI uploaded the names of 90 features as attachments. For the implementation of these features, I didn't use any pd.merge or df.join and most of the implementation are done using numpy. To make user-content features, I used scipy.sparse.lil_matrix to spare memory usage. I think my feature is not so good as other competitors (about 0.795 when using GBDT), but still improved score 0.001 ~ 0.002.<br>\nI made some tricky features (e.g. obtained by SVD), but it did not improve the score of NN model (improved GBDT model, though.), probably because the information of NN features includes that of such tricky features.</p>\n<h1>CV Strategy</h1>\n<p>My CV strategy is completely different from tito( <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> )'s one. First, I probed the number of new users in test set and I found there are about 7000 new users. as we know there is 2.5M rows in test set and we know the average length of the history of all users, we can estimate the number of rows of the new users and the number of rows of the existing users. After the calculation, I found</p>\n<pre><code>the number of rows of new users (i.e. user split): the number of rows of existing users (i.e. timeseries split) = 2 : 1.\n</code></pre>\n<p>So, For validation, I decided to use 1M rows for timeseries split and 2M rows for user split.<br>\nWhen making validation set of timeseries split, I was so careful that leakage can't happen. I mean, my training dataset and validation dataset never shares same task_container_id for a given user.</p>\n<h1>What worked</h1>\n<ol>\n<li>increase length from 100 to 400 worked.</li>\n<li>using normalized timedelta is better than digitized timedelta (mentioned in SAINT+ paper).</li>\n<li>concatenating embeddings is better than adding embedding, as mentioned in <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201798\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201798</a>. </li>\n<li>dropout = 0.2 is very important.</li>\n<li>StepLR with Adam is good. I trained my model for about 35 epoch with lr = 2e-3, then trained it for 1 epoch with lr = 2e-4.</li>\n</ol>\n<h1>What didn't work</h1>\n<ol>\n<li>Random masking of sequences didn't improve the score.</li>\n<li>bundle_id and normalized task_container_id is not needed for the input of LSTM.</li>\n<li>As I didn't make diverse models, weighted averaging is enough and blending using GBDT didn't work well.</li>\n<li>Full data training (without any validation data) didn't improve the score.</li>\n<li>In my implementation, SAINT didn't work well. In my opinion, it's logically very hard to implement SAINT without data leakage in training due to task_container_id.</li>\n</ol>\n<h1>Comments</h1>\n<p>I was very sad to see the private score bug problem of kaggle. Of course this competition is very great, but I think this competition could have been one of the most successful competition in kaggle without this problem.<br>\nAnyway, I really enjoyed my 4th kaggle competition. I'll continue kaggle and will surely become the winner in the next competition!</p>\n<h1>Thanks all, see you again!</h1>\n<p>P.S. The word <code>memory_mask</code> may be confusing because it is different from pytorch <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html#torch.nn.Transformer\" target=\"_blank\">Transformer</a>'s <code>memory_mask</code>.<br>\nThe word <code>memory_mask</code> is more like <code>mask</code> in pytorch's <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.TransformerEncoder.html\" target=\"_blank\">TransformerEncoder</a>.<br>\nThe reason my word is confusing is, I used Encoder-Decoder model at first, then gave up it and started using Encoder-only model, but I didn't change the function names. :(</p>",
  "messages": [
    {
      "id": 1146355,
      "postDate": "2021-01-09T18:13:42.940Z",
      "content": "<p>Thank you all teams who competed with me, all the people who participated in this competition, and the organizer that hosts such a great competition with the well-designed API!<br>\nCongrats <a href=\"https://www.kaggle.com/keetar\" target=\"_blank\">@keetar</a>, who defeats me and becomes the winner in this competition.</p>\n<p>I'm happy because it's my first time I get solo prize!</p>\n<p>It's my 4th kaggle competition and it was fun to compete with my past teammates ( <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a>, <a href=\"https://www.kaggle.com/pocketsuteado\" target=\"_blank\">@pocketsuteado</a>) and people who I competed with past competitions (e.g. <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>, <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>). <br>\nHere, I will explain the summary of my model and my features. </p>\n<p>I uploaded 2 kaggle notebooks for the explanation. As the ensemble is not so important in my solution, I will only explain my single model. </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution\" target=\"_blank\">6 similar models weighted average</a> : 0.817 public/0.818 private</li>\n<li><a href=\"https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold\" target=\"_blank\">single model</a> : 0.814 public/0.816 private</li>\n</ol>\n<h1>Models</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F762938%2F39ec58b701e6d37701cc794763d3a473%2F2nd_place.png?generation=1610213712807228&amp;alt=media\" alt=\"\"></p>\n<h2>Overview</h2>\n<p>my model is similar to SAKT, with 400 sequence length, 512 dimension, 4 nheads. <br>\nI don't use lecture information for the input of transformer model, which I guess is why I lost in this competition. For the query and key/value of SAKT-like model, I used LSTM-encoded features, whose input is as follows.</p>\n<ul>\n<li><p><strong>\"Query\" features</strong><br>\ncontent_id<br>\npart<br>\ntags<br>\nnormalized timedelta<br>\nnormalized log timestamp <br>\ncorrect answer<br>\ntask_container_id delta<br>\ncontent_type_id delta<br>\nnormalized absolute position </p></li>\n<li><p><strong>\"Memory\" features</strong><br>\nexplanation<br>\ncorrectness<br>\nnormalized elapsed time<br>\nuser_answer</p></li>\n</ul>\n<h2>Detailed Explanation of Training/Inference Process</h2>\n<p>I tried a very precise indexing/masking technique to avoid data leakage in training process, I guess which is partly why I became 2nd place. OK, suppose a very simple model of the sequence length = 5, and a task_container_id history of a specific user (without lecture) is like this.<br>\n<code>[0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 5, 8, 7, 7, 6]</code><br>\nAs there is 15 measurements in this user, I made 15/5 = 3 training samples and loss mask.</p>\n<p>(1) input task_container_id: <code>[pad, pad, pad, pad, pad, 0, 0, 0, 1, 1]</code> <br>\n(1) loss_mask: <code>[False, False, False, False, False, True, True, True, True, True]</code><br>\n(2) input task_container_id: <code>[0, 0, 0, 1, 1, 2, 2, 3, 3, 4]</code> <br>\n(2) loss_mask: <code>[False, False, False, False, False, True, True, True, True, True]</code><br>\n(3) input task_container_id: <code>[2, 2, 3, 3, 4, 5, 8, 7, 7, 6]</code> <br>\n(3) loss_mask: <code>[False, False, False, False, False, True, True, True, True, True]</code></p>\n<p>In this competition, the handling of the task_container_id is very important, as <strong>it is not allowed to use \"memory\" features of the same task container id</strong> to avoid leakage.<br>\nSo, after applying LSTM to the features, such fancy indexing is required to avoid the leakage for 3 training samples, where -1 means this position can't attend any position.</p>\n<p>(1) indices: <code>[-1, -1, -1, -1, -1, 4, 4, 4, 7, 7]</code><br>\n(2) indices: <code>[-1, -1, -1, 2, 2, 4, 4, 6, 6, 8]</code><br>\n(3) indices: <code>[-1, -1, 1, 1, 3, 4, 5, 6, 6, 8]</code></p>\n<p>To get this indices very fast in the training/inference phase, I wrote such cython (#1) function (vectorized implementation):</p>\n<pre><code>%%cython\nimport numpy as np\ncimport numpy as np\ncpdef np.ndarray[int] cget_memory_indices(np.ndarray task):\n    cdef Py_ssize_t n = task.shape[1]\n    cdef np.ndarray[int, ndim = 2] res = np.zeros_like(task, dtype = np.int32)\n    cdef np.ndarray[int] tmp_counter = np.full(task.shape[0], -1, dtype = np.int32)\n    cdef np.ndarray[int] u_counter = np.full(task.shape[0], task.shape[1] - 1, dtype = np.int32)\n    for i in range(n):\n        res[:, i] = u_counter\n        tmp_counter += 1\n        if i != n - 1:\n            mask = (task[:, i] != task[:, i + 1])\n            u_counter[mask] = tmp_counter[mask]\n    return res\n</code></pre>\n<p>After applying such fancy indexing to the output of LSTM features, I concatenated them with \"Query\" features and apply MLP, then we can get query for SAKT model. Then, I obtained key/value for SAKT model by applying MLP to query concatenated with \"Memory\" features. </p>\n<p>To train SAKT-like model, precise memory masking is also required to avoid leakage. I used <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html\" target=\"_blank\">3D attention mask</a> of torch.nn.MultiheadAttention. It should be noted <a href=\"https://pytorch.org/docs/stable/generated/torch.repeat_interleave.html\" target=\"_blank\">torch.repeat_interleave</a> must be leveraged to make 3D mask (batchsize * nhead, sequence length, sequence length). <br>\nThe memory mask for these 3 samples are like this. Here, 1 means True and 0 means False. Please remember the <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html\" target=\"_blank\">documentation</a> says </p>\n<pre><code>attn_mask ensure that position i is allowed to attend the unmasked positions. If a BoolTensor is provided, positions with True is not allowed to attend while False values will be unchanged.\n</code></pre>\n<p>(1) memory_mask:</p>\n<pre><code>array([[1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [1, 1, 1, 0, 0, 0, 0, 0, 1, 1],\n       [1, 1, 1, 0, 0, 0, 0, 0, 1, 1]])\n</code></pre>\n<p>(2) memory_mask:</p>\n<pre><code>array([[1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [0, 0, 0, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 1, 1, 0, 0, 0, 0, 0, 1]])\n</code></pre>\n<p>(3)  memory_mask:</p>\n<pre><code>array([[1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [0, 0, 1, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 1, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [1, 0, 0, 0, 0, 0, 1, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 1, 1, 0, 0, 0, 0, 0, 1]])\n</code></pre>\n<p>To make this memory mask very fast in the training phase, I wrote such cython (#2) function (not vectorized implementation):</p>\n<pre><code>%%cython\nimport numpy as np\ncimport numpy as np\ncpdef np.ndarray[int] cget_memory_mask(np.ndarray task, int n_length):\n    cdef Py_ssize_t n = task.shape[0]\n    cdef np.ndarray[int, ndim = 2] res = np.full((n, n), 1, dtype = np.int32)\n    cdef int tmp_counter = 0\n    cdef int u_counter = 0\n    for i in range(n):\n        tmp_counter += 1\n        if i == n - 1 or task[i] != task[i + 1]:\n            res[i - tmp_counter + 1 : i + 1, :u_counter] = 0\n            if u_counter == 0:\n                res[i - tmp_counter + 1 : i + 1, n - 1] = 0\n            if u_counter &gt; n_length:\n                res[i - tmp_counter + 1: i + 1, :(u_counter - n_length)] = 1\n            u_counter += tmp_counter\n            tmp_counter = 0\n    return res\n</code></pre>\n<p>In the inference phase, memory mask is not required. For more details, please look at my <a href=\"https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold\" target=\"_blank\">single model notebook</a>.</p>\n<h1>My Features</h1>\n<p>Transformer is great, but it suffers from a problem that the sequence length cannot be infinite. I mean, it cannot consider the information of very old samples. To tackle this problem and to leverage the lecture information, I made simple features and concatenated them with the output features of SAKT-like model, and applied MLP and sigmoid.<br>\nI uploaded the names of 90 features as attachments. For the implementation of these features, I didn't use any pd.merge or df.join and most of the implementation are done using numpy. To make user-content features, I used scipy.sparse.lil_matrix to spare memory usage. I think my feature is not so good as other competitors (about 0.795 when using GBDT), but still improved score 0.001 ~ 0.002.<br>\nI made some tricky features (e.g. obtained by SVD), but it did not improve the score of NN model (improved GBDT model, though.), probably because the information of NN features includes that of such tricky features.</p>\n<h1>CV Strategy</h1>\n<p>My CV strategy is completely different from tito( <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> )'s one. First, I probed the number of new users in test set and I found there are about 7000 new users. as we know there is 2.5M rows in test set and we know the average length of the history of all users, we can estimate the number of rows of the new users and the number of rows of the existing users. After the calculation, I found</p>\n<pre><code>the number of rows of new users (i.e. user split): the number of rows of existing users (i.e. timeseries split) = 2 : 1.\n</code></pre>\n<p>So, For validation, I decided to use 1M rows for timeseries split and 2M rows for user split.<br>\nWhen making validation set of timeseries split, I was so careful that leakage can't happen. I mean, my training dataset and validation dataset never shares same task_container_id for a given user.</p>\n<h1>What worked</h1>\n<ol>\n<li>increase length from 100 to 400 worked.</li>\n<li>using normalized timedelta is better than digitized timedelta (mentioned in SAINT+ paper).</li>\n<li>concatenating embeddings is better than adding embedding, as mentioned in <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201798\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201798</a>. </li>\n<li>dropout = 0.2 is very important.</li>\n<li>StepLR with Adam is good. I trained my model for about 35 epoch with lr = 2e-3, then trained it for 1 epoch with lr = 2e-4.</li>\n</ol>\n<h1>What didn't work</h1>\n<ol>\n<li>Random masking of sequences didn't improve the score.</li>\n<li>bundle_id and normalized task_container_id is not needed for the input of LSTM.</li>\n<li>As I didn't make diverse models, weighted averaging is enough and blending using GBDT didn't work well.</li>\n<li>Full data training (without any validation data) didn't improve the score.</li>\n<li>In my implementation, SAINT didn't work well. In my opinion, it's logically very hard to implement SAINT without data leakage in training due to task_container_id.</li>\n</ol>\n<h1>Comments</h1>\n<p>I was very sad to see the private score bug problem of kaggle. Of course this competition is very great, but I think this competition could have been one of the most successful competition in kaggle without this problem.<br>\nAnyway, I really enjoyed my 4th kaggle competition. I'll continue kaggle and will surely become the winner in the next competition!</p>\n<h1>Thanks all, see you again!</h1>\n<p>P.S. The word <code>memory_mask</code> may be confusing because it is different from pytorch <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html#torch.nn.Transformer\" target=\"_blank\">Transformer</a>'s <code>memory_mask</code>.<br>\nThe word <code>memory_mask</code> is more like <code>mask</code> in pytorch's <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.TransformerEncoder.html\" target=\"_blank\">TransformerEncoder</a>.<br>\nThe reason my word is confusing is, I used Encoder-Decoder model at first, then gave up it and started using Encoder-only model, but I didn't change the function names. :(</p>",
      "rawMarkdown": "Thank you all teams who competed with me, all the people who participated in this competition, and the organizer that hosts such a great competition with the well-designed API!\nCongrats @keetar, who defeats me and becomes the winner in this competition.\n\nI'm happy because it's my first time I get solo prize!\n\nIt's my 4th kaggle competition and it was fun to compete with my past teammates ( @nyanpn, @pocketsuteado) and people who I competed with past competitions (e.g. @aerdem4, @its7171). \nHere, I will explain the summary of my model and my features. \n\nI uploaded 2 kaggle notebooks for the explanation. As the ensemble is not so important in my solution, I will only explain my single model. \n1. [6 similar models weighted average](https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution) : 0.817 public/0.818 private\n2. [single model](https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold) : 0.814 public/0.816 private\n\n#Models \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F762938%2F39ec58b701e6d37701cc794763d3a473%2F2nd_place.png?generation=1610213712807228&alt=media)\n\n##Overview \nmy model is similar to SAKT, with 400 sequence length, 512 dimension, 4 nheads. \nI don't use lecture information for the input of transformer model, which I guess is why I lost in this competition. For the query and key/value of SAKT-like model, I used LSTM-encoded features, whose input is as follows.\n- **\"Query\" features**\ncontent_id\npart\ntags\nnormalized timedelta\nnormalized log timestamp \ncorrect answer\ntask_container_id delta\ncontent_type_id delta\nnormalized absolute position \n\n- **\"Memory\" features**\nexplanation\ncorrectness\nnormalized elapsed time\nuser_answer\n\n## Detailed Explanation of Training/Inference Process\nI tried a very precise indexing/masking technique to avoid data leakage in training process, I guess which is partly why I became 2nd place. OK, suppose a very simple model of the sequence length = 5, and a task_container_id history of a specific user (without lecture) is like this.\n`[0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 5, 8, 7, 7, 6]`\nAs there is 15 measurements in this user, I made 15/5 = 3 training samples and loss mask.\n\n(1) input task_container_id: `[pad, pad, pad, pad, pad, 0, 0, 0, 1, 1]` \n(1) loss_mask: `[False, False, False, False, False, True, True, True, True, True]`\n(2) input task_container_id: `[0, 0, 0, 1, 1, 2, 2, 3, 3, 4]` \n(2) loss_mask: `[False, False, False, False, False, True, True, True, True, True]`\n(3) input task_container_id: `[2, 2, 3, 3, 4, 5, 8, 7, 7, 6]` \n(3) loss_mask: `[False, False, False, False, False, True, True, True, True, True]`\n\nIn this competition, the handling of the task_container_id is very important, as **it is not allowed to use \"memory\" features of the same task container id** to avoid leakage.\nSo, after applying LSTM to the features, such fancy indexing is required to avoid the leakage for 3 training samples, where -1 means this position can't attend any position.\n\n(1) indices: `[-1, -1, -1, -1, -1, 4, 4, 4, 7, 7]`\n(2) indices: `[-1, -1, -1, 2, 2, 4, 4, 6, 6, 8]`\n(3) indices: `[-1, -1, 1, 1, 3, 4, 5, 6, 6, 8]`\n\nTo get this indices very fast in the training/inference phase, I wrote such cython (#1) function (vectorized implementation):\n\n```\n%%cython\nimport numpy as np\ncimport numpy as np\ncpdef np.ndarray[int] cget_memory_indices(np.ndarray task):\n    cdef Py_ssize_t n = task.shape[1]\n    cdef np.ndarray[int, ndim = 2] res = np.zeros_like(task, dtype = np.int32)\n    cdef np.ndarray[int] tmp_counter = np.full(task.shape[0], -1, dtype = np.int32)\n    cdef np.ndarray[int] u_counter = np.full(task.shape[0], task.shape[1] - 1, dtype = np.int32)\n    for i in range(n):\n        res[:, i] = u_counter\n        tmp_counter += 1\n        if i != n - 1:\n            mask = (task[:, i] != task[:, i + 1])\n            u_counter[mask] = tmp_counter[mask]\n    return res\n```\nAfter applying such fancy indexing to the output of LSTM features, I concatenated them with \"Query\" features and apply MLP, then we can get query for SAKT model. Then, I obtained key/value for SAKT model by applying MLP to query concatenated with \"Memory\" features. \n\nTo train SAKT-like model, precise memory masking is also required to avoid leakage. I used [3D attention mask](https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html) of torch.nn.MultiheadAttention. It should be noted [torch.repeat_interleave](https://pytorch.org/docs/stable/generated/torch.repeat_interleave.html) must be leveraged to make 3D mask (batchsize * nhead, sequence length, sequence length). \nThe memory mask for these 3 samples are like this. Here, 1 means True and 0 means False. Please remember the [documentation](https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html) says \n``` \nattn_mask ensure that position i is allowed to attend the unmasked positions. If a BoolTensor is provided, positions with True is not allowed to attend while False values will be unchanged.\n```\n(1) memory_mask:\n```\narray([[1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [1, 1, 1, 0, 0, 0, 0, 0, 1, 1],\n       [1, 1, 1, 0, 0, 0, 0, 0, 1, 1]])\n```\n(2) memory_mask:\n```\narray([[1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [0, 0, 0, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 1, 1, 0, 0, 0, 0, 0, 1]])\n```\n(3)  memory_mask:\n```\narray([[1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [0, 0, 1, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 1, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [1, 0, 0, 0, 0, 0, 1, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 1, 1, 0, 0, 0, 0, 0, 1]])\n```\n\nTo make this memory mask very fast in the training phase, I wrote such cython (#2) function (not vectorized implementation):\n```\n%%cython\nimport numpy as np\ncimport numpy as np\ncpdef np.ndarray[int] cget_memory_mask(np.ndarray task, int n_length):\n    cdef Py_ssize_t n = task.shape[0]\n    cdef np.ndarray[int, ndim = 2] res = np.full((n, n), 1, dtype = np.int32)\n    cdef int tmp_counter = 0\n    cdef int u_counter = 0\n    for i in range(n):\n        tmp_counter += 1\n        if i == n - 1 or task[i] != task[i + 1]:\n            res[i - tmp_counter + 1 : i + 1, :u_counter] = 0\n            if u_counter == 0:\n                res[i - tmp_counter + 1 : i + 1, n - 1] = 0\n            if u_counter > n_length:\n                res[i - tmp_counter + 1: i + 1, :(u_counter - n_length)] = 1\n            u_counter += tmp_counter\n            tmp_counter = 0\n    return res\n```\n\nIn the inference phase, memory mask is not required. For more details, please look at my [single model notebook](https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold).\n\n#My Features\nTransformer is great, but it suffers from a problem that the sequence length cannot be infinite. I mean, it cannot consider the information of very old samples. To tackle this problem and to leverage the lecture information, I made simple features and concatenated them with the output features of SAKT-like model, and applied MLP and sigmoid.\nI uploaded the names of 90 features as attachments. For the implementation of these features, I didn't use any pd.merge or df.join and most of the implementation are done using numpy. To make user-content features, I used scipy.sparse.lil_matrix to spare memory usage. I think my feature is not so good as other competitors (about 0.795 when using GBDT), but still improved score 0.001 ~ 0.002.\nI made some tricky features (e.g. obtained by SVD), but it did not improve the score of NN model (improved GBDT model, though.), probably because the information of NN features includes that of such tricky features.\n\n#CV Strategy\nMy CV strategy is completely different from tito( @its7171 )'s one. First, I probed the number of new users in test set and I found there are about 7000 new users. as we know there is 2.5M rows in test set and we know the average length of the history of all users, we can estimate the number of rows of the new users and the number of rows of the existing users. After the calculation, I found\n```\nthe number of rows of new users (i.e. user split): the number of rows of existing users (i.e. timeseries split) = 2 : 1.\n```\nSo, For validation, I decided to use 1M rows for timeseries split and 2M rows for user split.\nWhen making validation set of timeseries split, I was so careful that leakage can't happen. I mean, my training dataset and validation dataset never shares same task_container_id for a given user.\n\n#What worked \n1. increase length from 100 to 400 worked.\n2. using normalized timedelta is better than digitized timedelta (mentioned in SAINT+ paper).\n3. concatenating embeddings is better than adding embedding, as mentioned in https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201798. \n4. dropout = 0.2 is very important.\n5. StepLR with Adam is good. I trained my model for about 35 epoch with lr = 2e-3, then trained it for 1 epoch with lr = 2e-4.\n\n#What didn't work\n1. Random masking of sequences didn't improve the score.\n2. bundle_id and normalized task_container_id is not needed for the input of LSTM.\n3. As I didn't make diverse models, weighted averaging is enough and blending using GBDT didn't work well.\n4. Full data training (without any validation data) didn't improve the score.\n5. In my implementation, SAINT didn't work well. In my opinion, it's logically very hard to implement SAINT without data leakage in training due to task_container_id.\n\n#Comments\nI was very sad to see the private score bug problem of kaggle. Of course this competition is very great, but I think this competition could have been one of the most successful competition in kaggle without this problem.\nAnyway, I really enjoyed my 4th kaggle competition. I'll continue kaggle and will surely become the winner in the next competition!\n\n<h1>Thanks all, see you again!</h1>\n\n\nP.S. The word `memory_mask` may be confusing because it is different from pytorch [Transformer](https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html#torch.nn.Transformer)'s `memory_mask`.\nThe word `memory_mask` is more like `mask` in pytorch's [TransformerEncoder](https://pytorch.org/docs/stable/generated/torch.nn.TransformerEncoder.html).\nThe reason my word is confusing is, I used Encoder-Decoder model at first, then gave up it and started using Encoder-only model, but I didn't change the function names. :(\n",
      "votes": 206
    },
    {
      "id": 1146558,
      "postDate": "2021-01-09T21:16:22.747Z",
      "content": "<p>Good luck WINNING your next competition.</p>",
      "rawMarkdown": "Good luck WINNING your next competition.",
      "votes": 4
    },
    {
      "id": 1147926,
      "postDate": "2021-01-10T20:03:30.223Z",
      "content": "<p>Thanks for the nicely detailed write-up and congratz for the impressive finish!</p>",
      "rawMarkdown": "Thanks for the nicely detailed write-up and congratz for the impressive finish!",
      "votes": 1
    },
    {
      "id": 1146546,
      "postDate": "2021-01-09T21:03:17.123Z",
      "content": "<blockquote>\n  <p>In my implementation, SAINT didn't work well. In my opinion, it's logically very hard to implement SAINT without data leakage in training due to task_container_id.</p>\n</blockquote>\n<p>Couldn't you control that with <code>tgt_mask</code> from PyTorch's  <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html#torch.nn.Transformer\" target=\"_blank\">implementation</a> ?</p>",
      "rawMarkdown": "> In my implementation, SAINT didn't work well. In my opinion, it's logically very hard to implement SAINT without data leakage in training due to task_container_id.\n\nCouldn't you control that with `tgt_mask` from PyTorch's  [implementation](https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html#torch.nn.Transformer) ?",
      "votes": 1,
      "replies": [
        {
          "id": 1146561,
          "postDate": "2021-01-09T21:21:10.327Z",
          "content": "<p>Yes we can control <code>tgt_mask</code> in decoder. but if we use the same input of the decoder as SAINT+ paper in this image, I think the problem cannot be fixed just by controlling <code>tgt_mask</code>, because the <strong>query</strong> of decoder produces leakage (i.e. uses the information of the sample in the same task_container_id). In my understanding, the data leakage caused by <strong>query</strong> is unavoidable by controlling masking, while the data leakage caused by <strong>key/value</strong> is avoidable by masking.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F762938%2Fb29c0c24c02cb5b71707c2d378bc5a13%2Finput_of_saint.jpg?generation=1610226842294243&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Yes we can control `tgt_mask` in decoder. but if we use the same input of the decoder as SAINT+ paper in this image, I think the problem cannot be fixed just by controlling `tgt_mask`, because the **query** of decoder produces leakage (i.e. uses the information of the sample in the same task_container_id). In my understanding, the data leakage caused by **query** is unavoidable by controlling masking, while the data leakage caused by **key/value** is avoidable by masking.\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F762938%2Fb29c0c24c02cb5b71707c2d378bc5a13%2Finput_of_saint.jpg?generation=1610226842294243&alt=media)\n",
          "votes": 2
        },
        {
          "id": 1146572,
          "postDate": "2021-01-09T21:34:12.120Z",
          "content": "<p>If I understand what you are saying, you mean for questions in the same bundle that are under prediction (but during training), the corresponding input in the decoder will see the information (i.e. user correctness) in the same bundle (but prior the current position). However, while during inference, it can only see this information at the places before the current bundle.</p>\n<p>I am using the encoder-decoder approach, and during the inference time, I replace the values by the value obtained at the first question in the bundle under prediction. Basically, it treats the question in the bundle as separate questions, and use the history before them.</p>\n<p>Anyway, the best I can get is 0.802, so potentially there are some problems as you mentioned.</p>",
          "rawMarkdown": "If I understand what you are saying, you mean for questions in the same bundle that are under prediction (but during training), the corresponding input in the decoder will see the information (i.e. user correctness) in the same bundle (but prior the current position). However, while during inference, it can only see this information at the places before the current bundle.\n\nI am using the encoder-decoder approach, and during the inference time, I replace the values by the value obtained at the first question in the bundle under prediction. Basically, it treats the question in the bundle as separate questions, and use the history before them.\n\nAnyway, the best I can get is 0.802, so potentially there are some problems as you mentioned."
        },
        {
          "id": 1146587,
          "postDate": "2021-01-09T21:44:02.640Z",
          "content": "<p>Yes of course the leakage doesn't happen during inference, but I think it's problematic that leakage happens in training phase, because the model can learn meaningless patterns during training.</p>\n<p><code>I replace the values by the value obtained at the first question in the bundle under prediction.</code><br>\nHmm, I didn't notice this idea but this looks great. Why don't you try this idea in the training phase? I think it is possible by fancy indexing. </p>",
          "rawMarkdown": "Yes of course the leakage doesn't happen during inference, but I think it's problematic that leakage happens in training phase, because the model can learn meaningless patterns during training.\n\n `I replace the values by the value obtained at the first question in the bundle under prediction.`\nHmm, I didn't notice this idea but this looks great. Why don't you try this idea in the training phase? I think it is possible by fancy indexing. "
        },
        {
          "id": 1146606,
          "postDate": "2021-01-09T21:58:27.887Z",
          "content": "<p>I'm still not seeing which way the leak could go :?</p>",
          "rawMarkdown": "I'm still not seeing which way the leak could go :?"
        },
        {
          "id": 1146610,
          "postDate": "2021-01-09T22:02:29.533Z",
          "content": "<p>I didn't think of the training will be problematic. I am aware of some kind of inconsistency between decoder  training / inference input by this approach, but it is a comprising approach. Treating each question in the bundle under the current inference time step is not difficult, although still need to be careful.</p>\n<p>To do so in training, I feel it becomes more difficult. Because we are not just dealing the last bundle in the truncated sequence, we extract exact one question from each budle in that sequence. Maybe not so difficult, but I kind feel we also loss information that given by bundle. Anyway, without trying, I can't say if we avoid leakage or lossing info if we apply this approach to training</p>",
          "rawMarkdown": "I didn't think of the training will be problematic. I am aware of some kind of inconsistency between decoder  training / inference input by this approach, but it is a comprising approach. Treating each question in the bundle under the current inference time step is not difficult, although still need to be careful.\n\nTo do so in training, I feel it becomes more difficult. Because we are not just dealing the last bundle in the truncated sequence, we extract exact one question from each budle in that sequence. Maybe not so difficult, but I kind feel we also loss information that given by bundle. Anyway, without trying, I can't say if we avoid leakage or lossing info if we apply this approach to training",
          "votes": 1
        },
        {
          "id": 1146611,
          "postDate": "2021-01-09T22:05:18.623Z",
          "content": "<p><a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a><br>\nFor example, suppose task_container_id of a specific user is [0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 5, 8, 7, 7, 6]. Then, <br>\nposition [0, 1, 2] cannot use any information.<br>\nposition [3, 4] can use the information of [0, 1, 2].<br>\nposition [5, 6] can use the information of [0, 1, 2, 3, 4].<br>\nposition [7, 8] can use the information of [0, 1, 2, 3, 4, 5, 6].</p>\n<p>However, when simply using the input of the SAINT paper, position 1 can use the information (e.g. correctness) of position 0 and position 2 can use the information of position [0, 1]. I think this is problematic as the model uses the information that should not be utilized, during training.</p>",
          "rawMarkdown": "@bacicnikola\nFor example, suppose task_container_id of a specific user is [0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 5, 8, 7, 7, 6]. Then, \nposition [0, 1, 2] cannot use any information.\nposition [3, 4] can use the information of [0, 1, 2].\nposition [5, 6] can use the information of [0, 1, 2, 3, 4].\nposition [7, 8] can use the information of [0, 1, 2, 3, 4, 5, 6].\n\nHowever, when simply using the input of the SAINT paper, position 1 can use the information (e.g. correctness) of position 0 and position 2 can use the information of position [0, 1]. I think this is problematic as the model uses the information that should not be utilized, during training.",
          "votes": 3
        },
        {
          "id": 1146612,
          "postDate": "2021-01-09T22:06:17.623Z",
          "content": "<p>Not clear to me about what kind of leakage neither, but at least some form of inconsistency exist</p>",
          "rawMarkdown": "Not clear to me about what kind of leakage neither, but at least some form of inconsistency exist"
        },
        {
          "id": 1146617,
          "postDate": "2021-01-09T22:12:44.273Z",
          "content": "<p>Yes, it can be called \"incosistency\", as the problem doesn't happen at least in inference phase. </p>",
          "rawMarkdown": "Yes, it can be called \"incosistency\", as the problem doesn't happen at least in inference phase. ",
          "votes": 1
        },
        {
          "id": 1146618,
          "postDate": "2021-01-09T22:14:48.053Z",
          "content": "<p>To be clear, my approach mentioned above , for decoder, only applies to the last bundle in the truncated sequence. The bundles before it, even they appear in inference time, but already being predicted, the special processing is not applied to them anymore</p>",
          "rawMarkdown": "To be clear, my approach mentioned above , for decoder, only applies to the last bundle in the truncated sequence. The bundles before it, even they appear in inference time, but already being predicted, the special processing is not applied to them anymore",
          "votes": 1
        },
        {
          "id": 1146621,
          "postDate": "2021-01-09T22:19:53.660Z",
          "content": "<pre><code>To do so in training, I feel it becomes more difficult. Because we are not just dealing the last bundle in the truncated sequence, we extract exact one question from each budle in that sequence. Maybe not so difficult, but I kind feel we also loss information that given by bundle.\n</code></pre>\n<p>Now I understand what you mean. it seems the problem cannot be solved by simply using fancy indexing. Hmm…</p>",
          "rawMarkdown": "```\nTo do so in training, I feel it becomes more difficult. Because we are not just dealing the last bundle in the truncated sequence, we extract exact one question from each budle in that sequence. Maybe not so difficult, but I kind feel we also loss information that given by bundle.\n```\nNow I understand what you mean. it seems the problem cannot be solved by simply using fancy indexing. Hmm...",
          "votes": 1
        },
        {
          "id": 1146623,
          "postDate": "2021-01-09T22:20:17.523Z",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> <br>\nYes, I agree. But what confusses me is why tgt_mask can't control that. In example you mentioned, it should look like this:</p>\n<p>[0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 5, 8, 7, 7, 6]</p>\n<p>(1 mask, 0 no mask)</p>\n<p>1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 (for 0 position)<br>\n1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 (for 1 position)<br>\n1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 (for 2 position)<br>\n0 0 0 1 1 1 1 1 1 1 1 1 1 1 1 (for 3 position)<br>\n0 0 0 1 1 1 1 1 1 1 1 1 1 1 1 (for 4 position)<br>\n0 0 0 0 0 1 1 1 1 1 1 1 1 1 1 (for 5 position)<br>\n0 0 0 0 0 1 1 1 1 1 1 1 1 1 (for 6 position)<br>\netc.</p>\n<p>Although in my solution, I'm only using encoder part of the transformer, but that shouldn't make any difference.</p>",
          "rawMarkdown": "@mamasinkgs \nYes, I agree. But what confusses me is why tgt_mask can't control that. In example you mentioned, it should look like this:\n\n[0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 5, 8, 7, 7, 6]\n\n(1 mask, 0 no mask)\n\n1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 (for 0 position)\n1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 (for 1 position)\n1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 (for 2 position)\n0 0 0 1 1 1 1 1 1 1 1 1 1 1 1 (for 3 position)\n0 0 0 1 1 1 1 1 1 1 1 1 1 1 1 (for 4 position)\n0 0 0 0 0 1 1 1 1 1 1 1 1 1 1 (for 5 position)\n0 0 0 0 0 1 1 1 1 1 1 1 1 1 (for 6 position)\netc.\n\nAlthough in my solution, I'm only using encoder part of the transformer, but that shouldn't make any difference.\n"
        },
        {
          "id": 1146630,
          "postDate": "2021-01-09T22:28:16.987Z",
          "content": "<p><a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a> <br>\nWhen using encoder-only model, there is no problem. However, when using decoder of SAINT+ model, the sequence is <strong>shifted</strong> like the image above. So, by simply using tgt_mask you showed (actually the mask you showed is inherently same as the mask I showed in my solution), the problem cannot be solved.</p>",
          "rawMarkdown": "@bacicnikola \nWhen using encoder-only model, there is no problem. However, when using decoder of SAINT+ model, the sequence is **shifted** like the image above. So, by simply using tgt_mask you showed (actually the mask you showed is inherently same as the mask I showed in my solution), the problem cannot be solved.",
          "votes": 3
        },
        {
          "id": 1146633,
          "postDate": "2021-01-09T22:32:39.940Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1146640,
          "postDate": "2021-01-09T22:46:30.053Z",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a><br>\nYes, I see now. My implementation of this was only partly good because even though I did mask positions from the same task_container, responses were shifted to the right and question could always see previous response, which could be from the same task_container (since I used encoder only, position could always see itself)<br>\nI think I should've put responses into a encoder, and the rest of the features into decoder and control task_container with memory_mask.</p>\n<p>Thank you for your time!</p>",
          "rawMarkdown": "@mamasinkgs\nYes, I see now. My implementation of this was only partly good because even though I did mask positions from the same task_container, responses were shifted to the right and question could always see previous response, which could be from the same task_container (since I used encoder only, position could always see itself)\nI think I should've put responses into a encoder, and the rest of the features into decoder and control task_container with memory_mask.\n\nThank you for your time!",
          "votes": 2
        },
        {
          "id": 1146647,
          "postDate": "2021-01-09T22:57:28.953Z",
          "content": "<p><a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a> ， just a remark, the mask you shows for decoder also has an issue. For example , at a place, it can't look at the current position. Therefore, position embedding of that place becomes no meaning while predicting that place, if we use the usual way of position embedding</p>",
          "rawMarkdown": "@bacicnikola ， just a remark, the mask you shows for decoder also has an issue. For example , at a place, it can't look at the current position. Therefore, position embedding of that place becomes no meaning while predicting that place, if we use the usual way of position embedding"
        },
        {
          "id": 1146651,
          "postDate": "2021-01-09T23:06:13.917Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> Yes, you're right. But I couldn't get anything out of positions anyway.</p>",
          "rawMarkdown": "@yihdarshieh Yes, you're right. But I couldn't get anything out of positions anyway.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1146767,
      "postDate": "2021-01-10T03:13:26.980Z",
      "content": "<p>Big congrats to you on 2nd solo prize. Thanks for the detail writeup, very intelligence solution <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> </p>",
      "rawMarkdown": "Big congrats to you on 2nd solo prize. Thanks for the detail writeup, very intelligence solution @mamasinkgs ",
      "votes": 2
    },
    {
      "id": 1146439,
      "postDate": "2021-01-09T19:19:19.120Z",
      "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> , thank you for sharing. See you in the next one!</p>\n<p>About your CPython implementation for getting the masks, it is amazing, you must be an advanced Python programmer!<br>\nFor me, I use TensorFlow, and it is quite straightforward and fast to get these mask by using tf operations, like</p>\n<pre><code>def get_causal_attention_mask(nd, ns, dtype, only_before):\n    \"\"\"\n    1's in the lower triangle, counting from the lower right corner. Same as tf.matrix_band_part(tf.ones([nd, ns]),\n    -1, ns-nd), but doesn't produce garbage on TPUs.\n    \"\"\"\n\n    # Remark: Think `nd` as the number of queries and `ns` as the number of keys.\n    # In encoder-decoder case, the queries are the decoder features and the keys are the encoder features.\n\n    i = tf.range(nd)[:, tf.newaxis]  # repeat along dim 1\n    j = tf.range(ns) # repeat along dim 0 \n    m = i &gt;= (j - ns + nd) + tf.cast(only_before, dtype=tf.int32)\n\n    return tf.cast(m, dtype)\n\ndef get_attention_mask_from_timestamp_batch(timestamp_tensors, dtype, only_before):\n    \"\"\"\n    Args:\n        timestamp_tensors: 2-D tf.int32 tensor, representing a batch of sequences of non-decreasing timestamps.\n\n    Returns:\n        attention_mask: 3-D tf.int32 tensor of shape = [batch_size, query_len, key_len], consisting of 0 and 1.\n            Here `query_len` and `key_len` are actually `seq_len`. It should be reshpaed, when used to calculate \n            attention scores, to [batch_size, nb_attn_head, query_len, key_len].\n    \"\"\"\n\n    t = timestamp_tensors\n\n    batch_size = tf.math.reduce_sum(tf.ones_like(t[:, :1], dtype=tf.int32))\n    seq_len = tf.math.reduce_sum(tf.ones_like(t[:1, :], dtype=tf.int32))\n\n    x = tf.broadcast_to(t[:, :, tf.newaxis], shape=[batch_size, seq_len, seq_len]) # repeat along dim 2\n    y = tf.broadcast_to(t[:, tf.newaxis, :], shape=[batch_size, seq_len, seq_len]) + tf.cast(only_before, dtype=tf.int64) # repeat along dim 1\n\n    m =  x &gt;= y\n\n    return tf.cast(m, dtype)\n</code></pre>",
      "rawMarkdown": "@mamasinkgs , thank you for sharing. See you in the next one!\n\nAbout your CPython implementation for getting the masks, it is amazing, you must be an advanced Python programmer!\nFor me, I use TensorFlow, and it is quite straightforward and fast to get these mask by using tf operations, like\n\n```\ndef get_causal_attention_mask(nd, ns, dtype, only_before):\n    \"\"\"\n    1's in the lower triangle, counting from the lower right corner. Same as tf.matrix_band_part(tf.ones([nd, ns]),\n    -1, ns-nd), but doesn't produce garbage on TPUs.\n    \"\"\"\n    \n    # Remark: Think `nd` as the number of queries and `ns` as the number of keys.\n    # In encoder-decoder case, the queries are the decoder features and the keys are the encoder features.\n    \n    i = tf.range(nd)[:, tf.newaxis]  # repeat along dim 1\n    j = tf.range(ns) # repeat along dim 0 \n    m = i >= (j - ns + nd) + tf.cast(only_before, dtype=tf.int32)\n    \n    return tf.cast(m, dtype)\n\ndef get_attention_mask_from_timestamp_batch(timestamp_tensors, dtype, only_before):\n    \"\"\"\n    Args:\n        timestamp_tensors: 2-D tf.int32 tensor, representing a batch of sequences of non-decreasing timestamps.\n    \n    Returns:\n        attention_mask: 3-D tf.int32 tensor of shape = [batch_size, query_len, key_len], consisting of 0 and 1.\n            Here `query_len` and `key_len` are actually `seq_len`. It should be reshpaed, when used to calculate \n            attention scores, to [batch_size, nb_attn_head, query_len, key_len].\n    \"\"\"\n    \n    t = timestamp_tensors\n    \n    batch_size = tf.math.reduce_sum(tf.ones_like(t[:, :1], dtype=tf.int32))\n    seq_len = tf.math.reduce_sum(tf.ones_like(t[:1, :], dtype=tf.int32))\n    \n    x = tf.broadcast_to(t[:, :, tf.newaxis], shape=[batch_size, seq_len, seq_len]) # repeat along dim 2\n    y = tf.broadcast_to(t[:, tf.newaxis, :], shape=[batch_size, seq_len, seq_len]) + tf.cast(only_before, dtype=tf.int64) # repeat along dim 1\n    \n    m =  x >= y\n    \n    return tf.cast(m, dtype)\n```\n\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 1146459,
          "postDate": "2021-01-09T19:36:48.403Z",
          "content": "<p>Thanks! I'm not familiar with tensorflow, but I'll try it in future competitions.</p>",
          "rawMarkdown": "Thanks! I'm not familiar with tensorflow, but I'll try it in future competitions.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1590856,
      "postDate": "2021-11-21T17:37:10.020Z",
      "content": "<p>Hi, I have few questions</p>\n<ol>\n<li>Can you guide me how to run the code you provided? and how to get results from it.</li>\n<li>Did you published any research paper on this implementation. I want to understand your implementation in detail.</li>\n<li>At kaggle, I have seen your 3 submissions related to this competition. <br>\ni. <a href=\"https://www.kaggle.com/mamasinkgs/2nd-place-solution-for-hosts\" target=\"_blank\">https://www.kaggle.com/mamasinkgs/2nd-place-solution-for-hosts</a><br>\nii. <a href=\"https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution\" target=\"_blank\">https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution</a><br>\niii. <a href=\"https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold\" target=\"_blank\">https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold</a><br>\nI want to know which one is the actual and correct submission?</li>\n</ol>",
      "rawMarkdown": "Hi, I have few questions\n1. Can you guide me how to run the code you provided? and how to get results from it.\n2. Did you published any research paper on this implementation. I want to understand your implementation in detail.\n3. At kaggle, I have seen your 3 submissions related to this competition. \n   i. https://www.kaggle.com/mamasinkgs/2nd-place-solution-for-hosts\n  ii. https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution\n  iii. https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold\nI want to know which one is the actual and correct submission?\n",
      "replies": [
        {
          "id": 1591061,
          "postDate": "2021-11-21T22:53:26.467Z",
          "content": "<p>Hi Nimra, </p>\n<ol>\n<li>Please just run my notebook. I'm sorry I didn't provide the training code. </li>\n<li>Yes, please check <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/218148\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/218148</a></li>\n<li>I remember ii is the actual submission, which is the ensemble of 6 models. iii is the submission with single model.</li>\n</ol>",
          "rawMarkdown": "Hi Nimra, \n1.  Please just run my notebook. I'm sorry I didn't provide the training code. \n1. Yes, please check https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/218148\n1. I remember ii is the actual submission, which is the ensemble of 6 models. iii is the submission with single model.",
          "votes": 1
        },
        {
          "id": 1593182,
          "postDate": "2021-11-23T17:36:16.753Z",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> thanks for your reply.<br>\nActually I have run your notebook. Everything seems good, no errors. But from where can I see results? Do submission.csv file is getting updated every time I run the notebook?<br>\nAlso one more question, I am a newbie in machine learning and python. Can you suggest me what courses should I take or any thing else that can help me understand code?</p>",
          "rawMarkdown": "@mamasinkgs thanks for your reply.\nActually I have run your notebook. Everything seems good, no errors. But from where can I see results? Do submission.csv file is getting updated every time I run the notebook?\nAlso one more question, I am a newbie in machine learning and python. Can you suggest me what courses should I take or any thing else that can help me understand code?"
        }
      ]
    },
    {
      "id": 1158024,
      "postDate": "2021-01-18T10:35:01.017Z",
      "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> , Very Nice Approach of Using Transformer . Few Questions , I wanted to ask , if you don't mind sharing<br>\n1) How did you estimate New Users in Private Test set by probing ?<br>\n2) Is memory masking neccessary while training Model , what happens if We train Model without masking ?</p>",
      "rawMarkdown": "@mamasinkgs , Very Nice Approach of Using Transformer . Few Questions , I wanted to ask , if you don't mind sharing\n1) How did you estimate New Users in Private Test set by probing ?\n2) Is memory masking neccessary while training Model , what happens if We train Model without masking ?",
      "replies": [
        {
          "id": 1158203,
          "postDate": "2021-01-18T12:53:47.150Z",
          "content": "<p>1) I uploaded <a href=\"https://www.kaggle.com/mamasinkgs/submission-with-7000-limit?scriptVersionId=45368594\" target=\"_blank\">https://www.kaggle.com/mamasinkgs/submission-with-7000-limit?scriptVersionId=45368594</a>, so please check it. I tried binary search by changing the limit.<br>\n2) I thought models will learn some useless patterns. But, judging from others' solutions, precise masking is not necessarily needed. I think it depends on architectures, though.</p>",
          "rawMarkdown": "1) I uploaded https://www.kaggle.com/mamasinkgs/submission-with-7000-limit?scriptVersionId=45368594, so please check it. I tried binary search by changing the limit.\n2) I thought models will learn some useless patterns. But, judging from others' solutions, precise masking is not necessarily needed. I think it depends on architectures, though."
        },
        {
          "id": 1158346,
          "postDate": "2021-01-18T14:18:55.887Z",
          "content": "<p>Okay Thanks .</p>",
          "rawMarkdown": "Okay Thanks ."
        }
      ]
    },
    {
      "id": 1152889,
      "postDate": "2021-01-14T14:46:23.917Z",
      "content": "<p>Very nice approach, I was looking forward to this. A side question though, in this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195527#1071365\" target=\"_blank\">comment</a> it was inferred that you were not using a transformer. When did you shift to this? or was it a transformer all along?</p>",
      "rawMarkdown": "Very nice approach, I was looking forward to this. A side question though, in this [comment](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195527#1071365) it was inferred that you were not using a transformer. When did you shift to this? or was it a transformer all along?",
      "replies": [
        {
          "id": 1153011,
          "postDate": "2021-01-14T15:24:59.587Z",
          "content": "<p>Nice question! In the beginning of the competition, I used LightGBM as usual and got to 0.792 (1st place) soon. Then, I started using Transformer by reading this <a href=\"http://jalammar.github.io/illustrated-transformer/\" target=\"_blank\">blog</a>, soon I got to 0.801 using Transformer and 0.803 by ensembling GBDT and Transformer. I remember I wrote this comment then. (so, The fact is I was using both GBDT and Transformer). After that, my Transformer improved a lot and I decided to throw away GBDT and use my hand-crafted features for the additional input of Transformer. \"Is Attention All You Need?\" means I was wondering whether Transformer is better than GBDT when I wrote this comment.</p>",
          "rawMarkdown": "Nice question! In the beginning of the competition, I used LightGBM as usual and got to 0.792 (1st place) soon. Then, I started using Transformer by reading this [blog](http://jalammar.github.io/illustrated-transformer/), soon I got to 0.801 using Transformer and 0.803 by ensembling GBDT and Transformer. I remember I wrote this comment then. (so, The fact is I was using both GBDT and Transformer). After that, my Transformer improved a lot and I decided to throw away GBDT and use my hand-crafted features for the additional input of Transformer. \"Is Attention All You Need?\" means I was wondering whether Transformer is better than GBDT when I wrote this comment.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1150057,
      "postDate": "2021-01-12T10:28:19.493Z",
      "content": "<p>a dump question from a newbie. Which software do you use to draw that model structure graph? thank you !</p>",
      "rawMarkdown": "a dump question from a newbie. Which software do you use to draw that model structure graph? thank you !",
      "replies": [
        {
          "id": 1150308,
          "postDate": "2021-01-12T14:02:33.230Z",
          "content": "<p>I simply used powerpoint.</p>",
          "rawMarkdown": "I simply used powerpoint.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1148892,
      "postDate": "2021-01-11T13:07:43.350Z",
      "content": "<blockquote>\n  <p>I uploaded the names of 90 features as attachments.</p>\n</blockquote>\n<p>Where did you upload the attachments?</p>",
      "rawMarkdown": "> I uploaded the names of 90 features as attachments.\n\nWhere did you upload the attachments?",
      "replies": [
        {
          "id": 1148939,
          "postDate": "2021-01-11T13:44:16.993Z",
          "content": "<p>Thank you, I uploaded it now!</p>",
          "rawMarkdown": "Thank you, I uploaded it now!"
        }
      ]
    },
    {
      "id": 1148231,
      "postDate": "2021-01-11T02:46:41.373Z",
      "content": "<p>Congratulations and thank for detailed explanation!<br>\nMay I ask a question about how to deal with the tags features?<br>\nThank you.</p>",
      "rawMarkdown": "Congratulations and thank for detailed explanation!\nMay I ask a question about how to deal with the tags features?\nThank you.",
      "replies": [
        {
          "id": 1148245,
          "postDate": "2021-01-11T02:55:07.227Z",
          "content": "<p>Thank you! I used one-hot encoding (like multilabel problem) for tags and concatenate them with other features, then applied MLP.</p>",
          "rawMarkdown": "Thank you! I used one-hot encoding (like multilabel problem) for tags and concatenate them with other features, then applied MLP.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1146837,
      "postDate": "2021-01-10T05:09:53.260Z",
      "content": "<p>Congrats and thank you for your detailed explanation.<br>\nI have two questions about your solution.</p>\n<ul>\n<li>What is your cross-validation and model evaluating strategy?</li>\n<li>Did your task_container-aware masking improved CV score (and how much)?</li>\n</ul>\n<p>I think if we did not carefully split train/valid datasets and evaluate the model, CV score would be decreased when using task_container-aware masking, as the model cannot utilize the intra-container leakage you mentioned. </p>\n<p><a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">One of the major CV strategies in this competition</a> by <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> did not care task_container_id, though it split train/valid datasets by timestamp + random int.  [update] I confirmed that there are no such data at all.</p>",
      "rawMarkdown": "Congrats and thank you for your detailed explanation.\nI have two questions about your solution.\n- What is your cross-validation and model evaluating strategy?\n- Did your task_container-aware masking improved CV score (and how much)?\n\nI think if we did not carefully split train/valid datasets and evaluate the model, CV score would be decreased when using task_container-aware masking, as the model cannot utilize the intra-container leakage you mentioned. \n\n[One of the major CV strategies in this competition](https://www.kaggle.com/its7171/cv-strategy) by @its7171 did not care task_container_id, though it split train/valid datasets by timestamp + random int. ~~This CV strategy might split some task_containers into train/valid datasets, and might cause some intra-container leakage. (I should note that I used these datasets with good CV-LB relationship)~~ [update] I confirmed that there are no such data at all.",
      "replies": [
        {
          "id": 1146874,
          "postDate": "2021-01-10T06:03:54.940Z",
          "content": "<ul>\n<li><p>What is your cross-validation and model evaluating strategy?<br>\nMy CV strategy is completely different from tito's one. First, I probed the number of new users in test set and I found there is about 7000 new users. as we know there is 2.5M rows in test set and we know the average length of the history of all users, we can estimate the number of rows of the new users and the number of rows of the existing users. After the calculation, I found <br>\n<code># rows for user split (i.e. new users): # rows for timeseries split (i.e. existing users) = 2 : 1</code>. <br>\nSo, For validation, I decided to use 1M rows for timeseries split and 2M rows for user split. <br>\nWhen making validation set of timeseries split, I was so careful that inter-container leakage problem can't happen. I mean, my training dataset and validation dataset never shares same task_container_id for a specific user.</p></li>\n<li><p>Did your task_container-aware masking improved CV score (and how much)?<br>\nI don't know how much it improves CV/LB score, because I was really careful about inter-container leakage problem from the beginning and have never tried task_container-unaware masking and indexing (I hate leakage). So, there is no evidence that shows task_container-aware masking is important. That's just my guess from LB score.</p></li>\n</ul>",
          "rawMarkdown": "- What is your cross-validation and model evaluating strategy?\nMy CV strategy is completely different from tito's one. First, I probed the number of new users in test set and I found there is about 7000 new users. as we know there is 2.5M rows in test set and we know the average length of the history of all users, we can estimate the number of rows of the new users and the number of rows of the existing users. After the calculation, I found \n`# rows for user split (i.e. new users): # rows for timeseries split (i.e. existing users) = 2 : 1`. \nSo, For validation, I decided to use 1M rows for timeseries split and 2M rows for user split. \nWhen making validation set of timeseries split, I was so careful that inter-container leakage problem can't happen. I mean, my training dataset and validation dataset never shares same task_container_id for a specific user.\n\n- Did your task_container-aware masking improved CV score (and how much)?\nI don't know how much it improves CV/LB score, because I was really careful about inter-container leakage problem from the beginning and have never tried task_container-unaware masking and indexing (I hate leakage). So, there is no evidence that shows task_container-aware masking is important. That's just my guess from LB score.\n",
          "votes": 3
        },
        {
          "id": 1146949,
          "postDate": "2021-01-10T07:48:52.647Z",
          "content": "<p>Thank you for your clarification.<br>\nIt is so sophisticated CV strategy and very logical consequence.</p>",
          "rawMarkdown": "Thank you for your clarification.\nIt is so sophisticated CV strategy and very logical consequence."
        },
        {
          "id": 1147497,
          "postDate": "2021-01-10T14:47:40.403Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a>,</p>\n<blockquote>\n  <p>This CV strategy might split some task_containers into train/valid datasets</p>\n</blockquote>\n<p>Please check whether this kind of data is exists or not by your self if you think so.</p>\n<p>I believe there is no such data. :)</p>\n<p>Since:</p>\n<ul>\n<li>questions with same task have same timestamp</li>\n<li>train and validation are split after sorted by time</li>\n</ul>",
          "rawMarkdown": "Hi @tomooinubushi,\n\n> This CV strategy might split some task_containers into train/valid datasets\n\nPlease check whether this kind of data is exists or not by your self if you think so.\n\nI believe there is no such data. :)\n\nSince:\n* questions with same task have same timestamp\n* train and validation are split after sorted by time",
          "votes": 3
        },
        {
          "id": 1148181,
          "postDate": "2021-01-11T00:59:28.843Z",
          "content": "<p>I am sorry. I should have checked it before making a sloppy speculation.<br>\nI found there are no data which share the same timestamp or task_container_id between train/valid datasets. So, there are no intra-container leakage at all.<br>\nI found some of task_container_ids in valid dataset are lower than those in train dataset, which might be caused by known <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465\" target=\"_blank\">inconsistency between timestamp and task_container_id</a>. I am not sure how to deal with these data when creating attention masks.<br>\nAnyway, thank you for pointing out my mistake <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> and congrats to <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>. </p>",
          "rawMarkdown": "I am sorry. I should have checked it before making a sloppy speculation.\nI found there are no data which share the same timestamp or task_container_id between train/valid datasets. So, there are no intra-container leakage at all.\nI found some of task_container_ids in valid dataset are lower than those in train dataset, which might be caused by known [inconsistency between timestamp and task_container_id](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465). I am not sure how to deal with these data when creating attention masks.\nAnyway, thank you for pointing out my mistake @its7171 and congrats to @mamasinkgs. ",
          "votes": 1
        },
        {
          "id": 1148287,
          "postDate": "2021-01-11T03:48:54.867Z",
          "content": "<p>Thank you for your confirmation!</p>",
          "rawMarkdown": "Thank you for your confirmation!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1146747,
      "postDate": "2021-01-10T02:35:49.287Z",
      "content": "<p>thanks for sharing! a great job!  this is a big project!</p>",
      "rawMarkdown": "thanks for sharing! a great job!  this is a big project!"
    },
    {
      "id": 1146571,
      "postDate": "2021-01-09T21:34:07.693Z",
      "content": "<p><a href=\"https://www.kaggle.com/mamasisking\" target=\"_blank\">@mamasisking</a> , congratulations! Can you please also share info about the resources you put into the competition? hardware, amount of effort, time etc. </p>",
      "rawMarkdown": "@mamasisking , congratulations! Can you please also share info about the resources you put into the competition? hardware, amount of effort, time etc. ",
      "replies": [
        {
          "id": 1146590,
          "postDate": "2021-01-09T21:46:33.083Z",
          "content": "<p>I use V100 * 4 (VRAM 64GB) and training takes about 30h. I think I spent more than 500 hours for this competition. I did nothing other than this competition since 10/16.</p>",
          "rawMarkdown": "I use V100 * 4 (VRAM 64GB) and training takes about 30h. I think I spent more than 500 hours for this competition. I did nothing other than this competition since 10/16.",
          "votes": 9
        },
        {
          "id": 1147123,
          "postDate": "2021-01-10T09:55:45.060Z",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> You don't have a full-time job currently …? I can only work on this competition after my daily job …😭</p>",
          "rawMarkdown": "@mamasinkgs You don't have a full-time job currently ...? I can only work on this competition after my daily job ...😭",
          "votes": 1
        },
        {
          "id": 1147357,
          "postDate": "2021-01-10T13:05:35.330Z",
          "content": "<p>I'm a student 😄</p>",
          "rawMarkdown": "I'm a student 😄",
          "votes": 3
        },
        {
          "id": 1167076,
          "postDate": "2021-01-24T02:52:43.677Z",
          "content": "<p>Tesla V100 seems not cheap. That was a lot investment money and time wise!</p>",
          "rawMarkdown": "Tesla V100 seems not cheap. That was a lot investment money and time wise!"
        }
      ]
    },
    {
      "id": 1146503,
      "postDate": "2021-01-09T20:03:23.923Z",
      "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> , sorry to bother, but could you share your insight about this statement</p>\n<pre><code>In my implementation, SAINT didn't work well. In my opinion, it's logically very hard to implement SAINT without data leakage in training due to task_container_id.\n</code></pre>\n<p>Why you think <code>task_container_id</code> cause data leakage during training, and what kind of leakage?</p>",
      "rawMarkdown": "@mamasinkgs , sorry to bother, but could you share your insight about this statement\n\n```\nIn my implementation, SAINT didn't work well. In my opinion, it's logically very hard to implement SAINT without data leakage in training due to task_container_id.\n```\n\nWhy you think `task_container_id` cause data leakage during training, and what kind of leakage?",
      "replies": [
        {
          "id": 1146517,
          "postDate": "2021-01-09T20:14:54.737Z",
          "content": "<p>I mean, in my opinion, the input of the decoder of SAINT+ architecture cannot be implemented without any information loss or any data leakage, due to task_container_id. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F762938%2F8467453b525282d6754cbac4f2c9473e%2Finput_of_saint.jpg?generation=1610223284475779&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "I mean, in my opinion, the input of the decoder of SAINT+ architecture cannot be implemented without any information loss or any data leakage, due to task_container_id. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F762938%2F8467453b525282d6754cbac4f2c9473e%2Finput_of_saint.jpg?generation=1610223284475779&alt=media)",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1146558,
      "author_name": "Andrés Miguel Torrubia Sáez",
      "author_url": "",
      "post_date": "2021-01-09T21:16:22.747000",
      "content": "<p>Good luck WINNING your next competition.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1147926,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2021-01-10T20:03:30.223000",
      "content": "<p>Thanks for the nicely detailed write-up and congratz for the impressive finish!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1146546,
      "author_name": "Nikola Bacic",
      "author_url": "",
      "post_date": "2021-01-09T21:03:17.123000",
      "content": "<blockquote>\n  <p>In my implementation, SAINT didn't work well. In my opinion, it's logically very hard to implement SAINT without data leakage in training due to task_container_id.</p>\n</blockquote>\n<p>Couldn't you control that with <code>tgt_mask</code> from PyTorch's  <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html#torch.nn.Transformer\" target=\"_blank\">implementation</a> ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1146561,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-09T21:21:10.327000",
          "content": "<p>Yes we can control <code>tgt_mask</code> in decoder. but if we use the same input of the decoder as SAINT+ paper in this image, I think the problem cannot be fixed just by controlling <code>tgt_mask</code>, because the <strong>query</strong> of decoder produces leakage (i.e. uses the information of the sample in the same task_container_id). In my understanding, the data leakage caused by <strong>query</strong> is unavoidable by controlling masking, while the data leakage caused by <strong>key/value</strong> is avoidable by masking.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F762938%2Fb29c0c24c02cb5b71707c2d378bc5a13%2Finput_of_saint.jpg?generation=1610226842294243&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1146572,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-09T21:34:12.120000",
          "content": "<p>If I understand what you are saying, you mean for questions in the same bundle that are under prediction (but during training), the corresponding input in the decoder will see the information (i.e. user correctness) in the same bundle (but prior the current position). However, while during inference, it can only see this information at the places before the current bundle.</p>\n<p>I am using the encoder-decoder approach, and during the inference time, I replace the values by the value obtained at the first question in the bundle under prediction. Basically, it treats the question in the bundle as separate questions, and use the history before them.</p>\n<p>Anyway, the best I can get is 0.802, so potentially there are some problems as you mentioned.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1146587,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-09T21:44:02.640000",
          "content": "<p>Yes of course the leakage doesn't happen during inference, but I think it's problematic that leakage happens in training phase, because the model can learn meaningless patterns during training.</p>\n<p><code>I replace the values by the value obtained at the first question in the bundle under prediction.</code><br>\nHmm, I didn't notice this idea but this looks great. Why don't you try this idea in the training phase? I think it is possible by fancy indexing. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1146606,
          "author_name": "Nikola Bacic",
          "author_url": "",
          "post_date": "2021-01-09T21:58:27.887000",
          "content": "<p>I'm still not seeing which way the leak could go :?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1146610,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-09T22:02:29.533000",
          "content": "<p>I didn't think of the training will be problematic. I am aware of some kind of inconsistency between decoder  training / inference input by this approach, but it is a comprising approach. Treating each question in the bundle under the current inference time step is not difficult, although still need to be careful.</p>\n<p>To do so in training, I feel it becomes more difficult. Because we are not just dealing the last bundle in the truncated sequence, we extract exact one question from each budle in that sequence. Maybe not so difficult, but I kind feel we also loss information that given by bundle. Anyway, without trying, I can't say if we avoid leakage or lossing info if we apply this approach to training</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1146611,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-09T22:05:18.623000",
          "content": "<p><a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a><br>\nFor example, suppose task_container_id of a specific user is [0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 5, 8, 7, 7, 6]. Then, <br>\nposition [0, 1, 2] cannot use any information.<br>\nposition [3, 4] can use the information of [0, 1, 2].<br>\nposition [5, 6] can use the information of [0, 1, 2, 3, 4].<br>\nposition [7, 8] can use the information of [0, 1, 2, 3, 4, 5, 6].</p>\n<p>However, when simply using the input of the SAINT paper, position 1 can use the information (e.g. correctness) of position 0 and position 2 can use the information of position [0, 1]. I think this is problematic as the model uses the information that should not be utilized, during training.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1146612,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-09T22:06:17.623000",
          "content": "<p>Not clear to me about what kind of leakage neither, but at least some form of inconsistency exist</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1146617,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-09T22:12:44.273000",
          "content": "<p>Yes, it can be called \"incosistency\", as the problem doesn't happen at least in inference phase. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1146618,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-09T22:14:48.053000",
          "content": "<p>To be clear, my approach mentioned above , for decoder, only applies to the last bundle in the truncated sequence. The bundles before it, even they appear in inference time, but already being predicted, the special processing is not applied to them anymore</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1146621,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-09T22:19:53.660000",
          "content": "<pre><code>To do so in training, I feel it becomes more difficult. Because we are not just dealing the last bundle in the truncated sequence, we extract exact one question from each budle in that sequence. Maybe not so difficult, but I kind feel we also loss information that given by bundle.\n</code></pre>\n<p>Now I understand what you mean. it seems the problem cannot be solved by simply using fancy indexing. Hmm…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1146623,
          "author_name": "Nikola Bacic",
          "author_url": "",
          "post_date": "2021-01-09T22:20:17.523000",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> <br>\nYes, I agree. But what confusses me is why tgt_mask can't control that. In example you mentioned, it should look like this:</p>\n<p>[0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 5, 8, 7, 7, 6]</p>\n<p>(1 mask, 0 no mask)</p>\n<p>1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 (for 0 position)<br>\n1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 (for 1 position)<br>\n1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 (for 2 position)<br>\n0 0 0 1 1 1 1 1 1 1 1 1 1 1 1 (for 3 position)<br>\n0 0 0 1 1 1 1 1 1 1 1 1 1 1 1 (for 4 position)<br>\n0 0 0 0 0 1 1 1 1 1 1 1 1 1 1 (for 5 position)<br>\n0 0 0 0 0 1 1 1 1 1 1 1 1 1 (for 6 position)<br>\netc.</p>\n<p>Although in my solution, I'm only using encoder part of the transformer, but that shouldn't make any difference.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1146630,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-09T22:28:16.987000",
          "content": "<p><a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a> <br>\nWhen using encoder-only model, there is no problem. However, when using decoder of SAINT+ model, the sequence is <strong>shifted</strong> like the image above. So, by simply using tgt_mask you showed (actually the mask you showed is inherently same as the mask I showed in my solution), the problem cannot be solved.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1146633,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-09T22:32:39.940000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1146640,
          "author_name": "Nikola Bacic",
          "author_url": "",
          "post_date": "2021-01-09T22:46:30.053000",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a><br>\nYes, I see now. My implementation of this was only partly good because even though I did mask positions from the same task_container, responses were shifted to the right and question could always see previous response, which could be from the same task_container (since I used encoder only, position could always see itself)<br>\nI think I should've put responses into a encoder, and the rest of the features into decoder and control task_container with memory_mask.</p>\n<p>Thank you for your time!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1146647,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-09T22:57:28.953000",
          "content": "<p><a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a> ， just a remark, the mask you shows for decoder also has an issue. For example , at a place, it can't look at the current position. Therefore, position embedding of that place becomes no meaning while predicting that place, if we use the usual way of position embedding</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1146651,
          "author_name": "Nikola Bacic",
          "author_url": "",
          "post_date": "2021-01-09T23:06:13.917000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> Yes, you're right. But I couldn't get anything out of positions anyway.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1146767,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-01-10T03:13:26.980000",
      "content": "<p>Big congrats to you on 2nd solo prize. Thanks for the detail writeup, very intelligence solution <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1146439,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2021-01-09T19:19:19.120000",
      "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> , thank you for sharing. See you in the next one!</p>\n<p>About your CPython implementation for getting the masks, it is amazing, you must be an advanced Python programmer!<br>\nFor me, I use TensorFlow, and it is quite straightforward and fast to get these mask by using tf operations, like</p>\n<pre><code>def get_causal_attention_mask(nd, ns, dtype, only_before):\n    \"\"\"\n    1's in the lower triangle, counting from the lower right corner. Same as tf.matrix_band_part(tf.ones([nd, ns]),\n    -1, ns-nd), but doesn't produce garbage on TPUs.\n    \"\"\"\n\n    # Remark: Think `nd` as the number of queries and `ns` as the number of keys.\n    # In encoder-decoder case, the queries are the decoder features and the keys are the encoder features.\n\n    i = tf.range(nd)[:, tf.newaxis]  # repeat along dim 1\n    j = tf.range(ns) # repeat along dim 0 \n    m = i &gt;= (j - ns + nd) + tf.cast(only_before, dtype=tf.int32)\n\n    return tf.cast(m, dtype)\n\ndef get_attention_mask_from_timestamp_batch(timestamp_tensors, dtype, only_before):\n    \"\"\"\n    Args:\n        timestamp_tensors: 2-D tf.int32 tensor, representing a batch of sequences of non-decreasing timestamps.\n\n    Returns:\n        attention_mask: 3-D tf.int32 tensor of shape = [batch_size, query_len, key_len], consisting of 0 and 1.\n            Here `query_len` and `key_len` are actually `seq_len`. It should be reshpaed, when used to calculate \n            attention scores, to [batch_size, nb_attn_head, query_len, key_len].\n    \"\"\"\n\n    t = timestamp_tensors\n\n    batch_size = tf.math.reduce_sum(tf.ones_like(t[:, :1], dtype=tf.int32))\n    seq_len = tf.math.reduce_sum(tf.ones_like(t[:1, :], dtype=tf.int32))\n\n    x = tf.broadcast_to(t[:, :, tf.newaxis], shape=[batch_size, seq_len, seq_len]) # repeat along dim 2\n    y = tf.broadcast_to(t[:, tf.newaxis, :], shape=[batch_size, seq_len, seq_len]) + tf.cast(only_before, dtype=tf.int64) # repeat along dim 1\n\n    m =  x &gt;= y\n\n    return tf.cast(m, dtype)\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 1146459,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-09T19:36:48.403000",
          "content": "<p>Thanks! I'm not familiar with tensorflow, but I'll try it in future competitions.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1590856,
      "author_name": "Nimra Tassawar",
      "author_url": "",
      "post_date": "2021-11-21T17:37:10.020000",
      "content": "<p>Hi, I have few questions</p>\n<ol>\n<li>Can you guide me how to run the code you provided? and how to get results from it.</li>\n<li>Did you published any research paper on this implementation. I want to understand your implementation in detail.</li>\n<li>At kaggle, I have seen your 3 submissions related to this competition. <br>\ni. <a href=\"https://www.kaggle.com/mamasinkgs/2nd-place-solution-for-hosts\" target=\"_blank\">https://www.kaggle.com/mamasinkgs/2nd-place-solution-for-hosts</a><br>\nii. <a href=\"https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution\" target=\"_blank\">https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution</a><br>\niii. <a href=\"https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold\" target=\"_blank\">https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold</a><br>\nI want to know which one is the actual and correct submission?</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 1591061,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-11-21T22:53:26.467000",
          "content": "<p>Hi Nimra, </p>\n<ol>\n<li>Please just run my notebook. I'm sorry I didn't provide the training code. </li>\n<li>Yes, please check <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/218148\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/218148</a></li>\n<li>I remember ii is the actual submission, which is the ensemble of 6 models. iii is the submission with single model.</li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1593182,
          "author_name": "Nimra Tassawar",
          "author_url": "",
          "post_date": "2021-11-23T17:36:16.753000",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> thanks for your reply.<br>\nActually I have run your notebook. Everything seems good, no errors. But from where can I see results? Do submission.csv file is getting updated every time I run the notebook?<br>\nAlso one more question, I am a newbie in machine learning and python. Can you suggest me what courses should I take or any thing else that can help me understand code?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1158024,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2021-01-18T10:35:01.017000",
      "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> , Very Nice Approach of Using Transformer . Few Questions , I wanted to ask , if you don't mind sharing<br>\n1) How did you estimate New Users in Private Test set by probing ?<br>\n2) Is memory masking neccessary while training Model , what happens if We train Model without masking ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1158203,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-18T12:53:47.150000",
          "content": "<p>1) I uploaded <a href=\"https://www.kaggle.com/mamasinkgs/submission-with-7000-limit?scriptVersionId=45368594\" target=\"_blank\">https://www.kaggle.com/mamasinkgs/submission-with-7000-limit?scriptVersionId=45368594</a>, so please check it. I tried binary search by changing the limit.<br>\n2) I thought models will learn some useless patterns. But, judging from others' solutions, precise masking is not necessarily needed. I think it depends on architectures, though.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1158346,
          "author_name": "Athar Sayed",
          "author_url": "",
          "post_date": "2021-01-18T14:18:55.887000",
          "content": "<p>Okay Thanks .</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1152889,
      "author_name": "Wisso",
      "author_url": "",
      "post_date": "2021-01-14T14:46:23.917000",
      "content": "<p>Very nice approach, I was looking forward to this. A side question though, in this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195527#1071365\" target=\"_blank\">comment</a> it was inferred that you were not using a transformer. When did you shift to this? or was it a transformer all along?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1153011,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-14T15:24:59.587000",
          "content": "<p>Nice question! In the beginning of the competition, I used LightGBM as usual and got to 0.792 (1st place) soon. Then, I started using Transformer by reading this <a href=\"http://jalammar.github.io/illustrated-transformer/\" target=\"_blank\">blog</a>, soon I got to 0.801 using Transformer and 0.803 by ensembling GBDT and Transformer. I remember I wrote this comment then. (so, The fact is I was using both GBDT and Transformer). After that, my Transformer improved a lot and I decided to throw away GBDT and use my hand-crafted features for the additional input of Transformer. \"Is Attention All You Need?\" means I was wondering whether Transformer is better than GBDT when I wrote this comment.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1150057,
      "author_name": "yoyoyo",
      "author_url": "",
      "post_date": "2021-01-12T10:28:19.493000",
      "content": "<p>a dump question from a newbie. Which software do you use to draw that model structure graph? thank you !</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1150308,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-12T14:02:33.230000",
          "content": "<p>I simply used powerpoint.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1148892,
      "author_name": "Nya 🚀",
      "author_url": "",
      "post_date": "2021-01-11T13:07:43.350000",
      "content": "<blockquote>\n  <p>I uploaded the names of 90 features as attachments.</p>\n</blockquote>\n<p>Where did you upload the attachments?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1148939,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-11T13:44:16.993000",
          "content": "<p>Thank you, I uploaded it now!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1148231,
      "author_name": "Dean",
      "author_url": "",
      "post_date": "2021-01-11T02:46:41.373000",
      "content": "<p>Congratulations and thank for detailed explanation!<br>\nMay I ask a question about how to deal with the tags features?<br>\nThank you.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1148245,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-11T02:55:07.227000",
          "content": "<p>Thank you! I used one-hot encoding (like multilabel problem) for tags and concatenate them with other features, then applied MLP.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1146837,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2021-01-10T05:09:53.260000",
      "content": "<p>Congrats and thank you for your detailed explanation.<br>\nI have two questions about your solution.</p>\n<ul>\n<li>What is your cross-validation and model evaluating strategy?</li>\n<li>Did your task_container-aware masking improved CV score (and how much)?</li>\n</ul>\n<p>I think if we did not carefully split train/valid datasets and evaluate the model, CV score would be decreased when using task_container-aware masking, as the model cannot utilize the intra-container leakage you mentioned. </p>\n<p><a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">One of the major CV strategies in this competition</a> by <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> did not care task_container_id, though it split train/valid datasets by timestamp + random int.  [update] I confirmed that there are no such data at all.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1146874,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-10T06:03:54.940000",
          "content": "<ul>\n<li><p>What is your cross-validation and model evaluating strategy?<br>\nMy CV strategy is completely different from tito's one. First, I probed the number of new users in test set and I found there is about 7000 new users. as we know there is 2.5M rows in test set and we know the average length of the history of all users, we can estimate the number of rows of the new users and the number of rows of the existing users. After the calculation, I found <br>\n<code># rows for user split (i.e. new users): # rows for timeseries split (i.e. existing users) = 2 : 1</code>. <br>\nSo, For validation, I decided to use 1M rows for timeseries split and 2M rows for user split. <br>\nWhen making validation set of timeseries split, I was so careful that inter-container leakage problem can't happen. I mean, my training dataset and validation dataset never shares same task_container_id for a specific user.</p></li>\n<li><p>Did your task_container-aware masking improved CV score (and how much)?<br>\nI don't know how much it improves CV/LB score, because I was really careful about inter-container leakage problem from the beginning and have never tried task_container-unaware masking and indexing (I hate leakage). So, there is no evidence that shows task_container-aware masking is important. That's just my guess from LB score.</p></li>\n</ul>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1146949,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2021-01-10T07:48:52.647000",
          "content": "<p>Thank you for your clarification.<br>\nIt is so sophisticated CV strategy and very logical consequence.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1147497,
          "author_name": "tito",
          "author_url": "",
          "post_date": "2021-01-10T14:47:40.403000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a>,</p>\n<blockquote>\n  <p>This CV strategy might split some task_containers into train/valid datasets</p>\n</blockquote>\n<p>Please check whether this kind of data is exists or not by your self if you think so.</p>\n<p>I believe there is no such data. :)</p>\n<p>Since:</p>\n<ul>\n<li>questions with same task have same timestamp</li>\n<li>train and validation are split after sorted by time</li>\n</ul>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1148181,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2021-01-11T00:59:28.843000",
          "content": "<p>I am sorry. I should have checked it before making a sloppy speculation.<br>\nI found there are no data which share the same timestamp or task_container_id between train/valid datasets. So, there are no intra-container leakage at all.<br>\nI found some of task_container_ids in valid dataset are lower than those in train dataset, which might be caused by known <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465\" target=\"_blank\">inconsistency between timestamp and task_container_id</a>. I am not sure how to deal with these data when creating attention masks.<br>\nAnyway, thank you for pointing out my mistake <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> and congrats to <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a>. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1148287,
          "author_name": "tito",
          "author_url": "",
          "post_date": "2021-01-11T03:48:54.867000",
          "content": "<p>Thank you for your confirmation!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1146747,
      "author_name": "2981",
      "author_url": "",
      "post_date": "2021-01-10T02:35:49.287000",
      "content": "<p>thanks for sharing! a great job!  this is a big project!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1146571,
      "author_name": "kobi2000",
      "author_url": "",
      "post_date": "2021-01-09T21:34:07.693000",
      "content": "<p><a href=\"https://www.kaggle.com/mamasisking\" target=\"_blank\">@mamasisking</a> , congratulations! Can you please also share info about the resources you put into the competition? hardware, amount of effort, time etc. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1146590,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-09T21:46:33.083000",
          "content": "<p>I use V100 * 4 (VRAM 64GB) and training takes about 30h. I think I spent more than 500 hours for this competition. I did nothing other than this competition since 10/16.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1147123,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-10T09:55:45.060000",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> You don't have a full-time job currently …? I can only work on this competition after my daily job …😭</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1147357,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-10T13:05:35.330000",
          "content": "<p>I'm a student 😄</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1167076,
          "author_name": "Toddgm",
          "author_url": "",
          "post_date": "2021-01-24T02:52:43.677000",
          "content": "<p>Tesla V100 seems not cheap. That was a lot investment money and time wise!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1146503,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2021-01-09T20:03:23.923000",
      "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> , sorry to bother, but could you share your insight about this statement</p>\n<pre><code>In my implementation, SAINT didn't work well. In my opinion, it's logically very hard to implement SAINT without data leakage in training due to task_container_id.\n</code></pre>\n<p>Why you think <code>task_container_id</code> cause data leakage during training, and what kind of leakage?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1146517,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2021-01-09T20:14:54.737000",
          "content": "<p>I mean, in my opinion, the input of the decoder of SAINT+ architecture cannot be implemented without any information loss or any data leakage, due to task_container_id. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F762938%2F8467453b525282d6754cbac4f2c9473e%2Finput_of_saint.jpg?generation=1610223284475779&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1146355": "Thank you all teams who competed with me, all the people who participated in this competition, and the organizer that hosts such a great competition with the well-designed API!\nCongrats @keetar, who defeats me and becomes the winner in this competition.\n\nI'm happy because it's my first time I get solo prize!\n\nIt's my 4th kaggle competition and it was fun to compete with my past teammates ( @nyanpn, @pocketsuteado) and people who I competed with past competitions (e.g. @aerdem4, @its7171). \nHere, I will explain the summary of my model and my features. \n\nI uploaded 2 kaggle notebooks for the explanation. As the ensemble is not so important in my solution, I will only explain my single model. \n1. [6 similar models weighted average](https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution) : 0.817 public/0.818 private\n2. [single model](https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold) : 0.814 public/0.816 private\n\n#Models \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F762938%2F39ec58b701e6d37701cc794763d3a473%2F2nd_place.png?generation=1610213712807228&alt=media)\n\n##Overview \nmy model is similar to SAKT, with 400 sequence length, 512 dimension, 4 nheads. \nI don't use lecture information for the input of transformer model, which I guess is why I lost in this competition. For the query and key/value of SAKT-like model, I used LSTM-encoded features, whose input is as follows.\n- **\"Query\" features**\ncontent_id\npart\ntags\nnormalized timedelta\nnormalized log timestamp \ncorrect answer\ntask_container_id delta\ncontent_type_id delta\nnormalized absolute position \n\n- **\"Memory\" features**\nexplanation\ncorrectness\nnormalized elapsed time\nuser_answer\n\n## Detailed Explanation of Training/Inference Process\nI tried a very precise indexing/masking technique to avoid data leakage in training process, I guess which is partly why I became 2nd place. OK, suppose a very simple model of the sequence length = 5, and a task_container_id history of a specific user (without lecture) is like this.\n`[0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 5, 8, 7, 7, 6]`\nAs there is 15 measurements in this user, I made 15/5 = 3 training samples and loss mask.\n\n(1) input task_container_id: `[pad, pad, pad, pad, pad, 0, 0, 0, 1, 1]` \n(1) loss_mask: `[False, False, False, False, False, True, True, True, True, True]`\n(2) input task_container_id: `[0, 0, 0, 1, 1, 2, 2, 3, 3, 4]` \n(2) loss_mask: `[False, False, False, False, False, True, True, True, True, True]`\n(3) input task_container_id: `[2, 2, 3, 3, 4, 5, 8, 7, 7, 6]` \n(3) loss_mask: `[False, False, False, False, False, True, True, True, True, True]`\n\nIn this competition, the handling of the task_container_id is very important, as **it is not allowed to use \"memory\" features of the same task container id** to avoid leakage.\nSo, after applying LSTM to the features, such fancy indexing is required to avoid the leakage for 3 training samples, where -1 means this position can't attend any position.\n\n(1) indices: `[-1, -1, -1, -1, -1, 4, 4, 4, 7, 7]`\n(2) indices: `[-1, -1, -1, 2, 2, 4, 4, 6, 6, 8]`\n(3) indices: `[-1, -1, 1, 1, 3, 4, 5, 6, 6, 8]`\n\nTo get this indices very fast in the training/inference phase, I wrote such cython (#1) function (vectorized implementation):\n\n```\n%%cython\nimport numpy as np\ncimport numpy as np\ncpdef np.ndarray[int] cget_memory_indices(np.ndarray task):\n    cdef Py_ssize_t n = task.shape[1]\n    cdef np.ndarray[int, ndim = 2] res = np.zeros_like(task, dtype = np.int32)\n    cdef np.ndarray[int] tmp_counter = np.full(task.shape[0], -1, dtype = np.int32)\n    cdef np.ndarray[int] u_counter = np.full(task.shape[0], task.shape[1] - 1, dtype = np.int32)\n    for i in range(n):\n        res[:, i] = u_counter\n        tmp_counter += 1\n        if i != n - 1:\n            mask = (task[:, i] != task[:, i + 1])\n            u_counter[mask] = tmp_counter[mask]\n    return res\n```\nAfter applying such fancy indexing to the output of LSTM features, I concatenated them with \"Query\" features and apply MLP, then we can get query for SAKT model. Then, I obtained key/value for SAKT model by applying MLP to query concatenated with \"Memory\" features. \n\nTo train SAKT-like model, precise memory masking is also required to avoid leakage. I used [3D attention mask](https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html) of torch.nn.MultiheadAttention. It should be noted [torch.repeat_interleave](https://pytorch.org/docs/stable/generated/torch.repeat_interleave.html) must be leveraged to make 3D mask (batchsize * nhead, sequence length, sequence length). \nThe memory mask for these 3 samples are like this. Here, 1 means True and 0 means False. Please remember the [documentation](https://pytorch.org/docs/stable/generated/torch.nn.MultiheadAttention.html) says \n``` \nattn_mask ensure that position i is allowed to attend the unmasked positions. If a BoolTensor is provided, positions with True is not allowed to attend while False values will be unchanged.\n```\n(1) memory_mask:\n```\narray([[1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [1, 1, 1, 0, 0, 0, 0, 0, 1, 1],\n       [1, 1, 1, 0, 0, 0, 0, 0, 1, 1]])\n```\n(2) memory_mask:\n```\narray([[1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [0, 0, 0, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 1, 1, 0, 0, 0, 0, 0, 1]])\n```\n(3)  memory_mask:\n```\narray([[1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [1, 1, 1, 1, 1, 1, 1, 1, 1, 0],\n       [0, 0, 1, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 1, 1, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 1, 1, 1, 1, 1, 1],\n       [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],\n       [1, 0, 0, 0, 0, 0, 1, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 0, 0, 0, 0, 0, 1, 1, 1],\n       [1, 1, 1, 1, 0, 0, 0, 0, 0, 1]])\n```\n\nTo make this memory mask very fast in the training phase, I wrote such cython (#2) function (not vectorized implementation):\n```\n%%cython\nimport numpy as np\ncimport numpy as np\ncpdef np.ndarray[int] cget_memory_mask(np.ndarray task, int n_length):\n    cdef Py_ssize_t n = task.shape[0]\n    cdef np.ndarray[int, ndim = 2] res = np.full((n, n), 1, dtype = np.int32)\n    cdef int tmp_counter = 0\n    cdef int u_counter = 0\n    for i in range(n):\n        tmp_counter += 1\n        if i == n - 1 or task[i] != task[i + 1]:\n            res[i - tmp_counter + 1 : i + 1, :u_counter] = 0\n            if u_counter == 0:\n                res[i - tmp_counter + 1 : i + 1, n - 1] = 0\n            if u_counter > n_length:\n                res[i - tmp_counter + 1: i + 1, :(u_counter - n_length)] = 1\n            u_counter += tmp_counter\n            tmp_counter = 0\n    return res\n```\n\nIn the inference phase, memory mask is not required. For more details, please look at my [single model notebook](https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold).\n\n#My Features\nTransformer is great, but it suffers from a problem that the sequence length cannot be infinite. I mean, it cannot consider the information of very old samples. To tackle this problem and to leverage the lecture information, I made simple features and concatenated them with the output features of SAKT-like model, and applied MLP and sigmoid.\nI uploaded the names of 90 features as attachments. For the implementation of these features, I didn't use any pd.merge or df.join and most of the implementation are done using numpy. To make user-content features, I used scipy.sparse.lil_matrix to spare memory usage. I think my feature is not so good as other competitors (about 0.795 when using GBDT), but still improved score 0.001 ~ 0.002.\nI made some tricky features (e.g. obtained by SVD), but it did not improve the score of NN model (improved GBDT model, though.), probably because the information of NN features includes that of such tricky features.\n\n#CV Strategy\nMy CV strategy is completely different from tito( @its7171 )'s one. First, I probed the number of new users in test set and I found there are about 7000 new users. as we know there is 2.5M rows in test set and we know the average length of the history of all users, we can estimate the number of rows of the new users and the number of rows of the existing users. After the calculation, I found\n```\nthe number of rows of new users (i.e. user split): the number of rows of existing users (i.e. timeseries split) = 2 : 1.\n```\nSo, For validation, I decided to use 1M rows for timeseries split and 2M rows for user split.\nWhen making validation set of timeseries split, I was so careful that leakage can't happen. I mean, my training dataset and validation dataset never shares same task_container_id for a given user.\n\n#What worked \n1. increase length from 100 to 400 worked.\n2. using normalized timedelta is better than digitized timedelta (mentioned in SAINT+ paper).\n3. concatenating embeddings is better than adding embedding, as mentioned in https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201798. \n4. dropout = 0.2 is very important.\n5. StepLR with Adam is good. I trained my model for about 35 epoch with lr = 2e-3, then trained it for 1 epoch with lr = 2e-4.\n\n#What didn't work\n1. Random masking of sequences didn't improve the score.\n2. bundle_id and normalized task_container_id is not needed for the input of LSTM.\n3. As I didn't make diverse models, weighted averaging is enough and blending using GBDT didn't work well.\n4. Full data training (without any validation data) didn't improve the score.\n5. In my implementation, SAINT didn't work well. In my opinion, it's logically very hard to implement SAINT without data leakage in training due to task_container_id.\n\n#Comments\nI was very sad to see the private score bug problem of kaggle. Of course this competition is very great, but I think this competition could have been one of the most successful competition in kaggle without this problem.\nAnyway, I really enjoyed my 4th kaggle competition. I'll continue kaggle and will surely become the winner in the next competition!\n\n<h1>Thanks all, see you again!</h1>\n\n\nP.S. The word `memory_mask` may be confusing because it is different from pytorch [Transformer](https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html#torch.nn.Transformer)'s `memory_mask`.\nThe word `memory_mask` is more like `mask` in pytorch's [TransformerEncoder](https://pytorch.org/docs/stable/generated/torch.nn.TransformerEncoder.html).\nThe reason my word is confusing is, I used Encoder-Decoder model at first, then gave up it and started using Encoder-only model, but I didn't change the function names. :(\n",
    "1146558": "Good luck WINNING your next competition.",
    "1147926": "Thanks for the nicely detailed write-up and congratz for the impressive finish!",
    "1146546": "> In my implementation, SAINT didn't work well. In my opinion, it's logically very hard to implement SAINT without data leakage in training due to task_container_id.\n\nCouldn't you control that with `tgt_mask` from PyTorch's  [implementation](https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html#torch.nn.Transformer) ?",
    "1146767": "Big congrats to you on 2nd solo prize. Thanks for the detail writeup, very intelligence solution @mamasinkgs ",
    "1146439": "@mamasinkgs , thank you for sharing. See you in the next one!\n\nAbout your CPython implementation for getting the masks, it is amazing, you must be an advanced Python programmer!\nFor me, I use TensorFlow, and it is quite straightforward and fast to get these mask by using tf operations, like\n\n```\ndef get_causal_attention_mask(nd, ns, dtype, only_before):\n    \"\"\"\n    1's in the lower triangle, counting from the lower right corner. Same as tf.matrix_band_part(tf.ones([nd, ns]),\n    -1, ns-nd), but doesn't produce garbage on TPUs.\n    \"\"\"\n    \n    # Remark: Think `nd` as the number of queries and `ns` as the number of keys.\n    # In encoder-decoder case, the queries are the decoder features and the keys are the encoder features.\n    \n    i = tf.range(nd)[:, tf.newaxis]  # repeat along dim 1\n    j = tf.range(ns) # repeat along dim 0 \n    m = i >= (j - ns + nd) + tf.cast(only_before, dtype=tf.int32)\n    \n    return tf.cast(m, dtype)\n\ndef get_attention_mask_from_timestamp_batch(timestamp_tensors, dtype, only_before):\n    \"\"\"\n    Args:\n        timestamp_tensors: 2-D tf.int32 tensor, representing a batch of sequences of non-decreasing timestamps.\n    \n    Returns:\n        attention_mask: 3-D tf.int32 tensor of shape = [batch_size, query_len, key_len], consisting of 0 and 1.\n            Here `query_len` and `key_len` are actually `seq_len`. It should be reshpaed, when used to calculate \n            attention scores, to [batch_size, nb_attn_head, query_len, key_len].\n    \"\"\"\n    \n    t = timestamp_tensors\n    \n    batch_size = tf.math.reduce_sum(tf.ones_like(t[:, :1], dtype=tf.int32))\n    seq_len = tf.math.reduce_sum(tf.ones_like(t[:1, :], dtype=tf.int32))\n    \n    x = tf.broadcast_to(t[:, :, tf.newaxis], shape=[batch_size, seq_len, seq_len]) # repeat along dim 2\n    y = tf.broadcast_to(t[:, tf.newaxis, :], shape=[batch_size, seq_len, seq_len]) + tf.cast(only_before, dtype=tf.int64) # repeat along dim 1\n    \n    m =  x >= y\n    \n    return tf.cast(m, dtype)\n```\n\n\n",
    "1590856": "Hi, I have few questions\n1. Can you guide me how to run the code you provided? and how to get results from it.\n2. Did you published any research paper on this implementation. I want to understand your implementation in detail.\n3. At kaggle, I have seen your 3 submissions related to this competition. \n   i. https://www.kaggle.com/mamasinkgs/2nd-place-solution-for-hosts\n  ii. https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution\n  iii. https://www.kaggle.com/mamasinkgs/public-private-2nd-place-solution-single-fold\nI want to know which one is the actual and correct submission?\n",
    "1158024": "@mamasinkgs , Very Nice Approach of Using Transformer . Few Questions , I wanted to ask , if you don't mind sharing\n1) How did you estimate New Users in Private Test set by probing ?\n2) Is memory masking neccessary while training Model , what happens if We train Model without masking ?",
    "1152889": "Very nice approach, I was looking forward to this. A side question though, in this [comment](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195527#1071365) it was inferred that you were not using a transformer. When did you shift to this? or was it a transformer all along?",
    "1150057": "a dump question from a newbie. Which software do you use to draw that model structure graph? thank you !",
    "1148892": "> I uploaded the names of 90 features as attachments.\n\nWhere did you upload the attachments?",
    "1148231": "Congratulations and thank for detailed explanation!\nMay I ask a question about how to deal with the tags features?\nThank you.",
    "1146837": "Congrats and thank you for your detailed explanation.\nI have two questions about your solution.\n- What is your cross-validation and model evaluating strategy?\n- Did your task_container-aware masking improved CV score (and how much)?\n\nI think if we did not carefully split train/valid datasets and evaluate the model, CV score would be decreased when using task_container-aware masking, as the model cannot utilize the intra-container leakage you mentioned. \n\n[One of the major CV strategies in this competition](https://www.kaggle.com/its7171/cv-strategy) by @its7171 did not care task_container_id, though it split train/valid datasets by timestamp + random int. ~~This CV strategy might split some task_containers into train/valid datasets, and might cause some intra-container leakage. (I should note that I used these datasets with good CV-LB relationship)~~ [update] I confirmed that there are no such data at all.",
    "1146747": "thanks for sharing! a great job!  this is a big project!",
    "1146571": "@mamasisking , congratulations! Can you please also share info about the resources you put into the competition? hardware, amount of effort, time etc. ",
    "1146503": "@mamasinkgs , sorry to bother, but could you share your insight about this statement\n\n```\nIn my implementation, SAINT didn't work well. In my opinion, it's logically very hard to implement SAINT without data leakage in training due to task_container_id.\n```\n\nWhy you think `task_container_id` cause data leakage during training, and what kind of leakage?"
  }
}