{
  "id": 195632,
  "title": "SAINT benchmark",
  "url": "/competitions/riiid-test-answer-prediction/discussion/195632",
  "author_name": "MPWARE",
  "post_date": "2020-11-06T14:30:23.029000",
  "votes": 107,
  "comment_count": 381,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>Is someone able to reach <a href=\"https://arxiv.org/abs/2002.07033\" target=\"_blank\">SAINT</a> benchmark on RIIID dataset?<br>\nThey claim AUC around 0.78-0.79</p>\n<p>They use Exercice ID (content_id), Exercice category (could be question part), elapsed time, lag time  + (had explanation flag optionally) and answers.</p>\n<p><strong>Current status</strong> (sum up of all comments below) on <strong>2021/01/03</strong>:</p>\n<ul>\n<li>Pytorch <code>nn.Transformer</code> based implementation (not published yet)</li>\n<li>Sequences from around 334k users in training, 14k in validation</li>\n<li>My best local validation score: Fold1= 0.7903, LB=0.795</li>\n<li>Other best score from <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> LB=0.794</li>\n<li></li>\n<li>Inference time (private notebook) = Between 2h15min and 3h</li>\n</ul>\n<p>2 public notebooks implemented:</p>\n<ul>\n<li>Tensorflow: <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public</a></li>\n<li>Pytorch: <a href=\"https://www.kaggle.com/adityaecdrid/pytorch-demystifying-transformers\" target=\"_blank\">https://www.kaggle.com/adityaecdrid/pytorch-demystifying-transformers</a></li>\n</ul>\n<p>Other public notebooks related:</p>\n<ul>\n<li>SAKT (CV 0.745/LB 0.752): <a href=\"https://www.kaggle.com/mpware/sakt-fork\" target=\"_blank\">https://www.kaggle.com/mpware/sakt-fork</a></li>\n<li>SAKT (LB 0.751): <a href=\"https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing\" target=\"_blank\">https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing</a></li>\n<li>SAKT (LB 0.541) <a href=\"https://www.kaggle.com/leadbest/sakt-self-attentive-knowledge-tracing-submitter\" target=\"_blank\">https://www.kaggle.com/leadbest/sakt-self-attentive-knowledge-tracing-submitter</a></li>\n</ul>",
  "messages": [
    {
      "id": 1071126,
      "postDate": "2020-11-06T14:30:23.030Z",
      "content": "<p>Hi all,</p>\n<p>Is someone able to reach <a href=\"https://arxiv.org/abs/2002.07033\" target=\"_blank\">SAINT</a> benchmark on RIIID dataset?<br>\nThey claim AUC around 0.78-0.79</p>\n<p>They use Exercice ID (content_id), Exercice category (could be question part), elapsed time, lag time  + (had explanation flag optionally) and answers.</p>\n<p><strong>Current status</strong> (sum up of all comments below) on <strong>2021/01/03</strong>:</p>\n<ul>\n<li>Pytorch <code>nn.Transformer</code> based implementation (not published yet)</li>\n<li>Sequences from around 334k users in training, 14k in validation</li>\n<li>My best local validation score: Fold1= 0.7903, LB=0.795</li>\n<li>Other best score from <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> LB=0.794</li>\n<li></li>\n<li>Inference time (private notebook) = Between 2h15min and 3h</li>\n</ul>\n<p>2 public notebooks implemented:</p>\n<ul>\n<li>Tensorflow: <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public</a></li>\n<li>Pytorch: <a href=\"https://www.kaggle.com/adityaecdrid/pytorch-demystifying-transformers\" target=\"_blank\">https://www.kaggle.com/adityaecdrid/pytorch-demystifying-transformers</a></li>\n</ul>\n<p>Other public notebooks related:</p>\n<ul>\n<li>SAKT (CV 0.745/LB 0.752): <a href=\"https://www.kaggle.com/mpware/sakt-fork\" target=\"_blank\">https://www.kaggle.com/mpware/sakt-fork</a></li>\n<li>SAKT (LB 0.751): <a href=\"https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing\" target=\"_blank\">https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing</a></li>\n<li>SAKT (LB 0.541) <a href=\"https://www.kaggle.com/leadbest/sakt-self-attentive-knowledge-tracing-submitter\" target=\"_blank\">https://www.kaggle.com/leadbest/sakt-self-attentive-knowledge-tracing-submitter</a></li>\n</ul>",
      "rawMarkdown": "Hi all,\n\nIs someone able to reach [SAINT](https://arxiv.org/abs/2002.07033) benchmark on RIIID dataset?\nThey claim AUC around 0.78-0.79\n\nThey use Exercice ID (content_id), Exercice category (could be question part), elapsed time, lag time  + (had explanation flag optionally) and answers.\n\n**Current status** (sum up of all comments below) on **2021/01/03**:\n* Pytorch `nn.Transformer` based implementation (not published yet)\n* Sequences from around 334k users in training, 14k in validation\n* My best local validation score: Fold1= 0.7903, LB=0.795\n* Other best score from @claverru LB=0.794\n* ~~Padding mask issue for users with total interaction lower than sequence length~~\n* Inference time (private notebook) = Between 2h15min and 3h\n\n\n2 public notebooks implemented:\n* Tensorflow: https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\n* Pytorch: https://www.kaggle.com/adityaecdrid/pytorch-demystifying-transformers\n\nOther public notebooks related:\n* SAKT (CV 0.745/LB 0.752): https://www.kaggle.com/mpware/sakt-fork\n* SAKT (LB 0.751): https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing\n* SAKT (LB 0.541) https://www.kaggle.com/leadbest/sakt-self-attentive-knowledge-tracing-submitter\n \n",
      "votes": 106
    },
    {
      "id": 1107283,
      "postDate": "2020-12-09T15:03:30.890Z",
      "content": "<p>Hey guys, I finally got it. Still not trained with all the data, I will when I have more RAM (it is suposed to be coming). Single model very close to SAINT+:</p>\n<ul>\n<li>[Epochs: 24] loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768</li>\n<li>LB: 0.781</li>\n</ul>\n<p>Ask me anything about it. Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.</p>",
      "rawMarkdown": "Hey guys, I finally got it. Still not trained with all the data, I will when I have more RAM (it is suposed to be coming). Single model very close to SAINT+:\n- [Epochs: 24] loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768\n- LB: 0.781\n\nAsk me anything about it. Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.",
      "votes": 16,
      "replies": [
        {
          "id": 1107290,
          "postDate": "2020-12-09T15:10:02.763Z",
          "content": "<p>Wow! This is amazing score <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> ❤️🎉🎉! Your loss and auc gap is also really nice!</p>\n<blockquote>\n  <p>Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.</p>\n</blockquote>\n<p>Maybe, that's the crus, but can you give pointers for the same? Ty a lot! It's fine if we discuss after the comp as well :)</p>",
          "rawMarkdown": "Wow! This is amazing score @claverru ❤️🎉🎉! Your loss and auc gap is also really nice!\n\n>Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.\n\nMaybe, that's the crus, but can you give pointers for the same? Ty a lot! It's fine if we discuss after the comp as well :)",
          "votes": 1
        },
        {
          "id": 1107295,
          "postDate": "2020-12-09T15:12:48.367Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Hey good news for you! I'm picking user's sequence randomly on each epoch such as:<br>\nUser1: [Qn, Qn+100]<br>\nUser2: [Qm, Qm+100]<br>\n…<br>\nWhat about you?</p>",
          "rawMarkdown": "@claverru Hey good news for you! I'm picking user's sequence randomly on each epoch such as:\nUser1: [Qn, Qn+100]\nUser2: [Qm, Qm+100]\n...\nWhat about you?",
          "votes": 2
        },
        {
          "id": 1107297,
          "postDate": "2020-12-09T15:13:42.657Z",
          "content": "<blockquote>\n  <p>Hey guys, I finally got it. Still not trained with all the data, I will when I have more RAM (it is suposed to be coming). Single model very close to SAINT+:</p>\n  <ul>\n  <li>[Epochs: 24] loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768</li>\n  <li>LB: 0.781</li>\n  </ul>\n  <p>Ask me anything about it. Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.</p>\n</blockquote>\n<p>Great to see you up to LB, good job <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> !<br>\nI have similar results - however, in order to get 0.78, I need to </p>\n<ol>\n<li>reduce lr to smaller and train longer (lr = 1e-4 - 5e-5)</li>\n<li>use tag information</li>\n</ol>\n<p>Do you also need these 2 steps to get 0.78? And what kind of sampling strategy boost your LB score?</p>",
          "rawMarkdown": "> Hey guys, I finally got it. Still not trained with all the data, I will when I have more RAM (it is suposed to be coming). Single model very close to SAINT+:\n> - [Epochs: 24] loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768\n> - LB: 0.781\n> \n> Ask me anything about it. Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.\n\nGreat to see you up to LB, good job @claverru !\nI have similar results - however, in order to get 0.78, I need to \n   1. reduce lr to smaller and train longer (lr = 1e-4 - 5e-5)\n   2. use tag information\n\nDo you also need these 2 steps to get 0.78? And what kind of sampling strategy boost your LB score?",
          "votes": 2,
          "replies": [
            {
              "id": 1108935,
              "postDate": "2020-12-11T06:55:13.037Z",
              "content": "<blockquote>\n  <blockquote>\n    <p>Hey guys, I finally got it. Still not trained with all the data, I will when I have more RAM (it is suposed to be coming). Single model very close to SAINT+:</p>\n    <ul>\n    <li>[Epochs: 24] loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768</li>\n    <li>LB: 0.781</li>\n    </ul>\n    <p>Ask me anything about it. Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.</p>\n  </blockquote>\n  <p>Great to see you up to LB, good job <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> !<br>\n  I have similar results - however, in order to get 0.78, I need to </p>\n  <ol>\n  <li>reduce lr to smaller and train longer (lr = 1e-4 - 5e-5)</li>\n  <li>use tag information</li>\n  </ol>\n  <p>Do you also need these 2 steps to get 0.78? And what kind of sampling strategy boost your LB score?</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> <br>\nare you using plain bce loss or with some modification as i get very less loss for val in range 0.3sss for Auc you see.  may be mask bce ?<br>\nWat is custom Auc here</p>",
              "rawMarkdown": "> > Hey guys, I finally got it. Still not trained with all the data, I will when I have more RAM (it is suposed to be coming). Single model very close to SAINT+:\n> > - [Epochs: 24] loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768\n> > - LB: 0.781\n> > \n> > Ask me anything about it. Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.\n> \n> Great to see you up to LB, good job @claverru !\n> I have similar results - however, in order to get 0.78, I need to \n>    1. reduce lr to smaller and train longer (lr = 1e-4 - 5e-5)\n>    2. use tag information\n> \n> Do you also need these 2 steps to get 0.78? And what kind of sampling strategy boost your LB score?\n\n@yihdarshieh \nare you using plain bce loss or with some modification as i get very less loss for val in range 0.3sss for Auc you see.  may be mask bce ?\nWat is custom Auc here"
            }
          ]
        },
        {
          "id": 1107299,
          "postDate": "2020-12-09T15:18:22.860Z",
          "content": "<p>Nice! 👍</p>\n<p>Any key points other than a good sampling strategy to keep in mind? Plus what is it? I have seen that taking random sequences of (window length) from users with longer interactions helps.</p>\n<p>Does the size of d_model and num_layers impact the model significantly?</p>\n<p>I got 0.775 with SAINT+ in it's exact form with 2 num_layers and 128 d_model size.</p>",
          "rawMarkdown": "Nice! 👍\n\nAny key points other than a good sampling strategy to keep in mind? Plus what is it? I have seen that taking random sequences of (window length) from users with longer interactions helps.\n\nDoes the size of d_model and num_layers impact the model significantly?\n\nI got 0.775 with SAINT+ in it's exact form with 2 num_layers and 128 d_model size.",
          "votes": 7
        },
        {
          "id": 1107317,
          "postDate": "2020-12-09T15:35:32.383Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> I've seen that my models usually converge around 20-30 epochs on average. How much an improvement is seen on lowering lr and training for longer.</p>\n<p>Only using Kaggle GPU has really constrained testing for different hyperparameters, my current focus is on getting a good model and then optimizing in the last days of the competition.</p>",
          "rawMarkdown": "@yihdarshieh I've seen that my models usually converge around 20-30 epochs on average. How much an improvement is seen on lowering lr and training for longer.\n\nOnly using Kaggle GPU has really constrained testing for different hyperparameters, my current focus is on getting a good model and then optimizing in the last days of the competition.",
          "votes": 3
        },
        {
          "id": 1107319,
          "postDate": "2020-12-09T15:37:28.367Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> <a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a></p>\n<p>Sampling strategy:</p>\n<ul>\n<li>Take all user ids in training data.</li>\n<li>Compute their lengths and cap the result to 500.</li>\n<li>Scale the prior to make them sum 1 (probability distribution).</li>\n<li>Sample N ids with replacement with previous computed probabilities. N in my case is the same as different user ids you have in the training data. The result will have repeated ids.</li>\n<li>Take a random sequence for every id. It may lead to repeated sequences with low probability. For example, if you have id 115 repeated 7 times, you will take 7 random sequences for the user 115.</li>\n<li>Repeat each epoch.</li>\n</ul>\n<p>Why? </p>\n<ul>\n<li>If you take a random sequence for every user every epoch, you will overfit towards starter sequences, given that the majority of the users use the application only a few times. </li>\n<li>Cap to 500 before making the prob distribution denies super prolific users (outliers) from having a too big presence in the sampling.   </li>\n</ul>\n<p>Let me know if you need more details.</p>",
          "rawMarkdown": "@adityaecdrid @mpware @yihdarshieh @abdurrafae\n\nSampling strategy:\n- Take all user ids in training data.\n- Compute their lengths and cap the result to 500.\n- Scale the prior to make them sum 1 (probability distribution).\n- Sample N ids with replacement with previous computed probabilities. N in my case is the same as different user ids you have in the training data. The result will have repeated ids.\n- Take a random sequence for every id. It may lead to repeated sequences with low probability. For example, if you have id 115 repeated 7 times, you will take 7 random sequences for the user 115.\n- Repeat each epoch.\n\nWhy? \n\n- If you take a random sequence for every user every epoch, you will overfit towards starter sequences, given that the majority of the users use the application only a few times. \n- Cap to 500 before making the prob distribution denies super prolific users (outliers) from having a too big presence in the sampling.   \n\nLet me know if you need more details.",
          "votes": 19
        },
        {
          "id": 1107324,
          "postDate": "2020-12-09T15:40:32.897Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a></p>\n<p>I use every feature they use in SAINT+ paper with slightly modifications in them but not adding new ones. Learning rate schedule is exactly as described in SAINT+ (originally from Attention is All You Need).</p>",
          "rawMarkdown": "@yihdarshieh\n\nI use every feature they use in SAINT+ paper with slightly modifications in them but not adding new ones. Learning rate schedule is exactly as described in SAINT+ (originally from Attention is All You Need).",
          "votes": 3
        },
        {
          "id": 1107331,
          "postDate": "2020-12-09T15:43:34.243Z",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a></p>\n<p>I haven't tried to grow my model up. 2 encoder layers, 2 decoder layers and model size of 128.</p>\n<pre><code>train_ratio = 0.96\nwindows_size = 96\nepochs = 100\npatience = 3\nd_model = 128\nnum_heads = 4\nn_encoder_layers = 2\nn_decoder_layers = 2\nbatch_size = 256\n</code></pre>",
          "rawMarkdown": "@abdurrafae\n\nI haven't tried to grow my model up. 2 encoder layers, 2 decoder layers and model size of 128.\n\n```python\ntrain_ratio = 0.96\nwindows_size = 96\nepochs = 100\npatience = 3\nd_model = 128\nnum_heads = 4\nn_encoder_layers = 2\nn_decoder_layers = 2\nbatch_size = 256\n```",
          "votes": 4
        },
        {
          "id": 1107332,
          "postDate": "2020-12-09T15:44:59.897Z",
          "content": "<p>For all of you, if you get to implement the sampling strategy, let me know if it works for you. I have worked a lot (a lot) on it.</p>",
          "rawMarkdown": "For all of you, if you get to implement the sampling strategy, let me know if it works for you. I have worked a lot (a lot) on it.",
          "votes": 2
        },
        {
          "id": 1107338,
          "postDate": "2020-12-09T15:52:47.393Z",
          "content": "<p>Would you share the way you compute the lag time? It is still not clear to me how to compute it. Do you just use the current timestamp - the previous timestamp?</p>\n<p>If so:<br>\n    Q1: you will get 0 for question in a bundle (except for the 1st question in that bundle)??<br>\n    Q2: If a question follows a lecture, I think the above computation is not the lag time described in the paper. Or maybe you ignore all the lectures, so don't have care about lecture during computing lag time?</p>",
          "rawMarkdown": "Would you share the way you compute the lag time? It is still not clear to me how to compute it. Do you just use the current timestamp - the previous timestamp?\n\nIf so:\n    Q1: you will get 0 for question in a bundle (except for the 1st question in that bundle)??\n    Q2: If a question follows a lecture, I think the above computation is not the lag time described in the paper. Or maybe you ignore all the lectures, so don't have care about lecture during computing lag time?",
          "votes": 2
        },
        {
          "id": 1107391,
          "postDate": "2020-12-09T16:56:09.737Z",
          "content": "<p>That's one of the differences with SAINT+ I have. I don't compute lag as they describe it. I mean, I just compute the time between the start of two consecutive interactions, capping it to one day and scaling it [0, 1].  If you are feeding the model also with the other time feature, the model will agg that information as he pleases.</p>\n<p>I simply do:</p>\n<pre><code>user['timestamp'] = (user['timestamp'].diff().fillna(0)/8.64e+7).clip(upper=1)\n</code></pre>",
          "rawMarkdown": "That's one of the differences with SAINT+ I have. I don't compute lag as they describe it. I mean, I just compute the time between the start of two consecutive interactions, capping it to one day and scaling it [0, 1].  If you are feeding the model also with the other time feature, the model will agg that information as he pleases.\n\nI simply do:\n\n```python\nuser['timestamp'] = (user['timestamp'].diff().fillna(0)/8.64e+7).clip(upper=1)\n```",
          "votes": 6
        },
        {
          "id": 1107400,
          "postDate": "2020-12-09T17:05:12.500Z",
          "content": "<p>Thank you. Using smaller lr gives me from LB 0.778 to 0.781 (best CV epoch : 6 --&gt; 30). Not that much though.</p>",
          "rawMarkdown": "Thank you. Using smaller lr gives me from LB 0.778 to 0.781 (best CV epoch : 6 --> 30). Not that much though.",
          "votes": 2
        },
        {
          "id": 1107430,
          "postDate": "2020-12-09T17:25:37.503Z",
          "content": "<p>Hey mate what you said is true. I didn't even think about it. I'm probably getting zeroes for non-first questions in a bundle, though in inference time having the good diff.</p>\n<p>Anyway, there is a similar problem around which I find even worse. Imagine I have a sequence [A, B, C1, C2, C3], being C three questions in the same bundle. In training time I just let them as they are, but in inference time I do [A, B, C1], [A, B, C2], [A, B, C3]. Solving this I think could boost the performance by a lot, though I feel to lazy to set it up in training time.</p>",
          "rawMarkdown": "Hey mate what you said is true. I didn't even think about it. I'm probably getting zeroes for non-first questions in a bundle, though in inference time having the good diff.\n\nAnyway, there is a similar problem around which I find even worse. Imagine I have a sequence [A, B, C1, C2, C3], being C three questions in the same bundle. In training time I just let them as they are, but in inference time I do [A, B, C1], [A, B, C2], [A, B, C3]. Solving this I think could boost the performance by a lot, though I feel to lazy to set it up in training time.",
          "votes": 3
        },
        {
          "id": 1107459,
          "postDate": "2020-12-09T17:47:39.613Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> -- Yes, that is a major issue (in another way) I spent time to fix.</p>\n<p>Just like you, in training time, for me,  everything is in the normal/usual setting. I think it is too difficult to change the training time setting as in the inference time - unless you want to train with only the last few questions in the last bundle in each sequence. The model needs to see about 100 more times of the number of sequences to converge, which will be too time consuming.</p>\n<p>For inference, I do have something like <code>[A, B, C1], [A, B, C2], [A, B, C3].</code>, i.e. we don't use auto-regressive generation. Due to this, we need to keep the 2nd - the last question in the same bundle as the 1st question in the bundle. However, I keep the same questions (for prediction) for each user in a single sequence, but using special mask -  the 2nd question can't attend to the 1st question in the bundle during the inference time, etc. The same applies for the positional information. For example, I need to change [0, 1, 2, 3, 4] to [0, 1, 2, 2, 2].</p>",
          "rawMarkdown": "@claverru -- Yes, that is a major issue (in another way) I spent time to fix.\n\nJust like you, in training time, for me,  everything is in the normal/usual setting. I think it is too difficult to change the training time setting as in the inference time - unless you want to train with only the last few questions in the last bundle in each sequence. The model needs to see about 100 more times of the number of sequences to converge, which will be too time consuming.\n\nFor inference, I do have something like `[A, B, C1], [A, B, C2], [A, B, C3].`, i.e. we don't use auto-regressive generation. Due to this, we need to keep the 2nd - the last question in the same bundle as the 1st question in the bundle. However, I keep the same questions (for prediction) for each user in a single sequence, but using special mask -  the 2nd question can't attend to the 1st question in the bundle during the inference time, etc. The same applies for the positional information. For example, I need to change [0, 1, 2, 3, 4] to [0, 1, 2, 2, 2].",
          "votes": 2
        },
        {
          "id": 1107478,
          "postDate": "2020-12-09T18:13:41.480Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>, your sampling strategy seems very reasonable to me. However, it is somehow difficult for me to change the sampling - all the examples are store in tfrecord dataset, and using tf.data.Dataset as the input pipeline. I can do random subsequence selection, but not the user selection, seems the tf.data.Dataset will just yield the sequences from all users.</p>\n<p>The solution I have in mind is to calculate the probability (the one you calculated) for different users. And use it as a loss weight. For example, for a user with much shorter sequence - the probably you calculated is smaller. And it will be selected more often (if the proability is not applied to select sequences from different users) - but if I multiply the probably in loss calculation, it should have the same effect as <code>select few examples for that user using the smaller probability</code></p>",
          "rawMarkdown": "@claverru, your sampling strategy seems very reasonable to me. However, it is somehow difficult for me to change the sampling - all the examples are store in tfrecord dataset, and using tf.data.Dataset as the input pipeline. I can do random subsequence selection, but not the user selection, seems the tf.data.Dataset will just yield the sequences from all users.\n\nThe solution I have in mind is to calculate the probability (the one you calculated) for different users. And use it as a loss weight. For example, for a user with much shorter sequence - the probably you calculated is smaller. And it will be selected more often (if the proability is not applied to select sequences from different users) - but if I multiply the probably in loss calculation, it should have the same effect as `select few examples for that user using the smaller probability`\n\n",
          "votes": 2
        },
        {
          "id": 1107492,
          "postDate": "2020-12-09T18:24:24.917Z",
          "content": "<p>I have a couple of questions, Hope you guys won't mind answering them!</p>\n<p>First is regarding the sampling strategy shared by <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>. Is he attaching a probability for each user so that they get selected at random? Really clever idea, but not sure how to calculate that? Just by passing something to a distribution for a starter?</p>\n<p>Secondly, this</p>\n<blockquote>\n  <p>Imagine I have a sequence [A, B, C1, C2, C3], being C three questions in the same bundle. In training time I just let them as they are, but in inference time I do [A, B, C1], [A, B, C2], [A, B, C3]</p>\n</blockquote>\n<p>I am not sure i understood it. So, basically you guys are splitting the sequence of <code>[A, B, C1, C2, C3]</code> to <code>[A, B, C1], [A, B, C2], [A, B, C3]</code>. That's again very smart to do but I am not sure how can I append this useful insight to my inference without making preds on rows grouped by user_id let's say in inference?</p>\n<p>And then the below, need to think about this;</p>\n<blockquote>\n  <p>For inference, I do have something like [A, B, C1], [A, B, C2], [A, B, C3]., i.e. we don't use auto-regressive generation. Due to this, we need to keep the 2nd - the last question in the same bundle as the 1st question in the bundle. However, I keep the same questions (for prediction) for each user in a single sequence, but using special mask - the 2nd question can't attend to the 1st question in the bundle during the inference time, etc.</p>\n</blockquote>\n<p>This, i need to think over as of now as i guess, i am losing the context of this being discussed in past as well, in this or some other discussion as well.</p>",
          "rawMarkdown": "I have a couple of questions, Hope you guys won't mind answering them!\n\nFirst is regarding the sampling strategy shared by @claverru. Is he attaching a probability for each user so that they get selected at random? Really clever idea, but not sure how to calculate that? Just by passing something to a distribution for a starter?\n\nSecondly, this\n\n> Imagine I have a sequence [A, B, C1, C2, C3], being C three questions in the same bundle. In training time I just let them as they are, but in inference time I do [A, B, C1], [A, B, C2], [A, B, C3]\n\nI am not sure i understood it. So, basically you guys are splitting the sequence of `[A, B, C1, C2, C3]` to `[A, B, C1], [A, B, C2], [A, B, C3]`. That's again very smart to do but I am not sure how can I append this useful insight to my inference without making preds on rows grouped by user_id let's say in inference?\n\nAnd then the below, need to think about this;\n>For inference, I do have something like [A, B, C1], [A, B, C2], [A, B, C3]., i.e. we don't use auto-regressive generation. Due to this, we need to keep the 2nd - the last question in the same bundle as the 1st question in the bundle. However, I keep the same questions (for prediction) for each user in a single sequence, but using special mask - the 2nd question can't attend to the 1st question in the bundle during the inference time, etc.\n\n\nThis, i need to think over as of now as i guess, i am losing the context of this being discussed in past as well, in this or some other discussion as well."
        },
        {
          "id": 1107493,
          "postDate": "2020-12-09T18:24:48.780Z",
          "content": "<p>I think this competition requires equal amount of effort on architecture design and data management. It's all fun and enjoyable till to get into a situation where debugging for a week leads you no where and you ultimately revert the change you initially made. :/</p>",
          "rawMarkdown": "I think this competition requires equal amount of effort on architecture design and data management. It's all fun and enjoyable till to get into a situation where debugging for a week leads you no where and you ultimately revert the change you initially made. :/",
          "votes": 4
        },
        {
          "id": 1107508,
          "postDate": "2020-12-09T18:37:59.950Z",
          "content": "<p>Off topic but I would like to share this with you, avoid Kaggle Docker version 90 if you face time out, I've lost 3 days on it: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1107497\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1107497</a></p>",
          "rawMarkdown": "Off topic but I would like to share this with you, avoid Kaggle Docker version 90 if you face time out, I've lost 3 days on it: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1107497",
          "votes": 1
        },
        {
          "id": 1107520,
          "postDate": "2020-12-09T18:46:15.863Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> We were talking there about the differences between training time and inference time regarding to questions in the same bundle ID AKA same exam. You don't know their results until they are all finished, so we shouldn't be training as we do.</p>",
          "rawMarkdown": "@adityaecdrid We were talking there about the differences between training time and inference time regarding to questions in the same bundle ID AKA same exam. You don't know their results until they are all finished, so we shouldn't be training as we do.",
          "votes": 1
        },
        {
          "id": 1107524,
          "postDate": "2020-12-09T18:48:27.193Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> I'm just using my own Dockerfile trying to keep the same major versions as in the Kaggle image. I thank you anyways, I'll try to skip it when training here in Kaggle notebooks (this could be in fact be happening to me in Casava classification competition in which I'm training in Kaggle notebooks).</p>",
          "rawMarkdown": "@mpware I'm just using my own Dockerfile trying to keep the same major versions as in the Kaggle image. I thank you anyways, I'll try to skip it when training here in Kaggle notebooks (this could be in fact be happening to me in Casava classification competition in which I'm training in Kaggle notebooks)."
        },
        {
          "id": 1107531,
          "postDate": "2020-12-09T18:51:31.333Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> nvm you were talking about inference. I will have a look.</p>",
          "rawMarkdown": "@mpware nvm you were talking about inference. I will have a look."
        },
        {
          "id": 1107533,
          "postDate": "2020-12-09T18:56:52.707Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Sorry for the spam, but you were right, seems like I'm using v90. How can we rollback here? I noticed in fact that my running time increased with no apparently reason. </p>",
          "rawMarkdown": "@mpware Sorry for the spam, but you were right, seems like I'm using v90. How can we rollback here? I noticed in fact that my running time increased with no apparently reason. "
        },
        {
          "id": 1107536,
          "postDate": "2020-12-09T18:58:44.877Z",
          "content": "<p>Check the Env section in settings in interactive kernel mode, by default, i always keep it to the same version to which i made a sub with. (i.e. no upstream updates to the docker <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Fc95de2b124b677f4da391a2b77a9fa7d%2FScreenshot%202020-12-10%20at%2012.29.58%20AM.png?generation=1607540428537812&amp;alt=media\" alt=\"\">file)</p>",
          "rawMarkdown": "Check the Env section in settings in interactive kernel mode, by default, i always keep it to the same version to which i made a sub with. (i.e. no upstream updates to the docker ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Fc95de2b124b677f4da391a2b77a9fa7d%2FScreenshot%202020-12-10%20at%2012.29.58%20AM.png?generation=1607540428537812&alt=media)file)\n",
          "votes": 1
        },
        {
          "id": 1107540,
          "postDate": "2020-12-09T19:01:58.723Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> - I am not sure how other people doing it. The objective is: for each question in a bundle that we want to predict, we only used the history before the current question bundle. The reason: say for the 2nd question in the current bunlde to predict, some information for the 1st question in the same bundle is not available (for example, answer correctness). We use this position while prediction for the 2nd question.</p>\n<p>Ideally, we would like to use auto-regressive generation, i.e. we predict the results for the 1st question, and use it for prediction 2nd question etc. However it is time consuming and not so easy to implement in a efficient way (we have time limit for this competition). That's why we don't use this approach (at least, no notebook published for it).</p>\n<p>However, split to [A, B, C1], [A, B, C2], [A, B, C3], I don't like the idea - because it increase the number of sequences to compute the predictions.</p>\n<p>That's why I still keep [A, B, C1, C2, C3] , but modify some input information like positional info and make sure the attention mask also makes sense due to this constraint.</p>\n<p>If you don't do any of the above 2 approach, how do you perform inference? Unless you use encoder-only model. <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> and me use encoder-decoder architecture.</p>",
          "rawMarkdown": "@adityaecdrid - I am not sure how other people doing it. The objective is: for each question in a bundle that we want to predict, we only used the history before the current question bundle. The reason: say for the 2nd question in the current bunlde to predict, some information for the 1st question in the same bundle is not available (for example, answer correctness). We use this position while prediction for the 2nd question.\n\nIdeally, we would like to use auto-regressive generation, i.e. we predict the results for the 1st question, and use it for prediction 2nd question etc. However it is time consuming and not so easy to implement in a efficient way (we have time limit for this competition). That's why we don't use this approach (at least, no notebook published for it).\n\nHowever, split to [A, B, C1], [A, B, C2], [A, B, C3], I don't like the idea - because it increase the number of sequences to compute the predictions.\n\nThat's why I still keep [A, B, C1, C2, C3] , but modify some input information like positional info and make sure the attention mask also makes sense due to this constraint.\n\nIf you don't do any of the above 2 approach, how do you perform inference? Unless you use encoder-only model. @claverru and me use encoder-decoder architecture.",
          "votes": 4
        },
        {
          "id": 1107551,
          "postDate": "2020-12-09T19:09:08.083Z",
          "content": "<p>Thanks a lot for the clarifications! Now I can see where you guys are heading! Cool ideas :)</p>",
          "rawMarkdown": "Thanks a lot for the clarifications! Now I can see where you guys are heading! Cool ideas :)",
          "votes": 1
        },
        {
          "id": 1107555,
          "postDate": "2020-12-09T19:14:48.343Z",
          "content": "<p>Would you mind to share what is your current approach? BTW, great to see you up on LB also!</p>",
          "rawMarkdown": "Would you mind to share what is your current approach? BTW, great to see you up on LB also!",
          "votes": 1
        },
        {
          "id": 1107567,
          "postDate": "2020-12-09T19:25:25.687Z",
          "content": "<p>Well, I am way behind when it comes to SAINT, gave up on it for a week now; Switched to gbms with hand-crafted features and experiments on them as of now. (Current LB score is with a gbm, SAINT is at ~.74). But after seeing today's spike in discussions wrt SAINT, got some new inspirations, so I am going to try it a couple of more times!</p>\n<p>In my understanding, I am exactly doing this one for simplicity,</p>\n<blockquote>\n  <p>That's why I still keep [A, B, C1, C2, C3] , but modify some input information like positional info and make sure the attention mask also makes sense due to this constraint.</p>\n</blockquote>\n<p>For the above thing actually, I am <strong>not</strong> even changing my positional info (thanks for the tip!) and regarding the attention mask, I just (an hour back) found a bug while making inference, so I will get back hopefully with a better score than .74 in 1-2 days; 😊</p>",
          "rawMarkdown": "Well, I am way behind when it comes to SAINT, gave up on it for a week now; Switched to gbms with hand-crafted features and experiments on them as of now. (Current LB score is with a gbm, SAINT is at ~.74). But after seeing today's spike in discussions wrt SAINT, got some new inspirations, so I am going to try it a couple of more times!\n\nIn my understanding, I am exactly doing this one for simplicity,\n\n>That's why I still keep [A, B, C1, C2, C3] , but modify some input information like positional info and make sure the attention mask also makes sense due to this constraint.\n\nFor the above thing actually, I am **not** even changing my positional info (thanks for the tip!) and regarding the attention mask, I just (an hour back) found a bug while making inference, so I will get back hopefully with a better score than .74 in 1-2 days; 😊"
        },
        {
          "id": 1107572,
          "postDate": "2020-12-09T19:29:19.663Z",
          "content": "<p>BTW the next thing I'm going to work in is a feature like question_already_answered. This will give information beyond the sequence. Have any of you implemented this? Have it boosted your model?</p>",
          "rawMarkdown": "BTW the next thing I'm going to work in is a feature like question_already_answered. This will give information beyond the sequence. Have any of you implemented this? Have it boosted your model?",
          "votes": 1
        },
        {
          "id": 1107574,
          "postDate": "2020-12-09T19:33:48.710Z",
          "content": "<p>Oh, yes, in fact I have a feature that gives information about bundles, the task_container_id. Forgot to mention, I'm currently using it as embeddings with input [0, 10000] and adding them to the decoder (though I maybe should add them to both encoder and decoder).</p>\n<p>Here it is the snippet, hope it is useful:</p>\n<pre><code>e = tf.keras.layers.Add()([\n    content_emb,\n    part_emb\n])\n\nd = tf.keras.layers.Add()([\n    answered_correctly_emb,\n    task_container_emb,\n    time_features\n])\n\nfor _ in range(n_encoder_layers):\n    e = EncoderLayer(\n        d_model=d_model, \n        num_heads=num_heads, \n        dff=d_model*2, \n        rate=0.01,\n        rel_pos_enc=True\n    )(e, mask=mask)\n\nfor _ in range(n_decoder_layers):\n    d, _, _ = DecoderLayer(\n        d_model=d_model, \n        num_heads=num_heads, \n        dff=d_model*2, \n        rate=0.01, \n        rel_pos_enc=True\n    )(d, e, look_ahead_mask=mask, padding_mask=mask)\n</code></pre>",
          "rawMarkdown": "Oh, yes, in fact I have a feature that gives information about bundles, the task_container_id. Forgot to mention, I'm currently using it as embeddings with input [0, 10000] and adding them to the decoder (though I maybe should add them to both encoder and decoder).\n\nHere it is the snippet, hope it is useful:\n\n```python\ne = tf.keras.layers.Add()([\n    content_emb,\n    part_emb\n])\n\nd = tf.keras.layers.Add()([\n    answered_correctly_emb,\n    task_container_emb,\n    time_features\n])\n\nfor _ in range(n_encoder_layers):\n    e = EncoderLayer(\n        d_model=d_model, \n        num_heads=num_heads, \n        dff=d_model*2, \n        rate=0.01,\n        rel_pos_enc=True\n    )(e, mask=mask)\n\nfor _ in range(n_decoder_layers):\n    d, _, _ = DecoderLayer(\n        d_model=d_model, \n        num_heads=num_heads, \n        dff=d_model*2, \n        rate=0.01, \n        rel_pos_enc=True\n    )(d, e, look_ahead_mask=mask, padding_mask=mask)\n```",
          "votes": 4
        },
        {
          "id": 1107575,
          "postDate": "2020-12-09T19:33:50.533Z",
          "content": "<p>I plan to - but I might try pretraining the encoder with MLM loss first.<br>\nI am busy on preparing a TPU training notebook and will publish it before the end of week.<br>\nCurrently, I am running a few experiment like small model (you socre 0.781 with small model scare me already). Then I will try your sampling.</p>",
          "rawMarkdown": "I plan to - but I might try pretraining the encoder with MLM loss first.\nI am busy on preparing a TPU training notebook and will publish it before the end of week.\nCurrently, I am running a few experiment like small model (you socre 0.781 with small model scare me already). Then I will try your sampling.",
          "votes": 3
        },
        {
          "id": 1107577,
          "postDate": "2020-12-09T19:35:38.020Z",
          "content": "<p>I guess It's implemented and discussed <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266#1069347\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "I guess It's implemented and discussed [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266#1069347).",
          "votes": 1
        },
        {
          "id": 1107583,
          "postDate": "2020-12-09T19:42:00.063Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> - The task_container_id plays a role similar to absolute position - but a bit different it contains bundle information (so you do have some kind of absolute pos along with your relative pos ).</p>\n<p>I also tried absolute pos - but it gets lower socre (if trained at the same epochs) Considering you train up to 25 or more epochs, maybe it is good for me to try training more epochs also :) while using abs pos.</p>",
          "rawMarkdown": "@claverru - The task_container_id plays a role similar to absolute position - but a bit different it contains bundle information (so you do have some kind of absolute pos along with your relative pos ).\n\nI also tried absolute pos - but it gets lower socre (if trained at the same epochs) Considering you train up to 25 or more epochs, maybe it is good for me to try training more epochs also :) while using abs pos.",
          "votes": 1
        },
        {
          "id": 1107584,
          "postDate": "2020-12-09T19:42:07.483Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> Thanks, I will take a look.<br>\n<a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> I'm training at the moment a bigger model (3 enc, 3 dec, 256 d_model). I will share an update tomorrow when I submit it. </p>",
          "rawMarkdown": "@adityaecdrid Thanks, I will take a look.\n@yihdarshieh I'm training at the moment a bigger model (3 enc, 3 dec, 256 d_model). I will share an update tomorrow when I submit it. ",
          "votes": 2
        },
        {
          "id": 1107586,
          "postDate": "2020-12-09T19:43:23.753Z",
          "content": "<p>I train until EarlyStopping triggers, validating every 2 epochs, with a patience of 3 (6 epochs total).</p>",
          "rawMarkdown": "I train until EarlyStopping triggers, validating every 2 epochs, with a patience of 3 (6 epochs total)."
        },
        {
          "id": 1107590,
          "postDate": "2020-12-09T19:47:39.637Z",
          "content": "<p>Fuck <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>, have you implemented that bits thing shit? Sounds crazy.</p>",
          "rawMarkdown": "Fuck @adityaecdrid, have you implemented that bits thing shit? Sounds crazy."
        },
        {
          "id": 1107592,
          "postDate": "2020-12-09T19:50:56.313Z",
          "content": "<p>So at about which epoch you get the best CV?</p>",
          "rawMarkdown": "So at about which epoch you get the best CV?"
        },
        {
          "id": 1107595,
          "postDate": "2020-12-09T19:54:24.763Z",
          "content": "<p>The one I shared, 24th for my best model.</p>\n<p>24/100 loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768</p>",
          "rawMarkdown": "The one I shared, 24th for my best model.\n\n24/100 loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768",
          "votes": 1
        },
        {
          "id": 1107649,
          "postDate": "2020-12-09T20:47:57.073Z",
          "content": "<p>The only way to rollback is to upload your kernel code in on old kernel using the previous docker image. But you need at least one kernel with the old image.</p>",
          "rawMarkdown": "The only way to rollback is to upload your kernel code in on old kernel using the previous docker image. But you need at least one kernel with the old image.",
          "votes": 1
        },
        {
          "id": 1107733,
          "postDate": "2020-12-09T23:14:26.667Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a></p>\n<p>Before your sampling strategy, did you get the best CV in at a much earlier epoch? I tried smaller model (as yours), with the learning rate in SANIT(+) and batch 256 as yours, with absolute pos (not task_container_id, just the pos in the full history) --&gt; CV epoch 15 &gt; epoch 20 &gt; epoch 25, and all of them is worse than <code>smaller lr at epoch 10 without abs. pos</code>.</p>\n<p>I hope I can find the cause after I finished the TPU notebook - and definitely need try your sampling!</p>",
          "rawMarkdown": "@claverru\n\nBefore your sampling strategy, did you get the best CV in at a much earlier epoch? I tried smaller model (as yours), with the learning rate in SANIT(+) and batch 256 as yours, with absolute pos (not task_container_id, just the pos in the full history) --> CV epoch 15 > epoch 20 > epoch 25, and all of them is worse than `smaller lr at epoch 10 without abs. pos`.\n\nI hope I can find the cause after I finished the TPU notebook - and definitely need try your sampling!"
        },
        {
          "id": 1107886,
          "postDate": "2020-12-10T03:21:21.797Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> </p>\n<p>\"BTW the next thing I'm going to work in is a feature like question_already_answered. This will give information beyond the sequence. Have any of you implemented this? Have it boosted your model?\"</p>\n<p>I kept track of attempts for each user on each question on my tabular model and it boosted the score quite a bit. I've been curious how useful it would to a transformer. I would assume the attention likely captures that information.</p>",
          "rawMarkdown": "@claverru \n\n\"BTW the next thing I'm going to work in is a feature like question_already_answered. This will give information beyond the sequence. Have any of you implemented this? Have it boosted your model?\"\n\nI kept track of attempts for each user on each question on my tabular model and it boosted the score quite a bit. I've been curious how useful it would to a transformer. I would assume the attention likely captures that information.",
          "votes": 1
        },
        {
          "id": 1107992,
          "postDate": "2020-12-10T06:55:01.207Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a></p>\n<p>I must be too tired for the work on TPU notebook. I actually have a similar sampling as yours, as I mentioned earlier in another comment</p>\n<pre><code>I used to use random sample subsequence from users' history. But I found that potentially I overfit for users with much shorter sequences. For example, if a user has only 50 interaction. In each epoch, it will get the full history. While for a user having 1000 interaction, it only get one subsequence in a epoch.\n</code></pre>\n<p>The difference for my current sampling and yours:</p>\n<ul>\n<li><p>Instead of using probability, I use tf.data.Dataset to repeat user sequences. For example, if user X has a history of length <code>N</code>, it will be repeated for <code>N / W</code> times (<code>W</code> = window size). And I sample subsequences of length <code>W</code> from these <code>N / W</code> full sequences. For each epoch, I get about 1M sequences (while you get about 0.38M sequences). Therefore, your epoch 24 is somehow equivalent to my epoch 8. And it explains why we get the best CV at different epoch :)</p></li>\n<li><p>You have a control of <code>500</code> to avoid sampling too many times from particular users - while I don't do this. I will give it a try, since it should be quite easy to implement.</p></li>\n</ul>",
          "rawMarkdown": "@claverru\n\nI must be too tired for the work on TPU notebook. I actually have a similar sampling as yours, as I mentioned earlier in another comment\n\n```\nI used to use random sample subsequence from users' history. But I found that potentially I overfit for users with much shorter sequences. For example, if a user has only 50 interaction. In each epoch, it will get the full history. While for a user having 1000 interaction, it only get one subsequence in a epoch.\n```\n\nThe difference for my current sampling and yours:\n\n- Instead of using probability, I use tf.data.Dataset to repeat user sequences. For example, if user X has a history of length `N`, it will be repeated for `N / W` times (`W` = window size). And I sample subsequences of length `W` from these `N / W` full sequences. For each epoch, I get about 1M sequences (while you get about 0.38M sequences). Therefore, your epoch 24 is somehow equivalent to my epoch 8. And it explains why we get the best CV at different epoch :)\n\n- You have a control of `500` to avoid sampling too many times from particular users - while I don't do this. I will give it a try, since it should be quite easy to implement.\n",
          "votes": 3
        },
        {
          "id": 1108007,
          "postDate": "2020-12-10T07:11:14.350Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a></p>\n<p>Considering we use similar sampling (other than the upper limit of 500 or not), we have quite different results (with you model size, I get only 0.773), it makes thinking what's wrong in my model.</p>\n<p>I saw that in an earlier comment, you mentioned</p>\n<pre><code>80 M rows minus lectures (311567 users).\n</code></pre>\n<p>Is this the case at this moment? I do keep lecture, and pay attention to them. Maybe I should try not to pay attention to it - or even remove it (but it is not easy to change  my pipeline to remove lectures …).</p>\n<p>Do you have a reason not to use lectures?</p>",
          "rawMarkdown": "@claverru\n\nConsidering we use similar sampling (other than the upper limit of 500 or not), we have quite different results (with you model size, I get only 0.773), it makes thinking what's wrong in my model.\n\nI saw that in an earlier comment, you mentioned\n\n```\n80 M rows minus lectures (311567 users).\n```\n\nIs this the case at this moment? I do keep lecture, and pay attention to them. Maybe I should try not to pay attention to it - or even remove it (but it is not easy to change  my pipeline to remove lectures ...).\n\nDo you have a reason not to use lectures?",
          "votes": 1
        },
        {
          "id": 1108043,
          "postDate": "2020-12-10T08:01:25.010Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> I just woke up. Seems like we actually have a similar sampling strategy, the difference is that I may not see some of the users. If I understood correctly, you have 3 users like:</p>\n<p>[A, B, C, D]<br>\n[A]<br>\n[A, B]</p>\n<p>and do something like</p>\n<p>[A, B] [C, D]<br>\n[A, -]<br>\n[A, B]</p>\n<p>Don't you see any problem with that? You have 3 starter sequences even if they come from different users, which I pretended to smooth with my sampling strategy.</p>\n<p>On the other hand, I'm still with 80 M rows minus lectures (311567 users). But you just gave me an idea: in theory even if you predicted something idiotic for the lecture step, it should help the model in the way that it knows that <em>the user watched a lecture</em>, which is itself an important feature. So I think you are right there and I'm wrong. I will try that in the incoming days and come with the results.</p>\n<p>Finally, I just finished to train my bigger model, in which I also modified the way I compute the timestamp from an idea taken some comments above.</p>\n<pre><code>user['timestamp'] = (\n  user['timestamp'].diff().replace(0, method='ffill').fillna(0)/8.64e+7\n).clip(upper=1)\n</code></pre>\n<p>So far so good:</p>\n<p>[Epoch 28/100] loss: 0.5305 - custom_auc: 0.7850 - val_loss: 0.5290 - val_custom_auc: 0.7801</p>\n<p>I will make the submit and come with the results.</p>",
          "rawMarkdown": "Hi @yihdarshieh I just woke up. Seems like we actually have a similar sampling strategy, the difference is that I may not see some of the users. If I understood correctly, you have 3 users like:\n\n[A, B, C, D]\n[A]\n[A, B]\n\nand do something like\n\n[A, B] [C, D]\n[A, -]\n[A, B]\n\nDon't you see any problem with that? You have 3 starter sequences even if they come from different users, which I pretended to smooth with my sampling strategy.\n\nOn the other hand, I'm still with 80 M rows minus lectures (311567 users). But you just gave me an idea: in theory even if you predicted something idiotic for the lecture step, it should help the model in the way that it knows that _the user watched a lecture_, which is itself an important feature. So I think you are right there and I'm wrong. I will try that in the incoming days and come with the results.\n\nFinally, I just finished to train my bigger model, in which I also modified the way I compute the timestamp from an idea taken some comments above.\n\n```python\nuser['timestamp'] = (\n  user['timestamp'].diff().replace(0, method='ffill').fillna(0)/8.64e+7\n).clip(upper=1)\n```\nSo far so good:\n\n[Epoch 28/100] loss: 0.5305 - custom_auc: 0.7850 - val_loss: 0.5290 - val_custom_auc: 0.7801\n\nI will make the submit and come with the results.",
          "votes": 1
        },
        {
          "id": 1108081,
          "postDate": "2020-12-10T08:50:35.210Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> - you are right. I potentially sampled more starting sequence in each epoch, even they are from different users. (Actually I don't do [A, B, C, D] -&gt; [A, B] [C, D]. With your example, it is 2 subsequences from the 1st user, which could be selected from [A, B], [B, C] or [C, D])</p>",
          "rawMarkdown": "@claverru - you are right. I potentially sampled more starting sequence in each epoch, even they are from different users. (Actually I don't do [A, B, C, D] -> [A, B] [C, D]. With your example, it is 2 subsequences from the 1st user, which could be selected from [A, B], [B, C] or [C, D])",
          "votes": 1
        },
        {
          "id": 1108141,
          "postDate": "2020-12-10T10:20:20.957Z",
          "content": "<p>Then that's exactly how I had it before (a random subsequence for a user). The problem is that you don't have many subsequences to choose when the whole sequence is shorter than your window size.</p>",
          "rawMarkdown": "Then that's exactly how I had it before (a random subsequence for a user). The problem is that you don't have many subsequences to choose when the whole sequence is shorter than your window size."
        },
        {
          "id": 1108297,
          "postDate": "2020-12-10T14:07:09.173Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>, thanks for the sampling stategy for training, it's quite clear, I will try it too. And for validation, how do you do? As it should be always the same set (no random).</p>",
          "rawMarkdown": "@claverru, thanks for the sampling stategy for training, it's quite clear, I will try it too. And for validation, how do you do? As it should be always the same set (no random).",
          "votes": 1
        },
        {
          "id": 1108468,
          "postDate": "2020-12-10T17:24:03.857Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> </p>\n<p>I use every possible sequence for validation. [A, -, -, …], [A, B, -, …], …</p>\n<p>By the way, I got a submission error for my last model after waiting for 9h, though it doesn't say timeout. That sounds like what you told before.  WTF what can I do? I don't have any prior version enviroment. I've seen there's a v91 version, but I've checked the Dockerfile commit history and doesn't seem like they have fixed anything relevant.</p>",
          "rawMarkdown": "@mpware \n\nI use every possible sequence for validation. [A, -, -, ...], [A, B, -, ...], ...\n\nBy the way, I got a submission error for my last model after waiting for 9h, though it doesn't say timeout. That sounds like what you told before.  WTF what can I do? I don't have any prior version enviroment. I've seen there's a v91 version, but I've checked the Dockerfile commit history and doesn't seem like they have fixed anything relevant.",
          "votes": 1
        },
        {
          "id": 1108474,
          "postDate": "2020-12-10T17:43:51.090Z",
          "content": "<p>From my experience v90 works but slower than v89. If you don't have a kernel with v89 then you can try to fork one in the public kernels.</p>",
          "rawMarkdown": "From my experience v90 works but slower than v89. If you don't have a kernel with v89 then you can try to fork one in the public kernels.",
          "votes": 2
        },
        {
          "id": 1108518,
          "postDate": "2020-12-10T18:42:24.533Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> i too get similar CV but not same LBs..<br>\n could i be making mistake in passing response sequence to model</p>\n<pre><code>label = qa[1:]\n rt=qa[1:-1].copy()\n rt = np.append(np.zeros((1,)),rt)\n</code></pre>\n<p>during inference however i still case doubt  because if i know correctly model relies on previous exercise responses but during test time where do we have response for test q during that iteration.<br>\nAll i see in inferences so far that <br>\n<code>qa[1:] , q[2:].append( test questions)</code><br>\nwhere QA is series of Yes or No for correct ans from Train set/Valid set   but our question array has got question of Test data .</p>\n<pre><code>eg. E1 T1 T2 T3\nresp sequence is \n     E0,E1,E2,E3..  \n</code></pre>\n<p>i think m getting we append to last one or few test q given the history of train exercise responses</p>",
          "rawMarkdown": "@claverru i too get similar CV but not same LBs..\n could i be making mistake in passing response sequence to model\n```\nlabel = qa[1:]\n rt=qa[1:-1].copy()\n rt = np.append(np.zeros((1,)),rt)\n```\nduring inference however i still case doubt  because if i know correctly model relies on previous exercise responses but during test time where do we have response for test q during that iteration.\nAll i see in inferences so far that \n`qa[1:] , q[2:].append( test questions)`\nwhere QA is series of Yes or No for correct ans from Train set/Valid set   but our question array has got question of Test data .\n```\neg. E1 T1 T2 T3\nresp sequence is \n     E0,E1,E2,E3..  \n```\ni think m getting we append to last one or few test q given the history of train exercise responses\n"
        },
        {
          "id": 1109164,
          "postDate": "2020-12-11T11:27:47.703Z",
          "content": "<p>Okay it was not a Docker Image problem. I've been submitting with CPU all the time. I didn't know that I couldn't submit with GPU when doing \"Quick Save\" O.o (this is my first serious participation in a competition).</p>\n<p>With GPU, inference time was around 2 hours, and <strong>my new LB is 0.784</strong>. Now that I have a good pipeline and a decent correlation between CV and LB I think I can still improve my score with certain agility. </p>",
          "rawMarkdown": "Okay it was not a Docker Image problem. I've been submitting with CPU all the time. I didn't know that I couldn't submit with GPU when doing \"Quick Save\" O.o (this is my first serious participation in a competition).\n\nWith GPU, inference time was around 2 hours, and __my new LB is 0.784__. Now that I have a good pipeline and a decent correlation between CV and LB I think I can still improve my score with certain agility. ",
          "votes": 4
        },
        {
          "id": 1109171,
          "postDate": "2020-12-11T11:39:53.373Z",
          "content": "<p>You can edit the settings of the quick save as well and then probably we can use quick save as well…!<br>\nGreat LB! Keep Rising! 🎉🎊🎊</p>",
          "rawMarkdown": "You can edit the settings of the quick save as well and then probably we can use quick save as well...!\nGreat LB! Keep Rising! 🎉🎊🎊"
        },
        {
          "id": 1109238,
          "postDate": "2020-12-11T13:04:52.863Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>  I was doing it. I mark: \"Run with GPU for all sessions\", and have my GPU connected when submitting, but for some reason, it seems like you can only submit with GPU when doing \"Save and run all\".</p>\n<p>On the other hand, I'm grinding just thanks to the help received around, so thank you guys! I feel in debt and I'm trying to collaborate in the same proportion.</p>",
          "rawMarkdown": "@adityaecdrid  I was doing it. I mark: \"Run with GPU for all sessions\", and have my GPU connected when submitting, but for some reason, it seems like you can only submit with GPU when doing \"Save and run all\".\n\nOn the other hand, I'm grinding just thanks to the help received around, so thank you guys! I feel in debt and I'm trying to collaborate in the same proportion.",
          "votes": 1
        },
        {
          "id": 1109290,
          "postDate": "2020-12-11T14:08:43.090Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> great going  how much was score you got without updating status of test df with group ans ?<br>\nm able to get so far LB 74.8  without status update &amp; without masking  padded items during loss<br>\nm still trying to find what is holding me back :).. any advise appreciated . I joined the competition party week ago so still working to get basics right :)</p>",
          "rawMarkdown": "@claverru great going  how much was score you got without updating status of test df with group ans ?\nm able to get so far LB 74.8  without status update & without masking  padded items during loss\nm still trying to find what is holding me back :).. any advise appreciated . I joined the competition party week ago so still working to get basics right :)"
        },
        {
          "id": 1116239,
          "postDate": "2020-12-17T01:39:23.867Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Thanks for your sampling strategy and it works for me. With the sampling strategy and some feature engineering, I got 0.786 LB with a single SAINT+ model.</p>",
          "rawMarkdown": "@claverru Thanks for your sampling strategy and it works for me. With the sampling strategy and some feature engineering, I got 0.786 LB with a single SAINT+ model.",
          "votes": 2
        },
        {
          "id": 1122161,
          "postDate": "2020-12-22T08:35:53.913Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1131304,
          "postDate": "2020-12-29T16:40:38.877Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Thanks for your sampling strategy I was able to obtain CV of 0.765 (test set last 100 interactions) and LB 0.776  with the original Saint model (No time features). </p>\n<p>For those using this strategy how are you implementing it? I am sub classing from tf.keras.utils.Sequence but the GPU is now severely bottle necked.   </p>",
          "rawMarkdown": "@claverru Thanks for your sampling strategy I was able to obtain CV of 0.765 (test set last 100 interactions) and LB 0.776  with the original Saint model (No time features). \n\nFor those using this strategy how are you implementing it? I am sub classing from tf.keras.utils.Sequence but the GPU is now severely bottle necked.   "
        }
      ]
    },
    {
      "id": 1081936,
      "postDate": "2020-11-17T12:40:52.267Z",
      "content": "<p>I've tried several variant of saint, and only have a CV score to provide, no LB yet:<br>\nVariants : </p>\n<ul>\n<li>Saint++ implemented with the same parameters as in the papers and the same features : CV Roc AUC ~0.757</li>\n<li>Decoder architecture trained with both a next question prediction loss and a next correct prediction loss : CV Roc AUC ~0.762</li>\n<li>Encoder architecture (bidirectional) pretrained with a masked question modeling loss then finetuned with a sequence pair classification loss (seq 1 = question to predict, seq 2 = user history) : CV ROC AUC : ~0.77</li>\n<li>Replace the pure attention layers of transformers by LSTM + Attention : no performance improvement and instability during training.</li>\n</ul>\n<p>For each of these architectures I chose a sequence length of 128, a depth of model of 512 and a ffn representation size of 1024. Optimizer was adam with lr 3e-5 (more caused gradient explosion) and bs either 32 or 64.<br>\nNumber of layers per block was either 4 or 8</p>\n<p>Convergence is often decided during the first few epochs, but increase slitghly when continuing training (i've trained the encoder on MQM loss for about 8h on a RTX2070)</p>\n<p>Given the very low differences between the different architectures performances, i'd say that the architecture is not really impactfull for the modelisation (providing that you have a good one).</p>\n<p>So for next iteration i might focus more on how to build embeddings for the different sequences (right now i use on embedding layer for each input sequence and i do a summation)</p>\n<p>If you have any ideas of other variants to try or new features to include i'm open for discussion :)</p>\n<p>PS : For the ones trying to replicate the encoder results, the MQM loss need to be quickstart by first making a directionnal encoder, training for a few epochs with a next question prediction loss and then adding the look ahead masks and swithing to MQM loss (it fails to start when only on MQM loss)</p>",
      "rawMarkdown": "I've tried several variant of saint, and only have a CV score to provide, no LB yet:\nVariants : \n- Saint++ implemented with the same parameters as in the papers and the same features : CV Roc AUC ~0.757\n- Decoder architecture trained with both a next question prediction loss and a next correct prediction loss : CV Roc AUC ~0.762\n- Encoder architecture (bidirectional) pretrained with a masked question modeling loss then finetuned with a sequence pair classification loss (seq 1 = question to predict, seq 2 = user history) : CV ROC AUC : ~0.77\n- Replace the pure attention layers of transformers by LSTM + Attention : no performance improvement and instability during training.\n\nFor each of these architectures I chose a sequence length of 128, a depth of model of 512 and a ffn representation size of 1024. Optimizer was adam with lr 3e-5 (more caused gradient explosion) and bs either 32 or 64.\nNumber of layers per block was either 4 or 8\n\nConvergence is often decided during the first few epochs, but increase slitghly when continuing training (i've trained the encoder on MQM loss for about 8h on a RTX2070)\n\n\nGiven the very low differences between the different architectures performances, i'd say that the architecture is not really impactfull for the modelisation (providing that you have a good one).\n\nSo for next iteration i might focus more on how to build embeddings for the different sequences (right now i use on embedding layer for each input sequence and i do a summation)\n\nIf you have any ideas of other variants to try or new features to include i'm open for discussion :)\n\n\nPS : For the ones trying to replicate the encoder results, the MQM loss need to be quickstart by first making a directionnal encoder, training for a few epochs with a next question prediction loss and then adding the look ahead masks and swithing to MQM loss (it fails to start when only on MQM loss)\n\n",
      "votes": 13,
      "replies": [
        {
          "id": 1081967,
          "postDate": "2020-11-17T13:28:54.047Z",
          "content": "<p><a href=\"https://www.kaggle.com/rously\" target=\"_blank\">@rously</a> Thanks for your feedback, it's really interesting that you've achieved CV 0.757 with SAINT+ because it's quite similar for me, I'm not able to go beyond CV 0.76 with my implementation and with similar parameters as in paper (encoder + decoder, d_model=256, seq_len=100, nhead=8, 4 layers, batch size=256). LR with warmup (0 to 0.0008) works quite well for me and allows fast convergence after a few epochs.</p>\n<p>About:</p>\n<blockquote>\n  <p>Encoder architecture (bidirectional) pretrained with a masked question modeling loss then finetuned with a sequence pair classification loss (seq 1 = question to predict, seq 2 = user history) : CV ROC AUC : ~0.77</p>\n</blockquote>\n<p>masked question modeling = MQM, correct? Where does it apply? Only in loss? Or also in padding mask?</p>",
          "rawMarkdown": "@rously Thanks for your feedback, it's really interesting that you've achieved CV 0.757 with SAINT+ because it's quite similar for me, I'm not able to go beyond CV 0.76 with my implementation and with similar parameters as in paper (encoder + decoder, d_model=256, seq_len=100, nhead=8, 4 layers, batch size=256). LR with warmup (0 to 0.0008) works quite well for me and allows fast convergence after a few epochs.\n\nAbout:\n> Encoder architecture (bidirectional) pretrained with a masked question modeling loss then finetuned with a sequence pair classification loss (seq 1 = question to predict, seq 2 = user history) : CV ROC AUC : ~0.77\n\nmasked question modeling = MQM, correct? Where does it apply? Only in loss? Or also in padding mask?",
          "votes": 1
        },
        {
          "id": 1082438,
          "postDate": "2020-11-17T22:19:27.737Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <br>\nYep, MQM = masked question modeling loss, it's the same principle as mlm for language model, you hide 15% of the input tokens and you try to predict it. My intuition behind it was that only predicting the output of the user doesn't motivate the model enough to understand the pattern of sequences of questions, so i added this loss to do that.</p>\n<p>Basically a classical input will be constituted of several input sequence, but to understand it you can decompose it like that :</p>\n<p>[CLS] last_question_token [SEP] tok1 tok2 tok3 tok4 … [PAD] [PAD]…<br>\nThe MQM loss is only applied to the part of the sequence between the SEP and the PAD not anywhere else</p>",
          "rawMarkdown": "@mpware \nYep, MQM = masked question modeling loss, it's the same principle as mlm for language model, you hide 15% of the input tokens and you try to predict it. My intuition behind it was that only predicting the output of the user doesn't motivate the model enough to understand the pattern of sequences of questions, so i added this loss to do that.\n\nBasically a classical input will be constituted of several input sequence, but to understand it you can decompose it like that :\n\n[CLS] last_question_token [SEP] tok1 tok2 tok3 tok4 ... [PAD] [PAD]...\nThe MQM loss is only applied to the part of the sequence between the SEP and the PAD not anywhere else",
          "votes": 2
        },
        {
          "id": 1082516,
          "postDate": "2020-11-18T01:06:38.590Z",
          "content": "<p>This is really quite an interesting approach that one can try if time allows!</p>",
          "rawMarkdown": "This is really quite an interesting approach that one can try if time allows!"
        },
        {
          "id": 1093375,
          "postDate": "2020-11-27T16:52:15.793Z",
          "content": "<p><a href=\"https://www.kaggle.com/rously\" target=\"_blank\">@rously</a> masking of 15% of input tokens applies to both encoder and decoder inputs?</p>",
          "rawMarkdown": "@rously masking of 15% of input tokens applies to both encoder and decoder inputs?"
        },
        {
          "id": 1093735,
          "postDate": "2020-11-28T01:16:32.927Z",
          "content": "<p>In my case i was doing an only encoder architecture, but I guess that with an architecture encoder-decoder like saint you can pretrain the encoder with the masked loss as well.</p>\n<p>Havent had much time to experiment since my last post, but i noticed that using a mse or rankboost loss instead of crossentropy was giving slithly better result (my best one gives a CV of 77.3% with this)</p>",
          "rawMarkdown": "In my case i was doing an only encoder architecture, but I guess that with an architecture encoder-decoder like saint you can pretrain the encoder with the masked loss as well.\n\nHavent had much time to experiment since my last post, but i noticed that using a mse or rankboost loss instead of crossentropy was giving slithly better result (my best one gives a CV of 77.3% with this)"
        },
        {
          "id": 1108939,
          "postDate": "2020-12-11T07:00:13.740Z",
          "content": "<p><a href=\"https://www.kaggle.com/rously\" target=\"_blank\">@rously</a> <br>\nMQM= masked question modeling loss<br>\ndo you have pytorch impl repository for it ?</p>",
          "rawMarkdown": "@rously \nMQM= masked question modeling loss\ndo you have pytorch impl repository for it ?",
          "replies": [
            {
              "id": 1109339,
              "postDate": "2020-12-11T14:55:56.263Z",
              "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> <br>\nnope, i did the implementation in tensorflow. It is quite simple to do though with a generator, take your input sequence, generate an numpy array with numpy.random.choice([0,1],size = max_len, p = [1-r, r]) where r is your rate of masking, then you just have to make a multiplication to get the inputs and outputs of your network, just have to mask the zero tokens of the outputs in the loss calculation</p>",
              "rawMarkdown": "@jaideepvalani \nnope, i did the implementation in tensorflow. It is quite simple to do though with a generator, take your input sequence, generate an numpy array with numpy.random.choice([0,1],size = max_len, p = [1-r, r]) where r is your rate of masking, then you just have to make a multiplication to get the inputs and outputs of your network, just have to mask the zero tokens of the outputs in the loss calculation",
              "votes": 1
            },
            {
              "id": 1109350,
              "postDate": "2020-12-11T15:13:58.477Z",
              "content": "<p>By the way, some other ideas I tried but that did not improve the validation score:</p>\n<ul>\n<li>encode the input sequences with a tabnet architecture, doesn't improve the result and is a pain in the ass to train as tabnet requires a different learning rate than transformer</li>\n<li>modify the encoder-decoder merge by putting the encoder values as query and key and the decoder as value (this would make more sense to compare similar items)</li>\n<li>use a tabnet layer as classification, somehow it managed to peak forward in time through the batch normalization, so gave an AUC of 0.84 but did not generalize in pure inference</li>\n<li>increase batch size a lot (to 2048 thanks to TPU) did not help either</li>\n<li>use local attention windows (in order to look both at all history and at only the last 20 values)</li>\n</ul>\n<p>One thing that seems to help though is to replace the bce loss by a rank boost loss</p>\n<p>Perhaps with the sampling strategy from <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> these ideas might improve a bit</p>",
              "rawMarkdown": "By the way, some other ideas I tried but that did not improve the validation score:\n- encode the input sequences with a tabnet architecture, doesn't improve the result and is a pain in the ass to train as tabnet requires a different learning rate than transformer\n- modify the encoder-decoder merge by putting the encoder values as query and key and the decoder as value (this would make more sense to compare similar items)\n- use a tabnet layer as classification, somehow it managed to peak forward in time through the batch normalization, so gave an AUC of 0.84 but did not generalize in pure inference\n- increase batch size a lot (to 2048 thanks to TPU) did not help either\n- use local attention windows (in order to look both at all history and at only the last 20 values)\n\nOne thing that seems to help though is to replace the bce loss by a rank boost loss\n\nPerhaps with the sampling strategy from @claverru these ideas might improve a bit"
            }
          ]
        }
      ]
    },
    {
      "id": 1112659,
      "postDate": "2020-12-14T19:38:02.333Z",
      "content": "<p>PyTorch implementation of SAINT - <a href=\"https://github.com/arshadshk/SAINT-pytorch\" target=\"_blank\">https://github.com/arshadshk/SAINT-pytorch</a></p>\n<p>Let me know if any corrections or suggestions are there.<br>\n(Also have a look at SAKT-PyTorch <a href=\"https://github.com/arshadshk/SAKT-pytorch\" target=\"_blank\">https://github.com/arshadshk/SAKT-pytorch</a> )</p>",
      "rawMarkdown": "PyTorch implementation of SAINT - https://github.com/arshadshk/SAINT-pytorch\n\nLet me know if any corrections or suggestions are there.\n(Also have a look at SAKT-PyTorch https://github.com/arshadshk/SAKT-pytorch )",
      "votes": 11,
      "replies": [
        {
          "id": 1112800,
          "postDate": "2020-12-14T22:43:00.533Z",
          "content": "<p>Thanks for doing this. I'm two days into my own SAINT implementation, so it'd be nice to check my work against someone else's—considering there are no (other) public SAINT(+) implementation kernels or repos as far as I am aware. If you post your repo as a notebook, you can probably get some kaggle kernel points there as well 👍🏾.</p>",
          "rawMarkdown": "Thanks for doing this. I'm two days into my own SAINT implementation, so it'd be nice to check my work against someone else's—considering there are no (other) public SAINT(+) implementation kernels or repos as far as I am aware. If you post your repo as a notebook, you can probably get some kaggle kernel points there as well 👍🏾."
        },
        {
          "id": 1116146,
          "postDate": "2020-12-16T22:22:19.400Z",
          "content": "<p><a href=\"https://www.kaggle.com/arshad7\" target=\"_blank\">@arshad7</a> Can you please explain what is <code>total_in</code> input to the model ? I know readme says <code>Total number of unique interactions.</code> and for the random check you have it as 2.  Thank you. </p>",
          "rawMarkdown": "@arshad7 Can you please explain what is `total_in` input to the model ? I know readme says `Total number of unique interactions.` and for the random check you have it as 2.  Thank you. "
        },
        {
          "id": 1116253,
          "postDate": "2020-12-17T01:58:44.187Z",
          "content": "<p>I think it indicates the number of response class. Binay if total_in=2, i.e., 0 and 1. You can verify that by <code>in_de</code> from the code below.<br>\n<code>in_ex, in_cat, in_de = random_data(64, seq_len , total_ex, total_cat, total_in)</code></p>",
          "rawMarkdown": "I think it indicates the number of response class. Binay if total_in=2, i.e., 0 and 1. You can verify that by `in_de` from the code below.\n```in_ex, in_cat, in_de = random_data(64, seq_len , total_ex, total_cat, total_in)```\n",
          "votes": 1
        },
        {
          "id": 1116384,
          "postDate": "2020-12-17T06:02:30.477Z",
          "content": "<p><a href=\"https://www.kaggle.com/arshad7\" target=\"_blank\">@arshad7</a> Thank you for posting the repo! I just have one question, what are the three inputs in the forward pass? </p>",
          "rawMarkdown": "@arshad7 Thank you for posting the repo! I just have one question, what are the three inputs in the forward pass? "
        },
        {
          "id": 1116465,
          "postDate": "2020-12-17T07:35:37.350Z",
          "content": "<p><a href=\"https://www.kaggle.com/arshad7\" target=\"_blank\">@arshad7</a> I made use of it , refactored to separate encoder ,decider inputs and data input . What i find is skip version fair quite low in cv compared to non skip.<br>\n<a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> do you use skip version or non skip</p>",
          "rawMarkdown": "@arshad7 I made use of it , refactored to separate encoder ,decider inputs and data input . What i find is skip version fair quite low in cv compared to non skip.\n@mpware do you use skip version or non skip"
        },
        {
          "id": 1123601,
          "postDate": "2020-12-23T11:11:48.017Z",
          "content": "<p>I use skip version.</p>",
          "rawMarkdown": "I use skip version."
        },
        {
          "id": 1133723,
          "postDate": "2020-12-31T13:17:39.747Z",
          "content": "<p>I may be too late here. But I finally was able to make use of the SAINT implementation. But it's around CV 0.700. What was your CV/LB for this? I wonder if I have wrong dataloader implementation.</p>",
          "rawMarkdown": "I may be too late here. But I finally was able to make use of the SAINT implementation. But it's around CV 0.700. What was your CV/LB for this? I wonder if I have wrong dataloader implementation."
        },
        {
          "id": 1133755,
          "postDate": "2020-12-31T13:56:26.617Z",
          "content": "<p>Original saint I got CV 0.765 and LB 0.776. </p>",
          "rawMarkdown": "Original saint I got CV 0.765 and LB 0.776. ",
          "votes": 1
        },
        {
          "id": 1134178,
          "postDate": "2021-01-01T00:20:04.597Z",
          "content": "<p>Thank you for your reply. I must have something wrong in my code then.</p>",
          "rawMarkdown": "Thank you for your reply. I must have something wrong in my code then."
        },
        {
          "id": 1135225,
          "postDate": "2021-01-02T03:36:22.583Z",
          "content": "<p>Thanks. I got LB: 0.741 using <a href=\"https://github.com/arshadshk/SAINT-pytorch\" target=\"_blank\">https://github.com/arshadshk/SAINT-pytorch</a>.<br>\nI had a mistake on my dataloader implementation.</p>",
          "rawMarkdown": "Thanks. I got LB: 0.741 using https://github.com/arshadshk/SAINT-pytorch.\nI had a mistake on my dataloader implementation.",
          "votes": 1
        },
        {
          "id": 1136431,
          "postDate": "2021-01-03T04:23:09.743Z",
          "content": "<p>I get different model output after training and after loading a saved transformer model. I initially thought there must be some input processing bug, but then I directly passed tensors after training and after saving and loading the model, results are different.  I appreciate any  thoughts on what I could have missed. I used this model architecture - &nbsp;<a href=\"https://github.com/arshadshk/SAINT-pytorch\" target=\"_blank\">https://github.com/arshadshk/SAINT-pytorch</a>  Thank you! </p>\n<p>Edit: loading state_dict() is not working however loading full model works. don't know why 🤔</p>",
          "rawMarkdown": "I get different model output after training and after loading a saved transformer model. I initially thought there must be some input processing bug, but then I directly passed tensors after training and after saving and loading the model, results are different.  I appreciate any  thoughts on what I could have missed. I used this model architecture -  https://github.com/arshadshk/SAINT-pytorch  Thank you! \n\nEdit: loading state_dict() is not working however loading full model works. don't know why 🤔"
        }
      ]
    },
    {
      "id": 1128932,
      "postDate": "2020-12-27T22:28:18.320Z",
      "content": "<p>Gotta tell you guys, my last update brought me from 784 to 794. Only one detail got me there, related with timestamp (and I think it is still improvable). For those who haven't worked around this feature hard enough, I really recommend it.</p>",
      "rawMarkdown": "Gotta tell you guys, my last update brought me from 784 to 794. Only one detail got me there, related with timestamp (and I think it is still improvable). For those who haven't worked around this feature hard enough, I really recommend it.",
      "votes": 9,
      "replies": [
        {
          "id": 1128944,
          "postDate": "2020-12-27T22:56:48.733Z",
          "content": "<p>Great! You mean lag? Timestamp(N) - Timestamp(N-1)?<br>\nMy LB=0.791 is with SAINT+ paper + had_explanation. Nothing more.</p>",
          "rawMarkdown": "Great! You mean lag? Timestamp(N) - Timestamp(N-1)?\nMy LB=0.791 is with SAINT+ paper + had_explanation. Nothing more.",
          "votes": 3
        },
        {
          "id": 1128947,
          "postDate": "2020-12-27T23:07:08.190Z",
          "content": "<p>Yeah, lag it is, but you have to take it carefully.</p>",
          "rawMarkdown": "Yeah, lag it is, but you have to take it carefully.",
          "votes": 3
        },
        {
          "id": 1128951,
          "postDate": "2020-12-27T23:13:52.627Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> You don't know how glad I am to see you get higher in the LB - I checked your progress almost everyday - I thought you gave up but hoped you didn't because you have done so much!</p>\n<p>Again, good for you!</p>",
          "rawMarkdown": "@claverru You don't know how glad I am to see you get higher in the LB - I checked your progress almost everyday - I thought you gave up but hoped you didn't because you have done so much!\n\nAgain, good for you!",
          "votes": 3
        },
        {
          "id": 1128964,
          "postDate": "2020-12-28T00:04:38.233Z",
          "content": "<p>Thanks for reporting. I too had a jump from 780s to 790s (cv) by engineering the ts feature, and also agree there's plenty to be done with it in terms of stats surrounding categorical groupings.</p>\n<blockquote>\n  <p>Yeah, lag it is, but you have to take it carefully.</p>\n</blockquote>\n<p>I didn't do anything special to the variable though. Now you make me wonder if I should look at its values more closely 👀</p>",
          "rawMarkdown": "Thanks for reporting. I too had a jump from 780s to 790s (cv) by engineering the ts feature, and also agree there's plenty to be done with it in terms of stats surrounding categorical groupings.\n\n> Yeah, lag it is, but you have to take it carefully.\n\nI didn't do anything special to the variable though. Now you make me wonder if I should look at its values more closely 👀",
          "votes": 1
        },
        {
          "id": 1129013,
          "postDate": "2020-12-28T02:03:25.837Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Did you use continuous embedding or categorical embedding?</p>",
          "rawMarkdown": "@claverru Did you use continuous embedding or categorical embedding?\n"
        },
        {
          "id": 1129015,
          "postDate": "2020-12-28T02:08:46.640Z",
          "content": "<p>Wow! I thought you have left the competition for other one's after working hard enough in this one! Really cool and great going!</p>\n<p>For NNs h<strong>ow you pre-process your data matters a lot</strong>, I don't think the embedding/non_embedding of that will be a huge difference if used correctly</p>",
          "rawMarkdown": "Wow! I thought you have left the competition for other one's after working hard enough in this one! Really cool and great going!\n\nFor NNs h**ow you pre-process your data matters a lot**, I don't think the embedding/non_embedding of that will be a huge difference if used correctly"
        },
        {
          "id": 1129061,
          "postDate": "2020-12-28T03:47:14.653Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1129066,
          "postDate": "2020-12-28T04:03:50.133Z",
          "content": "<p>How much was wrt bad results </p>\n<p>And other thing 79.1 79.2 probably you would have expected more going by trend when we submit saint with cv 75 ,lb 776 or so . What is cv method you use default as in mpwares? </p>",
          "rawMarkdown": "How much was wrt bad results \n\nAnd other thing 79.1 79.2 probably you would have expected more going by trend when we submit saint with cv 75 ,lb 776 or so . What is cv method you use default as in mpwares? \n"
        },
        {
          "id": 1129367,
          "postDate": "2020-12-28T09:39:36.840Z",
          "content": "<p>Bad result is cv:0.65x. I not submitted. The cv method I used is same as popular public notebook.</p>",
          "rawMarkdown": " Bad result is cv:0.65x. I not submitted. The cv method I used is same as popular public notebook."
        },
        {
          "id": 1144675,
          "postDate": "2021-01-08T15:51:47.413Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> thank you for your words mate, I didn't read them until now.</p>",
          "rawMarkdown": "@yihdarshieh thank you for your words mate, I didn't read them until now.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1071786,
      "postDate": "2020-11-07T12:06:59.117Z",
      "content": "<p>I've read both SAINT and SAINT+ papers but some questions remain.</p>\n<p><strong>Input sequences</strong> should be something like:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fce0c6347798d2f33707b51b623e4a714%2Fsequences.png?generation=1604752912346221&amp;alt=media\" alt=\"\"></p>\n<p><strong>Start token(s):</strong><br>\n<em>The decoder takes O and another sequential input Re = [S,Re1,··· ,Rek−1] of response embeddings with the start token embedding S.</em></p>\n<p>Something like this should work for start token:</p>\n<pre><code>...\nresponse_size = 2\nself.response_embedding = nn.Embedding(response_size + 1, embedding_dim) # +1 to include start token\n...\nx_correctness = self.response_embedding(response_sequence)\n...\n# Add start token to correctness\nx_correctness = torch.roll(x_correctness, shifts=(0, 1, 0), dims=(0, 1, 0)) # Shift right the sequence\nx_correctness[:,0,:] = self.response_size # Start token\n</code></pre>\n<p>x_correctness is embedding of response sequence (0,1,1,1,0 …), response_size=2, so with token it will be (2,0,1,1,1,0 …)</p>\n<p><strong>Self attention mask:</strong></p>\n<pre><code># If a BoolTensor is provided, the positions with the value of True will be ignored while the position with the value of False will be unchanged.\n# tensor([[False,  True,  True,  True],\n#         [False, False,  True,  True],\n#         [False, False, False,  True],\n#         [False, False, False, False]])    \ndef generate_mask(self, size, diagonal=1):        \n    return torch.triu(torch.ones(size, size)==1, diagonal=diagonal)\n</code></pre>\n<p><strong>Learning rate:</strong><br>\nWarmup from 2000 to 4000 iterations works.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F9d3dc4354e1cb1ba653cfb8adba9e84a%2Flr.png?generation=1605009427277426&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I've read both SAINT and SAINT+ papers but some questions remain.\n\n**Input sequences** should be something like:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fce0c6347798d2f33707b51b623e4a714%2Fsequences.png?generation=1604752912346221&alt=media)\n\n**Start token(s):**\n*The decoder takes O and another sequential input Re = [S,Re1,··· ,Rek−1] of response embeddings with the start token embedding S.*\n\nSomething like this should work for start token:\n```\n...\nresponse_size = 2\nself.response_embedding = nn.Embedding(response_size + 1, embedding_dim) # +1 to include start token\n...\nx_correctness = self.response_embedding(response_sequence)\n...\n# Add start token to correctness\nx_correctness = torch.roll(x_correctness, shifts=(0, 1, 0), dims=(0, 1, 0)) # Shift right the sequence\nx_correctness[:,0,:] = self.response_size # Start token\n```\nx_correctness is embedding of response sequence (0,1,1,1,0 ...), response_size=2, so with token it will be (2,0,1,1,1,0 ...)\n\n\n**Self attention mask:**\n```\n# If a BoolTensor is provided, the positions with the value of True will be ignored while the position with the value of False will be unchanged.\n# tensor([[False,  True,  True,  True],\n#         [False, False,  True,  True],\n#         [False, False, False,  True],\n#         [False, False, False, False]])    \ndef generate_mask(self, size, diagonal=1):        \n    return torch.triu(torch.ones(size, size)==1, diagonal=diagonal)\n```\n\n**Learning rate:**\nWarmup from 2000 to 4000 iterations works.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F9d3dc4354e1cb1ba653cfb8adba9e84a%2Flr.png?generation=1605009427277426&alt=media)\n",
      "votes": 8,
      "replies": [
        {
          "id": 1071884,
          "postDate": "2020-11-07T14:15:42.897Z",
          "content": "<p>Nice work! Do we need the decoder here as well?</p>",
          "rawMarkdown": "Nice work! Do we need the decoder here as well?",
          "votes": 1
        },
        {
          "id": 1071920,
          "postDate": "2020-11-07T15:05:22.973Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> Yes, from the paper we need both encoder and decoder.</p>\n<pre><code># Transformer with default encoder/decoder        \nself.transformer = nn.Transformer(d_model=input_features_dim, \n                                  nhead=8, \n                                  num_encoder_layers= 6,\n                                  num_decoder_layers= 6, \n                                  dim_feedforward=2048, \n                                  dropout=0.1, \n                                  activation='relu', \n                                  custom_encoder = None,\n                                  custom_decoder = None)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fcbaeee6b56f3eeac7408c37d250ca3c8%2Ftransformer.png?generation=1604761231138288&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "@adityaecdrid Yes, from the paper we need both encoder and decoder.\n\n```\n# Transformer with default encoder/decoder        \nself.transformer = nn.Transformer(d_model=input_features_dim, \n                                  nhead=8, \n                                  num_encoder_layers= 6,\n                                  num_decoder_layers= 6, \n                                  dim_feedforward=2048, \n                                  dropout=0.1, \n                                  activation='relu', \n                                  custom_encoder = None,\n                                  custom_decoder = None)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fcbaeee6b56f3eeac7408c37d250ca3c8%2Ftransformer.png?generation=1604761231138288&alt=media)",
          "votes": 1,
          "replies": [
            {
              "id": 1108884,
              "postDate": "2020-12-11T05:39:03.937Z",
              "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> Yes, from the paper we need both encoder and decoder.</p>\n<pre><code># Transformer with default encoder/decoder        \nself.transformer = nn.Transformer(d_model=input_features_dim, \n                                  nhead=8, \n                                  num_encoder_layers= 6,\n                                  num_decoder_layers= 6, \n                                  dim_feedforward=2048, \n                                  dropout=0.1, \n                                  activation='relu', \n                                  custom_encoder = None,\n                                  custom_decoder = None)\n</code></pre>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fcbaeee6b56f3eeac7408c37d250ca3c8%2Ftransformer.png?generation=1604761231138288&amp;alt=media\" alt=\"\"></p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> about response shift<br>\n1) how about merely doing this in the dataset</p>\n<pre><code>target_id = q[1:]\n rt=qa[1:-1].copy()\n rt = np.append(np.zeros((1,)),rt)\n</code></pre>\n<p>so this should give prev ans to every current exercise. embeddings should get formed accordingly only ..  does it make a difference  compared to what you are doing after computing embeddings and then doing masking ?</p>\n<p>2) in dataset are you doing any shifting of inputs ?</p>",
              "rawMarkdown": "> @adityaecdrid Yes, from the paper we need both encoder and decoder.\n> \n> ```\n> # Transformer with default encoder/decoder        \n> self.transformer = nn.Transformer(d_model=input_features_dim, \n>                                   nhead=8, \n>                                   num_encoder_layers= 6,\n>                                   num_decoder_layers= 6, \n>                                   dim_feedforward=2048, \n>                                   dropout=0.1, \n>                                   activation='relu', \n>                                   custom_encoder = None,\n>                                   custom_decoder = None)\n> ```\n> \n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fcbaeee6b56f3eeac7408c37d250ca3c8%2Ftransformer.png?generation=1604761231138288&alt=media)\n\n@mpware about response shift\n1) how about merely doing this in the dataset\n \n   ```\ntarget_id = q[1:]\n rt=qa[1:-1].copy()\n rt = np.append(np.zeros((1,)),rt)\n```\nso this should give prev ans to every current exercise. embeddings should get formed accordingly only ..  does it make a difference  compared to what you are doing after computing embeddings and then doing masking ?\n\n2) in dataset are you doing any shifting of inputs ?"
            }
          ]
        },
        {
          "id": 1071932,
          "postDate": "2020-11-07T15:31:46.570Z",
          "content": "<p>Thanks MPWARE.  At inference time, when we will see unseen users, how can we pass in the i/p's that's needed? Plus we will have to maintain the seq for each un-seen user as well, right?<br>\nThanks for torch.roll as well!</p>",
          "rawMarkdown": "Thanks MPWARE.  At inference time, when we will see unseen users, how can we pass in the i/p's that's needed? Plus we will have to maintain the seq for each un-seen user as well, right?\nThanks for torch.roll as well!",
          "votes": 1
        },
        {
          "id": 1071946,
          "postDate": "2020-11-07T15:51:10.660Z",
          "content": "<p>During inference, we need to maintain the full sequences per user (i.e. the last 99 answers, questions,  …) then append the new question_id and our model will predict the next answer.<br>\nIt will be easy for groups including only one new question per user, but we know that we can have more than one question (between 1 to 5 if I'm not wrong according to <code>task_container_id</code> possible total questions) then we have several options: </p>\n<ul>\n<li>Consider our prediction as correct and append it to answer sequence and predict again with new question.</li>\n<li>Have another model trained to predict 2 answers, another to predict 3 answers …</li>\n<li>Or?</li>\n</ul>\n<p>Maintaining last 100 answers in memory for all users (with some flush to be memory friendly) is possible within the 9 hours inference time. You could implement it and simulate it with this nice <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">kernel</a> from <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> </p>\n<p>For unseen users, we won't have any previous answers but I believe we could use optional masks in our model to ignore padding we've added to reach sequence size for such new users.</p>",
          "rawMarkdown": "During inference, we need to maintain the full sequences per user (i.e. the last 99 answers, questions,  ...) then append the new question_id and our model will predict the next answer.\nIt will be easy for groups including only one new question per user, but we know that we can have more than one question (between 1 to 5 if I'm not wrong according to `task_container_id` possible total questions) then we have several options: \n- Consider our prediction as correct and append it to answer sequence and predict again with new question.\n- Have another model trained to predict 2 answers, another to predict 3 answers ...\n- Or?\n\nMaintaining last 100 answers in memory for all users (with some flush to be memory friendly) is possible within the 9 hours inference time. You could implement it and simulate it with this nice [kernel](https://www.kaggle.com/its7171/time-series-api-iter-test-emulator) from @its7171 \n\nFor unseen users, we won't have any previous answers but I believe we could use optional masks in our model to ignore padding we've added to reach sequence size for such new users.",
          "votes": 4
        },
        {
          "id": 1071950,
          "postDate": "2020-11-07T15:57:02.733Z",
          "content": "<blockquote>\n  <p>For unseen users, we won't have any previous answers but I believe we could use optional masks in our model to ignore padding we've added to reach sequence size for such new users.</p>\n</blockquote>\n<p>I see. Thanks for the tip; Have a lot of working to do.</p>\n<blockquote>\n  <p>Maintaining last 100 answers in memory for all users (with some flush to be memory friendly) is possible within the 9 hours inference time.</p>\n</blockquote>\n<p>Yep, last 100 is the window length as mentioned in the paper and I was using a deque for that but now torch.roll might come handy! </p>\n<p>Also, Are you simply taking the last 100 seq's / interactions or splitting let's say 1000 content_ids into 10 chunks?</p>",
          "rawMarkdown": ">For unseen users, we won't have any previous answers but I believe we could use optional masks in our model to ignore padding we've added to reach sequence size for such new users.\n\nI see. Thanks for the tip; Have a lot of working to do.\n\n>Maintaining last 100 answers in memory for all users (with some flush to be memory friendly) is possible within the 9 hours inference time.\n\nYep, last 100 is the window length as mentioned in the paper and I was using a deque for that but now torch.roll might come handy! \n\nAlso, Are you simply taking the last 100 seq's / interactions or splitting let's say 1000 content_ids into 10 chunks?",
          "votes": 1
        },
        {
          "id": 1071958,
          "postDate": "2020-11-07T16:14:21.930Z",
          "content": "<p>Just one dataframe in memory for all users, no chunks. </p>\n<p>But before being able to run inference with SAINT, we need to have SAINT model working first 😏. I've implemented it but it does not give results from the paper, it's stuck at score = 0.52. It's quite similar to SAKT results from this <a href=\"https://www.kaggle.com/leadbest/sakt-self-attentive-knowledge-tracing-submitter\" target=\"_blank\">notebook </a> attempt from <a href=\"https://www.kaggle.com/leadbest\" target=\"_blank\">@leadbest</a> (score = 0.54). SAKT should also perform better (paper claims 0.76). We might be doing something wrong or convergence conditions are not met.</p>\n<p><strong>Update</strong>: I'm sure something is wrong in my implementation has all OOF probabilities are equal to 0.68.</p>",
          "rawMarkdown": "Just one dataframe in memory for all users, no chunks. \n\nBut before being able to run inference with SAINT, we need to have SAINT model working first 😏. I've implemented it but it does not give results from the paper, it's stuck at score = 0.52. It's quite similar to SAKT results from this [notebook ](https://www.kaggle.com/leadbest/sakt-self-attentive-knowledge-tracing-submitter) attempt from @leadbest (score = 0.54). SAKT should also perform better (paper claims 0.76). We might be doing something wrong or convergence conditions are not met.\n\n**Update**: I'm sure something is wrong in my implementation has all OOF probabilities are equal to 0.68.",
          "votes": 2
        },
        {
          "id": 1072123,
          "postDate": "2020-11-07T20:26:24.217Z",
          "content": "<p>Reduce the size of model and it'll converge. Use 128 or 64 dimensions for d_model. That helped me overcome this issue.</p>",
          "rawMarkdown": "Reduce the size of model and it'll converge. Use 128 or 64 dimensions for d_model. That helped me overcome this issue.",
          "votes": 2
        },
        {
          "id": 1072174,
          "postDate": "2020-11-07T21:24:51.747Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> but I think root cause is something else. I've already tried 64 and 128 and it's quite the same.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F44b3ea36ff08c2e64d2720590f85ef4c%2Fbadtrain.png?generation=1604784242641607&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thanks @abdurrafae but I think root cause is something else. I've already tried 64 and 128 and it's quite the same.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F44b3ea36ff08c2e64d2720590f85ef4c%2Fbadtrain.png?generation=1604784242641607&alt=media)",
          "votes": 1
        },
        {
          "id": 1072250,
          "postDate": "2020-11-08T01:21:33.080Z",
          "content": "<p>The plots makes it look like models output is constant.</p>",
          "rawMarkdown": "The plots makes it look like models output is constant.",
          "votes": 1
        },
        {
          "id": 1072261,
          "postDate": "2020-11-08T02:02:16.347Z",
          "content": "<p>Thank you for referring the kernel. In the paper of Pardney and Karypis, SAKT claims auc of 0.824 in average. It assumes that all the correctness of the user-problem interactions are known in advance. So, I have to employ 1's as the fake correctness values in applying SAKT to this kernel. I've got 0.5x in submission, very low compared with the validation score of 0.9x, I think there are several reasons for this. </p>",
          "rawMarkdown": "Thank you for referring the kernel. In the paper of Pardney and Karypis, SAKT claims auc of 0.824 in average. It assumes that all the correctness of the user-problem interactions are known in advance. So, I have to employ 1's as the fake correctness values in applying SAKT to this kernel. I've got 0.5x in submission, very low compared with the validation score of 0.9x, I think there are several reasons for this. ",
          "votes": 2
        },
        {
          "id": 1072537,
          "postDate": "2020-11-08T11:57:31.023Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> Yes, output of predictions was constant and loss was not deceasing but I've made some progress. Root cause seems to be related to learning rate. After fine tuning it I've got better results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F15d4d8ba9beaf66dae6cee97d1661121%2Fbettertrain.png?generation=1604836526161499&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "@adityaecdrid Yes, output of predictions was constant and loss was not deceasing but I've made some progress. Root cause seems to be related to learning rate. After fine tuning it I've got better results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F15d4d8ba9beaf66dae6cee97d1661121%2Fbettertrain.png?generation=1604836526161499&alt=media)\n",
          "votes": 1
        },
        {
          "id": 1072544,
          "postDate": "2020-11-08T12:20:35.080Z",
          "content": "<p>🎊🎊🎊🎊🎉🎉🎉🎉🎉🎉 I had ~.66-.67 when i just used encoders. [no attention mask, just padding mask if needed] Not sure whether i did the evaluation correctly though 😅</p>",
          "rawMarkdown": "🎊🎊🎊🎊🎉🎉🎉🎉🎉🎉 I had ~.66-.67 when i just used encoders. [no attention mask, just padding mask if needed] Not sure whether i did the evaluation correctly though 😅",
          "votes": 1
        },
        {
          "id": 1072644,
          "postDate": "2020-11-08T14:46:20.913Z",
          "content": "<p>Now I'm able to reach 0.74 in validation.<br>\n<a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> You do need mask for self attention otherwise your model will also learn from future (that we don't want to).</p>",
          "rawMarkdown": "Now I'm able to reach 0.74 in validation.\n@adityaecdrid You do need mask for self attention otherwise your model will also learn from future (that we don't want to).",
          "votes": 1
        },
        {
          "id": 1072662,
          "postDate": "2020-11-08T15:08:01.857Z",
          "content": "<p>Wow! That's a good score. I will relook into them. Thanks for the inspiration &amp; Good luck with the inference pipeline! Don't forget to update us!</p>",
          "rawMarkdown": "Wow! That's a good score. I will relook into them. Thanks for the inspiration & Good luck with the inference pipeline! Don't forget to update us!",
          "votes": 1
        },
        {
          "id": 1073603,
          "postDate": "2020-11-09T17:39:36.403Z",
          "content": "<p>Everything implemented but I'm not able to reach the 0.79, only 0.748 on validation (train on 50% of data). Sequence size = 100, d_model=256. Maybe more data is needed…</p>",
          "rawMarkdown": "Everything implemented but I'm not able to reach the 0.79, only 0.748 on validation (train on 50% of data). Sequence size = 100, d_model=256. Maybe more data is needed...",
          "votes": 1
        },
        {
          "id": 1073606,
          "postDate": "2020-11-09T17:41:36.880Z",
          "content": "<p>Well, how's it on the LB :) Or on a simple lgbm with feature extracted from the same?</p>",
          "rawMarkdown": "Well, how's it on the LB :) Or on a simple lgbm with feature extracted from the same?",
          "votes": 1
        },
        {
          "id": 1073617,
          "postDate": "2020-11-09T17:51:05.260Z",
          "content": "<p>Not tested on LB (yet). I'm trying to have a solid local validation first. LB becomes quickly hypnotic and I don't want to focus on it for now. Compared to basic LightGBM (locally again), LGB gives 0.77.</p>",
          "rawMarkdown": "Not tested on LB (yet). I'm trying to have a solid local validation first. LB becomes quickly hypnotic and I don't want to focus on it for now. Compared to basic LightGBM (locally again), LGB gives 0.77.",
          "votes": 1
        },
        {
          "id": 1073619,
          "postDate": "2020-11-09T17:52:14.443Z",
          "content": "<p>Wow! That's a strong single model! Looking forward to learn from the same.</p>",
          "rawMarkdown": "Wow! That's a strong single model! Looking forward to learn from the same.",
          "votes": 1
        },
        {
          "id": 1073665,
          "postDate": "2020-11-09T19:05:33.673Z",
          "content": "<p>Hello mate, I've tried to comprehend all the thread. I wanted to add that depending on how you organize your input, maybe you don't need a triangular mask, just a padding one. Maybe what made you improve is not switching to only encoder, but to have only a padding mask.</p>",
          "rawMarkdown": "Hello mate, I've tried to comprehend all the thread. I wanted to add that depending on how you organize your input, maybe you don't need a triangular mask, just a padding one. Maybe what made you improve is not switching to only encoder, but to have only a padding mask.",
          "votes": 2
        },
        {
          "id": 1073676,
          "postDate": "2020-11-09T19:27:34.073Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> I've organized data to always have full real sequences (seq_len=100) in train/valid so I didn't need padding. The drawback is that I've much less data. I'm going to add padding support tomorrow. I will post results here.</p>\n<p>Did you also implement SAINT/SAINT+? What results do you get?</p>",
          "rawMarkdown": "@claverru I've organized data to always have full real sequences (seq_len=100) in train/valid so I didn't need padding. The drawback is that I've much less data. I'm going to add padding support tomorrow. I will post results here.\n\nDid you also implement SAINT/SAINT+? What results do you get?",
          "votes": 1
        },
        {
          "id": 1073745,
          "postDate": "2020-11-09T22:06:41.930Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> I'm trying to, but with slight modifications. </p>\n<p>Imagine I have a user story like this: [A, B, C, D, E] and I have a window size of 3. Then, what I would do is [[X, X, A], [X, A, B], [A, B, C], [B, C, D], [C, D, E]]. At the same time for every sequence I shift the answers for the input, and let them as they are for the targets. So in X everything is padded; in A I have a special SOS token for the answered_correctly (the rest features as they are); in B I have the answered_correctly corresponding to A (the rest features as they are); and so on.</p>\n<p>Also, instead of predicting a whole window, I just predict the next answered_correctly (one neuron).</p>\n<p>With this set up there's no target leakage so I don't need to look ahead mask the decoder.</p>\n<p>Though I'm afraid that I'm doing something wrong, because I still get better results with only the Encoder part. </p>\n<p>Also, I have a very small model (model dimension is less than 40 and sequence lenght is less than 70 in most experiments).</p>\n<p>My results are what you can see, ~0.758 on local validation and 0.763 on LB. I think that it has to be easy to beat the 0.775 with this approach, but for some reason it isn't working that well.</p>",
          "rawMarkdown": "@mpware I'm trying to, but with slight modifications. \n\nImagine I have a user story like this: [A, B, C, D, E] and I have a window size of 3. Then, what I would do is [[X, X, A], [X, A, B], [A, B, C], [B, C, D], [C, D, E]]. At the same time for every sequence I shift the answers for the input, and let them as they are for the targets. So in X everything is padded; in A I have a special SOS token for the answered_correctly (the rest features as they are); in B I have the answered_correctly corresponding to A (the rest features as they are); and so on.\n\nAlso, instead of predicting a whole window, I just predict the next answered_correctly (one neuron).\n\nWith this set up there's no target leakage so I don't need to look ahead mask the decoder.\n\nThough I'm afraid that I'm doing something wrong, because I still get better results with only the Encoder part. \n\nAlso, I have a very small model (model dimension is less than 40 and sequence lenght is less than 70 in most experiments).\n\nMy results are what you can see, ~0.758 on local validation and 0.763 on LB. I think that it has to be easy to beat the 0.775 with this approach, but for some reason it isn't working that well.\n",
          "votes": 3
        },
        {
          "id": 1073921,
          "postDate": "2020-11-10T03:28:02.117Z",
          "content": "<blockquote>\n  <p>Also, I have a very small model (model dimension is less than 40 and sequence lenght is less than 70 in most experiments).</p>\n</blockquote>\n<p>Can you increase the d_model to 256 and re-run?</p>\n<blockquote>\n  <p>I still get better results with only the Encoder part.</p>\n</blockquote>\n<p>But when you use only encoder, then you will have predictions for each and every seq, right? So how do you evaluate your model as your encoder only o/p will be <code>[BS, SEQ_LEN, 2]</code>? Can we simply take the <code>[:,:,1]</code> from the above mentioned tensor?</p>\n<p>My current setup is like this, (quite similar to <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>'s)</p>\n<p>I split the sequences of each user to 32 and pad where it's needed using src_key_padding_mask. (it's True where's there's a padding, otherwise False). I am not using any triangular mask.</p>\n<pre><code>def pad_seq(seq: List[int], max_batch_len: int = LAST_N, pad_value: int = True) -&gt; List[int]:\n    return seq + (max_batch_len - len(seq)) * [pad_value]\n</code></pre>\n<p>So as an e.g, for user_id 115, it's like this, (a deque for auto reduce to last 100)</p>\n<p><code>{'user_id': 115, 'content_id': deque([5692, 5716, 128, 7860, 7922, 156, 51, 50, 7896, 7863, 152, 104, 108, 7900, 7901, 7971, 25, 183, 7926, 7927, 4, 7984, 45, 185, 55, 7876, 6, 172, 7898, 175, 100, 7859, 57, 7948, 151, 167, 7897, 7882, 7962, 1278, 2065, 2064, 2063, 3363, 3365, 3364, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], maxlen=100), 'answered_correctly': deque([1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 1, 1, 1, 1, 0, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 1, 0, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], maxlen=100), 'task_container_id': deque([1, 2, 0, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 40, 40, 41, 41, 41, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], maxlen=100), 'part_id': deque([5, 5, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 2, 3, 3, 3, 4, 4, 4, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], maxlen=100), 'padded': deque([False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True], maxlen=100)}</code></p>\n<p>cc <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>. </p>\n<p>Ty!</p>",
          "rawMarkdown": ">Also, I have a very small model (model dimension is less than 40 and sequence lenght is less than 70 in most experiments).\n\nCan you increase the d_model to 256 and re-run?\n\n>I still get better results with only the Encoder part.\n\nBut when you use only encoder, then you will have predictions for each and every seq, right? So how do you evaluate your model as your encoder only o/p will be `[BS, SEQ_LEN, 2]`? Can we simply take the `[:,:,1]` from the above mentioned tensor?\n\nMy current setup is like this, (quite similar to @mpware's)\n\nI split the sequences of each user to 32 and pad where it's needed using src_key_padding_mask. (it's True where's there's a padding, otherwise False). I am not using any triangular mask.\n\n```\ndef pad_seq(seq: List[int], max_batch_len: int = LAST_N, pad_value: int = True) -> List[int]:\n    return seq + (max_batch_len - len(seq)) * [pad_value]\n```\n\nSo as an e.g, for user_id 115, it's like this, (a deque for auto reduce to last 100)\n\n```{'user_id': 115, 'content_id': deque([5692, 5716, 128, 7860, 7922, 156, 51, 50, 7896, 7863, 152, 104, 108, 7900, 7901, 7971, 25, 183, 7926, 7927, 4, 7984, 45, 185, 55, 7876, 6, 172, 7898, 175, 100, 7859, 57, 7948, 151, 167, 7897, 7882, 7962, 1278, 2065, 2064, 2063, 3363, 3365, 3364, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], maxlen=100), 'answered_correctly': deque([1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 1, 1, 1, 1, 0, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 1, 0, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], maxlen=100), 'task_container_id': deque([1, 2, 0, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 40, 40, 41, 41, 41, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], maxlen=100), 'part_id': deque([5, 5, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 2, 3, 3, 3, 4, 4, 4, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], maxlen=100), 'padded': deque([False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True], maxlen=100)}```\n\ncc @yihdarshieh. \n\nTy!",
          "votes": 2
        },
        {
          "id": 1074275,
          "postDate": "2020-11-10T13:07:25.757Z",
          "content": "<p>Hello mate,</p>\n<p>My input's shape is [batch_size, sequence_length, n_features] where those features are: the shifted answered_correctly, time features, question features, etc. Then I take every [:, :, feature] to create the embeddings. </p>\n<p>My output's shape is [batch_size, 1] (as well as the target), which means that for every sequence I have one output (0-1), that I then use to compute a binary logloss. </p>\n<p>I'm afraid that I can't increase my model size due to lack of computation power. My model is converging now, do you think that increasing the model size could improve the AUC from 0.763 to a relatively higher score?</p>\n<p>Also, I'm only using from 5M to 10M rows, depending on the experiment. And still some runs take several hours to complete.</p>",
          "rawMarkdown": "Hello mate,\n\nMy input's shape is [batch_size, sequence_length, n_features] where those features are: the shifted answered_correctly, time features, question features, etc. Then I take every [:, :, feature] to create the embeddings. \n\nMy output's shape is [batch_size, 1] (as well as the target), which means that for every sequence I have one output (0-1), that I then use to compute a binary logloss. \n\nI'm afraid that I can't increase my model size due to lack of computation power. My model is converging now, do you think that increasing the model size could improve the AUC from 0.763 to a relatively higher score?\n\nAlso, I'm only using from 5M to 10M rows, depending on the experiment. And still some runs take several hours to complete.",
          "votes": 3
        },
        {
          "id": 1074496,
          "postDate": "2020-11-10T18:10:42.197Z",
          "content": "<p>Interesting, need to check why it's not the same case for me. Regarding that position encoding,  isn't it true that position encodings are independent of input length?</p>",
          "rawMarkdown": "Interesting, need to check why it's not the same case for me. Regarding that position encoding,  isn't it true that position encodings are independent of input length?",
          "votes": 1
        },
        {
          "id": 1074544,
          "postDate": "2020-11-10T19:35:27.247Z",
          "content": "<p>I've tried multiple options there: </p>\n<ul>\n<li>Learnable sequence dependant position encoding.</li>\n<li>Non learnable sequence dependant position encoding.</li>\n<li>Learnable sequence non dependant position encoding.</li>\n<li>Non learnable sequence non dependant position encoding.</li>\n</ul>\n<p>I've seen no important difference yet TBH, so I'm sitcking to non learnable non dependant since theyre cheaper and easier to set up. Just a bunch of sinusoidal vectors with shape (sequence_length, d_model).</p>",
          "rawMarkdown": "I've tried multiple options there: \n- Learnable sequence dependant position encoding.\n- Non learnable sequence dependant position encoding.\n- Learnable sequence non dependant position encoding.\n- Non learnable sequence non dependant position encoding.\n\nI've seen no important difference yet TBH, so I'm sitcking to non learnable non dependant since theyre cheaper and easier to set up. Just a bunch of sinusoidal vectors with shape (sequence_length, d_model).",
          "votes": 3
        },
        {
          "id": 1074576,
          "postDate": "2020-11-10T20:11:05.287Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> In some papers I've read about KT and Transformer, some defined position encoding as a relative position (to a fixed sequence), so always 0, 1, 2, 3 … whatever the way you pick the sequence in the full time series. Sin/cos positional encoding (as in Attention is All you need) works for me, I've not tried direct 0, 1, 2, 3 …</p>",
          "rawMarkdown": "@claverru @adityaecdrid In some papers I've read about KT and Transformer, some defined position encoding as a relative position (to a fixed sequence), so always 0, 1, 2, 3 ... whatever the way you pick the sequence in the full time series. Sin/cos positional encoding (as in Attention is All you need) works for me, I've not tried direct 0, 1, 2, 3 ...",
          "votes": 2
        },
        {
          "id": 1074583,
          "postDate": "2020-11-10T20:23:44.007Z",
          "content": "<p>In a different discussion the current winner pointed out that the position encoding are actually learnable parameters (embedding vectors whose input is an integer and output is a vector of model dimension size). </p>\n<p>And what you said is true, in all those transformer based papers the position encoding is actually fixed - not increasing. Though, my intuition tells me that it shouldn't be as good as having an actual increasing position encoding for KT. Why?</p>\n<p>Imagine I have a sentence like <em>[..] and he said he was the […]</em>. You could actually complete that sentence with a high level of precision, even without knowing the beginning of the sentence. However, do you think a user will perform equally in the 1200th question than in the 70th even if their last sequence_length interactions were the same? </p>\n<p>My theory is that including that information should be relevant, and maybe decisive for those fighting for a 0.00x extra.</p>\n<p>Any thoughts?</p>",
          "rawMarkdown": "In a different discussion the current winner pointed out that the position encoding are actually learnable parameters (embedding vectors whose input is an integer and output is a vector of model dimension size). \n\nAnd what you said is true, in all those transformer based papers the position encoding is actually fixed - not increasing. Though, my intuition tells me that it shouldn't be as good as having an actual increasing position encoding for KT. Why?\n\nImagine I have a sentence like *[..] and he said he was the [...]*. You could actually complete that sentence with a high level of precision, even without knowing the beginning of the sentence. However, do you think a user will perform equally in the 1200th question than in the 70th even if their last sequence_length interactions were the same? \n\nMy theory is that including that information should be relevant, and maybe decisive for those fighting for a 0.00x extra.\n\nAny thoughts?",
          "votes": 4
        },
        {
          "id": 1074602,
          "postDate": "2020-11-10T21:00:21.573Z",
          "content": "<p>True but you could also add such information in another category like:<br>\nNewbie = Interactions with 0-100<br>\nNovice = Interactions with 0-500<br>\nMedium = Interactions with 500-2000<br>\n…<br>\nExpert = 5000+</p>\n<p>I did not try it yet but I will.</p>",
          "rawMarkdown": "True but you could also add such information in another category like:\nNewbie = Interactions with 0-100\nNovice = Interactions with 0-500\nMedium = Interactions with 500-2000\n...\nExpert = 5000+\n\nI did not try it yet but I will.",
          "votes": 2
        },
        {
          "id": 1074647,
          "postDate": "2020-11-10T22:35:56.640Z",
          "content": "<p>Nice idea! Take into account that a user is first a newbie, then a novice… don’t leak information from the future!</p>",
          "rawMarkdown": "Nice idea! Take into account that a user is first a newbie, then a novice... don’t leak information from the future!",
          "votes": 1
        },
        {
          "id": 1075080,
          "postDate": "2020-11-11T11:28:23.277Z",
          "content": "<p>I've seen in your notebook that you're using <code>task_container_id</code>, as it's increasing monotonically (except for some <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465\" target=\"_blank\">cases</a>) then it could also act as pseudo absolute sequence.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F0d6af8db417ea5f9393c11657ab345c7%2Ftask.png?generation=1605094256718118&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "I've seen in your notebook that you're using `task_container_id`, as it's increasing monotonically (except for some [cases](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465)) then it could also act as pseudo absolute sequence.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F0d6af8db417ea5f9393c11657ab345c7%2Ftask.png?generation=1605094256718118&alt=media)"
        },
        {
          "id": 1075147,
          "postDate": "2020-11-11T12:50:28.630Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> ,</p>\n<p>My current score on LB is a transformer, but I have difficulty to reproduce the results after adding more inference code for the validation. Something is really strange there, so my current LB might be kind of random. Hope I can find where the problem is. </p>",
          "rawMarkdown": "@adityaecdrid ,\n\nMy current score on LB is a transformer, but I have difficulty to reproduce the results after adding more inference code for the validation. Something is really strange there, so my current LB might be kind of random. Hope I can find where the problem is. ",
          "votes": 2
        },
        {
          "id": 1075152,
          "postDate": "2020-11-11T12:53:52.103Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> What is your score on local validation?</p>",
          "rawMarkdown": "@yihdarshieh What is your score on local validation?"
        },
        {
          "id": 1075198,
          "postDate": "2020-11-11T13:34:05.177Z",
          "content": "<p>Currently, I have only 3 CV measured, but it is for a smaller model, and they were run before the API update:</p>\n<pre><code>CV - 0.6651957631111145\nLB - 0.668\n\nCV - 0.6859979629516602\nLB - 0.690\n\nCV - 0.6827537417411804\nLB - 0.689\n\nCV - 0.679010808467865\nLB - 0.684\n</code></pre>\n<p>As I mentioned, I am still debugging the code because I am no longer to reproduce the same results as above. But the CV is tight to LB.</p>",
          "rawMarkdown": "Currently, I have only 3 CV measured, but it is for a smaller model, and they were run before the API update:\n\n\tCV - 0.6651957631111145\n\tLB - 0.668\n\n\tCV - 0.6859979629516602\n\tLB - 0.690\n\n\tCV - 0.6827537417411804\n\tLB - 0.689\n\n\tCV - 0.679010808467865\n\tLB - 0.684\n\nAs I mentioned, I am still debugging the code because I am no longer to reproduce the same results as above. But the CV is tight to LB.\n",
          "votes": 2
        },
        {
          "id": 1075203,
          "postDate": "2020-11-11T13:38:57.010Z",
          "content": "<p>I have added task_container_id as embeddings even if in the papers they don't say anything about it because I found a certainly jump in score. I also tried to normalize it (dividing it by 9999) and feed it as continuous variable to the network but didn't seem to work well.</p>",
          "rawMarkdown": "I have added task_container_id as embeddings even if in the papers they don't say anything about it because I found a certainly jump in score. I also tried to normalize it (dividing it by 9999) and feed it as continuous variable to the network but didn't seem to work well.",
          "votes": 3
        },
        {
          "id": 1076085,
          "postDate": "2020-11-12T08:53:46.213Z",
          "content": "<p>I saw your notebook, I am planning to release mine in PyTorch soon; Here's my question, your are doing a <strong>pooling</strong> o/p, that's why you are not seeing preds for every sequence element i feel as it <strong>reduces</strong> <code>(batch, steps, features)</code> (channel_last) to <code>(batch_size, features)</code>. Pl correct me if i misunderstood something.</p>\n<pre><code>    x = tf.keras.layers.GlobalAveragePooling1D()(x)\n    x = tf.keras.layers.Dropout(0.2)(x)    \n</code></pre>",
          "rawMarkdown": "I saw your notebook, I am planning to release mine in PyTorch soon; Here's my question, your are doing a **pooling** o/p, that's why you are not seeing preds for every sequence element i feel as it **reduces** `(batch, steps, features)` (channel_last) to `(batch_size, features)`. Pl correct me if i misunderstood something.\n\n```\n    x = tf.keras.layers.GlobalAveragePooling1D()(x)\n    x = tf.keras.layers.Dropout(0.2)(x)    \n```",
          "votes": 1
        },
        {
          "id": 1076107,
          "postDate": "2020-11-12T09:12:56.603Z",
          "content": "<p>That's it. The advantage is that I don't need lookahead mask. Maybe there's any drawback but I don't know.</p>",
          "rawMarkdown": "That's it. The advantage is that I don't need lookahead mask. Maybe there's any drawback but I don't know."
        },
        {
          "id": 1076179,
          "postDate": "2020-11-12T09:58:43.537Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid/pytorch-demystifying-transformers\" target=\"_blank\">Here</a> it's as promised. Please let me know my mistakes!</p>",
          "rawMarkdown": "[Here](https://www.kaggle.com/adityaecdrid/pytorch-demystifying-transformers) it's as promised. Please let me know my mistakes!",
          "votes": 3
        },
        {
          "id": 1076623,
          "postDate": "2020-11-12T17:42:09.390Z",
          "content": "<p>I've tried to implement left padding mask (for users with less total interactions than sequence length - i.e. like in inference) to use <code>nn.Transformer</code> but I'm facing to some conditions that leads to NaN tensor (and then NaN loss) to due to both self attention mask and padding mask. Nice example is provided here:<br>\n<a href=\"https://discuss.pytorch.org/t/how-to-add-padding-mask-to-nn-transformerencoder-module/63390/7\" target=\"_blank\">https://discuss.pytorch.org/t/how-to-add-padding-mask-to-nn-transformerencoder-module/63390/7</a></p>\n<p>Root cause is mask conditions that lead to the following result in <code>nn.MultiheadAttention</code>:</p>\n<pre><code>x = torch.Tensor([[[float(\"-inf\"), float(\"-inf\"), float(\"-inf\")]]])\nsoftmax = torch.nn.Softmax(dim=-1)\nsoftmax(x)\n...\ntensor([[[   nan,    nan,    nan]]])\n</code></pre>\n<p>Some discussions about this issue:<br>\n<a href=\"https://github.com/pytorch/pytorch/issues/41508\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/41508</a><br>\n<a href=\"https://github.com/pytorch/pytorch/pull/42323\" target=\"_blank\">https://github.com/pytorch/pytorch/pull/42323</a></p>",
          "rawMarkdown": "I've tried to implement left padding mask (for users with less total interactions than sequence length - i.e. like in inference) to use `nn.Transformer` but I'm facing to some conditions that leads to NaN tensor (and then NaN loss) to due to both self attention mask and padding mask. Nice example is provided here:\nhttps://discuss.pytorch.org/t/how-to-add-padding-mask-to-nn-transformerencoder-module/63390/7\n\nRoot cause is mask conditions that lead to the following result in `nn.MultiheadAttention`:\n```\nx = torch.Tensor([[[float(\"-inf\"), float(\"-inf\"), float(\"-inf\")]]])\nsoftmax = torch.nn.Softmax(dim=-1)\nsoftmax(x)\n...\ntensor([[[   nan,    nan,    nan]]])\n```\nSome discussions about this issue:\nhttps://github.com/pytorch/pytorch/issues/41508\nhttps://github.com/pytorch/pytorch/pull/42323",
          "votes": 2
        },
        {
          "id": 1076656,
          "postDate": "2020-11-12T18:06:49.403Z",
          "content": "<p>You are masking everything in the sequence mate!</p>\n<pre><code>&gt;&gt;&gt; x = torch.Tensor([[[float(\"-inf\"), float(\"-inf\"), float(\"-inf\"), 0.5]]])\n&gt;&gt;&gt; torch.nn.Softmax(dim=-1)(x)\ntensor([[[0., 0., 0., 1.]]])\n</code></pre>",
          "rawMarkdown": "You are masking everything in the sequence mate!\n\n```\n>>> x = torch.Tensor([[[float(\"-inf\"), float(\"-inf\"), float(\"-inf\"), 0.5]]])\n>>> torch.nn.Softmax(dim=-1)(x)\ntensor([[[0., 0., 0., 1.]]])\n```",
          "votes": 1
        },
        {
          "id": 1076668,
          "postDate": "2020-11-12T18:19:20.970Z",
          "content": "<p>Not in the input sequence (I've double checked) but within Attention matrix multiplication I've one (or more) rows with full <code>-inf</code>. Well, it's not a big problem for training as we can ignore users with small interactions but I'm wondering how it will behaves for new users in inference. </p>",
          "rawMarkdown": "Not in the input sequence (I've double checked) but within Attention matrix multiplication I've one (or more) rows with full `-inf`. Well, it's not a big problem for training as we can ignore users with small interactions but I'm wondering how it will behaves for new users in inference. ",
          "votes": 1
        },
        {
          "id": 1076710,
          "postDate": "2020-11-12T19:01:51.370Z",
          "content": "<p>That means you are probably masking a full row, thus having a full minus infinite row.</p>",
          "rawMarkdown": "That means you are probably masking a full row, thus having a full minus infinite row."
        }
      ]
    },
    {
      "id": 1142976,
      "postDate": "2021-01-07T17:43:09.117Z",
      "content": "<p>Sorry, I've been quiet since a few couple of days but I've been very busy with my team to try to improve our models.<br>\nLast 5 submissions have been sent so now competition is almost completed. I would like to thanks <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> <a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> and all other contributors to this thread. It really helped us to make transformer(s) model(s) train better and work. I won't share any secret right now but I think top teams and you guys have found similar \"things\" that made transformer better and better.<br>\nI wish you the best for private LB and for year 2021!</p>",
      "rawMarkdown": "Sorry, I've been quiet since a few couple of days but I've been very busy with my team to try to improve our models.\nLast 5 submissions have been sent so now competition is almost completed. I would like to thanks @adityaecdrid @yihdarshieh @claverru @abdurrafae @jaideepvalani and all other contributors to this thread. It really helped us to make transformer(s) model(s) train better and work. I won't share any secret right now but I think top teams and you guys have found similar \"things\" that made transformer better and better.\nI wish you the best for private LB and for year 2021!",
      "votes": 6,
      "replies": [
        {
          "id": 1143552,
          "postDate": "2021-01-08T00:22:58.553Z",
          "content": "<p>Great Work Guy's! A competition to remember ❤️ and next time will work insanely harder to have that top ~1-2% finish! </p>",
          "rawMarkdown": "Great Work Guy's! A competition to remember ❤️ and next time will work insanely harder to have that top ~1-2% finish! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1144673,
      "postDate": "2021-01-08T15:50:12.887Z",
      "content": "<p>I have shared my final solution here <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209793\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209793</a>. It has some tricks that weren't discussed in this thread, implemented during the last two weeks of competition. Grew me up from 0.794 to 0.800 in public LB, and I believe it could still have gotten a better score with more time to finetune.</p>",
      "rawMarkdown": "I have shared my final solution here [https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209793](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209793). It has some tricks that weren't discussed in this thread, implemented during the last two weeks of competition. Grew me up from 0.794 to 0.800 in public LB, and I believe it could still have gotten a better score with more time to finetune.",
      "votes": 3
    },
    {
      "id": 1110427,
      "postDate": "2020-12-12T18:16:20.857Z",
      "content": "<p>Here is update after unrelenting work of 5 days 77.1 SAINT score…Now will try some other stuff </p>",
      "rawMarkdown": "Here is update after unrelenting work of 5 days 77.1 SAINT score...Now will try some other stuff ",
      "votes": 3,
      "replies": [
        {
          "id": 1115292,
          "postDate": "2020-12-16T07:00:42.037Z",
          "content": "<p>Bravo! keep going Jaideep… see if some ensembles can take it further up by 1-2%age..</p>",
          "rawMarkdown": "Bravo! keep going Jaideep... see if some ensembles can take it further up by 1-2%age..",
          "votes": 1
        },
        {
          "id": 1134417,
          "postDate": "2021-01-01T09:30:34.067Z",
          "content": "<p>Congratulations on finally solving the submission problem</p>",
          "rawMarkdown": "Congratulations on finally solving the submission problem"
        },
        {
          "id": 1134429,
          "postDate": "2021-01-01T09:45:42.753Z",
          "content": "<p>yea..I had to team up w office colleague and tried multiple experiments…Finally it boils down to a browser issue :)<br>\nbut I guess it is too late now. Dont have time to experiment anything</p>",
          "rawMarkdown": "yea..I had to team up w office colleague and tried multiple experiments...Finally it boils down to a browser issue :)\nbut I guess it is too late now. Dont have time to experiment anything"
        },
        {
          "id": 1134750,
          "postDate": "2021-01-01T14:46:09.757Z",
          "content": "<p>Just simply enjoy the competition. I used to think that transformer can only be used in the field of NLP, growth of knowledge</p>",
          "rawMarkdown": "Just simply enjoy the competition. I used to think that transformer can only be used in the field of NLP, growth of knowledge",
          "votes": 1
        }
      ]
    },
    {
      "id": 1109542,
      "postDate": "2020-12-11T19:59:14.950Z",
      "content": "<p>I didn't contribute much in this thread, rather I asked a lot of your approaches.</p>\n<p>I just published a training / validation in TensorFlow with TPU (also works with GPU), and it works also on Colab (minimal change required). No competition submission pipeline is provided - it is (a lot, really a lot) personal effort to this competition, and providing it will also be unfair to those who work hard - you definitely know it.</p>\n<p>In particular, I implemented the auto-regressive prediction - but I found it doesn't perform better than <code>pretend each question (to be predicted during inference) in a question bundle as a single question</code>. And it doesn't work with TPU (with GPU, it is OK), so I didn't continue with it. </p>\n<p>As documentation - I have to say it is far beyond enough. I still need to focus on the competition, hope you could understand. I will add more during the time.</p>\n<p>However, I am currently not able to get higher. With <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> latest sampling strategy, I tried to train a model with his size. It didn't work at the 1st try, but after playing a bit of lr, I get a CV which is as good as a larger model whose LB is about 0.778. I guess the gap between 0.778 and 0.781 could be the different model design or other factors.</p>\n<p>I am currently train a lager model with the latest sampling strategy - thanks for TPU and Colab, I have quite resource to do experiments. </p>\n<p>Here it is <a href=\"https://www.kaggle.com/yihdarshieh/tpu-track-knowledge-states-of-1m-students\" target=\"_blank\">TPU - Track knowledge states of 1M+ students</a>.</p>\n<p>Thanks for your shares of your approaches - especially to <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>, <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> and <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>, among others.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "I didn't contribute much in this thread, rather I asked a lot of your approaches.\n\nI just published a training / validation in TensorFlow with TPU (also works with GPU), and it works also on Colab (minimal change required). No competition submission pipeline is provided - it is (a lot, really a lot) personal effort to this competition, and providing it will also be unfair to those who work hard - you definitely know it.\n\nIn particular, I implemented the auto-regressive prediction - but I found it doesn't perform better than `pretend each question (to be predicted during inference) in a question bundle as a single question`. And it doesn't work with TPU (with GPU, it is OK), so I didn't continue with it. \n\nAs documentation - I have to say it is far beyond enough. I still need to focus on the competition, hope you could understand. I will add more during the time.\n\nHowever, I am currently not able to get higher. With @claverru latest sampling strategy, I tried to train a model with his size. It didn't work at the 1st try, but after playing a bit of lr, I get a CV which is as good as a larger model whose LB is about 0.778. I guess the gap between 0.778 and 0.781 could be the different model design or other factors.\n\nI am currently train a lager model with the latest sampling strategy - thanks for TPU and Colab, I have quite resource to do experiments. \n\nHere it is [TPU - Track knowledge states of 1M+ students](https://www.kaggle.com/yihdarshieh/tpu-track-knowledge-states-of-1m-students).\n\nThanks for your shares of your approaches - especially to @claverru, @adityaecdrid and @mpware, among others.\n\nGood luck!",
      "votes": 3,
      "replies": [
        {
          "id": 1109600,
          "postDate": "2020-12-11T21:34:38.613Z",
          "content": "<blockquote>\n  <p>No competition submission pipeline is provided </p>\n</blockquote>\n<p>You mean inference code, right?</p>",
          "rawMarkdown": "> No competition submission pipeline is provided \n\nYou mean inference code, right?"
        },
        {
          "id": 1109604,
          "postDate": "2020-12-11T21:49:57.237Z",
          "content": "<p>Yes, no inference code  (there is a validation code, but it use the validation dataset stored in the tf record files, so the inference code for this part is not suitable for inference for the competition API setting)</p>",
          "rawMarkdown": "Yes, no inference code  (there is a validation code, but it use the validation dataset stored in the tf record files, so the inference code for this part is not suitable for inference for the competition API setting)"
        },
        {
          "id": 1109612,
          "postDate": "2020-12-11T21:58:18.213Z",
          "content": "<p>I see. Well, I agree with that (it would be bad to mess up LB with forked-and-run notebooks). I'm trying to write inference code for the first notebook from this post (PyTorch one), but some reason it produces repeating nearly-zero values. Have you encountered such thing?</p>",
          "rawMarkdown": "I see. Well, I agree with that (it would be bad to mess up LB with forked-and-run notebooks). I'm trying to write inference code for the first notebook from this post (PyTorch one), but some reason it produces repeating nearly-zero values. Have you encountered such thing?"
        },
        {
          "id": 1109613,
          "postDate": "2020-12-11T22:01:17.937Z",
          "content": "<p>I don't use that notebook. I built my own code from the beginning - and I can say the inference code need quite efforts to make it right. You can ask the author to share some tips.</p>",
          "rawMarkdown": "I don't use that notebook. I built my own code from the beginning - and I can say the inference code need quite efforts to make it right. You can ask the author to share some tips."
        },
        {
          "id": 1109616,
          "postDate": "2020-12-11T22:04:08.313Z",
          "content": "<p>Hmm, I see. Thanks for the reply.</p>",
          "rawMarkdown": "Hmm, I see. Thanks for the reply."
        },
        {
          "id": 1109645,
          "postDate": "2020-12-11T22:55:06.940Z",
          "content": "<p>Nice contribution mate!</p>",
          "rawMarkdown": "Nice contribution mate!"
        },
        {
          "id": 1115363,
          "postDate": "2020-12-16T08:37:37.510Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>  that is great what was you LB  and CV with basic Sampling strategy . <br>\nHow diff is that from claverru .. i couldnt figure out much .</p>\n<p>my latest lb  is 77.3 vs CV of 74.81 .Trying to narrow down the gap.</p>",
          "rawMarkdown": "@yihdarshieh  that is great what was you LB  and CV with basic Sampling strategy . \nHow diff is that from claverru .. i couldnt figure out much .\n\nmy latest lb  is 77.3 vs CV of 74.81 .Trying to narrow down the gap.\n\n"
        },
        {
          "id": 1119261,
          "postDate": "2020-12-19T21:10:03.087Z",
          "content": "<p>With that kernel, the best I can get is 0.781. For further improvement (if any), it won't be publish before the competition is finished.</p>\n<p>My sampling is random selection for each - but the nb of examples (each example is a sequence) from each user are different - but predetermined.</p>\n<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> sampling uses probability distribution for sampling from different users. So we still have different no. of sequences from different users, but that number is not predetermined, and in each epoch, some users might not been seen.</p>",
          "rawMarkdown": "With that kernel, the best I can get is 0.781. For further improvement (if any), it won't be publish before the competition is finished.\n\nMy sampling is random selection for each - but the nb of examples (each example is a sequence) from each user are different - but predetermined.\n\n@claverru sampling uses probability distribution for sampling from different users. So we still have different no. of sequences from different users, but that number is not predetermined, and in each epoch, some users might not been seen.",
          "votes": 1
        },
        {
          "id": 1120048,
          "postDate": "2020-12-20T15:05:01.513Z",
          "content": "<p>Ok , my best cv so far is 77.36 ,saint .will see how much lb does it gets . I suppose it should be less than 78 only ,higher cv don't get higher lb .</p>",
          "rawMarkdown": "Ok , my best cv so far is 77.36 ,saint .will see how much lb does it gets . I suppose it should be less than 78 only ,higher cv don't get higher lb ."
        }
      ]
    },
    {
      "id": 1092406,
      "postDate": "2020-11-26T19:26:03.690Z",
      "content": "<p>Hello guys, I've been trying to follow this discussion but I've been a little busy lately. I have one remaining question that I've not seen here answered. </p>\n<p>What do you do with users with more interactions than your input length? Do you roll a window? I've tried that and noticed it was (logically) overfitting towards those interactions that appear in multiple rolls, giving me an absurd AUC (~85%). </p>\n<p>Imagine a window size of 3, having a user with:<br>\n[A, B, C, D, E, F] interactions.</p>\n<p>If I roll a window I obtain [[A, B, C], [B, C, D], [C, D, E], [D, E, F]]. As you can see, for example, the interaction C appears multiple times.</p>\n<p>I've also been looking for this in NLP literature but I can't find an answer. Would you share your approach? Or a pointer to a paper/article/post?</p>",
      "rawMarkdown": "Hello guys, I've been trying to follow this discussion but I've been a little busy lately. I have one remaining question that I've not seen here answered. \n\nWhat do you do with users with more interactions than your input length? Do you roll a window? I've tried that and noticed it was (logically) overfitting towards those interactions that appear in multiple rolls, giving me an absurd AUC (~85%). \n\nImagine a window size of 3, having a user with:\n[A, B, C, D, E, F] interactions.\n\nIf I roll a window I obtain [[A, B, C], [B, C, D], [C, D, E], [D, E, F]]. As you can see, for example, the interaction C appears multiple times.\n\nI've also been looking for this in NLP literature but I can't find an answer. Would you share your approach? Or a pointer to a paper/article/post?",
      "votes": 3,
      "replies": [
        {
          "id": 1092415,
          "postDate": "2020-11-26T19:43:15.967Z",
          "content": "<p>Not sure what other folk's are doing, but as of now, I simply split them into multiple sequences of fixed length.(total // SEQ_LEN will be the all data you get by that way + some residual). </p>\n<p>I am really interested in knowing why Transformer's work actually without knowing the user_ids. (e.g.)</p>",
          "rawMarkdown": "Not sure what other folk's are doing, but as of now, I simply split them into multiple sequences of fixed length.(total // SEQ_LEN will be the all data you get by that way + some residual). \n\nI am really interested in knowing why Transformer's work actually without knowing the user_ids. (e.g.)",
          "votes": 1
        },
        {
          "id": 1092422,
          "postDate": "2020-11-26T19:54:37.507Z",
          "content": "<p>I used to use random sample subsequence from users' history. But I found that potentially I overfit for users with much shorter sequences. For example, if a user has only 50 interaction. In each epoch, it will get the full history. While for a user having 1000 interaction, it only get one subsequence in a epoch.</p>\n<p>For your question, I would say Adiitya's approach is standard. You could probably have some overlapping thought. Like [A, B, C, D], [C, D, E, F] etc.</p>",
          "rawMarkdown": "I used to use random sample subsequence from users' history. But I found that potentially I overfit for users with much shorter sequences. For example, if a user has only 50 interaction. In each epoch, it will get the full history. While for a user having 1000 interaction, it only get one subsequence in a epoch.\n\nFor your question, I would say Adiitya's approach is standard. You could probably have some overlapping thought. Like [A, B, C, D], [C, D, E, F] etc.",
          "votes": 2
        },
        {
          "id": 1092437,
          "postDate": "2020-11-26T20:17:53.017Z",
          "content": "<p>I see, I thought about that but didn't try it. If it is working for you I guess it's cool. If I get to set it up I will come to you with news. Thanks!</p>",
          "rawMarkdown": "I see, I thought about that but didn't try it. If it is working for you I guess it's cool. If I get to set it up I will come to you with news. Thanks!"
        },
        {
          "id": 1092442,
          "postDate": "2020-11-26T20:25:45.087Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> , The notion of user is based on the user's history. The user id is not essential, it is just a name. The real content is their history. The only problem of Transformer is that we can't use the full history , due to the timing and other resource constraint. I believe if we can use the full history (even not in this competition), we will get better results</p>",
          "rawMarkdown": "@adityaecdrid , The notion of user is based on the user's history. The user id is not essential, it is just a name. The real content is their history. The only problem of Transformer is that we can't use the full history , due to the timing and other resource constraint. I believe if we can use the full history (even not in this competition), we will get better results",
          "votes": 1
        },
        {
          "id": 1092503,
          "postDate": "2020-11-26T22:34:50.447Z",
          "content": "<p>Indeed, if you could use a full user history, it would be even better than having an ID. </p>",
          "rawMarkdown": "Indeed, if you could use a full user history, it would be even better than having an ID. "
        },
        {
          "id": 1092540,
          "postDate": "2020-11-27T00:20:43.800Z",
          "content": "<p>I trained on full history actually, an hour per epoch (~58-59mins) with 256 BS and a maximum of 5-7 epochs. (No apex as of now). But guess, there are bugs here and there, so :(</p>",
          "rawMarkdown": "I trained on full history actually, an hour per epoch (~58-59mins) with 256 BS and a maximum of 5-7 epochs. (No apex as of now). But guess, there are bugs here and there, so :("
        },
        {
          "id": 1093191,
          "postDate": "2020-11-27T14:27:48.153Z",
          "content": "<p>1 hour per epoch? It's 10 minutes for me for 96M rows. Did you index your pandas dataframe by user_id? I'm using something like this:</p>\n<pre><code>def __getitem__(self, idx):\n    if torch.is_tensor(idx):\n        idx = idx.tolist()\n\n    if self.subset == \"test\":\n        # For test, we need one question only to answer (with related user)\n        row = self.df.iloc[idx:idx+1,:]\n    else:\n        # For train/valid we need series per user\n        user_id = self.dex[idx]\n        row = self.df.loc[user_id]\n\n    sample =  self.get_sample(row)\n    return sample\n</code></pre>\n<p>Without such index it's super slow even with multiple workers.</p>",
          "rawMarkdown": "1 hour per epoch? It's 10 minutes for me for 96M rows. Did you index your pandas dataframe by user_id? I'm using something like this:\n\n```\ndef __getitem__(self, idx):\n\tif torch.is_tensor(idx):\n\t\tidx = idx.tolist()\n\t\n\tif self.subset == \"test\":\n\t\t# For test, we need one question only to answer (with related user)\n\t\trow = self.df.iloc[idx:idx+1,:]\n\telse:\n\t\t# For train/valid we need series per user\n\t\tuser_id = self.dex[idx]\n\t\trow = self.df.loc[user_id]\n\n\tsample =  self.get_sample(row)\n\treturn sample\n```\nWithout such index it's super slow even with multiple workers.",
          "votes": 4
        },
        {
          "id": 1093244,
          "postDate": "2020-11-27T15:04:13.877Z",
          "content": "<p>Ahh wow, that's pretty fast And I am not indexing the data-frame. Rather doing something incorrect it seems.</p>\n<p>I am not sure why it's slower for me; </p>\n<p>Here's my approach;<br>\nI prepared dataset in <a href=\"https://www.kaggle.com/adityaecdrid/pytorch-demystifying-transformers\" target=\"_blank\">this</a> way, except that's it's for all user data and those who have more than 100 SEQ's, i simply break them up. So my training <code>len(dataset)</code> is like 11_41_813. (with a BS of 256, that's like ~4.5K batches in total)</p>\n<pre><code># this is what i do when i prepare the dataset\ngrp.agg({\n    \"content_id\":list, \"answered_correctly\":list, \"task_container_id\":list,\n    \"part_id\":list, \"prior_question_elapsed_time\":list,\n</code></pre>\n<p>Actually I am detaching the loss etc from GPU to CPU, hence it's slow i believe, will remove that part and see it as well. (it's an anti-pattern but just for debugging etc)</p>\n<p>I will look into what you have suggested! Thanks for the tips.</p>",
          "rawMarkdown": "Ahh wow, that's pretty fast And I am not indexing the data-frame. Rather doing something incorrect it seems.\n\nI am not sure why it's slower for me; \n\nHere's my approach;\nI prepared dataset in [this](https://www.kaggle.com/adityaecdrid/pytorch-demystifying-transformers) way, except that's it's for all user data and those who have more than 100 SEQ's, i simply break them up. So my training `len(dataset)` is like 11_41_813. (with a BS of 256, that's like ~4.5K batches in total)\n\n```\n# this is what i do when i prepare the dataset\ngrp.agg({\n    \"content_id\":list, \"answered_correctly\":list, \"task_container_id\":list,\n    \"part_id\":list, \"prior_question_elapsed_time\":list,\n```\n\nActually I am detaching the loss etc from GPU to CPU, hence it's slow i believe, will remove that part and see it as well. (it's an anti-pattern but just for debugging etc)\n\nI will look into what you have suggested! Thanks for the tips.",
          "replies": [
            {
              "id": 1093254,
              "postDate": "2020-11-27T15:11:55.760Z",
              "content": "<p>To speed up you can also compute metric (add loss.item()) only every 50 batch iterations on training only. What does matter is the validation.</p>",
              "rawMarkdown": "To speed up you can also compute metric (add loss.item()) only every 50 batch iterations on training only. What does matter is the validation.",
              "votes": 1
            }
          ]
        },
        {
          "id": 1093328,
          "postDate": "2020-11-27T16:07:32.353Z",
          "content": "<p>Hmm interesting pseudo-code; So if i understand it correctly, you are fetching random rows from the user whoever is at <strong>idx</strong>; And then you maintain the LAST_SEQ or something in <code>self.get_sample</code>? Thanks for your help and tips!</p>",
          "rawMarkdown": "Hmm interesting pseudo-code; So if i understand it correctly, you are fetching random rows from the user whoever is at __idx__; And then you maintain the LAST_SEQ or something in `self.get_sample`? Thanks for your help and tips!"
        },
        {
          "id": 1094103,
          "postDate": "2020-11-28T10:20:02.723Z",
          "content": "<p>Wait what? 10 minutes for 100M rows per epoch? That just killed me. </p>",
          "rawMarkdown": "Wait what? 10 minutes for 100M rows per epoch? That just killed me. ",
          "votes": 3,
          "replies": [
            {
              "id": 1094155,
              "postDate": "2020-11-28T11:10:55.883Z",
              "content": "<p>14 minutes exactly, I'm running it on Colab/GPU (P100).<br>\nScreenshot below with in my current training with MSE loss:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Faaaf4526e1ad47e94f8eae2894f3f5df%2Ftrain_epoch.png?generation=1606561845720602&amp;alt=media\" alt=\"\"></p>\n<p>One epoch covers all 375k users but not all data for all users. It picks a sequence of 100 interactions (randomly) for each. So it's not 14min for 100M rows but 14min for around 37M rows. And all rows have <code>user_id</code> as index to speed up the rows retrievial once an user_id is picked. The dataframe is already sorted by user_id and timestamp, so the dataloader just needs to pad if needed and move data to GPU (4 workers). Also my train loop drops all useless move from GPU to CPU, it's quite important.</p>\n<p>Total RAM used during training: 8GB (all data prepared before and just loaded)<br>\nGPU RAM close to limit (16GB) with batch size = 256<br>\nAll tensors are <code>long</code> tensors</p>\n<p>I plan to try to run it on Kaggle kernel soon to see how it behaves.</p>",
              "rawMarkdown": "14 minutes exactly, I'm running it on Colab/GPU (P100).\nScreenshot below with in my current training with MSE loss:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Faaaf4526e1ad47e94f8eae2894f3f5df%2Ftrain_epoch.png?generation=1606561845720602&alt=media)\n\nOne epoch covers all 375k users but not all data for all users. It picks a sequence of 100 interactions (randomly) for each. So it's not 14min for 100M rows but 14min for around 37M rows. And all rows have `user_id` as index to speed up the rows retrievial once an user_id is picked. The dataframe is already sorted by user_id and timestamp, so the dataloader just needs to pad if needed and move data to GPU (4 workers). Also my train loop drops all useless move from GPU to CPU, it's quite important.\n\nTotal RAM used during training: 8GB (all data prepared before and just loaded)\nGPU RAM close to limit (16GB) with batch size = 256\nAll tensors are `long` tensors\n\nI plan to try to run it on Kaggle kernel soon to see how it behaves.",
              "votes": 2
            },
            {
              "id": 1094169,
              "postDate": "2020-11-28T11:23:41.387Z",
              "content": "<p>Beautiful! I am looking forward to your colab notebook link once we are done! And this explains the time_difference as well.</p>\n<blockquote>\n  <p>One epoch covers all 375k users but not all data for all users. It picks a sequence of 100 interactions (randomly) for each.</p>\n</blockquote>\n<p>What i am doing is all possible sequences for any user in any epoch capped at a SEQ_LEN as one row of training data from me. But your strategy is better i believe ❤️. Can't wait to go about doing this on V100's now!</p>\n<blockquote>\n  <p>Also my train loop drops all useless move from GPU to CPU, it's quite important.</p>\n</blockquote>\n<p>Very IMP!</p>\n<p>The next thing which i want to do is train an encoder first that can predict the next content_id and then add decoder etc to it if needed.</p>\n<p>Plus you are using MaskedBCE, only an encoder?</p>",
              "rawMarkdown": "Beautiful! I am looking forward to your colab notebook link once we are done! And this explains the time_difference as well.\n\n>One epoch covers all 375k users but not all data for all users. It picks a sequence of 100 interactions (randomly) for each.\n\nWhat i am doing is all possible sequences for any user in any epoch capped at a SEQ_LEN as one row of training data from me. But your strategy is better i believe ❤️. Can't wait to go about doing this on V100's now!\n\n> Also my train loop drops all useless move from GPU to CPU, it's quite important.\n\nVery IMP!\n\nThe next thing which i want to do is train an encoder first that can predict the next content_id and then add decoder etc to it if needed.\n\nPlus you are using MaskedBCE, only an encoder?",
              "votes": 1
            },
            {
              "id": 1094180,
              "postDate": "2020-11-28T11:34:24.887Z",
              "content": "<p>MaskedBCE is the name of my custom loss to workaround the padding mask issue I've with <code>nn.Transformer</code>. It allows to ignore padded parts.</p>",
              "rawMarkdown": "MaskedBCE is the name of my custom loss to workaround the padding mask issue I've with `nn.Transformer`. It allows to ignore padded parts.",
              "votes": 1
            },
            {
              "id": 1094356,
              "postDate": "2020-11-28T14:41:24.530Z",
              "content": "<p>But I saw MSE on the screenshot?</p>",
              "rawMarkdown": "But I saw MSE on the screenshot?",
              "votes": 2
            },
            {
              "id": 1094494,
              "postDate": "2020-11-28T16:58:47.617Z",
              "content": "<p>Yes, my MaskedBCE includes either BCE or MSE (so I should have named it differently).</p>",
              "rawMarkdown": "Yes, my MaskedBCE includes either BCE or MSE (so I should have named it differently).",
              "votes": 1
            },
            {
              "id": 1094497,
              "postDate": "2020-11-28T17:02:50.440Z",
              "content": "<p>Would you mind to share your CV/LB for MSE / BCE losses? I am busy to debug and haven't been able to try extra configuration yet …</p>",
              "rawMarkdown": "Would you mind to share your CV/LB for MSE / BCE losses? I am busy to debug and haven't been able to try extra configuration yet ..."
            },
            {
              "id": 1094520,
              "postDate": "2020-11-28T17:24:41.233Z",
              "content": "<p>MSE loss should raise eyebrows..! Trying to think why it's for 😅</p>",
              "rawMarkdown": "MSE loss should raise eyebrows..! Trying to think why it's for 😅"
            },
            {
              "id": 1094537,
              "postDate": "2020-11-28T17:42:53.180Z",
              "content": "<p>Sure, for fold1: <br>\nBCE:  Loss=0.5627, AUC 0.7685, LB=0.775<br>\nMSE:  Loss=0.1907, AUC 0.7687, LB=0.768</p>",
              "rawMarkdown": "Sure, for fold1: \nBCE:  Loss=0.5627, AUC 0.7685, LB=0.775\nMSE:  Loss=0.1907, AUC 0.7687, LB=0.768",
              "votes": 3
            },
            {
              "id": 1094632,
              "postDate": "2020-11-28T19:17:06.310Z",
              "content": "<p>Your results are a bit odd, usually i get a better score with mse than with BCE (between +0.5 and +1 pts in AUC) did you think about rescaling you labels (for exemple 0 to 0 and 1 to 10)? This might help your mse loss to perform better (for values between zero and 1 the mse loss is very low as square diminish the value between zero and one)</p>",
              "rawMarkdown": "Your results are a bit odd, usually i get a better score with mse than with BCE (between +0.5 and +1 pts in AUC) did you think about rescaling you labels (for exemple 0 to 0 and 1 to 10)? This might help your mse loss to perform better (for values between zero and 1 the mse loss is very low as square diminish the value between zero and one)",
              "votes": 1
            },
            {
              "id": 1094681,
              "postDate": "2020-11-28T20:10:58.327Z",
              "content": "<p>Nope, I will try now. Thanks.</p>",
              "rawMarkdown": "Nope, I will try now. Thanks."
            },
            {
              "id": 1096094,
              "postDate": "2020-11-30T08:01:07.643Z",
              "content": "<p><a href=\"https://www.kaggle.com/rously\" target=\"_blank\">@rously</a> I've tried to scale up the labels to 0-10, moved loss to MSE, removed the sigmoid and scale down labels/predictions to 0-1 just before metrics but I get similar results. MSE loss: 19.16, AUC: 0.766</p>",
              "rawMarkdown": "@rously I've tried to scale up the labels to 0-10, moved loss to MSE, removed the sigmoid and scale down labels/predictions to 0-1 just before metrics but I get similar results. MSE loss: 19.16, AUC: 0.766"
            }
          ]
        },
        {
          "id": 1094147,
          "postDate": "2020-11-28T11:04:56.293Z",
          "content": "<p>Ya, I am also dead now 🤐🥺.</p>",
          "rawMarkdown": "Ya, I am also dead now 🤐🥺.",
          "votes": 1
        },
        {
          "id": 1094258,
          "postDate": "2020-11-28T12:51:16.983Z",
          "content": "<p>if you use the TPUs on collab you can even reduce the training time at about 1.5min per epoch, with the whole dataset seen at each epochs.</p>\n<p>Took me a bit of time to understand how to use properly TPU though<br>\n(you need to put your data on tfrecord format on a private google cloud storage and connect it to the collab env …)</p>",
          "rawMarkdown": "if you use the TPUs on collab you can even reduce the training time at about 1.5min per epoch, with the whole dataset seen at each epochs.\n\nTook me a bit of time to understand how to use properly TPU though\n(you need to put your data on tfrecord format on a private google cloud storage and connect it to the collab env ...)",
          "votes": 1
        },
        {
          "id": 1094278,
          "postDate": "2020-11-28T13:07:14.517Z",
          "content": "<p>Is that even worth it?</p>",
          "rawMarkdown": "Is that even worth it?",
          "votes": 1
        },
        {
          "id": 1094325,
          "postDate": "2020-11-28T14:10:16.013Z",
          "content": "<p>TFRecord isn't needed if you use PyTorch but it's tricky to get it working. If you are on Colab Pro and have sufficient RAM, it can be done but will take some time. Nonetheless 14-15 mins is fine for one epoch. <a href=\"https://www.kaggle.com/adityaecdrid/simple-xlmr-tpu-pytorch\" target=\"_blank\">ref</a>. </p>\n<p>Just a warning, It's not straightforward to do it and there are stability issues.</p>",
          "rawMarkdown": "TFRecord isn't needed if you use PyTorch but it's tricky to get it working. If you are on Colab Pro and have sufficient RAM, it can be done but will take some time. Nonetheless 14-15 mins is fine for one epoch. [ref](https://www.kaggle.com/adityaecdrid/simple-xlmr-tpu-pytorch). \n\nJust a warning, It's not straightforward to do it and there are stability issues."
        },
        {
          "id": 1094625,
          "postDate": "2020-11-28T19:14:22.453Z",
          "content": "<p>You can also bypass tfrecords on tensorflow if you can fit your whole dataset in ram, but my dataset processed takes more than 10go in total, so it is a bit tricky to use as is ^^</p>",
          "rawMarkdown": "You can also bypass tfrecords on tensorflow if you can fit your whole dataset in ram, but my dataset processed takes more than 10go in total, so it is a bit tricky to use as is ^^",
          "votes": 1
        },
        {
          "id": 1104665,
          "postDate": "2020-12-07T07:07:56.327Z",
          "content": "<p>'window size-1' dummy padding at beginning and end?</p>",
          "rawMarkdown": "'window size-1' dummy padding at beginning and end?"
        }
      ]
    },
    {
      "id": 1091985,
      "postDate": "2020-11-26T12:52:42.800Z",
      "content": "<p>The best I can get so far CV=0.764 and LB=0.775<br>\nIt required more memory optimizations to make it works within 3 hours.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fbcf2f62ad526a9a1f64d31409d0e37c7%2Finference3.png?generation=1606395072707510&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The best I can get so far CV=0.764 and LB=0.775\nIt required more memory optimizations to make it works within 3 hours.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fbcf2f62ad526a9a1f64d31409d0e37c7%2Finference3.png?generation=1606395072707510&alt=media)",
      "votes": 3,
      "replies": [
        {
          "id": 1092032,
          "postDate": "2020-11-26T13:42:46.757Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> , for LB, do you train on your validation dataset before submitting? Currently I can't pass 0.770, I am not sure if I should try to train on the validation dataset before submitting ….</p>",
          "rawMarkdown": "@mpware , for LB, do you train on your validation dataset before submitting? Currently I can't pass 0.770, I am not sure if I should try to train on the validation dataset before submitting ....",
          "votes": 2,
          "replies": [
            {
              "id": 1092299,
              "postDate": "2020-11-26T17:20:27.993Z",
              "content": "<p>No, just on train fold (1 fold currently). I'm going to try with another fold and ensemble both.</p>",
              "rawMarkdown": "No, just on train fold (1 fold currently). I'm going to try with another fold and ensemble both.",
              "votes": 2
            }
          ]
        },
        {
          "id": 1092041,
          "postDate": "2020-11-26T13:50:03.303Z",
          "content": "<p>my LB score can't pass 0.77 either</p>",
          "rawMarkdown": "my LB score can't pass 0.77 either",
          "votes": 1,
          "replies": [
            {
              "id": 1109396,
              "postDate": "2020-12-11T16:16:51.627Z",
              "content": "<blockquote>\n  <p>my LB score can't pass 0.77 either<br>\n  <a href=\"https://www.kaggle.com/yangxiaoshuai\" target=\"_blank\">@yangxiaoshuai</a>  how are  you passing inputs.. <br>\n  i think there are two possible approaches. <br>\n  q[1:],qa[1:] then shifting embeddings<br>\n  or q[1:],qa[:-1] to take care of shifting of embeddings.</p>\n</blockquote>",
              "rawMarkdown": "> my LB score can't pass 0.77 either\n@yangxiaoshuai  how are  you passing inputs.. \ni think there are two possible approaches. \nq[1:],qa[1:] then shifting embeddings\nor q[1:],qa[:-1] to take care of shifting of embeddings."
            }
          ]
        },
        {
          "id": 1092078,
          "postDate": "2020-11-26T14:24:06.773Z",
          "content": "<p>I cannot beat .74 🥺; there are some bugs for sure and my CV is higher than LB when i use SAINT. <br>\nAny tips on this would be helpful?</p>\n<p>There are some bugs which i am aware of and trying to find a fix for them. But why do you want to make it work in 3 hours? We have 9 hours for GPU's submission.</p>",
          "rawMarkdown": "I cannot beat .74 🥺; there are some bugs for sure and my CV is higher than LB when i use SAINT. \nAny tips on this would be helpful?\n\nThere are some bugs which i am aware of and trying to find a fix for them. But why do you want to make it work in 3 hours? We have 9 hours for GPU's submission.",
          "votes": 2,
          "replies": [
            {
              "id": 1092306,
              "postDate": "2020-11-26T17:24:54.257Z",
              "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> </p>\n<blockquote>\n  <p>But why do you want to make it work in 3 hours?</p>\n</blockquote>\n<p>Good question 👍 <br>\nIt's not a secret but I plan to ensemble it with other models so I need to make each as fast as possible. If SAINT+ takes 9 hours then I'm done with it.</p>\n<p>Currently I'm stuck with LGB models, I cannot train them anymore, too much features and not enough memory.</p>",
              "rawMarkdown": "@adityaecdrid \n\n> But why do you want to make it work in 3 hours?\n\nGood question 👍 \nIt's not a secret but I plan to ensemble it with other models so I need to make each as fast as possible. If SAINT+ takes 9 hours then I'm done with it.\n\nCurrently I'm stuck with LGB models, I cannot train them anymore, too much features and not enough memory.",
              "votes": 1
            },
            {
              "id": 1092308,
              "postDate": "2020-11-26T17:27:22.697Z",
              "content": "<p>Ya, trying to squeeze out as much as possible so that i can have 2-3 models at-least;</p>\n<blockquote>\n  <p>Currently I'm stuck with LGB models, I cannot train them anymore, too much features and not enough memory.</p>\n</blockquote>\n<p>I am just curious, have you tried only running on 10M rows with SAINT's feats into your lgbm as well?</p>\n<p>Also, try to avoid \"object\" dtype, i saw my RAM exploding because if it; (I am sure you are aware of this, just leaving what i found)<br>\nTy!</p>",
              "rawMarkdown": "Ya, trying to squeeze out as much as possible so that i can have 2-3 models at-least;\n\n>Currently I'm stuck with LGB models, I cannot train them anymore, too much features and not enough memory.\n\nI am just curious, have you tried only running on 10M rows with SAINT's feats into your lgbm as well?\n\nAlso, try to avoid \"object\" dtype, i saw my RAM exploding because if it; (I am sure you are aware of this, just leaving what i found)\nTy!",
              "votes": 1
            },
            {
              "id": 1092318,
              "postDate": "2020-11-26T17:33:31.487Z",
              "content": "<blockquote>\n  <p>I am just curious, have you tried only running on 10M rows with SAINT's feats into your lgbm as well?</p>\n</blockquote>\n<p>Nope, not yet.</p>\n<blockquote>\n  <p>Also, try to avoid \"object\" dtype, i saw my RAM exploding because if it; (I am sure you are aware of this, just leaving what i found)</p>\n</blockquote>\n<p>Yes, it's not easy to optimize everything, for some structure I had to use built-in types and for other numpy to balance memory/compute performances.</p>",
              "rawMarkdown": "> I am just curious, have you tried only running on 10M rows with SAINT's feats into your lgbm as well?\n\nNope, not yet.\n\n> Also, try to avoid \"object\" dtype, i saw my RAM exploding because if it; (I am sure you are aware of this, just leaving what i found)\n\nYes, it's not easy to optimize everything, for some structure I had to use built-in types and for other numpy to balance memory/compute performances.\n"
            },
            {
              "id": 1092357,
              "postDate": "2020-11-26T18:20:49.713Z",
              "content": "<p>Yep, it's more of a software challenge than ML :) (60-40%) (personal sentiments)</p>",
              "rawMarkdown": "Yep, it's more of a software challenge than ML :) (60-40%) (personal sentiments)",
              "votes": 1
            },
            {
              "id": 1093179,
              "postDate": "2020-11-27T14:19:01.580Z",
              "content": "<blockquote>\n  <p>Currently I'm stuck with LGB models, I cannot train them anymore, too much features and not enough memory.</p>\n</blockquote>\n<p>I cannot train on my lapi (16 gigs) for more than 8M rows now with 15 feats :(</p>",
              "rawMarkdown": ">Currently I'm stuck with LGB models, I cannot train them anymore, too much features and not enough memory.\n\nI cannot train on my lapi (16 gigs) for more than 8M rows now with 15 feats :("
            }
          ]
        },
        {
          "id": 1092671,
          "postDate": "2020-11-27T04:42:00.263Z",
          "content": "<p>Similar here. Using 1kw training samples, my transformer model achieve 0.766 auc, while with 2kw training training samples, the LB AUC is 0.769. Haven't trained models using more data yet. But curious  about how to achieve AUC more than 0.79+ as the paper stated… And as for the inference time, I found it's very unstable. My submit time for transformer models change between 1hour and 3 hours(even for the same model, the submission time is changing). Similar things happen for LGB models submission…</p>",
          "rawMarkdown": "Similar here. Using 1kw training samples, my transformer model achieve 0.766 auc, while with 2kw training training samples, the LB AUC is 0.769. Haven't trained models using more data yet. But curious  about how to achieve AUC more than 0.79+ as the paper stated... And as for the inference time, I found it's very unstable. My submit time for transformer models change between 1hour and 3 hours(even for the same model, the submission time is changing). Similar things happen for LGB models submission...",
          "votes": 1
        },
        {
          "id": 1092963,
          "postDate": "2020-11-27T10:37:53.010Z",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> Inference time is also unstable for me, for the same model/inference code, it's between 2h15 to 3h. BTW: Which loss are you using?</p>",
          "rawMarkdown": "@lihaorocky Inference time is also unstable for me, for the same model/inference code, it's between 2h15 to 3h. BTW: Which loss are you using?"
        },
        {
          "id": 1093331,
          "postDate": "2020-11-27T16:09:36.173Z",
          "content": "<p>Currently it's nn.CrossEntropyLoss, but I'm planning tuning it also the overall structure after my work with GBDT is done.</p>",
          "rawMarkdown": "Currently it's nn.CrossEntropyLoss, but I'm planning tuning it also the overall structure after my work with GBDT is done.",
          "votes": 1
        },
        {
          "id": 1093350,
          "postDate": "2020-11-27T16:28:45.457Z",
          "content": "<p>within 3 hours is so fast, use greedy decoding?<br>\nI'm struggling with inference time…</p>",
          "rawMarkdown": "within 3 hours is so fast, use greedy decoding?\nI'm struggling with inference time..."
        }
      ]
    },
    {
      "id": 1074914,
      "postDate": "2020-11-11T08:27:13.503Z",
      "content": "<p>I have opened a Notebook with a light version of a Transformer Encoder in TensorFlow. Feel free to check it out to gather some ideas. <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">NOTEBOOK</a></p>",
      "rawMarkdown": "I have opened a Notebook with a light version of a Transformer Encoder in TensorFlow. Feel free to check it out to gather some ideas. [NOTEBOOK](https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public)",
      "votes": 3
    },
    {
      "id": 1096744,
      "postDate": "2020-11-30T18:33:56.887Z",
      "content": "<p>I've forked a SAKT <a href=\"https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing\" target=\"_blank\">kernel</a> which is super fast (train + test) within 3 hours. Thanks <a href=\"https://www.kaggle.com/wangsg\" target=\"_blank\">@wangsg</a> for the baseline.</p>\n<p>It shows a simple way to index data to speed up dataloading. It uses almost all data.<br>\nI've added random sequence picking + simple train/valid split + score fixes.</p>\n<p>SAKT is not SAINT as it uses an encoder only but it's a good place to start.</p>",
      "rawMarkdown": "I've forked a SAKT [kernel](https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing) which is super fast (train + test) within 3 hours. Thanks @wangsg for the baseline.\n\nIt shows a simple way to index data to speed up dataloading. It uses almost all data.\nI've added random sequence picking + simple train/valid split + score fixes.\n\nSAKT is not SAINT as it uses an encoder only but it's a good place to start.",
      "votes": 4,
      "replies": [
        {
          "id": 1096787,
          "postDate": "2020-11-30T19:23:19.197Z",
          "content": "<p>This is beautiful! Just curious, Have you plotted attention weights as well!</p>",
          "rawMarkdown": "This is beautiful! Just curious, Have you plotted attention weights as well!"
        },
        {
          "id": 1100122,
          "postDate": "2020-12-02T21:26:11.463Z",
          "content": "<p>That SAKT kernel has been super useful, but it doesn't use close to nearly all the data. It uses one 100 max sequence from every user. But for some users that's less than 1 percent of their interactions. It ends up using only about a quarter of the potential data.</p>",
          "rawMarkdown": "That SAKT kernel has been super useful, but it doesn't use close to nearly all the data. It uses one 100 max sequence from every user. But for some users that's less than 1 percent of their interactions. It ends up using only about a quarter of the potential data.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1085471,
      "postDate": "2020-11-20T22:35:42.523Z",
      "content": "<p>Another result with masked loss for padding (as I'm not able to have padding mask working in Transformer) and fixed total questions length (thanks <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> for the good catch!)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F71b24b4a24b4e6875a95c45c28cf7a29%2Finference2.png?generation=1605911666424581&amp;alt=media\" alt=\"\"> <br>\nI think we can do better.</p>",
      "rawMarkdown": "Another result with masked loss for padding (as I'm not able to have padding mask working in Transformer) and fixed total questions length (thanks @adityaecdrid for the good catch!)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F71b24b4a24b4e6875a95c45c28cf7a29%2Finference2.png?generation=1605911666424581&alt=media) \nI think we can do better.",
      "votes": 4,
      "replies": [
        {
          "id": 1085626,
          "postDate": "2020-11-21T03:16:41.407Z",
          "content": "<p>Great Job Pal! Godspeed :)</p>\n<blockquote>\n  <p>masked loss for padding</p>\n</blockquote>\n<p>I am not sure I understood this part; Is it that you trained it just like we do it in NLP when using Bert etc with a MLM loss?</p>\n<p>Your input sequence could be something like <code>[CLS] last_5_tokens [SEP] remaining_tokens [PAD]…</code><br>\nand something similar to MLM loss on <code>remaining_tokens</code>? In that case, How are you passing other items e.g. elapsed_time etc and dealing with the fact that Bert restricts to 512 as max_seq_len?</p>",
          "rawMarkdown": "Great Job Pal! Godspeed :)\n\n>masked loss for padding\n\nI am not sure I understood this part; Is it that you trained it just like we do it in NLP when using Bert etc with a MLM loss?\n\nYour input sequence could be something like `[CLS] last_5_tokens [SEP] remaining_tokens [PAD]…`\nand something similar to MLM loss on `remaining_tokens`? In that case, How are you passing other items e.g. elapsed_time etc and dealing with the fact that Bert restricts to 512 as max_seq_len?"
        }
      ]
    },
    {
      "id": 1078159,
      "postDate": "2020-11-14T13:03:24.867Z",
      "content": "<p>Some additional results depending on train data size (sequence length = 100)</p>\n<ul>\n<li>33% of data: Local CV= 0.738</li>\n<li>50% of data: Local CV = 0.749</li>\n<li>90% of data: Local CV = 0.757</li>\n<li>95% of data: Local CV = 0.760 (update)</li>\n</ul>",
      "rawMarkdown": "Some additional results depending on train data size (sequence length = 100)\n- 33% of data: Local CV= 0.738\n- 50% of data: Local CV = 0.749\n- 90% of data: Local CV = 0.757\n- 95% of data: Local CV = 0.760 (update)",
      "votes": 4,
      "replies": [
        {
          "id": 1079062,
          "postDate": "2020-11-15T15:18:31.337Z",
          "content": "<p>Hey mate, congrats on your LB update!</p>",
          "rawMarkdown": "Hey mate, congrats on your LB update!",
          "votes": 2
        },
        {
          "id": 1079561,
          "postDate": "2020-11-16T07:59:46.240Z",
          "content": "<p>My current LB is with LGB not with transformer yet. I'm working on inference kernel with a trained transformer, I need to optimize it as it will exceed the 9h time limit. I should have first result this week.<br>\nOne difficulty is around the padding for new user.</p>",
          "rawMarkdown": "My current LB is with LGB not with transformer yet. I'm working on inference kernel with a trained transformer, I need to optimize it as it will exceed the 9h time limit. I should have first result this week.\nOne difficulty is around the padding for new user.",
          "votes": 2
        },
        {
          "id": 1079590,
          "postDate": "2020-11-16T09:13:11.703Z",
          "content": "<p>Oh! Then could you tell us when it's done?</p>",
          "rawMarkdown": "Oh! Then could you tell us when it's done?",
          "votes": 3
        },
        {
          "id": 1079595,
          "postDate": "2020-11-16T09:17:08.810Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1081047,
          "postDate": "2020-11-16T18:47:42.317Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Currently it does not fit within 9h. I need to try to optimize it.</p>\n<p><strong>Update</strong>:  I've finally reached acceptable inference time, I would like to share the pain points and some solutions:</p>\n<ul>\n<li>History: Keep last 100 interactions <strong>per user in memory</strong></li>\n<li>Build test sequence <strong>on-fly</strong> from current test_df and history for each user</li>\n<li>Forget pandas to concat or basic clean operations, move to plain <strong>list or numpy</strong>.</li>\n<li>Avoid workers &gt; 0 in Pytorch dataloader, batches are small due to API and for some reasons concurrent workers introduce large overhead.</li>\n</ul>\n<p>First test submission inference time: 5h</p>",
          "rawMarkdown": "@claverru Currently it does not fit within 9h. I need to try to optimize it.\n\n**Update**:  I've finally reached acceptable inference time, I would like to share the pain points and some solutions:\n-  History: Keep last 100 interactions **per user in memory**\n-  Build test sequence **on-fly** from current test_df and history for each user\n-  Forget pandas to concat or basic clean operations, move to plain **list or numpy**.\n- Avoid workers > 0 in Pytorch dataloader, batches are small due to API and for some reasons concurrent workers introduce large overhead.\n\nFirst test submission inference time: 5h\n",
          "votes": 5
        },
        {
          "id": 1082265,
          "postDate": "2020-11-17T18:38:54.173Z",
          "content": "<p>After all optimizations, inference time = <strong>3h</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F6febdde32456300991e2b4ac63a07c41%2Finference.png?generation=1605638316021117&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "After all optimizations, inference time = **3h**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F6febdde32456300991e2b4ac63a07c41%2Finference.png?generation=1605638316021117&alt=media)",
          "votes": 5
        },
        {
          "id": 1082268,
          "postDate": "2020-11-17T18:41:06.070Z",
          "content": "<p>Wow <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> ❤️🎊🎉; You are gonna rock the LB soon! Congratulations! Now i can go back and start working again on the transformers! As it's possible to do it 😅</p>",
          "rawMarkdown": "Wow @mpware ❤️🎊🎉; You are gonna rock the LB soon! Congratulations! Now i can go back and start working again on the transformers! As it's possible to do it 😅",
          "votes": 2
        },
        {
          "id": 1082276,
          "postDate": "2020-11-17T18:48:05.783Z",
          "content": "<p>Few doubts I have,</p>\n<ul>\n<li>History: Keep last 100 interactions per user in memory</li>\n</ul>\n<p>Did you use a deque for this ?</p>\n<ul>\n<li>Build test sequence on-fly from current test_df and history for each user</li>\n</ul>\n<p>Append operation on the deque would take care of it automatically i feel (but you have to use your own collate_fn)</p>\n<ul>\n<li>You are doing things that's mentioned in the paper more or less like triangular masking etc?</li>\n</ul>\n<p>Ty!</p>",
          "rawMarkdown": "Few doubts I have,\n\n\n- History: Keep last 100 interactions per user in memory\n\nDid you use a deque for this ?\n\n- Build test sequence on-fly from current test_df and history for each user\n\nAppend operation on the deque would take care of it automatically i feel (but you have to use your own collate_fn)\n\n- You are doing things that's mentioned in the paper more or less like triangular masking etc?\n\nTy!",
          "votes": 1
        },
        {
          "id": 1082278,
          "postDate": "2020-11-17T18:53:31.470Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> would you mind to share your model archeticture and model size? I have difficulty to go beyond 0.765 ….</p>\n<p>For example, do you use encoder-decoder? If so, how do you do inference when submitting?<br>\nBecause it seems we need to perform inference at several timestamps in each test batch, and it is kind time consuming, no?</p>",
          "rawMarkdown": "@mpware would you mind to share your model archeticture and model size? I have difficulty to go beyond 0.765 ....\n\nFor example, do you use encoder-decoder? If so, how do you do inference when submitting?\nBecause it seems we need to perform inference at several timestamps in each test batch, and it is kind time consuming, no?",
          "votes": 1
        },
        {
          "id": 1082293,
          "postDate": "2020-11-17T19:15:14.757Z",
          "content": "<p>Nice <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> mate, you really deserve. Also, I was wondering if any of you guys could help me out to understand something.</p>\n<p>I've been training with an only output per sequence, as you could have read before in this discussion. I'm trying to switch to a sequence output, having [A, B, C] with shifted answered_correctly as input and predicting answered_correctly for [A, B, C].</p>\n<p>My main concern is: what if you have the sequence [F, G, H]? Predicting the answered_correctly for F wouldn't make sense since in that sequence you dont have previous interactions.</p>\n<p>How do you train those interactions above the sequence length?</p>\n<p>EDIT: They don't say anything about this in the reference papers.</p>",
          "rawMarkdown": "Nice @mpware mate, you really deserve. Also, I was wondering if any of you guys could help me out to understand something.\n\nI've been training with an only output per sequence, as you could have read before in this discussion. I'm trying to switch to a sequence output, having [A, B, C] with shifted answered_correctly as input and predicting answered_correctly for [A, B, C].\n\nMy main concern is: what if you have the sequence [F, G, H]? Predicting the answered_correctly for F wouldn't make sense since in that sequence you dont have previous interactions.\n\nHow do you train those interactions above the sequence length?\n\nEDIT: They don't say anything about this in the reference papers.",
          "votes": 1
        },
        {
          "id": 1082297,
          "postDate": "2020-11-17T19:19:58.043Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> - not clear to me about your question. What the relationship between [A, B, C] and [F, G, H] …?</p>",
          "rawMarkdown": "@claverru - not clear to me about your question. What the relationship between [A, B, C] and [F, G, H] ...?"
        },
        {
          "id": 1082301,
          "postDate": "2020-11-17T19:23:59.020Z",
          "content": "<blockquote>\n  <p>My main concern is: what if you have the sequence [F, G, H]? Predicting the answered_correctly for F wouldn't make sense since in that sequence you dont have previous interactions.</p>\n</blockquote>\n<p>I guess, you are worrying about unseen users for which we have no history, right?<br>\nIn that case, I would default the preds to 0.67 for the first interaction. [keeping things simple, as i don't have a transformer working 😅]</p>",
          "rawMarkdown": ">My main concern is: what if you have the sequence [F, G, H]? Predicting the answered_correctly for F wouldn't make sense since in that sequence you dont have previous interactions.\n\nI guess, you are worrying about unseen users for which we have no history, right?\nIn that case, I would default the preds to 0.67 for the first interaction. [keeping things simple, as i don't have a transformer working 😅]"
        },
        {
          "id": 1082306,
          "postDate": "2020-11-17T19:30:39.180Z",
          "content": "<p>if the concern is the new user, than the model will learn the distribution of the answer correction for each question, independent of the users.</p>\n<p>Just like the prediction the first word in a corpus.</p>",
          "rawMarkdown": "if the concern is the new user, than the model will learn the distribution of the answer correction for each question, independent of the users.\n\nJust like the prediction the first word in a corpus.",
          "votes": 1
        },
        {
          "id": 1082312,
          "postDate": "2020-11-17T19:37:45.830Z",
          "content": "<p>I'll try to explain again. Imagine you have 200 interactions for a user, and a window size/sequence length of 50. If you try to input [100:150] for example, the output for the first element in the sequence makes no sense, since in that specific sequence you don't have any prior interactions, but the user actually interacted 100 times before that.</p>",
          "rawMarkdown": "I'll try to explain again. Imagine you have 200 interactions for a user, and a window size/sequence length of 50. If you try to input [100:150] for example, the output for the first element in the sequence makes no sense, since in that specific sequence you don't have any prior interactions, but the user actually interacted 100 times before that.\n\n"
        },
        {
          "id": 1082314,
          "postDate": "2020-11-17T19:41:16.443Z",
          "content": "<p>Have into account that you output the same number of elements than the input has. The last element of the output makes sense, but the rest of them don't (except for the first sequence length interactions).</p>",
          "rawMarkdown": "Have into account that you output the same number of elements than the input has. The last element of the output makes sense, but the rest of them don't (except for the first sequence length interactions)."
        },
        {
          "id": 1082322,
          "postDate": "2020-11-17T19:48:48.623Z",
          "content": "<p>This only applies for the training phase. In the inference, you just take whatever element in the output you have to. The question is: what is the utility of training N to N instead of N to 1?</p>",
          "rawMarkdown": "This only applies for the training phase. In the inference, you just take whatever element in the output you have to. The question is: what is the utility of training N to N instead of N to 1?"
        },
        {
          "id": 1082330,
          "postDate": "2020-11-17T19:55:40.287Z",
          "content": "<p>I don't have a good answer for it. If we use window size, that's it. I think it as <code>a model use the previous fixed length of history to predict things</code>.</p>\n<p>If you want, a simple add-on will be use the absolute pos embedding --&gt; say the 1000th interaction. So at least you have something global in user history added into the model.</p>",
          "rawMarkdown": "I don't have a good answer for it. If we use window size, that's it. I think it as `a model use the previous fixed length of history to predict things`.\n\nIf you want, a simple add-on will be use the absolute pos embedding --> say the 1000th interaction. So at least you have something global in user history added into the model."
        },
        {
          "id": 1082339,
          "postDate": "2020-11-17T20:09:43.127Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> No need to use dequeue, just use of built-in python list like:</p>\n<pre><code>my_list = [] # Start with empty list\nmy_list.extend([1,2,3]) # now list is [1,2,3]\nmy_list.extend([4,5,6,7,8,9,10,11]) # now list is [1,2,3,4,5,6,7,8,9,10,11]\nmy_list = my_list[-10:] # now list contains only the last 10 items: [2,3,4,5,6,7,8,9,10,11]\n</code></pre>\n<p>Maintain a list for each user and for each items (question_id, answer, elapsed time …) with the last 100 interactions and it will fit into memory and it will be much faster than <code>np.append(...)</code> or <code>pd.concat(...)</code> by an order of magnitude!</p>\n<p>I'm almost following the SAINT/SAINT+ papers. I've added PRIOR_QUESTION_HAD_EXPLANATION has a new feature and I'm using cos/sin positional encoding. My sequence length is 100. Embedding dimension is 256, so my batches are (BS, 100, 256) and output of my model is (BS, 100) followed by <code>nn.BCEWithLogitsLoss()</code> to compute loss and then <code>torch.sigmoid()</code> to get probabilities.</p>\n<p>Currently I'm not using padding mask at all, I'm padding with existing sequence from another user. For self attention, I'm using the triu mask.</p>",
          "rawMarkdown": "@adityaecdrid No need to use dequeue, just use of built-in python list like:\n\n```\nmy_list = [] # Start with empty list\nmy_list.extend([1,2,3]) # now list is [1,2,3]\nmy_list.extend([4,5,6,7,8,9,10,11]) # now list is [1,2,3,4,5,6,7,8,9,10,11]\nmy_list = my_list[-10:] # now list contains only the last 10 items: [2,3,4,5,6,7,8,9,10,11]\n```\nMaintain a list for each user and for each items (question_id, answer, elapsed time ...) with the last 100 interactions and it will fit into memory and it will be much faster than `np.append(...)` or `pd.concat(...)` by an order of magnitude!\n\nI'm almost following the SAINT/SAINT+ papers. I've added PRIOR_QUESTION_HAD_EXPLANATION has a new feature and I'm using cos/sin positional encoding. My sequence length is 100. Embedding dimension is 256, so my batches are (BS, 100, 256) and output of my model is (BS, 100) followed by `nn.BCEWithLogitsLoss()` to compute loss and then `torch.sigmoid()` to get probabilities.\n\nCurrently I'm not using padding mask at all, I'm padding with existing sequence from another user. For self attention, I'm using the triu mask.",
          "votes": 4
        },
        {
          "id": 1082348,
          "postDate": "2020-11-17T20:17:34.923Z",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> . Do you perform several inferences in each test bacth (because there are potentially multiple questions in each bundle )?</p>",
          "rawMarkdown": "Thanks, @mpware . Do you perform several inferences in each test bacth (because there are potentially multiple questions in each bundle )?",
          "votes": 2
        },
        {
          "id": 1082354,
          "postDate": "2020-11-17T20:25:45.187Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> My model encoder + decoder, configuration is:</p>\n<pre><code>class raw_conf:\n\n    mtype = \"SAINT\"\n    backbone = \"transformer\" \n\n    pad_mode = \"random\" # i.e. pad with data from another user\n    seq_len = 100\n    embedding_dim = 256 # embed_dim must be divisible by num_heads\n    exercices_id_size = 13782\n    exercices_part_size = 7\n    response_size = 2 \n    elapsed_time_size = 300 \n    lag_time_size = 720 # It was 1440 in SAINT paper.\n    explanation_size = 2\n    position_encoding_enabled = True\n\n    # Model\n    nhead = 8\n    num_encoder_layers = 4\n    num_decoder_layers = 4\n    dim_feedforward = 2048\n    dropout = 0.1\n    activation = None\n</code></pre>\n<p>For inference, assuming an API batch with 16 users and 24 questions, I build - on-fly with user's history in memory - a tensor with (24, 100) and after prediction I get (24,1) by returning only the last answer (predition). The on-fly part is tricky (pick history + padding) and must be super fast to fit the 9h time limit.</p>\n<pre><code>def per_model_predict(model, X_test, features_cols=None):\n    test_dataset = RIIIDDataset(X_test, conf, None, subset=\"test\")\n    # Data loader must be super fast, the bottleneck is here.\n    test_loader = DataLoader(test_dataset, batch_size=conf.BATCH_SIZE, shuffle=False, num_workers=0, drop_last = False, pin_memory=False)\n    # Predict is done here, no bottleneck here.\n    _, y_prob = test_loop_fn(conf, test_loader, None, model, conf.L_DEVICE, verbose=False) # (BS, 100)\n    return y_prob[:, -1]\n</code></pre>",
          "rawMarkdown": "@yihdarshieh My model encoder + decoder, configuration is:\n\n```\nclass raw_conf:\n\n    mtype = \"SAINT\"\n    backbone = \"transformer\" \n\n    pad_mode = \"random\" # i.e. pad with data from another user\n    seq_len = 100\n    embedding_dim = 256 # embed_dim must be divisible by num_heads\n    exercices_id_size = 13782\n    exercices_part_size = 7\n    response_size = 2 \n    elapsed_time_size = 300 \n    lag_time_size = 720 # It was 1440 in SAINT paper.\n    explanation_size = 2\n    position_encoding_enabled = True\n\n    # Model\n    nhead = 8\n    num_encoder_layers = 4\n    num_decoder_layers = 4\n    dim_feedforward = 2048\n    dropout = 0.1\n    activation = None\n```\n\nFor inference, assuming an API batch with 16 users and 24 questions, I build - on-fly with user's history in memory - a tensor with (24, 100) and after prediction I get (24,1) by returning only the last answer (predition). The on-fly part is tricky (pick history + padding) and must be super fast to fit the 9h time limit.\n\n```\ndef per_model_predict(model, X_test, features_cols=None):\n    test_dataset = RIIIDDataset(X_test, conf, None, subset=\"test\")\n    # Data loader must be super fast, the bottleneck is here.\n    test_loader = DataLoader(test_dataset, batch_size=conf.BATCH_SIZE, shuffle=False, num_workers=0, drop_last = False, pin_memory=False)\n    # Predict is done here, no bottleneck here.\n    _, y_prob = test_loop_fn(conf, test_loader, None, model, conf.L_DEVICE, verbose=False) # (BS, 100)\n    return y_prob[:, -1]\n```",
          "votes": 5
        },
        {
          "id": 1082359,
          "postDate": "2020-11-17T20:36:47.310Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> About:</p>\n<blockquote>\n  <p>If you try to input [100:150] for example, the output for the first element in the sequence makes no sense, since in that specific sequence you don't have any prior interactions, but the user actually interacted 100 times before that.</p>\n</blockquote>\n<p>You can ignore it in the loss by computing BCE only from 101 to 150.</p>",
          "rawMarkdown": "@claverru About:\n> If you try to input [100:150] for example, the output for the first element in the sequence makes no sense, since in that specific sequence you don't have any prior interactions, but the user actually interacted 100 times before that.\n\nYou can ignore it in the loss by computing BCE only from 101 to 150."
        },
        {
          "id": 1082361,
          "postDate": "2020-11-17T20:41:44.790Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a></p>\n<blockquote>\n  <p>Do you perform several inferences in each test batch (because there are potentially multiple questions in each bundle )?</p>\n</blockquote>\n<p>Yes. If we've 3 questions in one task container like [Q1,Q2,Q3] for one user then it becomes a tensor of (3, 100) and then 3 independant predictions for this user. All the timestamps for Q1, Q2, Q3 are the same. I don't use the prediction from Q1 to predict Q2. I only use the history for this user each time.</p>",
          "rawMarkdown": "@yihdarshieh\n\n> Do you perform several inferences in each test batch (because there are potentially multiple questions in each bundle )?\n\nYes. If we've 3 questions in one task container like [Q1,Q2,Q3] for one user then it becomes a tensor of (3, 100) and then 3 independant predictions for this user. All the timestamps for Q1, Q2, Q3 are the same. I don't use the prediction from Q1 to predict Q2. I only use the history for this user each time.",
          "votes": 1
        },
        {
          "id": 1082381,
          "postDate": "2020-11-17T21:09:06.680Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>, thank you very much, very clear explanation of the approach. Keep going!</p>",
          "rawMarkdown": "@mpware, thank you very much, very clear explanation of the approach. Keep going!",
          "votes": 1
        },
        {
          "id": 1082384,
          "postDate": "2020-11-17T21:11:55.900Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Why are you using sinusoidal position encoding? I'm noticing that my model struggles to converge when using learnable position embeddings. Did you have the same problem?</p>",
          "rawMarkdown": "@mpware Why are you using sinusoidal position encoding? I'm noticing that my model struggles to converge when using learnable position embeddings. Did you have the same problem?"
        },
        {
          "id": 1082400,
          "postDate": "2020-11-17T21:38:11.970Z",
          "content": "<p>Sinuasoidal position encoding gives better performance on my training. With \"Noam\" LR warm up and LR always below 0.0008 I don't have convergence issue.</p>",
          "rawMarkdown": "Sinuasoidal position encoding gives better performance on my training. With \"Noam\" LR warm up and LR always below 0.0008 I don't have convergence issue.",
          "votes": 4
        },
        {
          "id": 1082529,
          "postDate": "2020-11-18T01:35:06.740Z",
          "content": "<p>Thanks for the detailed responses;</p>\n<blockquote>\n  <p>I'm almost following the SAINT/SAINT+ papers. I've added PRIOR_QUESTION_HAD_EXPLANATION has a new feature and I'm using cos/sin positional encoding. My sequence length is 100. Embedding dimension is 256, so my batches are (BS, 100, 256) and output of my model is (BS, 100) followed by nn.BCEWithLogitsLoss() to compute loss and then torch.sigmoid() to get probabilities.</p>\n</blockquote>\n<p>Shouldn't the Transformer's decoder output's shape be (T, N, E) as per docs where T is the target sequence length, N is the batch size, E is the feature number (embedding_dims)?</p>\n<p>Is your target sequence length defined as 1 as on the output side, we don't need a sequence right? Sorry, but i am still very confused about this 🤐. Would appreciate is someone can help with the same.  If it's 1, then we don't need any mask's on decoder side, right? <br>\nTy! </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F15f4aeda8965f6cce896f8d988f7bf16%2FScreenshot%202020-11-18%20at%207.04.17%20AM.png?generation=1605663277214065&amp;alt=media\" alt=\"Decoder_SAINT+\"></p>\n<blockquote>\n  <p>Sinuasoidal position encoding gives better performance on my training.</p>\n</blockquote>\n<p>This might be better actually as it's position agnostic and can easily extrapolate to sequences longer than our training sequences len.</p>",
          "rawMarkdown": "Thanks for the detailed responses;\n\n>I'm almost following the SAINT/SAINT+ papers. I've added PRIOR_QUESTION_HAD_EXPLANATION has a new feature and I'm using cos/sin positional encoding. My sequence length is 100. Embedding dimension is 256, so my batches are (BS, 100, 256) and output of my model is (BS, 100) followed by nn.BCEWithLogitsLoss() to compute loss and then torch.sigmoid() to get probabilities.\n\nShouldn't the Transformer's decoder output's shape be (T, N, E) as per docs where T is the target sequence length, N is the batch size, E is the feature number (embedding_dims)?\n\nIs your target sequence length defined as 1 as on the output side, we don't need a sequence right? Sorry, but i am still very confused about this 🤐. Would appreciate is someone can help with the same.  If it's 1, then we don't need any mask's on decoder side, right? \nTy! \n\n![Decoder_SAINT+](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F15f4aeda8965f6cce896f8d988f7bf16%2FScreenshot%202020-11-18%20at%207.04.17%20AM.png?generation=1605663277214065&alt=media)\n\n> Sinuasoidal position encoding gives better performance on my training.\n\nThis might be better actually as it's position agnostic and can easily extrapolate to sequences longer than our training sequences len."
        },
        {
          "id": 1082800,
          "postDate": "2020-11-18T08:33:28.613Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> About:</p>\n<blockquote>\n  <p>Shouldn't the Transformer's decoder output's shape be (T, N, E) </p>\n</blockquote>\n<p>Yes, it is so you have to make sure you're passing the correct input and getting the correct output. This is what I did with <code>transpose</code> before/after using <code>nn.Transformer</code>:</p>\n<pre><code>x_position_exercices = x_position_exercices.transpose(1,0) # (seq_len, BS, embedding_dim)\nx_position_responses = x_position_responses.transpose(1,0) # (seq_len, BS, embedding_dim)\nx_transformer = self.transformer(src=x_position_exercices, tgt=x_position_responses, src_mask=src_mask, tgt_mask=tgt_mask, memory_mask=mem_mask) # (seq_len, BS, embedding_dim)\nx_transformer = x_transformer.transpose(1,0) # (BS, seq_len, embedding_dim)\n</code></pre>\n<p>My output is (BS, 100), the model tries to predict each answer of the 100, not only the last one. I think predicting only the last one will also work (as it's only what we need for inference). I'm going to try it too.</p>",
          "rawMarkdown": "@adityaecdrid About:\n\n> Shouldn't the Transformer's decoder output's shape be (T, N, E) \n\nYes, it is so you have to make sure you're passing the correct input and getting the correct output. This is what I did with `transpose ` before/after using `nn.Transformer`:\n\n```\nx_position_exercices = x_position_exercices.transpose(1,0) # (seq_len, BS, embedding_dim)\nx_position_responses = x_position_responses.transpose(1,0) # (seq_len, BS, embedding_dim)\nx_transformer = self.transformer(src=x_position_exercices, tgt=x_position_responses, src_mask=src_mask, tgt_mask=tgt_mask, memory_mask=mem_mask) # (seq_len, BS, embedding_dim)\nx_transformer = x_transformer.transpose(1,0) # (BS, seq_len, embedding_dim)\n```\n\nMy output is (BS, 100), the model tries to predict each answer of the 100, not only the last one. I think predicting only the last one will also work (as it's only what we need for inference). I'm going to try it too.",
          "votes": 2
        },
        {
          "id": 1082953,
          "postDate": "2020-11-18T12:30:47.747Z",
          "content": "<p>Thanks a lot! It's quite helpful; Keep galloping on the LB :)</p>\n<p>Edit -&gt; You are using mem_mask as well, nice; Need to check that as well 😅</p>",
          "rawMarkdown": "Thanks a lot! It's quite helpful; Keep galloping on the LB :)\n\nEdit -> You are using mem_mask as well, nice; Need to check that as well 😅",
          "replies": [
            {
              "id": 1082968,
              "postDate": "2020-11-18T12:58:17.367Z",
              "content": "<p>And the more data you include, the more iterations you've per epoch and the less epochs you need. For example:</p>\n<pre><code>Fold 1 train users: 289764 valid users: 12452\nFold 1 train size: (86298794, 6) valid size: (863478, 6)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Ff1950fc7b58b4e1e56b64091a29cd8ce%2Ftrain_0.758.png?generation=1605704249682412&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "And the more data you include, the more iterations you've per epoch and the less epochs you need. For example:\n\n```\nFold 1 train users: 289764 valid users: 12452\nFold 1 train size: (86298794, 6) valid size: (863478, 6)\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Ff1950fc7b58b4e1e56b64091a29cd8ce%2Ftrain_0.758.png?generation=1605704249682412&alt=media)",
              "votes": 2
            },
            {
              "id": 1082969,
              "postDate": "2020-11-18T13:00:43.143Z",
              "content": "<p>Looking forward for a fork-able version of the same! Hehe; You can actually create more data this way by swapping content_ids in and out from different user attempts as it doesn't depends on the user. Not sure, maybe you can try to \"fake\" the data but it might impact the performance.</p>",
              "rawMarkdown": "Looking forward for a fork-able version of the same! Hehe; You can actually create more data this way by swapping content_ids in and out from different user attempts as it doesn't depends on the user. Not sure, maybe you can try to \"fake\" the data but it might impact the performance.",
              "votes": 1
            },
            {
              "id": 1082978,
              "postDate": "2020-11-18T13:06:46.640Z",
              "content": "<p>I'm not sure to publish my kernel now as it provides score in medals zone. Moreover, I'm using around 90% of data so It won't load in Kaggle.</p>",
              "rawMarkdown": "I'm not sure to publish my kernel now as it provides score in medals zone. Moreover, I'm using around 90% of data so It won't load in Kaggle.",
              "votes": 2
            },
            {
              "id": 1082983,
              "postDate": "2020-11-18T13:11:51.887Z",
              "content": "<p>I completely agree, (was just kidding) you shouldn't do it. You have already provided enough information for everyone to implement their own now; GL!</p>",
              "rawMarkdown": "I completely agree, (was just kidding) you shouldn't do it. You have already provided enough information for everyone to implement their own now; GL!",
              "votes": 1
            },
            {
              "id": 1082994,
              "postDate": "2020-11-18T13:17:51.473Z",
              "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>  <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> What is mem_mask? I never heard about it for transformer.</p>",
              "rawMarkdown": "@adityaecdrid  @mpware What is mem_mask? I never heard about it for transformer."
            },
            {
              "id": 1082999,
              "postDate": "2020-11-18T13:23:35.193Z",
              "content": "<p>It's an attention mask for output of the decoder:<br>\n<a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html</a><br>\n<a href=\"https://discuss.pytorch.org/t/memory-mask-in-nn-transformer/55230\" target=\"_blank\">https://discuss.pytorch.org/t/memory-mask-in-nn-transformer/55230</a></p>",
              "rawMarkdown": "It's an attention mask for output of the decoder:\nhttps://pytorch.org/docs/stable/generated/torch.nn.Transformer.html\nhttps://discuss.pytorch.org/t/memory-mask-in-nn-transformer/55230\n",
              "votes": 2
            },
            {
              "id": 1083387,
              "postDate": "2020-11-18T23:17:43.277Z",
              "content": "<p>Thanks, I don't use pytorch, and i just call the mem mask as decoder to encoder attention mask😄</p>",
              "rawMarkdown": " Thanks, I don't use pytorch, and i just call the mem mask as decoder to encoder attention mask😄"
            }
          ]
        },
        {
          "id": 1083372,
          "postDate": "2020-11-18T22:41:23.803Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1083664,
          "postDate": "2020-11-19T08:20:16.460Z",
          "content": "<blockquote>\n  <p>exercices_id_size = 13782</p>\n</blockquote>\n<p>I am just trying to combine so many tips here and there for the NN's; <br>\nIt seems that you are training on lecture videos as well from the config's?</p>\n<p>Secondly, are you training it in an auto-regressive fashion? I don't think we are as we aren't generating anything here.</p>\n<p>From the Attention Is All You Need Paper,</p>\n<blockquote>\n  <p>At each step the model is auto-regressive[10], consuming the previously generated symbols as additional input when generating the next.</p>\n</blockquote>",
          "rawMarkdown": ">    exercices_id_size = 13782\n\nI am just trying to combine so many tips here and there for the NN's; \nIt seems that you are training on lecture videos as well from the config's?\n\nSecondly, are you training it in an auto-regressive fashion? I don't think we are as we aren't generating anything here.\n\nFrom the Attention Is All You Need Paper,\n>At each step the model is auto-regressive[10], consuming the previously generated symbols as additional input when generating the next.",
          "votes": 1,
          "replies": [
            {
              "id": 1083680,
              "postDate": "2020-11-19T08:32:07.520Z",
              "content": "<p>Yes, you're correct there are only 13523 questions and I planned to use lectures later (but I did not yet). Good catch! So it should be:</p>\n<p><code>exercices_id_size = 13523 # 13782</code></p>",
              "rawMarkdown": "Yes, you're correct there are only 13523 questions and I planned to use lectures later (but I did not yet). Good catch! So it should be:\n\n`exercices_id_size = 13523 # 13782`\n",
              "votes": 2
            },
            {
              "id": 1083686,
              "postDate": "2020-11-19T08:38:44.550Z",
              "content": "<p>One more I have for you, there's a start token for everything on the decoder side, right? [Cs, Ps, ETs, LTs]</p>",
              "rawMarkdown": "One more I have for you, there's a start token for everything on the decoder side, right? [Cs, Ps, ETs, LTs]"
            },
            {
              "id": 1083694,
              "postDate": "2020-11-19T08:43:33.757Z",
              "content": "<p>Yes on all, I'm not sure of the best way to do it, so I did the following:</p>\n<pre><code># Add start token to correctness\nx_correctness = torch.roll(x_correctness, shifts=(0, 1, 0), dims=(0, 1, 0)) # Shift right the sequence\nx_correctness[:,0,:] = self.response_size # Start token\n</code></pre>\n<p>People with NL skills should know the best way, maybe directly embedding with padding index parameter=0.</p>",
              "rawMarkdown": "Yes on all, I'm not sure of the best way to do it, so I did the following:\n\n```\n# Add start token to correctness\nx_correctness = torch.roll(x_correctness, shifts=(0, 1, 0), dims=(0, 1, 0)) # Shift right the sequence\nx_correctness[:,0,:] = self.response_size # Start token\n```\n\nPeople with NL skills should know the best way, maybe directly embedding with padding index parameter=0.\n\n",
              "votes": 2
            },
            {
              "id": 1083697,
              "postDate": "2020-11-19T08:54:57.330Z",
              "content": "<p>Hi, do you add start token to the truncated sequence, or to the full history sequence (up to the current time) before the truncation? I mean, we might have [  Q100, Q101, ….] and [R100, R101] …., and if you add start token at this truncated sequence, we will have [START, R100, …]. Not sure if this is the best, but I don't think it doesn't hurt much.</p>",
              "rawMarkdown": "Hi, do you add start token to the truncated sequence, or to the full history sequence (up to the current time) before the truncation? I mean, we might have [  Q100, Q101, ....] and [R100, R101] ...., and if you add start token at this truncated sequence, we will have [START, R100, ...]. Not sure if this is the best, but I don't think it doesn't hurt much."
            },
            {
              "id": 1083781,
              "postDate": "2020-11-19T11:28:44.163Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 1085711,
          "postDate": "2020-11-21T06:19:47.493Z",
          "content": "<p><code>For inference, assuming an API batch with 16 users and 24 questions, I build - on-fly with user's history in memory - a tensor with (24, 100) and after prediction I get (24,1) by returning only the last answer (predition). The on-fly part is tricky (pick history + padding) and must be super fast to fit the 9h time limit.</code></p>\n<p>Can you share, how much it took for the sample example_test.csv? It's taking ~3 secs for all 4 iters of the sample test_set but it's not going to pass through that time limit of 9 hours as it has to be less than 0.55/iter but it's .75/iter for me.</p>",
          "rawMarkdown": "`For inference, assuming an API batch with 16 users and 24 questions, I build - on-fly with user's history in memory - a tensor with (24, 100) and after prediction I get (24,1) by returning only the last answer (predition). The on-fly part is tricky (pick history + padding) and must be super fast to fit the 9h time limit.`\n\nCan you share, how much it took for the sample example_test.csv? It's taking ~3 secs for all 4 iters of the sample test_set but it's not going to pass through that time limit of 9 hours as it has to be less than 0.55/iter but it's .75/iter for me."
        },
        {
          "id": 1085791,
          "postDate": "2020-11-21T07:43:08.203Z",
          "content": "<blockquote>\n  <p>For inference, assuming an API batch with 16 users and 24 questions, I build - on-fly with user's history in memory - a tensor with (24, 100) and after prediction I get (24,1) by returning only the last answer (predition). The on-fly part is tricky (pick history + padding) and must be super fast to fit the 9h time limit.</p>\n</blockquote>\n<p>I guess my init draft of inferencing with fixed total questions length is working (atleast on the dummy data),  <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> If i understand correctly, wouldn't this mean that you will be returning the pred for padded_token as well sometimes if the seq_len isn't equal to the max_value?</p>",
          "rawMarkdown": ">For inference, assuming an API batch with 16 users and 24 questions, I build - on-fly with user's history in memory - a tensor with (24, 100) and after prediction I get (24,1) by returning only the last answer (predition). The on-fly part is tricky (pick history + padding) and must be super fast to fit the 9h time limit.\n\nI guess my init draft of inferencing with fixed total questions length is working (atleast on the dummy data),  @mpware If i understand correctly, wouldn't this mean that you will be returning the pred for padded_token as well sometimes if the seq_len isn't equal to the max_value?"
        },
        {
          "id": 1085914,
          "postDate": "2020-11-21T09:20:02.067Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> would you mind to share how you calculate lag time and how to use it. As mentioned in <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197528\" target=\"_blank\">this thread</a>, for the decoder, suppose we have Qn and R(n-1) and we want to predict R(n), we use the elapsed time for answering Q(n-1), that is the <code>priior_question_elapsed_time</code> given at R(n).</p>\n<p>For lag time, I think it would be similar: We need to use the lag time between the ending of answering Q(n-2) and the stating time for answering Q(n-1). But I have some trouble to calculate this using tensors. (otherwise I have to build this information directly in the datasets and convert them to tensors later …)</p>",
          "rawMarkdown": "@mpware would you mind to share how you calculate lag time and how to use it. As mentioned in [this thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197528), for the decoder, suppose we have Qn and R(n-1) and we want to predict R(n), we use the elapsed time for answering Q(n-1), that is the `priior_question_elapsed_time` given at R(n).\n\nFor lag time, I think it would be similar: We need to use the lag time between the ending of answering Q(n-2) and the stating time for answering Q(n-1). But I have some trouble to calculate this using tensors. (otherwise I have to build this information directly in the datasets and convert them to tensors later ...)",
          "votes": 1,
          "replies": [
            {
              "id": 1085933,
              "postDate": "2020-11-21T09:35:29.047Z",
              "content": "<p>Start with something simple, timestamp(N) - timestamp(N-1), even if not the exact definition it' s a good estimation.</p>",
              "rawMarkdown": "Start with something simple, timestamp(N) - timestamp(N-1), even if not the exact definition it' s a good estimation.",
              "votes": 3
            },
            {
              "id": 1085959,
              "postDate": "2020-11-21T09:54:09.127Z",
              "content": "<p>Yes, I might over-complicates a lot of things :) I will give it a try first. Thanks</p>",
              "rawMarkdown": "Yes, I might over-complicates a lot of things :) I will give it a try first. Thanks"
            },
            {
              "id": 1085962,
              "postDate": "2020-11-21T09:56:24.717Z",
              "content": "<p>If we do this, we have to take care of some outliers as well i believe. difference b/w my last logout today and my next login time can bee quite large.</p>",
              "rawMarkdown": "If we do this, we have to take care of some outliers as well i believe. difference b/w my last logout today and my next login time can bee quite large."
            },
            {
              "id": 1085988,
              "postDate": "2020-11-21T10:11:13.800Z",
              "content": "<p>they use a maximal lag time of 300 seconds and categorical features for it. So it shouldn't be a real problem for large lag time (will be truncated anyway)</p>",
              "rawMarkdown": "they use a maximal lag time of 300 seconds and categorical features for it. So it shouldn't be a real problem for large lag time (will be truncated anyway)",
              "votes": 1
            },
            {
              "id": 1128442,
              "postDate": "2020-12-27T12:51:21.317Z",
              "content": "<blockquote>\n  <p>they use a maximal lag time of 300 seconds and categorical features for it. So it shouldn't be a real problem for large lag time (will be truncated anyway)</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> <br>\n1)to what max time lag time can be truncated is it 1400 minutes.. or 1400*60 sec?<br>\n2) by categorical embedding means each lag time is given a time bucket ? if it falls in some range </p>",
              "rawMarkdown": "> they use a maximal lag time of 300 seconds and categorical features for it. So it shouldn't be a real problem for large lag time (will be truncated anyway)\n\n@yihdarshieh \n1)to what max time lag time can be truncated is it 1400 minutes.. or 1400*60 sec?\n2) by categorical embedding means each lag time is given a time bucket ? if it falls in some range "
            }
          ]
        },
        {
          "id": 1108534,
          "postDate": "2020-12-10T19:04:38.120Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <br>\nhow do you pass memory mask… i currently do simply this,where future mask is simple np.triu<br>\n<code>att_mask = future_mask(x.size(0)).to(device)</code><br>\n<code>x1=self.transformer_decoder(response,att_output,tgt_mask=resp_attn,memory_mask=att_mask)</code></p>",
          "rawMarkdown": "@mpware \nhow do you pass memory mask... i currently do simply this,where future mask is simple np.triu\n`att_mask = future_mask(x.size(0)).to(device)`\n`x1=self.transformer_decoder(response,att_output,tgt_mask=resp_attn,memory_mask=att_mask)`",
          "replies": [
            {
              "id": 1108541,
              "postDate": "2020-12-10T19:14:16.387Z",
              "content": "<p>I'm using <code>nn.Transformer</code> directly. There is a parameter dedicated for mem_mask.<br>\n<a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html</a></p>",
              "rawMarkdown": "I'm using ` nn.Transformer` directly. There is a parameter dedicated for mem_mask.\nhttps://pytorch.org/docs/stable/generated/torch.nn.Transformer.html\n"
            }
          ]
        },
        {
          "id": 1108675,
          "postDate": "2020-12-10T22:24:29.047Z",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> :</p>\n<p>Finally, there is minimal change in my code to having your sampling stragegy, something like</p>\n<pre><code>            ds = ds.flat_map(\n                lambda x: tf.data.Dataset.from_tensors(x).repeat(\n                    tf.cast(tf.random.uniform(shape=[]) &lt; self.training_sample_prob_table.lookup(x['user_id']), tf.int64)\n                )\n            )\n</code></pre>\n<p>Just repeat the (full) sequences of each user, depending on a probably pre-computed in a table like</p>\n<pre><code>        initializer = tf.lookup.KeyValueTensorInitializer(user_id_tensor, sample_prob_tensor)\n        training_sample_prob_table = tf.lookup.StaticHashTable(initializer, default_value=0.0)\n        self.training_sample_prob_table = training_sample_prob_table\n</code></pre>\n<p>I don't know yet how much it helps for me, because I am still not able to get 0.78 with your model size. I don't know what causes it:</p>\n<pre><code>- I use position embedding (learnable) for pos in [0, 1, ... window_size - 1]. I guess you use `sinusoidal` embedding?\n\n- You use softmax and I use sigmoid.\n\n- You use lag time, I don't\n</code></pre>",
          "rawMarkdown": "@claverru :\n\nFinally, there is minimal change in my code to having your sampling stragegy, something like\n\n```\n            ds = ds.flat_map(\n                lambda x: tf.data.Dataset.from_tensors(x).repeat(\n                    tf.cast(tf.random.uniform(shape=[]) < self.training_sample_prob_table.lookup(x['user_id']), tf.int64)\n                )\n            )\n```\nJust repeat the (full) sequences of each user, depending on a probably pre-computed in a table like\n\n```\n        initializer = tf.lookup.KeyValueTensorInitializer(user_id_tensor, sample_prob_tensor)\n        training_sample_prob_table = tf.lookup.StaticHashTable(initializer, default_value=0.0)\n        self.training_sample_prob_table = training_sample_prob_table\n```\n\nI don't know yet how much it helps for me, because I am still not able to get 0.78 with your model size. I don't know what causes it:\n\n    - I use position embedding (learnable) for pos in [0, 1, ... window_size - 1]. I guess you use `sinusoidal` embedding?\n\n    - You use softmax and I use sigmoid.\n\n    - You use lag time, I don't"
        },
        {
          "id": 1108896,
          "postDate": "2020-12-11T05:50:29.493Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>  </p>\n<pre><code>src – the sequence to the encoder (required).\n\ntgt – the sequence to the decoder (required).\n\nsrc_mask – the additive mask for the src sequence (optional).\n\ntgt_mask – the additive mask for the tgt sequence (optional).\n\nmemory_mask – the additive mask for the encoder output (optional).\n</code></pre>\n<p>what i meant was for every mask  input we can pass this way ?</p>\n<pre><code>src_mask = future_mask(x.size(0)).to(device) #using triu\ntgt_mask=  future_mask(response.size(0)).to(device) #using triu\nmemory_mask =future_mask(x.size(0)).to(device) #using triu\n</code></pre>",
          "rawMarkdown": "@mpware  \n```\nsrc – the sequence to the encoder (required).\n\ntgt – the sequence to the decoder (required).\n\nsrc_mask – the additive mask for the src sequence (optional).\n\ntgt_mask – the additive mask for the tgt sequence (optional).\n\nmemory_mask – the additive mask for the encoder output (optional).\n```\n\nwhat i meant was for every mask  input we can pass this way ?\n\n```\nsrc_mask = future_mask(x.size(0)).to(device) #using triu\ntgt_mask=  future_mask(response.size(0)).to(device) #using triu\nmemory_mask =future_mask(x.size(0)).to(device) #using triu\n\n```\n\n \n "
        }
      ]
    },
    {
      "id": 1134850,
      "postDate": "2021-01-01T16:24:31.470Z",
      "content": "<p>I'm in the mix with a SAINT implementation! 0.778 LB with just content_id so far. Super stoked to improve on that. </p>\n<p>I do have a question related to the paper though, if anyone cares to enlighten me. In the SAINT, paper they state that the decoder takes in a sequential input R of response embeddings with start token embedding S. As of now, my implementation doesn't have a unique starting token embedding. It's just padded with zeros. Is this much of an issue? Can anyone share how they're dealing with this start token embedding? I'm not sure I understand that part.</p>",
      "rawMarkdown": "I'm in the mix with a SAINT implementation! 0.778 LB with just content_id so far. Super stoked to improve on that. \n\nI do have a question related to the paper though, if anyone cares to enlighten me. In the SAINT, paper they state that the decoder takes in a sequential input R of response embeddings with start token embedding S. As of now, my implementation doesn't have a unique starting token embedding. It's just padded with zeros. Is this much of an issue? Can anyone share how they're dealing with this start token embedding? I'm not sure I understand that part.",
      "votes": 1,
      "replies": [
        {
          "id": 1134882,
          "postDate": "2021-01-01T17:09:09.987Z",
          "content": "<p>Congrats!</p>\n<p>I haven't heard the term 'stoked' since I left Northern California :-).</p>\n<p>Regarding the starter tokens on the decoder side—imo, <strong>none</strong> of them are necessary except, perhaps, on the 'correctness embedding'. I can't say more than that though =] we're getting near the finish line.</p>",
          "rawMarkdown": "Congrats!\n\nI haven't heard the term 'stoked' since I left Northern California :-).\n\nRegarding the starter tokens on the decoder side—imo, **none** of them are necessary except, perhaps, on the 'correctness embedding'. I can't say more than that though =] we're getting near the finish line.",
          "votes": 1
        },
        {
          "id": 1134969,
          "postDate": "2021-01-01T18:46:28.487Z",
          "content": "<p>I incremented by 1 on the answered_correctly field and use (0 for padding, 1 for answered incorrectly, 2 for correct, 3 for start token)</p>\n<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> Good point around not needing the starter tokens, I had overlooked that point thanks :)</p>",
          "rawMarkdown": "I incremented by 1 on the answered_correctly field and use (0 for padding, 1 for answered incorrectly, 2 for correct, 3 for start token)\n\n@authman Good point around not needing the starter tokens, I had overlooked that point thanks :)"
        }
      ]
    },
    {
      "id": 1107907,
      "postDate": "2020-12-10T04:30:34.847Z",
      "content": "<p>Alright, so here's one of my concern which i am not able to help find answer to, Let's assume your batch looks like, (0 represents the padding token) (assume that next question you are going to attempt is the last non_zero value) (6,8,12,19,25,25)</p>\n<pre><code>tensor([[ 1,  2,  3,  4,  5,  6],\n        [ 7,  8,  0,  0,  0,  0],\n        [10, 11, 12,  0,  0,  0],\n        [14, 15, 16, 17, 18, 19],\n        [20, 21, 22, 23, 24, 25],\n        20, 21, 22, 23, 24, 25]])\n</code></pre>\n<p>So in this case, the src_mask i.e. the self_attention mask in the encoder side, that shouldn't simply be a upper-traingular matrix, right?</p>\n<p>Like if we give mask like the below, then that's incorrect, right?</p>\n<pre><code>tensor([[0., -inf, -inf, -inf, -inf, -inf],\n        [0., 0., -inf, -inf, -inf, -inf],\n        [0., 0., 0., -inf, -inf, -inf],\n        [0., 0., 0., 0., -inf, -inf],\n        [0., 0., 0., 0., 0., -inf],\n        [0., 0., 0., 0., 0., 0.]])\n</code></pre>\n<p>And it should actually be like this, right? (basically i would call it a length mask 😅)</p>\n<pre><code>tensor([[0., 0., 0., 0., 0., -inf],\n        [0., -inf, -inf, -inf, -inf, -inf],\n        [0., 0., -inf, -inf, -inf, -inf],\n        [0., 0., 0., 0., 0., -inf],\n        [0., 0., 0., 0., 0., -inf],\n        [0., 0., 0., 0., 0., -inf]])\n</code></pre>\n<p>Is my understanding apt?</p>",
      "rawMarkdown": "Alright, so here's one of my concern which i am not able to help find answer to, Let's assume your batch looks like, (0 represents the padding token) (assume that next question you are going to attempt is the last non_zero value) (6,8,12,19,25,25)\n\n```\ntensor([[ 1,  2,  3,  4,  5,  6],\n        [ 7,  8,  0,  0,  0,  0],\n        [10, 11, 12,  0,  0,  0],\n        [14, 15, 16, 17, 18, 19],\n        [20, 21, 22, 23, 24, 25],\n        20, 21, 22, 23, 24, 25]])\n```\nSo in this case, the src_mask i.e. the self_attention mask in the encoder side, that shouldn't simply be a upper-traingular matrix, right?\n\nLike if we give mask like the below, then that's incorrect, right?\n```\ntensor([[0., -inf, -inf, -inf, -inf, -inf],\n        [0., 0., -inf, -inf, -inf, -inf],\n        [0., 0., 0., -inf, -inf, -inf],\n        [0., 0., 0., 0., -inf, -inf],\n        [0., 0., 0., 0., 0., -inf],\n        [0., 0., 0., 0., 0., 0.]])\n```\n\nAnd it should actually be like this, right? (basically i would call it a length mask 😅)\n\n```\ntensor([[0., 0., 0., 0., 0., -inf],\n        [0., -inf, -inf, -inf, -inf, -inf],\n        [0., 0., -inf, -inf, -inf, -inf],\n        [0., 0., 0., 0., 0., -inf],\n        [0., 0., 0., 0., 0., -inf],\n        [0., 0., 0., 0., 0., -inf]])\n```\n\nIs my understanding apt?",
      "votes": 1,
      "replies": [
        {
          "id": 1107933,
          "postDate": "2020-12-10T05:10:58.473Z",
          "content": "<p>No, your understanding is not correct. The mask is not applied on this dimension, instead, it is applied to every attention head.<br>\nThe upper-triangular mask will make sure at each step we only look at the past.<br>\nIn you above example, every position needs a different mask. When you combine all the masks for all steps, you will get a upper-triangular matrix.</p>",
          "rawMarkdown": "No, your understanding is not correct. The mask is not applied on this dimension, instead, it is applied to every attention head.\nThe upper-triangular mask will make sure at each step we only look at the past.\nIn you above example, every position needs a different mask. When you combine all the masks for all steps, you will get a upper-triangular matrix.",
          "votes": 2
        },
        {
          "id": 1107968,
          "postDate": "2020-12-10T06:14:11.953Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> for the clarifications!  Then it will remain the same as tgt_mask as well? (look-ahead-mask from decoder point of view) as there also we need to prevent looking ahead at future tokens pretty much.</p>",
          "rawMarkdown": "Thanks @frankpanxj for the clarifications!  Then it will remain the same as tgt_mask as well? (look-ahead-mask from decoder point of view) as there also we need to prevent looking ahead at future tokens pretty much."
        },
        {
          "id": 1107983,
          "postDate": "2020-12-10T06:34:46.527Z",
          "content": "<p>That is a different thing, although they are all called \"mask\". The target mask is only to filter out the steps we care about, i.e. excluding padding and lectures in this case.</p>",
          "rawMarkdown": "That is a different thing, although they are all called \"mask\". The target mask is only to filter out the steps we care about, i.e. excluding padding and lectures in this case."
        },
        {
          "id": 1107989,
          "postDate": "2020-12-10T06:44:15.427Z",
          "content": "<p>You don't use lecture information? (i.e. no attention to it?)</p>",
          "rawMarkdown": "You don't use lecture information? (i.e. no attention to it?)"
        },
        {
          "id": 1107993,
          "postDate": "2020-12-10T06:58:34.677Z",
          "content": "<p>I use lecture information. What I mean here is that you might want to exclude lectures when calculating the loss and AUC. It won't be a very big problem if you don't, other than causing some distortion to the metrics.</p>",
          "rawMarkdown": "I use lecture information. What I mean here is that you might want to exclude lectures when calculating the loss and AUC. It won't be a very big problem if you don't, other than causing some distortion to the metrics.",
          "votes": 2
        },
        {
          "id": 1109058,
          "postDate": "2020-12-11T09:39:55.617Z",
          "content": "<p><a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a>  i dont use lectures and pass the mask this way  ,where future mask is np.triu.<br>\nI do this as i see in Architecture that every input to MultiAttn head is through Triu mask.</p>\n<pre><code>att_mask = future_mask(x.size(0)).to(device)\nresp_attn= future_mask(response.size(0)).to(device)\n        x1=self.transformer (src=x,tgt=response ,\n                             src_mask=att_mask,\n                               tgt_mask=resp_attn,\n                               memory_mask=att_mask)\n</code></pre>\n<p>Let me know if am following the paper correctly. </p>",
          "rawMarkdown": "@frankpanxj  i dont use lectures and pass the mask this way  ,where future mask is np.triu.\nI do this as i see in Architecture that every input to MultiAttn head is through Triu mask.\n```\natt_mask = future_mask(x.size(0)).to(device)\nresp_attn= future_mask(response.size(0)).to(device)\n        x1=self.transformer (src=x,tgt=response ,\n                             src_mask=att_mask,\n                               tgt_mask=resp_attn,\n                               memory_mask=att_mask)\n```\nLet me know if am following the paper correctly. ",
          "votes": 2
        },
        {
          "id": 1109234,
          "postDate": "2020-12-11T12:59:44.987Z",
          "content": "<p>Actually I just realized you guys are talking about the Transformer Module of Torch <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>. So what I said about tgt_mask above was not correct.</p>\n<p>To my understanding src_mask and tgt_mask should all be triu matrix. <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> your implementation looks good to me, as long as size(0) is the sequence length (not batch size) here. (Since the length of x and tgt is the same in this case, maybe you can just use att_mask for all three masks?)</p>",
          "rawMarkdown": "Actually I just realized you guys are talking about the Transformer Module of Torch @jaideepvalani @adityaecdrid. So what I said about tgt_mask above was not correct.\n\nTo my understanding src_mask and tgt_mask should all be triu matrix. @jaideepvalani your implementation looks good to me, as long as size(0) is the sequence length (not batch size) here. (Since the length of x and tgt is the same in this case, maybe you can just use att_mask for all three masks?)"
        },
        {
          "id": 1109303,
          "postDate": "2020-12-11T14:17:29.713Z",
          "content": "<p>I think src[tgt]_key_padding_mask can be used for individual sequence length masking in the transformer module of torch.</p>",
          "rawMarkdown": "I think src[tgt]_key_padding_mask can be used for individual sequence length masking in the transformer module of torch."
        },
        {
          "id": 1117041,
          "postDate": "2020-12-17T17:08:29.307Z",
          "content": "<p>This was a great sub-thread conversation.</p>\n<p>Along the same lines, I am interested in how you all are handling <em>validation</em>. In my current scheme, I have a number of users who are only in train and not val, and a number of users who are in both. Since I don't have user_id embeddings, it makes the most sense that users in both should have their historical data available (fed to the transformer model) during val; but for validation sake, I don't want the validation loss/metric to include those historical values. Is anyone handling this using double masks, or how are you all dealing with that?</p>",
          "rawMarkdown": "This was a great sub-thread conversation.\n\nAlong the same lines, I am interested in how you all are handling *validation*. In my current scheme, I have a number of users who are only in train and not val, and a number of users who are in both. Since I don't have user_id embeddings, it makes the most sense that users in both should have their historical data available (fed to the transformer model) during val; but for validation sake, I don't want the validation loss/metric to include those historical values. Is anyone handling this using double masks, or how are you all dealing with that?",
          "votes": 2
        },
        {
          "id": 1119990,
          "postDate": "2020-12-20T14:18:18.647Z",
          "content": "<p>I just use 95% users to train, 5% users for validation (only take the latest seqence length of history for validation). This is certainly not ideal as some of the targets in validation can only rely on none to very short history to predict, even if longer history exists. Therefore the local validation score will be lower than LB score.  But my submissions suggest higher local validation score always means higher LB score, so in practice this might be good enough.</p>",
          "rawMarkdown": "I just use 95% users to train, 5% users for validation (only take the latest seqence length of history for validation). This is certainly not ideal as some of the targets in validation can only rely on none to very short history to predict, even if longer history exists. Therefore the local validation score will be lower than LB score.  But my submissions suggest higher local validation score always means higher LB score, so in practice this might be good enough.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1096342,
      "postDate": "2020-11-30T12:29:36.033Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4570913%2Fb5c717170357fb281614b150c232e4a0%2Fxy2.png?generation=1606666420284171&amp;alt=media\" alt=\"\"><br>\nHi, this graph is AUC for given position in a 128 sequence length of a transformer model.<br>\nI think we can interpret the slope as the ability of the model to take into account the past …</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4570913%2Fb5c717170357fb281614b150c232e4a0%2Fxy2.png?generation=1606666420284171&alt=media)\nHi, this graph is AUC for given position in a 128 sequence length of a transformer model.\nI think we can interpret the slope as the ability of the model to take into account the past ...\n ",
      "votes": 1,
      "replies": [
        {
          "id": 1096685,
          "postDate": "2020-11-30T17:36:05.120Z",
          "content": "<p>Is this for train data or validation?</p>\n<p>I have a similar trajectory as well.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2555700%2Fb25a23e4905be16922936b5f320e1b1d%2FScreenshot%202020-11-30%20223258.png?generation=1606757745828460&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Is this for train data or validation?\n\nI have a similar trajectory as well.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2555700%2Fb25a23e4905be16922936b5f320e1b1d%2FScreenshot%202020-11-30%20223258.png?generation=1606757745828460&alt=media)\n"
        },
        {
          "id": 1096702,
          "postDate": "2020-11-30T17:50:39.300Z",
          "content": "<p>Validation on about 2% of the dataset and you ?</p>",
          "rawMarkdown": "Validation on about 2% of the dataset and you ?"
        },
        {
          "id": 1096914,
          "postDate": "2020-11-30T21:46:50.850Z",
          "content": "<p>Validation, 3% of users. I run validation for complete users sequences (max 128).</p>",
          "rawMarkdown": "Validation, 3% of users. I run validation for complete users sequences (max 128)."
        }
      ]
    },
    {
      "id": 1082028,
      "postDate": "2020-11-17T14:24:40.857Z",
      "content": "<p>I am struggling to make my lgbm inference properly for 4 days now, so forget about DL models inference here for now; (to save your time on more stronger features that can help your lgbm models);  Another reason is, one small mistake and the whole thing will blow, as it's quite unforgivable in nature</p>\n<p>Simple idea is to use the feature vectors from them as inputs to your lgbm to get started where you know you won't be breaking anything for sure! (as that should have some visible improvement on your lgbm model)</p>",
      "rawMarkdown": "I am struggling to make my lgbm inference properly for 4 days now, so forget about DL models inference here for now; (to save your time on more stronger features that can help your lgbm models);  Another reason is, one small mistake and the whole thing will blow, as it's quite unforgivable in nature\n\nSimple idea is to use the feature vectors from them as inputs to your lgbm to get started where you know you won't be breaking anything for sure! (as that should have some visible improvement on your lgbm model)",
      "votes": 1
    },
    {
      "id": 1138890,
      "postDate": "2021-01-05T03:34:30.850Z",
      "content": "<p>Not a great implementation, but as promised in my old kernel, here's the one which has <a href=\"https://www.kaggle.com/adityaecdrid/fork-of-saint-inference-ea970c\" target=\"_blank\">SAINT's Inference</a>. It's buggy somewhere (~.739 on LB )and I am not able to fix it on my own, so making it public, hoping it will help someone for sure!</p>\n<p>NB, there are bugs in it, so use it by analysing the same, the code's for reference only.</p>\n<p>Thanks for all your shares!</p>",
      "rawMarkdown": "Not a great implementation, but as promised in my old kernel, here's the one which has [SAINT's Inference](https://www.kaggle.com/adityaecdrid/fork-of-saint-inference-ea970c). It's buggy somewhere (~.739 on LB )and I am not able to fix it on my own, so making it public, hoping it will help someone for sure!\n\nNB, there are bugs in it, so use it by analysing the same, the code's for reference only.\n\nThanks for all your shares!",
      "votes": 2
    },
    {
      "id": 1133430,
      "postDate": "2020-12-31T07:59:11.803Z",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> Is there any chance you share a bit of the secret that boosts your LB from 0.798 to 0.806? That's very amazing, you are doing so great!</p>",
      "rawMarkdown": "@abdurrafae Is there any chance you share a bit of the secret that boosts your LB from 0.798 to 0.806? That's very amazing, you are doing so great!",
      "votes": 2,
      "replies": [
        {
          "id": 1133475,
          "postDate": "2020-12-31T08:46:01.947Z",
          "content": "<p>Here it's not easy to climb, the team is working hard. You also made solid progress! with transformer too I guess.</p>",
          "rawMarkdown": "Here it's not easy to climb, the team is working hard. You also made solid progress! with transformer too I guess.",
          "votes": 1
        },
        {
          "id": 1133533,
          "postDate": "2020-12-31T09:40:18.523Z",
          "content": "<p>Lastest update 79.3 is saint plus. we are working  towards making it better<br>\n<a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>  did you use Lagtime , what is way u computed it  ts2-ts1 ?</p>",
          "rawMarkdown": "Lastest update 79.3 is saint plus. we are working  towards making it better\n@mpware  did you use Lagtime , what is way u computed it  ts2-ts1 ?",
          "votes": -2,
          "replies": [
            {
              "id": 1133577,
              "postDate": "2020-12-31T10:44:40.370Z",
              "content": "<p>Yes, timestamp(N) - timestamp(N-1). It's not exactly the same definition as in SAINT+ paper but an approximation</p>",
              "rawMarkdown": "Yes, timestamp(N) - timestamp(N-1). It's not exactly the same definition as in SAINT+ paper but an approximation"
            }
          ]
        },
        {
          "id": 1133583,
          "postDate": "2020-12-31T10:54:34.727Z",
          "content": "<p>Thanks, I'm just fixed a few bugs (2 to be precise) :) <br>\nNo change in the model architecture or pipeline.</p>",
          "rawMarkdown": "Thanks, I'm just fixed a few bugs (2 to be precise) :) \nNo change in the model architecture or pipeline.",
          "votes": 4
        },
        {
          "id": 1133664,
          "postDate": "2020-12-31T12:17:00.807Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Yes, difficult … and yes, worked (and will work) hard, very hard …😂😆</p>",
          "rawMarkdown": "@mpware Yes, difficult ... and yes, worked (and will work) hard, very hard ...😂😆"
        },
        {
          "id": 1133753,
          "postDate": "2020-12-31T13:53:33.353Z",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> if possible, could you share a bit, other than the features used in the original SAINT / SAINT+, what are (some of) the extra things you have added to your model / features?</p>\n<p>I am surprised that you could get LB <code>0.792</code> with <code>d_model=128 with n_layer=2</code>. If you have no extra things in your model / features, I feel I must have done something wrong … (I also tried to debug, but the inputs for my training/validation has no difference to my submission pipeline … i.e. no bug found anymore)</p>\n<p>I hope I can hear a bit from you - at least give me some hope 😄</p>",
          "rawMarkdown": "@abdurrafae if possible, could you share a bit, other than the features used in the original SAINT / SAINT+, what are (some of) the extra things you have added to your model / features?\n\nI am surprised that you could get LB `0.792` with `d_model=128 with n_layer=2`. If you have no extra things in your model / features, I feel I must have done something wrong ... (I also tried to debug, but the inputs for my training/validation has no difference to my submission pipeline ... i.e. no bug found anymore)\n\nI hope I can hear a bit from you - at least give me some hope 😄"
        },
        {
          "id": 1133781,
          "postDate": "2020-12-31T14:23:30.593Z",
          "content": "<p>I reached 0.775 with pure saint iirc. I added the time features to get up to 0.788, fixed a few bugs and got 0.792. Increasing model size from there took me to 0.798 and fixing a few more bugs helped me reach 0.806. Now these bugs were present all along the way so not sure what to make of it.</p>\n<p>I think you should have 0.79 just by using SAINT+ with 128 d_model and  2 n_layers.</p>",
          "rawMarkdown": "I reached 0.775 with pure saint iirc. I added the time features to get up to 0.788, fixed a few bugs and got 0.792. Increasing model size from there took me to 0.798 and fixing a few more bugs helped me reach 0.806. Now these bugs were present all along the way so not sure what to make of it.\n\nI think you should have 0.79 just by using SAINT+ with 128 d_model and  2 n_layers.",
          "votes": 8
        },
        {
          "id": 1133788,
          "postDate": "2020-12-31T14:26:17.910Z",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> Do you use all the data for training? And how long does it take you to train one epoch?</p>",
          "rawMarkdown": "@abdurrafae Do you use all the data for training? And how long does it take you to train one epoch?",
          "votes": 2
        },
        {
          "id": 1133800,
          "postDate": "2020-12-31T14:38:28.557Z",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> Thank you! One last question (at least for today), what's is your CV strategy, and what's your validation dataset size?</p>",
          "rawMarkdown": "@abdurrafae Thank you! One last question (at least for today), what's is your CV strategy, and what's your validation dataset size?",
          "votes": 1
        },
        {
          "id": 1133818,
          "postDate": "2020-12-31T15:02:16.700Z",
          "content": "<p>only use kaggle gpu?amazing</p>",
          "rawMarkdown": "only use kaggle gpu?amazing"
        },
        {
          "id": 1133819,
          "postDate": "2020-12-31T15:04:24.983Z",
          "content": "<p>Just keeping a few users (3~4%) unseen for validation. </p>",
          "rawMarkdown": "Just keeping a few users (3~4%) unseen for validation. "
        },
        {
          "id": 1133822,
          "postDate": "2020-12-31T15:05:48.417Z",
          "content": "<p>Yeah, and TPUs.<br>\nLearned how to use TPUs 2 weeks ago. Got inspired by <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> :)</p>",
          "rawMarkdown": "Yeah, and TPUs.\nLearned how to use TPUs 2 weeks ago. Got inspired by @yihdarshieh :)",
          "votes": 2
        },
        {
          "id": 1133827,
          "postDate": "2020-12-31T15:10:43.270Z",
          "content": "<p>Looking forward to your proposal after the competition end.Can I know the parameters of your pure saint model?</p>",
          "rawMarkdown": "Looking forward to your proposal after the competition end.Can I know the parameters of your pure saint model?"
        },
        {
          "id": 1133830,
          "postDate": "2020-12-31T15:13:21.390Z",
          "content": "<p>d_model 128 and n_layers 2</p>\n<p>Edit: I'm using a modified version now with higher dimensions and layers.</p>",
          "rawMarkdown": "d_model 128 and n_layers 2\n\nEdit: I'm using a modified version now with higher dimensions and layers."
        },
        {
          "id": 1133863,
          "postDate": "2020-12-31T15:36:25.747Z",
          "content": "<p>is it as heavy as mentioned in paper. I feel  512 dim, 6 /6  thats an  overkill given the resource every one has.</p>",
          "rawMarkdown": "is it as heavy as mentioned in paper. I feel  512 dim, 6 /6  thats an  overkill given the resource every one has."
        },
        {
          "id": 1133866,
          "postDate": "2020-12-31T15:37:17.490Z",
          "content": "<p>I definitely helped a DL monster to grow further 😆</p>",
          "rawMarkdown": "I definitely helped a DL monster to grow further 😆"
        },
        {
          "id": 1133870,
          "postDate": "2020-12-31T15:38:40.537Z",
          "content": "<p>So you only validate on unseen users? I have keep users with some training history and some new users</p>",
          "rawMarkdown": "So you only validate on unseen users? I have keep users with some training history and some new users"
        },
        {
          "id": 1133876,
          "postDate": "2020-12-31T15:43:45.190Z",
          "content": "<p>Hope my TPU notebooks get much more votes once you give your <code>thank you speech</code> after winning 😊</p>",
          "rawMarkdown": "Hope my TPU notebooks get much more votes once you give your `thank you speech` after winning 😊",
          "votes": 2
        },
        {
          "id": 1133917,
          "postDate": "2020-12-31T16:35:41Z",
          "content": "<p><a href=\"https://www.kaggle.com/jjaideepvalani\" target=\"_blank\">@jjaideepvalani</a></p>\n<p>It's currently at 512 with 4 layers but I think it overfits a bit. I'll try with 256. Thing is the time difference in training with 256 vs 512 is not more than ~60 secs per epoch so I go with 512 while training.</p>\n<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> Amen to that :)</p>",
          "rawMarkdown": "@jjaideepvalani\n\nIt's currently at 512 with 4 layers but I think it overfits a bit. I'll try with 256. Thing is the time difference in training with 256 vs 512 is not more than ~60 secs per epoch so I go with 512 while training.\n\n@yihdarshieh Amen to that :)",
          "votes": 2
        },
        {
          "id": 1133932,
          "postDate": "2020-12-31T16:59:57.437Z",
          "content": "<p>My goal of Saint model is to achieve the effect of the paper, I am just a novice, and cost too much time in lgb, I can not achieve a high score。</p>",
          "rawMarkdown": "My goal of Saint model is to achieve the effect of the paper, I am just a novice, and cost too much time in lgb, I can not achieve a high score。"
        },
        {
          "id": 1134939,
          "postDate": "2021-01-01T18:05:32.597Z",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> will we have neg lag times ,apart from the ones that correspond to timestep -0 ?<br>\nif yes around how many ?</p>",
          "rawMarkdown": "@abdurrafae will we have neg lag times ,apart from the ones that correspond to timestep -0 ?\nif yes around how many ?",
          "votes": -3
        },
        {
          "id": 1134956,
          "postDate": "2021-01-01T18:21:11.023Z",
          "content": "<p>lag time all above 0,</p>",
          "rawMarkdown": "lag time all above 0,"
        },
        {
          "id": 1134965,
          "postDate": "2021-01-01T18:36:55.250Z",
          "content": "<p>I can't comment on the time feature for now</p>",
          "rawMarkdown": "I can't comment on the time feature for now"
        }
      ]
    },
    {
      "id": 1142307,
      "postDate": "2021-01-07T09:59:32.260Z",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>  congrats for  continuing to be in gold hunt.<br>\nLooks like after hard efforts we might miss by few miles :)</p>\n<p>Any thing one can think around to go beyond .806 .  :)</p>",
      "rawMarkdown": "@mpware  congrats for  continuing to be in gold hunt.\nLooks like after hard efforts we might miss by few miles :)\n\nAny thing one can think around to go beyond .806 .  :)\n",
      "replies": [
        {
          "id": 1142313,
          "postDate": "2021-01-07T10:02:55.740Z",
          "content": "<p>Better data usage </p>",
          "rawMarkdown": "Better data usage "
        },
        {
          "id": 1142320,
          "postDate": "2021-01-07T10:08:48.490Z",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> Your team have great work and have good rank.<br>\nIf you don't mind. How about your single model?</p>",
          "rawMarkdown": "@jaideepvalani Your team have great work and have good rank.\nIf you don't mind. How about your single model?"
        },
        {
          "id": 1142821,
          "postDate": "2021-01-07T15:58:53.963Z",
          "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a>  around same as ensemble only.. close to what  bro abdur  had got earlier few days ago. <br>\nBtw any way to fasten up the inference. I put optimized merges only, but i think too much time goes in status updates. using np.append.</p>\n<p>I regret that i lost my resources so couldnt contribute to experiments but was trying to compensate best by doing analysis and coming up with new features , improving scores</p>",
          "rawMarkdown": "@m10515009  around same as ensemble only.. close to what  bro abdur  had got earlier few days ago. \nBtw any way to fasten up the inference. I put optimized merges only, but i think too much time goes in status updates. using np.append.\n\nI regret that i lost my resources so couldnt contribute to experiments but was trying to compensate best by doing analysis and coming up with new features , improving scores"
        }
      ]
    },
    {
      "id": 1103856,
      "postDate": "2020-12-06T11:19:10.600Z",
      "content": "<p>Any of you still working this way?</p>",
      "rawMarkdown": "Any of you still working this way?",
      "votes": 2,
      "replies": [
        {
          "id": 1103878,
          "postDate": "2020-12-06T11:50:29.150Z",
          "content": "<p>Okay, I think I'm sharing a full description of my current best model since I have taken some ideas from this discussion:</p>\n<ul>\n<li>CV: 0.7610, LB: 0.767</li>\n<li>Model<ul>\n<li>2 Encoder Layers, 2 Decoder Layers</li>\n<li>Model dimension: 128</li>\n<li>Number of heads: 4</li>\n<li>Sequence Length: 96</li>\n<li>Padding mask and look-ahead mask in both encoder and decoder.</li>\n<li>Padding-masked loss.</li></ul></li>\n<li>Features (differences with SAINT+):<ul>\n<li>No absolute position encoding. Instead relative position in the MHA layers.</li>\n<li>Using task_container_id embeddings which I add to the decoder.</li>\n<li>I use lag (scaled 0-1), elapsed_time (scaled 0-1) and prior_question_had_explanation as continous features (they all together in a dense layer that transform to model dimension size, activated with tanh). It is added to the decoder.</li></ul></li>\n<li>Training procedure:<ul>\n<li>Batch size: 512</li>\n<li>Optimizer: Adam (beta_1=0.9, beta_2=0.999, epsilon=1e-8) with a LR scheduler as stated in SAINT+ paper.</li>\n<li>Loss: Categorical crossentropy (2 output neurons activated with softmax).</li>\n<li>AUC computed taking only into account last non-padded element in every sequence.</li></ul></li>\n<li>Input flow:<ul>\n<li>80 M rows minus lectures (311567 users).</li>\n<li>For training, random sample a sequence per user per epoch, having a linear increasing probability (the last sequence in a user has double the chances to be selected than the first one).</li>\n<li>For validation I use every possible sequence for 4% of the users. </li></ul></li>\n</ul>",
          "rawMarkdown": "Okay, I think I'm sharing a full description of my current best model since I have taken some ideas from this discussion:\n- CV: 0.7610, LB: 0.767\n- Model\n  - 2 Encoder Layers, 2 Decoder Layers\n  - Model dimension: 128\n  - Number of heads: 4\n  - Sequence Length: 96\n  - Padding mask and look-ahead mask in both encoder and decoder.\n  - Padding-masked loss.\n- Features (differences with SAINT+):\n  - No absolute position encoding. Instead relative position in the MHA layers.\n  - Using task_container_id embeddings which I add to the decoder.\n  - I use lag (scaled 0-1), elapsed_time (scaled 0-1) and prior_question_had_explanation as continous features (they all together in a dense layer that transform to model dimension size, activated with tanh). It is added to the decoder.\n- Training procedure:\n  - Batch size: 512\n  - Optimizer: Adam (beta_1=0.9, beta_2=0.999, epsilon=1e-8) with a LR scheduler as stated in SAINT+ paper.\n  - Loss: Categorical crossentropy (2 output neurons activated with softmax).\n  - AUC computed taking only into account last non-padded element in every sequence.\n- Input flow:\n  - 80 M rows minus lectures (311567 users).\n  - For training, random sample a sequence per user per epoch, having a linear increasing probability (the last sequence in a user has double the chances to be selected than the first one).\n  - For validation I use every possible sequence for 4% of the users. ",
          "votes": 11
        },
        {
          "id": 1104350,
          "postDate": "2020-12-06T20:51:51.400Z",
          "content": "<p>For me num_layers = 2 has been more effective, the model diverges (AUC gets stuck around 0.63) if I increase this even to 3.<br>\nSimilar is the case with d_model, if I increase this from 128, the model again diverges. </p>\n<p>CV: 0.775 (No LB for this)</p>\n<p>Maybe it's an issue with my implementation.</p>",
          "rawMarkdown": "For me num_layers = 2 has been more effective, the model diverges (AUC gets stuck around 0.63) if I increase this even to 3.\nSimilar is the case with d_model, if I increase this from 128, the model again diverges. \n\nCV: 0.775 (No LB for this)\n\nMaybe it's an issue with my implementation.",
          "votes": 1
        },
        {
          "id": 1104406,
          "postDate": "2020-12-06T22:47:13.907Z",
          "content": "<p>Mine still converges when bigger but no needed. With my current size is actually overfitting in a few epochs, and training keeps improving over time (hence showing the model can still learn). I still don't understand how can we obtain the metris reported by SAINT+ neither why they use such a big model (4 encoder layers, 4 decoder layers, 512 model dimension).</p>",
          "rawMarkdown": "Mine still converges when bigger but no needed. With my current size is actually overfitting in a few epochs, and training keeps improving over time (hence showing the model can still learn). I still don't understand how can we obtain the metris reported by SAINT+ neither why they use such a big model (4 encoder layers, 4 decoder layers, 512 model dimension)."
        },
        {
          "id": 1106063,
          "postDate": "2020-12-08T13:26:40.173Z",
          "content": "<p>Still on it but fighting with memory/time out issues since a few days. Fighting for both training and inference!</p>",
          "rawMarkdown": "Still on it but fighting with memory/time out issues since a few days. Fighting for both training and inference!",
          "votes": 3
        },
        {
          "id": 1106077,
          "postDate": "2020-12-08T13:50:18.117Z",
          "content": "<blockquote>\n  <p>AUC computed taking only into account last non-padded element in every sequence</p>\n</blockquote>\n<p>Hmm, that's interesting. Have you tried computing it for all non-padded elements?</p>\n<blockquote>\n  <p>Loss: Categorical crossentropy (2 output neurons activated with softmax).</p>\n</blockquote>\n<p>In paper, they used a sigmoid at the end. Plus, if we use ignore_index in the loss for the padded element, that's should bring in the same effevt, right? (wrt NN.CrossEntropy, torch).</p>\n<p>Also, I use 4 encoder/decoder with 8 Attention heads and hidden dim as 256/512. 2048 is too high for me.</p>",
          "rawMarkdown": ">AUC computed taking only into account last non-padded element in every sequence\n\nHmm, that's interesting. Have you tried computing it for all non-padded elements?\n\n>Loss: Categorical crossentropy (2 output neurons activated with softmax).\n\nIn paper, they used a sigmoid at the end. Plus, if we use ignore_index in the loss for the padded element, that's should bring in the same effevt, right? (wrt NN.CrossEntropy, torch).\n\nAlso, I use 4 encoder/decoder with 8 Attention heads and hidden dim as 256/512. 2048 is too high for me.",
          "votes": 2
        },
        {
          "id": 1106230,
          "postDate": "2020-12-08T16:09:49.697Z",
          "content": "<p>I have tried, yes, but I found it to be less accurate.</p>\n<p>I think it should be practically the same, I read some guys pointing out that 2-softmax gives better results than 1-sigmoid but can't confirm.</p>\n<p>That model you describe is too big I think, do you really need it?</p>",
          "rawMarkdown": "I have tried, yes, but I found it to be less accurate.\n\nI think it should be practically the same, I read some guys pointing out that 2-softmax gives better results than 1-sigmoid but can't confirm.\n\nThat model you describe is too big I think, do you really need it?"
        },
        {
          "id": 1107298,
          "postDate": "2020-12-09T15:15:54.613Z",
          "content": "<blockquote>\n  <p>That model you describe is too big I think, do you really need it?</p>\n</blockquote>\n<p>Well, i don't need it because a 2 layer model also scores the same almost for me. Was just trying to mimic the paper, that's it. Great to see you up on LB 🎊🎊</p>",
          "rawMarkdown": ">That model you describe is too big I think, do you really need it?\n\nWell, i don't need it because a 2 layer model also scores the same almost for me. Was just trying to mimic the paper, that's it. Great to see you up on LB 🎊🎊"
        },
        {
          "id": 1119389,
          "postDate": "2020-12-20T02:21:08.490Z",
          "content": "<p>Hi claverru, I trained a four-layer SAINT+, with final LB 0.78. But I find it's quite hard to blend this model with LGBM. It always leads to Submission Scoring Error. I suppose it is due to running over time.  Do you meet with any similar problems?  </p>",
          "rawMarkdown": "Hi claverru, I trained a four-layer SAINT+, with final LB 0.78. But I find it's quite hard to blend this model with LGBM. It always leads to Submission Scoring Error. I suppose it is due to running over time.  Do you meet with any similar problems?  "
        }
      ]
    },
    {
      "id": 1100139,
      "postDate": "2020-12-02T21:56:04.503Z",
      "content": "<p>MPWARE, without divulging too much information, can you tell me if you're adding engineered features to your SAINT model? I'm particularly interested in engineered continuous features. I'm amazed at how well SAKT does with just the content_id as input. It does better than a tabular model I've been working on with several engineered features. I've just been wondering how many of these engineered features are being captured indirectly by the sequences, and how useful it would be to add them to my SAKT model. Cheers </p>",
      "rawMarkdown": "MPWARE, without divulging too much information, can you tell me if you're adding engineered features to your SAINT model? I'm particularly interested in engineered continuous features. I'm amazed at how well SAKT does with just the content_id as input. It does better than a tabular model I've been working on with several engineered features. I've just been wondering how many of these engineered features are being captured indirectly by the sequences, and how useful it would be to add them to my SAKT model. Cheers ",
      "votes": 2,
      "replies": [
        {
          "id": 1100145,
          "postDate": "2020-12-02T22:05:03.633Z",
          "content": "<p><a href=\"https://www.kaggle.com/gannonreynolds\" target=\"_blank\">@gannonreynolds</a> Zero engineered features in my SAINT implementation. I've tried to add <code>tags1</code> and <code>task_container_id</code>  related to questions but it did not help. </p>\n<blockquote>\n  <p>I'm amazed at how well SAKT does with just the content_id as input</p>\n</blockquote>\n<p>Yes, that's the magic of Transformer. It learns hidden patterns from sequences (it seems that Transformers could overperform some CNN models for computer vision). It's promising. <br>\nCurrently I'm trying TabNet with and without engineered features but I'm not able to beat any of my other models.</p>",
          "rawMarkdown": "@gannonreynolds Zero engineered features in my SAINT implementation. I've tried to add `tags1` and `task_container_id`  related to questions but it did not help. \n\n> I'm amazed at how well SAKT does with just the content_id as input\n\nYes, that's the magic of Transformer. It learns hidden patterns from sequences (it seems that Transformers could overperform some CNN models for computer vision). It's promising. \nCurrently I'm trying TabNet with and without engineered features but I'm not able to beat any of my other models.",
          "votes": 4
        },
        {
          "id": 1100169,
          "postDate": "2020-12-02T22:23:08.390Z",
          "content": "<p>Wow that's really impressive, thanks for the intel. I had become pretty uninspired just running into a wall with the performance of my tabular model, but this thread has re-sparked my enthusiasm. Exciting stuff</p>",
          "rawMarkdown": "Wow that's really impressive, thanks for the intel. I had become pretty uninspired just running into a wall with the performance of my tabular model, but this thread has re-sparked my enthusiasm. Exciting stuff",
          "votes": 1
        }
      ]
    },
    {
      "id": 1086735,
      "postDate": "2020-11-22T02:15:46.330Z",
      "content": "<p></p>\n<p>Had a first successful submission at .708;</p>\n<p>Edit -&gt; Current best is at .740 ~2.5-3 hours inference.</p>",
      "rawMarkdown": "~~So, My inference design finally works ~2.5-3 hours but the performance isn't quite expected. (~.67) 🥺. So i was looking for tips as to how can one debug the same except printing tensors 😅? Ty!\n~~\n\nHad a first successful submission at .708;\n\nEdit -> Current best is at .740 ~2.5-3 hours inference.",
      "votes": 2
    },
    {
      "id": 1141340,
      "postDate": "2021-01-06T16:35:58.273Z",
      "content": "<p>This competition is so stressful and frustrating with these technical constraints. I hope my last submission will work …</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F371132%2F657af4d0a68945a185d57f1198dbe893%2FScreenshot_2021-01-06%20TensorBoard(1).png?generation=1609950875720205&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "This competition is so stressful and frustrating with these technical constraints. I hope my last submission will work ...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F371132%2F657af4d0a68945a185d57f1198dbe893%2FScreenshot_2021-01-06%20TensorBoard(1).png?generation=1609950875720205&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 1141365,
          "postDate": "2021-01-06T16:51:36.053Z",
          "content": "<p>Mine keeps giving submission errors after 4 hours GPU time and 4 hours of submission. If anyone has a clue do let me know. It is some silly mistake but I am doing so many things at last min that I just am gave up on this error</p>\n<p>calculate lagtime list for the test_df received<br>\nnote that max_timestamp_u_dict has the last timestamp for each user</p>\n<p>prev_test_df = test_df.copy()<br>\nquestion_len=len( test_df[test_df['content_type_id'] == 0])<br>\nlagtime_list = np.zeros(question_len, dtype=np.float32)<br>\ni=0</p>\n<p>for j, (user_id,content_type_id,timestamp,content_id) in enumerate(zip(test_df['user_id'].values,test_df['content_type_id'].values,test_df['timestamp'].values, test_df['content_id'].values)):<br>\nif(content_type_id==0):<br>\nif user_id in max_timestamp_u_dict['max_time_stamp'].keys():<br>\nlagtime_list[i]=int((timestamp - max_timestamp_u_dict['max_time_stamp'][user_id]) /(1000360010))<br>\nmax_timestamp_u_dict['max_time_stamp'][user_id] = timestamp<br>\nelse:<br>\nlagtime_list[i] = int(lagtime_mean)<br>\nmax_timestamp_u_dict['max_time_stamp'].update({user_id:timestamp})<br>\ni=i+1</p>\n<p>Now lagtime is in a list. We just have to assign it when processing the next test_df as below:<br>\nprev_test_df = prev_test_df[prev_test_df.content_type_id == False].reset_index(drop=True)<br>\nprev_test_df[\"lagtime\"]=lagtime_list<br>\nprev_test_df['lagtime'].fillna(int(lagtime_mean), inplace=True)<br>\nprev_test_df.lagtime=prev_test_df.lagtime.astype('int')</p>\n<p>dump the state<br>\nprev_group = prev_test_df[['user_id', 'content_id', 'answered_correctly', 'lagtime']].groupby('user_id').apply(lambda r: (r['content_id'].values,r['answered_correctly'].values,r['lagtime'].values))</p>\n<p>Somehow this gives scoring error and my GPU time is out<br>\nAny hints, clues appreciated</p>",
          "rawMarkdown": "Mine keeps giving submission errors after 4 hours GPU time and 4 hours of submission. If anyone has a clue do let me know. It is some silly mistake but I am doing so many things at last min that I just am gave up on this error\n\ncalculate lagtime list for the test_df received\nnote that max_timestamp_u_dict has the last timestamp for each user\n\nprev_test_df = test_df.copy()\nquestion_len=len( test_df[test_df['content_type_id'] == 0])\nlagtime_list = np.zeros(question_len, dtype=np.float32)\ni=0\n\nfor j, (user_id,content_type_id,timestamp,content_id) in enumerate(zip(test_df['user_id'].values,test_df['content_type_id'].values,test_df['timestamp'].values, test_df['content_id'].values)):\nif(content_type_id==0):\nif user_id in max_timestamp_u_dict['max_time_stamp'].keys():\nlagtime_list[i]=int((timestamp - max_timestamp_u_dict['max_time_stamp'][user_id]) /(1000360010))\nmax_timestamp_u_dict['max_time_stamp'][user_id] = timestamp\nelse:\nlagtime_list[i] = int(lagtime_mean)\nmax_timestamp_u_dict['max_time_stamp'].update({user_id:timestamp})\ni=i+1\n\nNow lagtime is in a list. We just have to assign it when processing the next test_df as below:\nprev_test_df = prev_test_df[prev_test_df.content_type_id == False].reset_index(drop=True)\nprev_test_df[\"lagtime\"]=lagtime_list\nprev_test_df['lagtime'].fillna(int(lagtime_mean), inplace=True)\nprev_test_df.lagtime=prev_test_df.lagtime.astype('int')\n\ndump the state\nprev_group = prev_test_df[['user_id', 'content_id', 'answered_correctly', 'lagtime']].groupby('user_id').apply(lambda r: (r['content_id'].values,r['answered_correctly'].values,r['lagtime'].values))\n\nSomehow this gives scoring error and my GPU time is out\nAny hints, clues appreciated"
        },
        {
          "id": 1141460,
          "postDate": "2021-01-06T18:12:25.363Z",
          "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> Have you tried running it against the CV script from <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a> ? That was useful for me because you can test your inference against data without having to submit it. So if it errors out, you know why. Or if it doesn't error out, in my experience, the error probably has something to do with how new users are handled.</p>\n<p>Also, if you wrap your code blocks in <code></code> for comments, it makes it much easier to read, for anyone trying to help.</p>",
          "rawMarkdown": "@allohvk Have you tried running it against the CV script from https://www.kaggle.com/its7171/cv-strategy ? That was useful for me because you can test your inference against data without having to submit it. So if it errors out, you know why. Or if it doesn't error out, in my experience, the error probably has something to do with how new users are handled.\n\nAlso, if you wrap your code blocks in ```  ``` for comments, it makes it much easier to read, for anyone trying to help."
        },
        {
          "id": 1141480,
          "postDate": "2021-01-06T18:21:45.917Z",
          "content": "<p><a href=\"https://www.kaggle.com/gannonreynolds\" target=\"_blank\">@gannonreynolds</a> Thanks for the update. I havent taken a look at that. Will do so tomorrow if time permits.</p>",
          "rawMarkdown": "@gannonreynolds Thanks for the update. I havent taken a look at that. Will do so tomorrow if time permits."
        },
        {
          "id": 1141517,
          "postDate": "2021-01-06T18:35:33.597Z",
          "content": "<p>you may need TPU if you have used 40hours gpu</p>",
          "rawMarkdown": "you may need TPU if you have used 40hours gpu"
        }
      ]
    },
    {
      "id": 1140053,
      "postDate": "2021-01-05T19:17:52.897Z",
      "content": "<p>without time features 0.778, however it's too late and add time feature reduce score.didnt find the reason yet</p>",
      "rawMarkdown": "without time features 0.778, however it's too late and add time feature reduce score.didnt find the reason yet"
    },
    {
      "id": 1138465,
      "postDate": "2021-01-04T17:49:22.060Z",
      "content": "<p>I'm curious about the usefulness of the exercise categories used in the SAINT paper. I would imagine the embeddings of the content ids would do a better job at finding the relationships between questions better than a human deciding what the tags should be. Am I off base in thinking that? Has anyone compared models run with the provided tags and question details vs without?</p>\n<p>I understand if no one wants to discuss SAINT implementations this close to the end of the comp. I'm just curious and have very limited GPU time left to test it myself.</p>",
      "rawMarkdown": "I'm curious about the usefulness of the exercise categories used in the SAINT paper. I would imagine the embeddings of the content ids would do a better job at finding the relationships between questions better than a human deciding what the tags should be. Am I off base in thinking that? Has anyone compared models run with the provided tags and question details vs without?\n\nI understand if no one wants to discuss SAINT implementations this close to the end of the comp. I'm just curious and have very limited GPU time left to test it myself.",
      "replies": [
        {
          "id": 1138501,
          "postDate": "2021-01-04T18:22:30.133Z",
          "content": "<p>The category is just the TOEIC part. Sure, the question_id embeddings are more useful, but the TOEIC part (category) also provides signal. Think of it as a type of question clustering. Like \"Reading\", or \"Listening\", or \"Comprehension\", etc.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F2ee06e85eca1af9fdb28c38be8b44eeb%2FScreen%20Shot%202021-01-04%20at%2012.20.58%20PM.png?generation=1609784547132431&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "The category is just the TOEIC part. Sure, the question_id embeddings are more useful, but the TOEIC part (category) also provides signal. Think of it as a type of question clustering. Like \"Reading\", or \"Listening\", or \"Comprehension\", etc.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F2ee06e85eca1af9fdb28c38be8b44eeb%2FScreen%20Shot%202021-01-04%20at%2012.20.58%20PM.png?generation=1609784547132431&alt=media)",
          "votes": 1
        },
        {
          "id": 1138521,
          "postDate": "2021-01-04T18:43:07.927Z",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> Thanks for the response. I imagine it doesn't hurt to include it, if you have the time and computing resources to do so. I've just been wondering how useful it is. If a question only has one part (category), the relevant information about that category should be learned by the embedding of that question id right? I would imagine the embedding learns the part of the question implicitly, and with more nuance, since it's being represented by a higher dimensional tensor. </p>\n<p>I'm still pretty new to the idea of embeddings, so I'm trying to learn what exactly their capabilities and limitations are.</p>",
          "rawMarkdown": "@authman Thanks for the response. I imagine it doesn't hurt to include it, if you have the time and computing resources to do so. I've just been wondering how useful it is. If a question only has one part (category), the relevant information about that category should be learned by the embedding of that question id right? I would imagine the embedding learns the part of the question implicitly, and with more nuance, since it's being represented by a higher dimensional tensor. \n\nI'm still pretty new to the idea of embeddings, so I'm trying to learn what exactly their capabilities and limitations are.",
          "votes": 1
        },
        {
          "id": 1138530,
          "postDate": "2021-01-04T18:49:14Z",
          "content": "<p>At the end of the day, DL is just finding statistical correlations within the data to the point of over-fitting and then early stopping before validation performance gets any worse. It can also be proven that given enough hidden units, a single layer MLP can approximate any function =]. But sometimes the 'art' is about finding helpful ways to assist or even coax the net into behaving nicely. Adding things like part embeddings, as a form a clustering, help the net learn better to associate questions within the same part (perhaps the user sucks at the skills necessary for some part and so user-responses belonging to it suffer), even though as you mention, an exercise embedding at the content_id level can/should/does do this as well. Best to try w/ and w/o the part embedding and you'll see it'll help improve validation score before over-fitting occurs =].</p>\n<p>Easy way to think about it—imagine you have only 10 / 100M responses to a particular question. With a 128 or 256 content_id embedding, these few samples might not be sufficient for the net to properly learn to associate those questions with the appropriate part, without our assistance. In our dataset, we have a number of questions that are even just asked a single time!</p>",
          "rawMarkdown": "At the end of the day, DL is just finding statistical correlations within the data to the point of over-fitting and then early stopping before validation performance gets any worse. It can also be proven that given enough hidden units, a single layer MLP can approximate any function =]. But sometimes the 'art' is about finding helpful ways to assist or even coax the net into behaving nicely. Adding things like part embeddings, as a form a clustering, help the net learn better to associate questions within the same part (perhaps the user sucks at the skills necessary for some part and so user-responses belonging to it suffer), even though as you mention, an exercise embedding at the content_id level can/should/does do this as well. Best to try w/ and w/o the part embedding and you'll see it'll help improve validation score before over-fitting occurs =].\n\nEasy way to think about it—imagine you have only 10 / 100M responses to a particular question. With a 128 or 256 content_id embedding, these few samples might not be sufficient for the net to properly learn to associate those questions with the appropriate part, without our assistance. In our dataset, we have a number of questions that are even just asked a single time!",
          "votes": 2
        },
        {
          "id": 1138538,
          "postDate": "2021-01-04T18:55:52.577Z",
          "content": "<p>That's some really interesting food for thought. Thanks <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> :)</p>",
          "rawMarkdown": "That's some really interesting food for thought. Thanks @authman :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1133918,
      "postDate": "2020-12-31T16:35:58.110Z",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> does  LR scheduler type make a difference ,</p>",
      "rawMarkdown": "@abdurrafae does  LR scheduler type make a difference ,",
      "replies": [
        {
          "id": 1133939,
          "postDate": "2020-12-31T17:09:29.180Z",
          "content": "<p>I've only used the Noam LR so can't compare.</p>",
          "rawMarkdown": "I've only used the Noam LR so can't compare."
        }
      ]
    },
    {
      "id": 1128432,
      "postDate": "2020-12-27T12:41:03.187Z",
      "content": "<p>My Saint Benchmark<br>\nCV 0.789 LB 0.783</p>",
      "rawMarkdown": "My Saint Benchmark\nCV 0.789 LB 0.783\n"
    },
    {
      "id": 1123628,
      "postDate": "2020-12-23T11:36:32.517Z",
      "content": "<p>what a cool work!</p>",
      "rawMarkdown": "what a cool work!"
    },
    {
      "id": 1122162,
      "postDate": "2020-12-22T08:36:59.323Z",
      "content": "<p>Dose anyone know how to convert prior_elasped_time to continuous embedding as input of Transformer?</p>",
      "rawMarkdown": "Dose anyone know how to convert prior_elasped_time to continuous embedding as input of Transformer?",
      "replies": [
        {
          "id": 1122185,
          "postDate": "2020-12-22T08:52:41.020Z",
          "content": "<p>Try this:</p>\n<pre><code>if elapsed_time_cat is True:\n    elapsed_time_embedding = nn.Embedding(elapsed_time_size, embedding_dim)\nelse:\n    elapsed_time_embedding = nn.Linear(1, embedding_dim, bias=False)\n</code></pre>",
          "rawMarkdown": "Try this:\n\n```\nif elapsed_time_cat is True:\n\telapsed_time_embedding = nn.Embedding(elapsed_time_size, embedding_dim)\nelse:\n\telapsed_time_embedding = nn.Linear(1, embedding_dim, bias=False)\n```",
          "votes": 2
        },
        {
          "id": 1122194,
          "postDate": "2020-12-22T08:56:52.120Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>  very thank you!!</p>",
          "rawMarkdown": "@mpware  very thank you!!"
        },
        {
          "id": 1122689,
          "postDate": "2020-12-22T16:15:35.640Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <br>\nMy Transformer model LB : 0.783 without time features. In first epoch, the AUC of validation set is 0.762. However, when I add continuous embedding(prior_elasped_time), the first epoch become 0.622 and AUC improve very slow during training. Do you have the same experience? </p>",
          "rawMarkdown": "Hi @mpware \nMy Transformer model LB : 0.783 without time features. In first epoch, the AUC of validation set is 0.762. However, when I add continuous embedding(prior_elasped_time), the first epoch become 0.622 and AUC improve very slow during training. Do you have the same experience? ",
          "replies": [
            {
              "id": 1122711,
              "postDate": "2020-12-22T16:39:32.227Z",
              "content": "<p>I've never had 0.762 on first epoch but 0.62 is closer to what I got.</p>",
              "rawMarkdown": "I've never had 0.762 on first epoch but 0.62 is closer to what I got.",
              "votes": 1
            }
          ]
        },
        {
          "id": 1122747,
          "postDate": "2020-12-22T17:17:29.293Z",
          "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a> <br>\nHow much was your LB /CV without  any Time. with just Part/Exercise/Response. <br>\nI get 77.3 as LB without Part vs 74.81 cv , when i add Part i get  very good CV of 77.4 but LB not very high. <br>\nAlso if you added part did you face any sub scoring error at first. </p>",
          "rawMarkdown": "@m10515009 \nHow much was your LB /CV without  any Time. with just Part/Exercise/Response. \nI get 77.3 as LB without Part vs 74.81 cv , when i add Part i get  very good CV of 77.4 but LB not very high. \nAlso if you added part did you face any sub scoring error at first. "
        },
        {
          "id": 1122753,
          "postDate": "2020-12-22T17:23:50.847Z",
          "content": "<p>Take this with a grain of salt, as I haven't written an inference pipeline yet so I can't test against LB… but I am using the same val others have said works great. At the end of my first epoch, my transformer model which makes use of two ts features, hits 0.7733 validation AUC. That seems in-line with what you are seeing <a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F50b80bad9c8c6ab11b35a8fbfbc29376%2Fepoch1.png?generation=1608656664787102&amp;alt=media\" alt=\"\"></p>\n<blockquote>\n  <p>Do you have the same experience? </p>\n</blockquote>\n<p>However I have to add that in order for your question to be meaningful when comparing to others, a few additional details are necessary to share, such as: your batch size, optimizer, and/or lr. In my case, I am consuming batches of 512 through Adam with 1e-3 LR. You?</p>",
          "rawMarkdown": "Take this with a grain of salt, as I haven't written an inference pipeline yet so I can't test against LB... but I am using the same val others have said works great. At the end of my first epoch, my transformer model which makes use of two ts features, hits 0.7733 validation AUC. That seems in-line with what you are seeing @m10515009.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F50b80bad9c8c6ab11b35a8fbfbc29376%2Fepoch1.png?generation=1608656664787102&alt=media)\n\n> Do you have the same experience? \n\nHowever I have to add that in order for your question to be meaningful when comparing to others, a few additional details are necessary to share, such as: your batch size, optimizer, and/or lr. In my case, I am consuming batches of 512 through Adam with 1e-3 LR. You?"
        },
        {
          "id": 1122792,
          "postDate": "2020-12-22T17:53:10.610Z",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>  have u included Part Info to you model.  what is TS.. tgt id/response ? What is Optim u use,  i use simple Adam with 1 cycle LR.</p>",
          "rawMarkdown": "@authman  have u included Part Info to you model.  what is TS.. tgt id/response ? What is Optim u use,  i use simple Adam with 1 cycle LR.\n\n",
          "votes": 1
        },
        {
          "id": 1122872,
          "postDate": "2020-12-22T18:50:47.967Z",
          "content": "<ul>\n<li>parts included</li>\n<li>tags not included</li>\n<li>ts = timestamp</li>\n<li>adam</li>\n<li>no scheduling</li>\n</ul>",
          "rawMarkdown": "- parts included\n- tags not included\n- ts = timestamp\n- adam\n- no scheduling"
        },
        {
          "id": 1123175,
          "postDate": "2020-12-23T01:54:39.550Z",
          "content": "<p>It seems that the model has some problems?<br>\nThe AUC of the validation set is not improvement.</p>\n<pre><code>epoch - 0 train_loss - 0.5145 train_auc - 0.602 val_loss - 0.4125 val_auc - 0.623 time=362.28s\nepoch - 1 train_loss - 0.5068 train_auc - 0.615 val_loss - 0.4115 val_auc - 0.630 time=361.66s\nepoch - 2 train_loss - 0.5050 train_auc - 0.619 val_loss - 0.4112 val_auc - 0.627 time=361.42s\nepoch - 3 train_loss - 0.5040 train_auc - 0.622 val_loss - 0.4090 val_auc - 0.628 time=361.26s\nepoch - 4 train_loss - 0.5033 train_auc - 0.623 val_loss - 0.4087 val_auc - 0.628 time=360.97s\nepoch - 5 train_loss - 0.5026 train_auc - 0.625 val_loss - 0.4082 val_auc - 0.631 time=360.81s\nepoch - 6 train_loss - 0.5021 train_auc - 0.626 val_loss - 0.4080 val_auc - 0.630 time=360.54s\nepoch - 7 train_loss - 0.5017 train_auc - 0.627 val_loss - 0.4080 val_auc - 0.631 time=360.71s\nepoch - 8 train_loss - 0.5013 train_auc - 0.628 val_loss - 0.4078 val_auc - 0.633 time=359.89s\nepoch - 9 train_loss - 0.5009 train_auc - 0.629 val_loss - 0.4083 val_auc - 0.630 time=359.87s\nepoch - 10 train_loss - 0.5006 train_auc - 0.630 val_loss - 0.4079 val_auc - 0.627 time=359.87s\nepoch - 11 train_loss - 0.5003 train_auc - 0.631 val_loss - 0.4080 val_auc - 0.627 time=359.65s\nepoch - 12 train_loss - 0.5000 train_auc - 0.632 val_loss - 0.4079 val_auc - 0.629 time=359.48s\nepoch - 13 train_loss - 0.4997 train_auc - 0.632 val_loss - 0.4083 val_auc - 0.631 time=358.99s\nepoch - 14 train_loss - 0.4994 train_auc - 0.633 val_loss - 0.4079 val_auc - 0.628 time=359.33s\nepoch - 15 train_loss - 0.4992 train_auc - 0.633 val_loss - 0.4084 val_auc - 0.631 time=359.09s\nepoch - 16 train_loss - 0.4989 train_auc - 0.634 val_loss - 0.4079 val_auc - 0.628 time=358.95s\nepoch - 17 train_loss - 0.4987 train_auc - 0.635 val_loss - 0.4080 val_auc - 0.628 time=359.72s\nepoch - 18 train_loss - 0.4984 train_auc - 0.635 val_loss - 0.4084 val_auc - 0.628 time=360.06s\nepoch - 19 train_loss - 0.4981 train_auc - 0.636 val_loss - 0.4080 val_auc - 0.630 time=359.68s\nepoch - 20 train_loss - 0.4978 train_auc - 0.637 val_loss - 0.4081 val_auc - 0.628 time=358.90s\n</code></pre>\n<p>The parameters are as follows<br>\nbatch size: 512<br>\nd_model : 256<br>\nlr :1e-3  (Adam)</p>\n<hr>\n<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> yes, CV: 0.783/ LB: 0.784  with just Part/Exercise/Response.<br>\n<a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Do you have the same problem by using continuous embedding?</p>",
          "rawMarkdown": "It seems that the model has some problems?\nThe AUC of the validation set is not improvement.\n```\nepoch - 0 train_loss - 0.5145 train_auc - 0.602 val_loss - 0.4125 val_auc - 0.623 time=362.28s\nepoch - 1 train_loss - 0.5068 train_auc - 0.615 val_loss - 0.4115 val_auc - 0.630 time=361.66s\nepoch - 2 train_loss - 0.5050 train_auc - 0.619 val_loss - 0.4112 val_auc - 0.627 time=361.42s\nepoch - 3 train_loss - 0.5040 train_auc - 0.622 val_loss - 0.4090 val_auc - 0.628 time=361.26s\nepoch - 4 train_loss - 0.5033 train_auc - 0.623 val_loss - 0.4087 val_auc - 0.628 time=360.97s\nepoch - 5 train_loss - 0.5026 train_auc - 0.625 val_loss - 0.4082 val_auc - 0.631 time=360.81s\nepoch - 6 train_loss - 0.5021 train_auc - 0.626 val_loss - 0.4080 val_auc - 0.630 time=360.54s\nepoch - 7 train_loss - 0.5017 train_auc - 0.627 val_loss - 0.4080 val_auc - 0.631 time=360.71s\nepoch - 8 train_loss - 0.5013 train_auc - 0.628 val_loss - 0.4078 val_auc - 0.633 time=359.89s\nepoch - 9 train_loss - 0.5009 train_auc - 0.629 val_loss - 0.4083 val_auc - 0.630 time=359.87s\nepoch - 10 train_loss - 0.5006 train_auc - 0.630 val_loss - 0.4079 val_auc - 0.627 time=359.87s\nepoch - 11 train_loss - 0.5003 train_auc - 0.631 val_loss - 0.4080 val_auc - 0.627 time=359.65s\nepoch - 12 train_loss - 0.5000 train_auc - 0.632 val_loss - 0.4079 val_auc - 0.629 time=359.48s\nepoch - 13 train_loss - 0.4997 train_auc - 0.632 val_loss - 0.4083 val_auc - 0.631 time=358.99s\nepoch - 14 train_loss - 0.4994 train_auc - 0.633 val_loss - 0.4079 val_auc - 0.628 time=359.33s\nepoch - 15 train_loss - 0.4992 train_auc - 0.633 val_loss - 0.4084 val_auc - 0.631 time=359.09s\nepoch - 16 train_loss - 0.4989 train_auc - 0.634 val_loss - 0.4079 val_auc - 0.628 time=358.95s\nepoch - 17 train_loss - 0.4987 train_auc - 0.635 val_loss - 0.4080 val_auc - 0.628 time=359.72s\nepoch - 18 train_loss - 0.4984 train_auc - 0.635 val_loss - 0.4084 val_auc - 0.628 time=360.06s\nepoch - 19 train_loss - 0.4981 train_auc - 0.636 val_loss - 0.4080 val_auc - 0.630 time=359.68s\nepoch - 20 train_loss - 0.4978 train_auc - 0.637 val_loss - 0.4081 val_auc - 0.628 time=358.90s\n```\nThe parameters are as follows\nbatch size: 512\nd_model : 256\nlr :1e-3  (Adam)\n\n----\n\n@jaideepvalani yes, CV: 0.783/ LB: 0.784  with just Part/Exercise/Response.\n@mpware Do you have the same problem by using continuous embedding?"
        },
        {
          "id": 1123251,
          "postDate": "2020-12-23T04:42:22.387Z",
          "content": "<p>Reduce the learning rate (5e-4 or 1e-4) and it should work. Transformers are notorious for being too dependent on a good lr.</p>",
          "rawMarkdown": "Reduce the learning rate (5e-4 or 1e-4) and it should work. Transformers are notorious for being too dependent on a good lr.",
          "votes": 1
        },
        {
          "id": 1123285,
          "postDate": "2020-12-23T05:20:22.527Z",
          "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a>  thanks.. do you initialize the embedding weights also ?</p>",
          "rawMarkdown": "@m10515009  thanks.. do you initialize the embedding weights also ?"
        },
        {
          "id": 1123778,
          "postDate": "2020-12-23T13:48:07.677Z",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> I train the model from scratch.</p>",
          "rawMarkdown": "@jaideepvalani I train the model from scratch."
        },
        {
          "id": 1124578,
          "postDate": "2020-12-24T04:12:57.107Z",
          "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a>  you cv with part is impressive … m not sure what mistake i could be making but i could get only 77.3 cv with part and lb around that. <br>\nwas wondering if you also got similar results in first.<br>\nmy model is simple<br>\n4 heads,2 layers, 128 to 256 dim.  sampling default as used in public kernels</p>",
          "rawMarkdown": "@m10515009  you cv with part is impressive ... m not sure what mistake i could be making but i could get only 77.3 cv with part and lb around that. \nwas wondering if you also got similar results in first.\nmy model is simple\n4 heads,2 layers, 128 to 256 dim.  sampling default as used in public kernels"
        },
        {
          "id": 1135374,
          "postDate": "2021-01-02T07:50:01.863Z",
          "content": "<p>Did you solve the problem? I added lagtime to the decoder, but the results were worse. Maybe I need to try discontinuous embedding</p>",
          "rawMarkdown": "Did you solve the problem? I added lagtime to the decoder, but the results were worse. Maybe I need to try discontinuous embedding"
        }
      ]
    },
    {
      "id": 1119945,
      "postDate": "2020-12-20T13:26:58.133Z",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> , your team's current score LB 0.797, is it a score obtained from combining LGBM with your transformer's model (LB=0.783)?</p>",
      "rawMarkdown": "@mpware , your team's current score LB 0.797, is it a score obtained from combining LGBM with your transformer's model (LB=0.783)?"
    },
    {
      "id": 1097973,
      "postDate": "2020-12-01T11:28:15.390Z",
      "content": "<p>So i was checking my work, it seems that in SAINT they apply masking for every Multi-Head Attention layers in decoder step? Is my understanding correct?</p>",
      "rawMarkdown": "So i was checking my work, it seems that in SAINT they apply masking for every Multi-Head Attention layers in decoder step? Is my understanding correct?",
      "replies": [
        {
          "id": 1098062,
          "postDate": "2020-12-01T12:22:50.817Z",
          "content": "<p>Yes, triu mask on each.</p>",
          "rawMarkdown": "Yes, triu mask on each.",
          "votes": 1
        },
        {
          "id": 1098082,
          "postDate": "2020-12-01T12:30:27.560Z",
          "content": "<p>Does pytorch's masks behave that way in decoder phase? I am not sure about that as it's not what the paper \"Attention Is All you Need\" does. So confused 🤐</p>",
          "rawMarkdown": "Does pytorch's masks behave that way in decoder phase? I am not sure about that as it's not what the paper \"Attention Is All you Need\" does. So confused 🤐"
        },
        {
          "id": 1098166,
          "postDate": "2020-12-01T13:30:47.450Z",
          "content": "<p>They are applying encoding mask also in the encoder because questions are also conditionals to time. In the original Transformer, you have a whole sentence you want to translate. Here, if you didn't decode-mask in the encoder, you would be leaking questions from the future.</p>",
          "rawMarkdown": "They are applying encoding mask also in the encoder because questions are also conditionals to time. In the original Transformer, you have a whole sentence you want to translate. Here, if you didn't decode-mask in the encoder, you would be leaking questions from the future.",
          "votes": 2
        },
        {
          "id": 1098238,
          "postDate": "2020-12-01T14:09:16.583Z",
          "content": "<p>Hmm, I am doing that, have verified the same by manually printing the same after attention. Not sure why i can't pass .74 mark 🥺. Thanks!</p>",
          "rawMarkdown": "Hmm, I am doing that, have verified the same by manually printing the same after attention. Not sure why i can't pass .74 mark 🥺. Thanks!"
        },
        {
          "id": 1098575,
          "postDate": "2020-12-01T17:54:25.457Z",
          "content": "<p>I'm also stuck after reviewing every piece in my code and having tried different approaches. For some reason I can't get my model to improve. I have even adjusted my input flow as <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> suggested, being able to train a whole epoch over 80M rows in ~5 min (random sampling a sequence for every user every epoch). </p>",
          "rawMarkdown": "I'm also stuck after reviewing every piece in my code and having tried different approaches. For some reason I can't get my model to improve. I have even adjusted my input flow as @mpware suggested, being able to train a whole epoch over 80M rows in ~5 min (random sampling a sequence for every user every epoch). ",
          "votes": 1
        },
        {
          "id": 1098639,
          "postDate": "2020-12-01T18:45:34.897Z",
          "content": "<p>I also reviewed a lot of times in my code :( - always found something was wrong each time</p>",
          "rawMarkdown": "I also reviewed a lot of times in my code :( - always found something was wrong each time",
          "votes": 1
        },
        {
          "id": 1101879,
          "postDate": "2020-12-04T11:11:49.110Z",
          "content": "<p>Great going <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> ! I am about to give up over SAINT; Will just try once more from scratch without referring to any old snips :(</p>",
          "rawMarkdown": "Great going @yihdarshieh ! I am about to give up over SAINT; Will just try once more from scratch without referring to any old snips :("
        }
      ]
    },
    {
      "id": 1097346,
      "postDate": "2020-12-01T03:03:09.067Z",
      "content": "<p>If you don't mind, can anyone reveal how high the score can be achieved with a single nn model?👀👀👀</p>",
      "rawMarkdown": "If you don't mind, can anyone reveal how high the score can be achieved with a single nn model?👀👀👀"
    },
    {
      "id": 1096962,
      "postDate": "2020-11-30T22:48:46.227Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> thanks for this information and the time and effort put into it.</p>",
      "rawMarkdown": "Hi @mpware thanks for this information and the time and effort put into it."
    },
    {
      "id": 1094094,
      "postDate": "2020-11-28T10:11:56.640Z",
      "content": "<p>I have one extra concern in addition to time: how much memory is required for full data training in Transformer?</p>",
      "rawMarkdown": "I have one extra concern in addition to time: how much memory is required for full data training in Transformer?",
      "replies": [
        {
          "id": 1094101,
          "postDate": "2020-11-28T10:17:04.413Z",
          "content": "<p>For me, 10% of the data is used to occupy about 2GB of memory space in GPU. I don’t know what the step size the author takes</p>",
          "rawMarkdown": "For me, 10% of the data is used to occupy about 2GB of memory space in GPU. I don’t know what the step size the author takes"
        }
      ]
    },
    {
      "id": 1093402,
      "postDate": "2020-11-27T17:23:36.967Z",
      "content": "<p>I am not sure where there's a bug in my pipeline for me but it seems but content_id's sequence can be quite similar at times… (check user_id's 204790744 and 189703047)</p>\n<p>Attached is what I see,</p>\n<pre><code>{'user_id': 189703047, 'content_id': [7900, 7876, 175, 1278, 2065, 2064, 2063, 3364, 3365, 3363, 2948, 2946, 2947, 2595, 2594, 2593, 4492, 4120, 4696, 6116, 6173, 6370##, 6911, 6910, 6909, 6908, 7219, 7218, 7216, 7217]}\n\n{'user_id': 204790744, 'content_id': [7900, 7876, 175, 1278, 2064, 2063, 2065, 3365, 3364, 3363, 2946, 2947, 2948, 2595, 2594, 2593, 4492, 4120, 4696, 6116, 6173, 6370\", 6879, 6880, 6878, 6877, 7219, 7216, 7218, 7217, 4237, 5675, 6276, 9063, 4740, 5154, 6439, 9163, 5328, 4537]}\n</code></pre>",
      "rawMarkdown": "I am not sure where there's a bug in my pipeline for me but it seems but content_id's sequence can be quite similar at times... (check user_id's 204790744 and 189703047)\n\nAttached is what I see,\n\n```\n{'user_id': 189703047, 'content_id': [7900, 7876, 175, 1278, 2065, 2064, 2063, 3364, 3365, 3363, 2948, 2946, 2947, 2595, 2594, 2593, 4492, 4120, 4696, 6116, 6173, 6370##, 6911, 6910, 6909, 6908, 7219, 7218, 7216, 7217]}\n\n{'user_id': 204790744, 'content_id': [7900, 7876, 175, 1278, 2064, 2063, 2065, 3365, 3364, 3363, 2946, 2947, 2948, 2595, 2594, 2593, 4492, 4120, 4696, 6116, 6173, 6370\", 6879, 6880, 6878, 6877, 7219, 7216, 7218, 7217, 4237, 5675, 6276, 9063, 4740, 5154, 6439, 9163, 5328, 4537]}\n```",
      "replies": [
        {
          "id": 1100422,
          "postDate": "2020-12-03T04:43:34.783Z",
          "content": "<p>I think this is just because first 30 questions for some participants were exactly the same.<br>\nMany of these participants stop answering after these 30 questions.</p>",
          "rawMarkdown": "I think this is just because first 30 questions for some participants were exactly the same.\nMany of these participants stop answering after these 30 questions.",
          "votes": 2
        },
        {
          "id": 1100457,
          "postDate": "2020-12-03T05:12:36.203Z",
          "content": "<p>Yep! I will drop all such records while training and then see what happens! (keep them during inference/test_time)</p>",
          "rawMarkdown": "Yep! I will drop all such records while training and then see what happens! (keep them during inference/test_time)"
        }
      ]
    },
    {
      "id": 1083824,
      "postDate": "2020-11-19T12:29:34.437Z",
      "content": "<p>Can you share <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> as to what's your accuracy stats? For me, it's pretty bad and AUC is around .67 for SAINT+ on unseen users. [WIP] Ty!</p>",
      "rawMarkdown": "Can you share @mpware as to what's your accuracy stats? For me, it's pretty bad and AUC is around .67 for SAINT+ on unseen users. [WIP] Ty!",
      "replies": [
        {
          "id": 1083870,
          "postDate": "2020-11-19T13:30:14.420Z",
          "content": "<p>You have them in the previous screenshot I've posted (for validation).</p>",
          "rawMarkdown": "You have them in the previous screenshot I've posted (for validation).",
          "votes": 1
        },
        {
          "id": 1083880,
          "postDate": "2020-11-19T13:41:51.013Z",
          "content": "<p>Ahh, I see it now; It's great for you!</p>",
          "rawMarkdown": "Ahh, I see it now; It's great for you!",
          "replies": [
            {
              "id": 1083885,
              "postDate": "2020-11-19T13:51:26.177Z",
              "content": "<p>To be more precise:</p>\n<pre><code>Fold 1 train users: 334532 valid users: 14731\nFold 1 train size: (86990390, 6) valid size: (917403, 6)\nvalid_auc - 0.7585, valid_accuracy - 0.743, valid_loss - 0.5247\n</code></pre>\n<p>For validation, I'm using sequence from users with 20 to 100 interactions. I need to add more because it means it only validates on first 100 interactions.</p>",
              "rawMarkdown": "To be more precise:\n\n```\nFold 1 train users: 334532 valid users: 14731\nFold 1 train size: (86990390, 6) valid size: (917403, 6)\nvalid_auc - 0.7585, valid_accuracy - 0.743, valid_loss - 0.5247\n```\n\nFor validation, I'm using sequence from users with 20 to 100 interactions. I need to add more because it means it only validates on first 100 interactions.",
              "votes": 1
            },
            {
              "id": 1085595,
              "postDate": "2020-11-21T02:33:39.110Z",
              "content": "<p>So, I got it to training with few hacky stuffs and few things which are little incorrect but making it work for inference is driving me crazy :( </p>\n<p>No Lag Time as of now and CV is off by 0.005/0.006</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F3376210d62b68b36e83f00fb7d6316ec%2FScreenshot%202020-11-21%20at%208.02.46%20AM.png?generation=1605926010430181&amp;alt=media\" alt=\"\">.</p>",
              "rawMarkdown": "So, I got it to training with few hacky stuffs and few things which are little incorrect but making it work for inference is driving me crazy :( \n\nNo Lag Time as of now and CV is off by 0.005/0.006\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F3376210d62b68b36e83f00fb7d6316ec%2FScreenshot%202020-11-21%20at%208.02.46%20AM.png?generation=1605926010430181&alt=media).\n",
              "votes": 2
            },
            {
              "id": 1085865,
              "postDate": "2020-11-21T08:49:38.607Z",
              "content": "<p>Your training score looks good now. For inference it took me one day of optimization/pipeline re-design to make it works within the time limit. You've to identify where your code is slow (timeit should help) and then find solutions to speed up until acceptable. The solution for me was to drop all my pandas code.</p>",
              "rawMarkdown": "Your training score looks good now. For inference it took me one day of optimization/pipeline re-design to make it works within the time limit. You've to identify where your code is slow (timeit should help) and then find solutions to speed up until acceptable. The solution for me was to drop all my pandas code.",
              "votes": 2
            },
            {
              "id": 1086071,
              "postDate": "2020-11-21T10:42:39Z",
              "content": "<p>Thanks Kaggle for this;  </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Fa6ff55df170ca971c2bb7b0a12b4692d%2FScreenshot%202020-11-21%20at%204.08.35%20PM.png?generation=1605955161875797&amp;alt=media\" alt=\" \"></p>\n<blockquote>\n  <p>The solution for me was to drop all my pandas code.</p>\n</blockquote>\n<p>Yep, no pandas for me as well except for filtering the col once and a map, can be circumvented if needed!</p>\n<p>It should be fast enough as for 4 iters it's just 1.03 secs, right?  But it's failing so it doesn't matter. 🤐😑🤕🥺. Hard limit someone reported was .55 secs/iters  ~2.2 secs for 4 iters. Anything more than this will result in Timeout.</p>\n<p>NB, It's a dummy model with 4 layers and 32 hidden dims, that explains the fastness i am seeing.😁</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F26b23e6d7794f784c05b43e6ddd6aa30%2FScreenshot%202020-11-21%20at%204.10.00%20PM.png?generation=1605955234630126&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Thanks Kaggle for this;  \n\n![ ](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Fa6ff55df170ca971c2bb7b0a12b4692d%2FScreenshot%202020-11-21%20at%204.08.35%20PM.png?generation=1605955161875797&alt=media)\n\n>The solution for me was to drop all my pandas code.\n\nYep, no pandas for me as well except for filtering the col once and a map, can be circumvented if needed!\n\nIt should be fast enough as for 4 iters it's just 1.03 secs, right?  But it's failing so it doesn't matter. 🤐😑🤕🥺. Hard limit someone reported was .55 secs/iters  ~2.2 secs for 4 iters. Anything more than this will result in Timeout.\n\nNB, It's a dummy model with 4 layers and 32 hidden dims, that explains the fastness i am seeing.😁\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F26b23e6d7794f784c05b43e6ddd6aa30%2FScreenshot%202020-11-21%20at%204.10.00%20PM.png?generation=1605955234630126&alt=media)",
              "votes": 1
            },
            {
              "id": 1086078,
              "postDate": "2020-11-21T10:49:19.803Z",
              "content": "<p>Have you measured your model inference time per batch (about batch size 30~40)?</p>",
              "rawMarkdown": "Have you measured your model inference time per batch (about batch size 30~40)?"
            },
            {
              "id": 1086080,
              "postDate": "2020-11-21T10:52:02.230Z",
              "content": "<p>You got error after 9 hours or earlier? Might also be a bug in the code ….</p>",
              "rawMarkdown": "You got error after 9 hours or earlier? Might also be a bug in the code ...."
            },
            {
              "id": 1086085,
              "postDate": "2020-11-21T10:56:59.260Z",
              "content": "<p>Yep, way less than 9 hours, so logic issue as of now..</p>\n<blockquote>\n  <p>Have you measured your model inference time per batch (about batch size 30~40)?</p>\n</blockquote>\n<p>Not quite accurate, but it's around 30-40 msecs, only making predictions.</p>",
              "rawMarkdown": "Yep, way less than 9 hours, so logic issue as of now..\n\n>Have you measured your model inference time per batch (about batch size 30~40)?\n\nNot quite accurate, but it's around 30-40 msecs, only making predictions."
            },
            {
              "id": 1086147,
              "postDate": "2020-11-21T12:36:58.803Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 1087072,
              "postDate": "2020-11-22T10:08:15.260Z",
              "content": "<p>Isn't a great score but I am happy about it as well. After 2 days of white board coding (and sleepless nights), another day to write the inference pipeline and 3 failed submissions wrt Transformer's;  Now, I will try to close in the bugs I am aware of.</p>\n<p>Would just like to add one more tip here, if we do things correctly with a proper design first and dry runs in our minds as to how/what all we need to take care of, the solution will reveal itself pretty much.</p>\n<blockquote>\n  <p>Have you measured your model inference time per batch (about batch size 30~40)?</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>, It's pretty fast on GPUs. &lt;2-2.5 hours for me.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Fa14d06d57a62981f8d91d1234b9c227f%2FScreenshot%202020-11-22%20at%203.35.30%20PM.png?generation=1606039568359206&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Isn't a great score but I am happy about it as well. After 2 days of white board coding (and sleepless nights), another day to write the inference pipeline and 3 failed submissions wrt Transformer's;  Now, I will try to close in the bugs I am aware of.\n\nWould just like to add one more tip here, if we do things correctly with a proper design first and dry runs in our minds as to how/what all we need to take care of, the solution will reveal itself pretty much.\n\n>Have you measured your model inference time per batch (about batch size 30~40)?\n\n@yihdarshieh, It's pretty fast on GPUs. <2-2.5 hours for me.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Fa14d06d57a62981f8d91d1234b9c227f%2FScreenshot%202020-11-22%20at%203.35.30%20PM.png?generation=1606039568359206&alt=media)",
              "votes": 1
            }
          ]
        },
        {
          "id": 1108903,
          "postDate": "2020-12-11T06:04:18.300Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> why your val loss is high.. i get arund .3…  using simple bce loss with logits</p>",
          "rawMarkdown": "@adityaecdrid why your val loss is high.. i get arund .3...  using simple bce loss with logits"
        }
      ]
    },
    {
      "id": 1142042,
      "postDate": "2021-01-07T05:23:00.030Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1087159,
      "postDate": "2020-11-22T12:19:50.060Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1084568,
      "postDate": "2020-11-20T06:55:38.887Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1095395,
      "postDate": "2020-11-29T15:08:07.370Z",
      "content": "<p>thank you for sharing!</p>",
      "rawMarkdown": "thank you for sharing!",
      "votes": -1
    },
    {
      "id": 1124810,
      "postDate": "2020-12-24T08:08:16.417Z",
      "content": "<p>Thanks for sharing and updating! </p>",
      "rawMarkdown": "Thanks for sharing and updating! "
    }
  ],
  "comments": [
    {
      "id": 1107283,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2020-12-09T15:03:30.890000",
      "content": "<p>Hey guys, I finally got it. Still not trained with all the data, I will when I have more RAM (it is suposed to be coming). Single model very close to SAINT+:</p>\n<ul>\n<li>[Epochs: 24] loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768</li>\n<li>LB: 0.781</li>\n</ul>\n<p>Ask me anything about it. Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.</p>",
      "votes": 16,
      "replies": [
        {
          "id": 1107290,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-09T15:10:02.763000",
          "content": "<p>Wow! This is amazing score <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> ❤️🎉🎉! Your loss and auc gap is also really nice!</p>\n<blockquote>\n  <p>Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.</p>\n</blockquote>\n<p>Maybe, that's the crus, but can you give pointers for the same? Ty a lot! It's fine if we discuss after the comp as well :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107295,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-09T15:12:48.367000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Hey good news for you! I'm picking user's sequence randomly on each epoch such as:<br>\nUser1: [Qn, Qn+100]<br>\nUser2: [Qm, Qm+100]<br>\n…<br>\nWhat about you?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1107297,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-09T15:13:42.657000",
          "content": "<blockquote>\n  <p>Hey guys, I finally got it. Still not trained with all the data, I will when I have more RAM (it is suposed to be coming). Single model very close to SAINT+:</p>\n  <ul>\n  <li>[Epochs: 24] loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768</li>\n  <li>LB: 0.781</li>\n  </ul>\n  <p>Ask me anything about it. Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.</p>\n</blockquote>\n<p>Great to see you up to LB, good job <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> !<br>\nI have similar results - however, in order to get 0.78, I need to </p>\n<ol>\n<li>reduce lr to smaller and train longer (lr = 1e-4 - 5e-5)</li>\n<li>use tag information</li>\n</ol>\n<p>Do you also need these 2 steps to get 0.78? And what kind of sampling strategy boost your LB score?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 1108935,
              "author_name": "Jaideep",
              "author_url": "",
              "post_date": "2020-12-11T06:55:13.037000",
              "content": "<blockquote>\n  <blockquote>\n    <p>Hey guys, I finally got it. Still not trained with all the data, I will when I have more RAM (it is suposed to be coming). Single model very close to SAINT+:</p>\n    <ul>\n    <li>[Epochs: 24] loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768</li>\n    <li>LB: 0.781</li>\n    </ul>\n    <p>Ask me anything about it. Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.</p>\n  </blockquote>\n  <p>Great to see you up to LB, good job <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> !<br>\n  I have similar results - however, in order to get 0.78, I need to </p>\n  <ol>\n  <li>reduce lr to smaller and train longer (lr = 1e-4 - 5e-5)</li>\n  <li>use tag information</li>\n  </ol>\n  <p>Do you also need these 2 steps to get 0.78? And what kind of sampling strategy boost your LB score?</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> <br>\nare you using plain bce loss or with some modification as i get very less loss for val in range 0.3sss for Auc you see.  may be mask bce ?<br>\nWat is custom Auc here</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1107299,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-09T15:18:22.860000",
          "content": "<p>Nice! 👍</p>\n<p>Any key points other than a good sampling strategy to keep in mind? Plus what is it? I have seen that taking random sequences of (window length) from users with longer interactions helps.</p>\n<p>Does the size of d_model and num_layers impact the model significantly?</p>\n<p>I got 0.775 with SAINT+ in it's exact form with 2 num_layers and 128 d_model size.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1107317,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-09T15:35:32.383000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> I've seen that my models usually converge around 20-30 epochs on average. How much an improvement is seen on lowering lr and training for longer.</p>\n<p>Only using Kaggle GPU has really constrained testing for different hyperparameters, my current focus is on getting a good model and then optimizing in the last days of the competition.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1107319,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T15:37:28.367000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> <a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a></p>\n<p>Sampling strategy:</p>\n<ul>\n<li>Take all user ids in training data.</li>\n<li>Compute their lengths and cap the result to 500.</li>\n<li>Scale the prior to make them sum 1 (probability distribution).</li>\n<li>Sample N ids with replacement with previous computed probabilities. N in my case is the same as different user ids you have in the training data. The result will have repeated ids.</li>\n<li>Take a random sequence for every id. It may lead to repeated sequences with low probability. For example, if you have id 115 repeated 7 times, you will take 7 random sequences for the user 115.</li>\n<li>Repeat each epoch.</li>\n</ul>\n<p>Why? </p>\n<ul>\n<li>If you take a random sequence for every user every epoch, you will overfit towards starter sequences, given that the majority of the users use the application only a few times. </li>\n<li>Cap to 500 before making the prob distribution denies super prolific users (outliers) from having a too big presence in the sampling.   </li>\n</ul>\n<p>Let me know if you need more details.</p>",
          "votes": 19,
          "replies": []
        },
        {
          "id": 1107324,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T15:40:32.897000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a></p>\n<p>I use every feature they use in SAINT+ paper with slightly modifications in them but not adding new ones. Learning rate schedule is exactly as described in SAINT+ (originally from Attention is All You Need).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1107331,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T15:43:34.243000",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a></p>\n<p>I haven't tried to grow my model up. 2 encoder layers, 2 decoder layers and model size of 128.</p>\n<pre><code>train_ratio = 0.96\nwindows_size = 96\nepochs = 100\npatience = 3\nd_model = 128\nnum_heads = 4\nn_encoder_layers = 2\nn_decoder_layers = 2\nbatch_size = 256\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1107332,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T15:44:59.897000",
          "content": "<p>For all of you, if you get to implement the sampling strategy, let me know if it works for you. I have worked a lot (a lot) on it.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1107338,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-09T15:52:47.393000",
          "content": "<p>Would you share the way you compute the lag time? It is still not clear to me how to compute it. Do you just use the current timestamp - the previous timestamp?</p>\n<p>If so:<br>\n    Q1: you will get 0 for question in a bundle (except for the 1st question in that bundle)??<br>\n    Q2: If a question follows a lecture, I think the above computation is not the lag time described in the paper. Or maybe you ignore all the lectures, so don't have care about lecture during computing lag time?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1107391,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T16:56:09.737000",
          "content": "<p>That's one of the differences with SAINT+ I have. I don't compute lag as they describe it. I mean, I just compute the time between the start of two consecutive interactions, capping it to one day and scaling it [0, 1].  If you are feeding the model also with the other time feature, the model will agg that information as he pleases.</p>\n<p>I simply do:</p>\n<pre><code>user['timestamp'] = (user['timestamp'].diff().fillna(0)/8.64e+7).clip(upper=1)\n</code></pre>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1107400,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-09T17:05:12.500000",
          "content": "<p>Thank you. Using smaller lr gives me from LB 0.778 to 0.781 (best CV epoch : 6 --&gt; 30). Not that much though.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1107430,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T17:25:37.503000",
          "content": "<p>Hey mate what you said is true. I didn't even think about it. I'm probably getting zeroes for non-first questions in a bundle, though in inference time having the good diff.</p>\n<p>Anyway, there is a similar problem around which I find even worse. Imagine I have a sequence [A, B, C1, C2, C3], being C three questions in the same bundle. In training time I just let them as they are, but in inference time I do [A, B, C1], [A, B, C2], [A, B, C3]. Solving this I think could boost the performance by a lot, though I feel to lazy to set it up in training time.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1107459,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-09T17:47:39.613000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> -- Yes, that is a major issue (in another way) I spent time to fix.</p>\n<p>Just like you, in training time, for me,  everything is in the normal/usual setting. I think it is too difficult to change the training time setting as in the inference time - unless you want to train with only the last few questions in the last bundle in each sequence. The model needs to see about 100 more times of the number of sequences to converge, which will be too time consuming.</p>\n<p>For inference, I do have something like <code>[A, B, C1], [A, B, C2], [A, B, C3].</code>, i.e. we don't use auto-regressive generation. Due to this, we need to keep the 2nd - the last question in the same bundle as the 1st question in the bundle. However, I keep the same questions (for prediction) for each user in a single sequence, but using special mask -  the 2nd question can't attend to the 1st question in the bundle during the inference time, etc. The same applies for the positional information. For example, I need to change [0, 1, 2, 3, 4] to [0, 1, 2, 2, 2].</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1107478,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-09T18:13:41.480000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>, your sampling strategy seems very reasonable to me. However, it is somehow difficult for me to change the sampling - all the examples are store in tfrecord dataset, and using tf.data.Dataset as the input pipeline. I can do random subsequence selection, but not the user selection, seems the tf.data.Dataset will just yield the sequences from all users.</p>\n<p>The solution I have in mind is to calculate the probability (the one you calculated) for different users. And use it as a loss weight. For example, for a user with much shorter sequence - the probably you calculated is smaller. And it will be selected more often (if the proability is not applied to select sequences from different users) - but if I multiply the probably in loss calculation, it should have the same effect as <code>select few examples for that user using the smaller probability</code></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1107492,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-09T18:24:24.917000",
          "content": "<p>I have a couple of questions, Hope you guys won't mind answering them!</p>\n<p>First is regarding the sampling strategy shared by <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>. Is he attaching a probability for each user so that they get selected at random? Really clever idea, but not sure how to calculate that? Just by passing something to a distribution for a starter?</p>\n<p>Secondly, this</p>\n<blockquote>\n  <p>Imagine I have a sequence [A, B, C1, C2, C3], being C three questions in the same bundle. In training time I just let them as they are, but in inference time I do [A, B, C1], [A, B, C2], [A, B, C3]</p>\n</blockquote>\n<p>I am not sure i understood it. So, basically you guys are splitting the sequence of <code>[A, B, C1, C2, C3]</code> to <code>[A, B, C1], [A, B, C2], [A, B, C3]</code>. That's again very smart to do but I am not sure how can I append this useful insight to my inference without making preds on rows grouped by user_id let's say in inference?</p>\n<p>And then the below, need to think about this;</p>\n<blockquote>\n  <p>For inference, I do have something like [A, B, C1], [A, B, C2], [A, B, C3]., i.e. we don't use auto-regressive generation. Due to this, we need to keep the 2nd - the last question in the same bundle as the 1st question in the bundle. However, I keep the same questions (for prediction) for each user in a single sequence, but using special mask - the 2nd question can't attend to the 1st question in the bundle during the inference time, etc.</p>\n</blockquote>\n<p>This, i need to think over as of now as i guess, i am losing the context of this being discussed in past as well, in this or some other discussion as well.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107493,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-09T18:24:48.780000",
          "content": "<p>I think this competition requires equal amount of effort on architecture design and data management. It's all fun and enjoyable till to get into a situation where debugging for a week leads you no where and you ultimately revert the change you initially made. :/</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1107508,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-09T18:37:59.950000",
          "content": "<p>Off topic but I would like to share this with you, avoid Kaggle Docker version 90 if you face time out, I've lost 3 days on it: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1107497\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1107497</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107520,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T18:46:15.863000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> We were talking there about the differences between training time and inference time regarding to questions in the same bundle ID AKA same exam. You don't know their results until they are all finished, so we shouldn't be training as we do.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107524,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T18:48:27.193000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> I'm just using my own Dockerfile trying to keep the same major versions as in the Kaggle image. I thank you anyways, I'll try to skip it when training here in Kaggle notebooks (this could be in fact be happening to me in Casava classification competition in which I'm training in Kaggle notebooks).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107531,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T18:51:31.333000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> nvm you were talking about inference. I will have a look.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107533,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T18:56:52.707000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Sorry for the spam, but you were right, seems like I'm using v90. How can we rollback here? I noticed in fact that my running time increased with no apparently reason. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107536,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-09T18:58:44.877000",
          "content": "<p>Check the Env section in settings in interactive kernel mode, by default, i always keep it to the same version to which i made a sub with. (i.e. no upstream updates to the docker <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Fc95de2b124b677f4da391a2b77a9fa7d%2FScreenshot%202020-12-10%20at%2012.29.58%20AM.png?generation=1607540428537812&amp;alt=media\" alt=\"\">file)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107540,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-09T19:01:58.723000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> - I am not sure how other people doing it. The objective is: for each question in a bundle that we want to predict, we only used the history before the current question bundle. The reason: say for the 2nd question in the current bunlde to predict, some information for the 1st question in the same bundle is not available (for example, answer correctness). We use this position while prediction for the 2nd question.</p>\n<p>Ideally, we would like to use auto-regressive generation, i.e. we predict the results for the 1st question, and use it for prediction 2nd question etc. However it is time consuming and not so easy to implement in a efficient way (we have time limit for this competition). That's why we don't use this approach (at least, no notebook published for it).</p>\n<p>However, split to [A, B, C1], [A, B, C2], [A, B, C3], I don't like the idea - because it increase the number of sequences to compute the predictions.</p>\n<p>That's why I still keep [A, B, C1, C2, C3] , but modify some input information like positional info and make sure the attention mask also makes sense due to this constraint.</p>\n<p>If you don't do any of the above 2 approach, how do you perform inference? Unless you use encoder-only model. <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> and me use encoder-decoder architecture.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1107551,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-09T19:09:08.083000",
          "content": "<p>Thanks a lot for the clarifications! Now I can see where you guys are heading! Cool ideas :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107555,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-09T19:14:48.343000",
          "content": "<p>Would you mind to share what is your current approach? BTW, great to see you up on LB also!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107567,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-09T19:25:25.687000",
          "content": "<p>Well, I am way behind when it comes to SAINT, gave up on it for a week now; Switched to gbms with hand-crafted features and experiments on them as of now. (Current LB score is with a gbm, SAINT is at ~.74). But after seeing today's spike in discussions wrt SAINT, got some new inspirations, so I am going to try it a couple of more times!</p>\n<p>In my understanding, I am exactly doing this one for simplicity,</p>\n<blockquote>\n  <p>That's why I still keep [A, B, C1, C2, C3] , but modify some input information like positional info and make sure the attention mask also makes sense due to this constraint.</p>\n</blockquote>\n<p>For the above thing actually, I am <strong>not</strong> even changing my positional info (thanks for the tip!) and regarding the attention mask, I just (an hour back) found a bug while making inference, so I will get back hopefully with a better score than .74 in 1-2 days; 😊</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107572,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T19:29:19.663000",
          "content": "<p>BTW the next thing I'm going to work in is a feature like question_already_answered. This will give information beyond the sequence. Have any of you implemented this? Have it boosted your model?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107574,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T19:33:48.710000",
          "content": "<p>Oh, yes, in fact I have a feature that gives information about bundles, the task_container_id. Forgot to mention, I'm currently using it as embeddings with input [0, 10000] and adding them to the decoder (though I maybe should add them to both encoder and decoder).</p>\n<p>Here it is the snippet, hope it is useful:</p>\n<pre><code>e = tf.keras.layers.Add()([\n    content_emb,\n    part_emb\n])\n\nd = tf.keras.layers.Add()([\n    answered_correctly_emb,\n    task_container_emb,\n    time_features\n])\n\nfor _ in range(n_encoder_layers):\n    e = EncoderLayer(\n        d_model=d_model, \n        num_heads=num_heads, \n        dff=d_model*2, \n        rate=0.01,\n        rel_pos_enc=True\n    )(e, mask=mask)\n\nfor _ in range(n_decoder_layers):\n    d, _, _ = DecoderLayer(\n        d_model=d_model, \n        num_heads=num_heads, \n        dff=d_model*2, \n        rate=0.01, \n        rel_pos_enc=True\n    )(d, e, look_ahead_mask=mask, padding_mask=mask)\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1107575,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-09T19:33:50.533000",
          "content": "<p>I plan to - but I might try pretraining the encoder with MLM loss first.<br>\nI am busy on preparing a TPU training notebook and will publish it before the end of week.<br>\nCurrently, I am running a few experiment like small model (you socre 0.781 with small model scare me already). Then I will try your sampling.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1107577,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-09T19:35:38.020000",
          "content": "<p>I guess It's implemented and discussed <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266#1069347\" target=\"_blank\">here</a>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107583,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-09T19:42:00.063000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> - The task_container_id plays a role similar to absolute position - but a bit different it contains bundle information (so you do have some kind of absolute pos along with your relative pos ).</p>\n<p>I also tried absolute pos - but it gets lower socre (if trained at the same epochs) Considering you train up to 25 or more epochs, maybe it is good for me to try training more epochs also :) while using abs pos.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107584,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T19:42:07.483000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> Thanks, I will take a look.<br>\n<a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> I'm training at the moment a bigger model (3 enc, 3 dec, 256 d_model). I will share an update tomorrow when I submit it. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1107586,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T19:43:23.753000",
          "content": "<p>I train until EarlyStopping triggers, validating every 2 epochs, with a patience of 3 (6 epochs total).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107590,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T19:47:39.637000",
          "content": "<p>Fuck <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>, have you implemented that bits thing shit? Sounds crazy.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107592,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-09T19:50:56.313000",
          "content": "<p>So at about which epoch you get the best CV?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107595,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-09T19:54:24.763000",
          "content": "<p>The one I shared, 24th for my best model.</p>\n<p>24/100 loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107649,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-09T20:47:57.073000",
          "content": "<p>The only way to rollback is to upload your kernel code in on old kernel using the previous docker image. But you need at least one kernel with the old image.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107733,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-09T23:14:26.667000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a></p>\n<p>Before your sampling strategy, did you get the best CV in at a much earlier epoch? I tried smaller model (as yours), with the learning rate in SANIT(+) and batch 256 as yours, with absolute pos (not task_container_id, just the pos in the full history) --&gt; CV epoch 15 &gt; epoch 20 &gt; epoch 25, and all of them is worse than <code>smaller lr at epoch 10 without abs. pos</code>.</p>\n<p>I hope I can find the cause after I finished the TPU notebook - and definitely need try your sampling!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107886,
          "author_name": "Gannon Reynolds",
          "author_url": "",
          "post_date": "2020-12-10T03:21:21.797000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> </p>\n<p>\"BTW the next thing I'm going to work in is a feature like question_already_answered. This will give information beyond the sequence. Have any of you implemented this? Have it boosted your model?\"</p>\n<p>I kept track of attempts for each user on each question on my tabular model and it boosted the score quite a bit. I've been curious how useful it would to a transformer. I would assume the attention likely captures that information.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107992,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-10T06:55:01.207000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a></p>\n<p>I must be too tired for the work on TPU notebook. I actually have a similar sampling as yours, as I mentioned earlier in another comment</p>\n<pre><code>I used to use random sample subsequence from users' history. But I found that potentially I overfit for users with much shorter sequences. For example, if a user has only 50 interaction. In each epoch, it will get the full history. While for a user having 1000 interaction, it only get one subsequence in a epoch.\n</code></pre>\n<p>The difference for my current sampling and yours:</p>\n<ul>\n<li><p>Instead of using probability, I use tf.data.Dataset to repeat user sequences. For example, if user X has a history of length <code>N</code>, it will be repeated for <code>N / W</code> times (<code>W</code> = window size). And I sample subsequences of length <code>W</code> from these <code>N / W</code> full sequences. For each epoch, I get about 1M sequences (while you get about 0.38M sequences). Therefore, your epoch 24 is somehow equivalent to my epoch 8. And it explains why we get the best CV at different epoch :)</p></li>\n<li><p>You have a control of <code>500</code> to avoid sampling too many times from particular users - while I don't do this. I will give it a try, since it should be quite easy to implement.</p></li>\n</ul>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1108007,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-10T07:11:14.350000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a></p>\n<p>Considering we use similar sampling (other than the upper limit of 500 or not), we have quite different results (with you model size, I get only 0.773), it makes thinking what's wrong in my model.</p>\n<p>I saw that in an earlier comment, you mentioned</p>\n<pre><code>80 M rows minus lectures (311567 users).\n</code></pre>\n<p>Is this the case at this moment? I do keep lecture, and pay attention to them. Maybe I should try not to pay attention to it - or even remove it (but it is not easy to change  my pipeline to remove lectures …).</p>\n<p>Do you have a reason not to use lectures?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1108043,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-10T08:01:25.010000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> I just woke up. Seems like we actually have a similar sampling strategy, the difference is that I may not see some of the users. If I understood correctly, you have 3 users like:</p>\n<p>[A, B, C, D]<br>\n[A]<br>\n[A, B]</p>\n<p>and do something like</p>\n<p>[A, B] [C, D]<br>\n[A, -]<br>\n[A, B]</p>\n<p>Don't you see any problem with that? You have 3 starter sequences even if they come from different users, which I pretended to smooth with my sampling strategy.</p>\n<p>On the other hand, I'm still with 80 M rows minus lectures (311567 users). But you just gave me an idea: in theory even if you predicted something idiotic for the lecture step, it should help the model in the way that it knows that <em>the user watched a lecture</em>, which is itself an important feature. So I think you are right there and I'm wrong. I will try that in the incoming days and come with the results.</p>\n<p>Finally, I just finished to train my bigger model, in which I also modified the way I compute the timestamp from an idea taken some comments above.</p>\n<pre><code>user['timestamp'] = (\n  user['timestamp'].diff().replace(0, method='ffill').fillna(0)/8.64e+7\n).clip(upper=1)\n</code></pre>\n<p>So far so good:</p>\n<p>[Epoch 28/100] loss: 0.5305 - custom_auc: 0.7850 - val_loss: 0.5290 - val_custom_auc: 0.7801</p>\n<p>I will make the submit and come with the results.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1108081,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-10T08:50:35.210000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> - you are right. I potentially sampled more starting sequence in each epoch, even they are from different users. (Actually I don't do [A, B, C, D] -&gt; [A, B] [C, D]. With your example, it is 2 subsequences from the 1st user, which could be selected from [A, B], [B, C] or [C, D])</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1108141,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-10T10:20:20.957000",
          "content": "<p>Then that's exactly how I had it before (a random subsequence for a user). The problem is that you don't have many subsequences to choose when the whole sequence is shorter than your window size.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1108297,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-10T14:07:09.173000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>, thanks for the sampling stategy for training, it's quite clear, I will try it too. And for validation, how do you do? As it should be always the same set (no random).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1108468,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-10T17:24:03.857000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> </p>\n<p>I use every possible sequence for validation. [A, -, -, …], [A, B, -, …], …</p>\n<p>By the way, I got a submission error for my last model after waiting for 9h, though it doesn't say timeout. That sounds like what you told before.  WTF what can I do? I don't have any prior version enviroment. I've seen there's a v91 version, but I've checked the Dockerfile commit history and doesn't seem like they have fixed anything relevant.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1108474,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-10T17:43:51.090000",
          "content": "<p>From my experience v90 works but slower than v89. If you don't have a kernel with v89 then you can try to fork one in the public kernels.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1108518,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-10T18:42:24.533000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> i too get similar CV but not same LBs..<br>\n could i be making mistake in passing response sequence to model</p>\n<pre><code>label = qa[1:]\n rt=qa[1:-1].copy()\n rt = np.append(np.zeros((1,)),rt)\n</code></pre>\n<p>during inference however i still case doubt  because if i know correctly model relies on previous exercise responses but during test time where do we have response for test q during that iteration.<br>\nAll i see in inferences so far that <br>\n<code>qa[1:] , q[2:].append( test questions)</code><br>\nwhere QA is series of Yes or No for correct ans from Train set/Valid set   but our question array has got question of Test data .</p>\n<pre><code>eg. E1 T1 T2 T3\nresp sequence is \n     E0,E1,E2,E3..  \n</code></pre>\n<p>i think m getting we append to last one or few test q given the history of train exercise responses</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1109164,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-11T11:27:47.703000",
          "content": "<p>Okay it was not a Docker Image problem. I've been submitting with CPU all the time. I didn't know that I couldn't submit with GPU when doing \"Quick Save\" O.o (this is my first serious participation in a competition).</p>\n<p>With GPU, inference time was around 2 hours, and <strong>my new LB is 0.784</strong>. Now that I have a good pipeline and a decent correlation between CV and LB I think I can still improve my score with certain agility. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1109171,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-11T11:39:53.373000",
          "content": "<p>You can edit the settings of the quick save as well and then probably we can use quick save as well…!<br>\nGreat LB! Keep Rising! 🎉🎊🎊</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1109238,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-11T13:04:52.863000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>  I was doing it. I mark: \"Run with GPU for all sessions\", and have my GPU connected when submitting, but for some reason, it seems like you can only submit with GPU when doing \"Save and run all\".</p>\n<p>On the other hand, I'm grinding just thanks to the help received around, so thank you guys! I feel in debt and I'm trying to collaborate in the same proportion.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1109290,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-11T14:08:43.090000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> great going  how much was score you got without updating status of test df with group ans ?<br>\nm able to get so far LB 74.8  without status update &amp; without masking  padded items during loss<br>\nm still trying to find what is holding me back :).. any advise appreciated . I joined the competition party week ago so still working to get basics right :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1116239,
          "author_name": "jojo",
          "author_url": "",
          "post_date": "2020-12-17T01:39:23.867000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Thanks for your sampling strategy and it works for me. With the sampling strategy and some feature engineering, I got 0.786 LB with a single SAINT+ model.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1122161,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-22T08:35:53.913000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1131304,
          "author_name": "Darren Lahr",
          "author_url": "",
          "post_date": "2020-12-29T16:40:38.877000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Thanks for your sampling strategy I was able to obtain CV of 0.765 (test set last 100 interactions) and LB 0.776  with the original Saint model (No time features). </p>\n<p>For those using this strategy how are you implementing it? I am sub classing from tf.keras.utils.Sequence but the GPU is now severely bottle necked.   </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1081936,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-17T12:40:52.267000",
      "content": "<p>I've tried several variant of saint, and only have a CV score to provide, no LB yet:<br>\nVariants : </p>\n<ul>\n<li>Saint++ implemented with the same parameters as in the papers and the same features : CV Roc AUC ~0.757</li>\n<li>Decoder architecture trained with both a next question prediction loss and a next correct prediction loss : CV Roc AUC ~0.762</li>\n<li>Encoder architecture (bidirectional) pretrained with a masked question modeling loss then finetuned with a sequence pair classification loss (seq 1 = question to predict, seq 2 = user history) : CV ROC AUC : ~0.77</li>\n<li>Replace the pure attention layers of transformers by LSTM + Attention : no performance improvement and instability during training.</li>\n</ul>\n<p>For each of these architectures I chose a sequence length of 128, a depth of model of 512 and a ffn representation size of 1024. Optimizer was adam with lr 3e-5 (more caused gradient explosion) and bs either 32 or 64.<br>\nNumber of layers per block was either 4 or 8</p>\n<p>Convergence is often decided during the first few epochs, but increase slitghly when continuing training (i've trained the encoder on MQM loss for about 8h on a RTX2070)</p>\n<p>Given the very low differences between the different architectures performances, i'd say that the architecture is not really impactfull for the modelisation (providing that you have a good one).</p>\n<p>So for next iteration i might focus more on how to build embeddings for the different sequences (right now i use on embedding layer for each input sequence and i do a summation)</p>\n<p>If you have any ideas of other variants to try or new features to include i'm open for discussion :)</p>\n<p>PS : For the ones trying to replicate the encoder results, the MQM loss need to be quickstart by first making a directionnal encoder, training for a few epochs with a next question prediction loss and then adding the look ahead masks and swithing to MQM loss (it fails to start when only on MQM loss)</p>",
      "votes": 13,
      "replies": [
        {
          "id": 1081967,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-17T13:28:54.047000",
          "content": "<p><a href=\"https://www.kaggle.com/rously\" target=\"_blank\">@rously</a> Thanks for your feedback, it's really interesting that you've achieved CV 0.757 with SAINT+ because it's quite similar for me, I'm not able to go beyond CV 0.76 with my implementation and with similar parameters as in paper (encoder + decoder, d_model=256, seq_len=100, nhead=8, 4 layers, batch size=256). LR with warmup (0 to 0.0008) works quite well for me and allows fast convergence after a few epochs.</p>\n<p>About:</p>\n<blockquote>\n  <p>Encoder architecture (bidirectional) pretrained with a masked question modeling loss then finetuned with a sequence pair classification loss (seq 1 = question to predict, seq 2 = user history) : CV ROC AUC : ~0.77</p>\n</blockquote>\n<p>masked question modeling = MQM, correct? Where does it apply? Only in loss? Or also in padding mask?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1082438,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-17T22:19:27.737000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <br>\nYep, MQM = masked question modeling loss, it's the same principle as mlm for language model, you hide 15% of the input tokens and you try to predict it. My intuition behind it was that only predicting the output of the user doesn't motivate the model enough to understand the pattern of sequences of questions, so i added this loss to do that.</p>\n<p>Basically a classical input will be constituted of several input sequence, but to understand it you can decompose it like that :</p>\n<p>[CLS] last_question_token [SEP] tok1 tok2 tok3 tok4 … [PAD] [PAD]…<br>\nThe MQM loss is only applied to the part of the sequence between the SEP and the PAD not anywhere else</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1082516,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-18T01:06:38.590000",
          "content": "<p>This is really quite an interesting approach that one can try if time allows!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1093375,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-27T16:52:15.793000",
          "content": "<p><a href=\"https://www.kaggle.com/rously\" target=\"_blank\">@rously</a> masking of 15% of input tokens applies to both encoder and decoder inputs?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1093735,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-28T01:16:32.927000",
          "content": "<p>In my case i was doing an only encoder architecture, but I guess that with an architecture encoder-decoder like saint you can pretrain the encoder with the masked loss as well.</p>\n<p>Havent had much time to experiment since my last post, but i noticed that using a mse or rankboost loss instead of crossentropy was giving slithly better result (my best one gives a CV of 77.3% with this)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1108939,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-11T07:00:13.740000",
          "content": "<p><a href=\"https://www.kaggle.com/rously\" target=\"_blank\">@rously</a> <br>\nMQM= masked question modeling loss<br>\ndo you have pytorch impl repository for it ?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1109339,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-12-11T14:55:56.263000",
              "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> <br>\nnope, i did the implementation in tensorflow. It is quite simple to do though with a generator, take your input sequence, generate an numpy array with numpy.random.choice([0,1],size = max_len, p = [1-r, r]) where r is your rate of masking, then you just have to make a multiplication to get the inputs and outputs of your network, just have to mask the zero tokens of the outputs in the loss calculation</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1109350,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-12-11T15:13:58.477000",
              "content": "<p>By the way, some other ideas I tried but that did not improve the validation score:</p>\n<ul>\n<li>encode the input sequences with a tabnet architecture, doesn't improve the result and is a pain in the ass to train as tabnet requires a different learning rate than transformer</li>\n<li>modify the encoder-decoder merge by putting the encoder values as query and key and the decoder as value (this would make more sense to compare similar items)</li>\n<li>use a tabnet layer as classification, somehow it managed to peak forward in time through the batch normalization, so gave an AUC of 0.84 but did not generalize in pure inference</li>\n<li>increase batch size a lot (to 2048 thanks to TPU) did not help either</li>\n<li>use local attention windows (in order to look both at all history and at only the last 20 values)</li>\n</ul>\n<p>One thing that seems to help though is to replace the bce loss by a rank boost loss</p>\n<p>Perhaps with the sampling strategy from <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> these ideas might improve a bit</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1112659,
      "author_name": "Arshad Shaikh",
      "author_url": "",
      "post_date": "2020-12-14T19:38:02.333000",
      "content": "<p>PyTorch implementation of SAINT - <a href=\"https://github.com/arshadshk/SAINT-pytorch\" target=\"_blank\">https://github.com/arshadshk/SAINT-pytorch</a></p>\n<p>Let me know if any corrections or suggestions are there.<br>\n(Also have a look at SAKT-PyTorch <a href=\"https://github.com/arshadshk/SAKT-pytorch\" target=\"_blank\">https://github.com/arshadshk/SAKT-pytorch</a> )</p>",
      "votes": 11,
      "replies": [
        {
          "id": 1112800,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-14T22:43:00.533000",
          "content": "<p>Thanks for doing this. I'm two days into my own SAINT implementation, so it'd be nice to check my work against someone else's—considering there are no (other) public SAINT(+) implementation kernels or repos as far as I am aware. If you post your repo as a notebook, you can probably get some kaggle kernel points there as well 👍🏾.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1116146,
          "author_name": "RB",
          "author_url": "",
          "post_date": "2020-12-16T22:22:19.400000",
          "content": "<p><a href=\"https://www.kaggle.com/arshad7\" target=\"_blank\">@arshad7</a> Can you please explain what is <code>total_in</code> input to the model ? I know readme says <code>Total number of unique interactions.</code> and for the random check you have it as 2.  Thank you. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1116253,
          "author_name": "Ryosuke Horiuchi",
          "author_url": "",
          "post_date": "2020-12-17T01:58:44.187000",
          "content": "<p>I think it indicates the number of response class. Binay if total_in=2, i.e., 0 and 1. You can verify that by <code>in_de</code> from the code below.<br>\n<code>in_ex, in_cat, in_de = random_data(64, seq_len , total_ex, total_cat, total_in)</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1116384,
          "author_name": "Cakey",
          "author_url": "",
          "post_date": "2020-12-17T06:02:30.477000",
          "content": "<p><a href=\"https://www.kaggle.com/arshad7\" target=\"_blank\">@arshad7</a> Thank you for posting the repo! I just have one question, what are the three inputs in the forward pass? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1116465,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-17T07:35:37.350000",
          "content": "<p><a href=\"https://www.kaggle.com/arshad7\" target=\"_blank\">@arshad7</a> I made use of it , refactored to separate encoder ,decider inputs and data input . What i find is skip version fair quite low in cv compared to non skip.<br>\n<a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> do you use skip version or non skip</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1123601,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-23T11:11:48.017000",
          "content": "<p>I use skip version.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133723,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "2020-12-31T13:17:39.747000",
          "content": "<p>I may be too late here. But I finally was able to make use of the SAINT implementation. But it's around CV 0.700. What was your CV/LB for this? I wonder if I have wrong dataloader implementation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133755,
          "author_name": "Darren Lahr",
          "author_url": "",
          "post_date": "2020-12-31T13:56:26.617000",
          "content": "<p>Original saint I got CV 0.765 and LB 0.776. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1134178,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "2021-01-01T00:20:04.597000",
          "content": "<p>Thank you for your reply. I must have something wrong in my code then.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135225,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "2021-01-02T03:36:22.583000",
          "content": "<p>Thanks. I got LB: 0.741 using <a href=\"https://github.com/arshadshk/SAINT-pytorch\" target=\"_blank\">https://github.com/arshadshk/SAINT-pytorch</a>.<br>\nI had a mistake on my dataloader implementation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1136431,
          "author_name": "RB",
          "author_url": "",
          "post_date": "2021-01-03T04:23:09.743000",
          "content": "<p>I get different model output after training and after loading a saved transformer model. I initially thought there must be some input processing bug, but then I directly passed tensors after training and after saving and loading the model, results are different.  I appreciate any  thoughts on what I could have missed. I used this model architecture - &nbsp;<a href=\"https://github.com/arshadshk/SAINT-pytorch\" target=\"_blank\">https://github.com/arshadshk/SAINT-pytorch</a>  Thank you! </p>\n<p>Edit: loading state_dict() is not working however loading full model works. don't know why 🤔</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1128932,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2020-12-27T22:28:18.320000",
      "content": "<p>Gotta tell you guys, my last update brought me from 784 to 794. Only one detail got me there, related with timestamp (and I think it is still improvable). For those who haven't worked around this feature hard enough, I really recommend it.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1128944,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-27T22:56:48.733000",
          "content": "<p>Great! You mean lag? Timestamp(N) - Timestamp(N-1)?<br>\nMy LB=0.791 is with SAINT+ paper + had_explanation. Nothing more.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1128947,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-27T23:07:08.190000",
          "content": "<p>Yeah, lag it is, but you have to take it carefully.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1128951,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-27T23:13:52.627000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> You don't know how glad I am to see you get higher in the LB - I checked your progress almost everyday - I thought you gave up but hoped you didn't because you have done so much!</p>\n<p>Again, good for you!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1128964,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-28T00:04:38.233000",
          "content": "<p>Thanks for reporting. I too had a jump from 780s to 790s (cv) by engineering the ts feature, and also agree there's plenty to be done with it in terms of stats surrounding categorical groupings.</p>\n<blockquote>\n  <p>Yeah, lag it is, but you have to take it carefully.</p>\n</blockquote>\n<p>I didn't do anything special to the variable though. Now you make me wonder if I should look at its values more closely 👀</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1129013,
          "author_name": "Dean",
          "author_url": "",
          "post_date": "2020-12-28T02:03:25.837000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Did you use continuous embedding or categorical embedding?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129015,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-28T02:08:46.640000",
          "content": "<p>Wow! I thought you have left the competition for other one's after working hard enough in this one! Really cool and great going!</p>\n<p>For NNs h<strong>ow you pre-process your data matters a lot</strong>, I don't think the embedding/non_embedding of that will be a huge difference if used correctly</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129061,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-28T03:47:14.653000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129066,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-28T04:03:50.133000",
          "content": "<p>How much was wrt bad results </p>\n<p>And other thing 79.1 79.2 probably you would have expected more going by trend when we submit saint with cv 75 ,lb 776 or so . What is cv method you use default as in mpwares? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1129367,
          "author_name": "Dean",
          "author_url": "",
          "post_date": "2020-12-28T09:39:36.840000",
          "content": "<p>Bad result is cv:0.65x. I not submitted. The cv method I used is same as popular public notebook.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1144675,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2021-01-08T15:51:47.413000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> thank you for your words mate, I didn't read them until now.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1071786,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-11-07T12:06:59.117000",
      "content": "<p>I've read both SAINT and SAINT+ papers but some questions remain.</p>\n<p><strong>Input sequences</strong> should be something like:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fce0c6347798d2f33707b51b623e4a714%2Fsequences.png?generation=1604752912346221&amp;alt=media\" alt=\"\"></p>\n<p><strong>Start token(s):</strong><br>\n<em>The decoder takes O and another sequential input Re = [S,Re1,··· ,Rek−1] of response embeddings with the start token embedding S.</em></p>\n<p>Something like this should work for start token:</p>\n<pre><code>...\nresponse_size = 2\nself.response_embedding = nn.Embedding(response_size + 1, embedding_dim) # +1 to include start token\n...\nx_correctness = self.response_embedding(response_sequence)\n...\n# Add start token to correctness\nx_correctness = torch.roll(x_correctness, shifts=(0, 1, 0), dims=(0, 1, 0)) # Shift right the sequence\nx_correctness[:,0,:] = self.response_size # Start token\n</code></pre>\n<p>x_correctness is embedding of response sequence (0,1,1,1,0 …), response_size=2, so with token it will be (2,0,1,1,1,0 …)</p>\n<p><strong>Self attention mask:</strong></p>\n<pre><code># If a BoolTensor is provided, the positions with the value of True will be ignored while the position with the value of False will be unchanged.\n# tensor([[False,  True,  True,  True],\n#         [False, False,  True,  True],\n#         [False, False, False,  True],\n#         [False, False, False, False]])    \ndef generate_mask(self, size, diagonal=1):        \n    return torch.triu(torch.ones(size, size)==1, diagonal=diagonal)\n</code></pre>\n<p><strong>Learning rate:</strong><br>\nWarmup from 2000 to 4000 iterations works.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F9d3dc4354e1cb1ba653cfb8adba9e84a%2Flr.png?generation=1605009427277426&amp;alt=media\" alt=\"\"></p>",
      "votes": 8,
      "replies": [
        {
          "id": 1071884,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-07T14:15:42.897000",
          "content": "<p>Nice work! Do we need the decoder here as well?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1071920,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-07T15:05:22.973000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> Yes, from the paper we need both encoder and decoder.</p>\n<pre><code># Transformer with default encoder/decoder        \nself.transformer = nn.Transformer(d_model=input_features_dim, \n                                  nhead=8, \n                                  num_encoder_layers= 6,\n                                  num_decoder_layers= 6, \n                                  dim_feedforward=2048, \n                                  dropout=0.1, \n                                  activation='relu', \n                                  custom_encoder = None,\n                                  custom_decoder = None)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fcbaeee6b56f3eeac7408c37d250ca3c8%2Ftransformer.png?generation=1604761231138288&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": [
            {
              "id": 1108884,
              "author_name": "Jaideep",
              "author_url": "",
              "post_date": "2020-12-11T05:39:03.937000",
              "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> Yes, from the paper we need both encoder and decoder.</p>\n<pre><code># Transformer with default encoder/decoder        \nself.transformer = nn.Transformer(d_model=input_features_dim, \n                                  nhead=8, \n                                  num_encoder_layers= 6,\n                                  num_decoder_layers= 6, \n                                  dim_feedforward=2048, \n                                  dropout=0.1, \n                                  activation='relu', \n                                  custom_encoder = None,\n                                  custom_decoder = None)\n</code></pre>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fcbaeee6b56f3eeac7408c37d250ca3c8%2Ftransformer.png?generation=1604761231138288&amp;alt=media\" alt=\"\"></p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> about response shift<br>\n1) how about merely doing this in the dataset</p>\n<pre><code>target_id = q[1:]\n rt=qa[1:-1].copy()\n rt = np.append(np.zeros((1,)),rt)\n</code></pre>\n<p>so this should give prev ans to every current exercise. embeddings should get formed accordingly only ..  does it make a difference  compared to what you are doing after computing embeddings and then doing masking ?</p>\n<p>2) in dataset are you doing any shifting of inputs ?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1071932,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-07T15:31:46.570000",
          "content": "<p>Thanks MPWARE.  At inference time, when we will see unseen users, how can we pass in the i/p's that's needed? Plus we will have to maintain the seq for each un-seen user as well, right?<br>\nThanks for torch.roll as well!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1071946,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-07T15:51:10.660000",
          "content": "<p>During inference, we need to maintain the full sequences per user (i.e. the last 99 answers, questions,  …) then append the new question_id and our model will predict the next answer.<br>\nIt will be easy for groups including only one new question per user, but we know that we can have more than one question (between 1 to 5 if I'm not wrong according to <code>task_container_id</code> possible total questions) then we have several options: </p>\n<ul>\n<li>Consider our prediction as correct and append it to answer sequence and predict again with new question.</li>\n<li>Have another model trained to predict 2 answers, another to predict 3 answers …</li>\n<li>Or?</li>\n</ul>\n<p>Maintaining last 100 answers in memory for all users (with some flush to be memory friendly) is possible within the 9 hours inference time. You could implement it and simulate it with this nice <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">kernel</a> from <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> </p>\n<p>For unseen users, we won't have any previous answers but I believe we could use optional masks in our model to ignore padding we've added to reach sequence size for such new users.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1071950,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-07T15:57:02.733000",
          "content": "<blockquote>\n  <p>For unseen users, we won't have any previous answers but I believe we could use optional masks in our model to ignore padding we've added to reach sequence size for such new users.</p>\n</blockquote>\n<p>I see. Thanks for the tip; Have a lot of working to do.</p>\n<blockquote>\n  <p>Maintaining last 100 answers in memory for all users (with some flush to be memory friendly) is possible within the 9 hours inference time.</p>\n</blockquote>\n<p>Yep, last 100 is the window length as mentioned in the paper and I was using a deque for that but now torch.roll might come handy! </p>\n<p>Also, Are you simply taking the last 100 seq's / interactions or splitting let's say 1000 content_ids into 10 chunks?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1071958,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-07T16:14:21.930000",
          "content": "<p>Just one dataframe in memory for all users, no chunks. </p>\n<p>But before being able to run inference with SAINT, we need to have SAINT model working first 😏. I've implemented it but it does not give results from the paper, it's stuck at score = 0.52. It's quite similar to SAKT results from this <a href=\"https://www.kaggle.com/leadbest/sakt-self-attentive-knowledge-tracing-submitter\" target=\"_blank\">notebook </a> attempt from <a href=\"https://www.kaggle.com/leadbest\" target=\"_blank\">@leadbest</a> (score = 0.54). SAKT should also perform better (paper claims 0.76). We might be doing something wrong or convergence conditions are not met.</p>\n<p><strong>Update</strong>: I'm sure something is wrong in my implementation has all OOF probabilities are equal to 0.68.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1072123,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-11-07T20:26:24.217000",
          "content": "<p>Reduce the size of model and it'll converge. Use 128 or 64 dimensions for d_model. That helped me overcome this issue.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1072174,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-07T21:24:51.747000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> but I think root cause is something else. I've already tried 64 and 128 and it's quite the same.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F44b3ea36ff08c2e64d2720590f85ef4c%2Fbadtrain.png?generation=1604784242641607&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1072250,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-08T01:21:33.080000",
          "content": "<p>The plots makes it look like models output is constant.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1072261,
          "author_name": "HDKIM",
          "author_url": "",
          "post_date": "2020-11-08T02:02:16.347000",
          "content": "<p>Thank you for referring the kernel. In the paper of Pardney and Karypis, SAKT claims auc of 0.824 in average. It assumes that all the correctness of the user-problem interactions are known in advance. So, I have to employ 1's as the fake correctness values in applying SAKT to this kernel. I've got 0.5x in submission, very low compared with the validation score of 0.9x, I think there are several reasons for this. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1072537,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-08T11:57:31.023000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> Yes, output of predictions was constant and loss was not deceasing but I've made some progress. Root cause seems to be related to learning rate. After fine tuning it I've got better results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F15d4d8ba9beaf66dae6cee97d1661121%2Fbettertrain.png?generation=1604836526161499&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1072544,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-08T12:20:35.080000",
          "content": "<p>🎊🎊🎊🎊🎉🎉🎉🎉🎉🎉 I had ~.66-.67 when i just used encoders. [no attention mask, just padding mask if needed] Not sure whether i did the evaluation correctly though 😅</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1072644,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-08T14:46:20.913000",
          "content": "<p>Now I'm able to reach 0.74 in validation.<br>\n<a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> You do need mask for self attention otherwise your model will also learn from future (that we don't want to).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1072662,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-08T15:08:01.857000",
          "content": "<p>Wow! That's a good score. I will relook into them. Thanks for the inspiration &amp; Good luck with the inference pipeline! Don't forget to update us!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1073603,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-09T17:39:36.403000",
          "content": "<p>Everything implemented but I'm not able to reach the 0.79, only 0.748 on validation (train on 50% of data). Sequence size = 100, d_model=256. Maybe more data is needed…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1073606,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-09T17:41:36.880000",
          "content": "<p>Well, how's it on the LB :) Or on a simple lgbm with feature extracted from the same?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1073617,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-09T17:51:05.260000",
          "content": "<p>Not tested on LB (yet). I'm trying to have a solid local validation first. LB becomes quickly hypnotic and I don't want to focus on it for now. Compared to basic LightGBM (locally again), LGB gives 0.77.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1073619,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-09T17:52:14.443000",
          "content": "<p>Wow! That's a strong single model! Looking forward to learn from the same.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1073665,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-09T19:05:33.673000",
          "content": "<p>Hello mate, I've tried to comprehend all the thread. I wanted to add that depending on how you organize your input, maybe you don't need a triangular mask, just a padding one. Maybe what made you improve is not switching to only encoder, but to have only a padding mask.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1073676,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-09T19:27:34.073000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> I've organized data to always have full real sequences (seq_len=100) in train/valid so I didn't need padding. The drawback is that I've much less data. I'm going to add padding support tomorrow. I will post results here.</p>\n<p>Did you also implement SAINT/SAINT+? What results do you get?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1073745,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-09T22:06:41.930000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> I'm trying to, but with slight modifications. </p>\n<p>Imagine I have a user story like this: [A, B, C, D, E] and I have a window size of 3. Then, what I would do is [[X, X, A], [X, A, B], [A, B, C], [B, C, D], [C, D, E]]. At the same time for every sequence I shift the answers for the input, and let them as they are for the targets. So in X everything is padded; in A I have a special SOS token for the answered_correctly (the rest features as they are); in B I have the answered_correctly corresponding to A (the rest features as they are); and so on.</p>\n<p>Also, instead of predicting a whole window, I just predict the next answered_correctly (one neuron).</p>\n<p>With this set up there's no target leakage so I don't need to look ahead mask the decoder.</p>\n<p>Though I'm afraid that I'm doing something wrong, because I still get better results with only the Encoder part. </p>\n<p>Also, I have a very small model (model dimension is less than 40 and sequence lenght is less than 70 in most experiments).</p>\n<p>My results are what you can see, ~0.758 on local validation and 0.763 on LB. I think that it has to be easy to beat the 0.775 with this approach, but for some reason it isn't working that well.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1073921,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-10T03:28:02.117000",
          "content": "<blockquote>\n  <p>Also, I have a very small model (model dimension is less than 40 and sequence lenght is less than 70 in most experiments).</p>\n</blockquote>\n<p>Can you increase the d_model to 256 and re-run?</p>\n<blockquote>\n  <p>I still get better results with only the Encoder part.</p>\n</blockquote>\n<p>But when you use only encoder, then you will have predictions for each and every seq, right? So how do you evaluate your model as your encoder only o/p will be <code>[BS, SEQ_LEN, 2]</code>? Can we simply take the <code>[:,:,1]</code> from the above mentioned tensor?</p>\n<p>My current setup is like this, (quite similar to <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>'s)</p>\n<p>I split the sequences of each user to 32 and pad where it's needed using src_key_padding_mask. (it's True where's there's a padding, otherwise False). I am not using any triangular mask.</p>\n<pre><code>def pad_seq(seq: List[int], max_batch_len: int = LAST_N, pad_value: int = True) -&gt; List[int]:\n    return seq + (max_batch_len - len(seq)) * [pad_value]\n</code></pre>\n<p>So as an e.g, for user_id 115, it's like this, (a deque for auto reduce to last 100)</p>\n<p><code>{'user_id': 115, 'content_id': deque([5692, 5716, 128, 7860, 7922, 156, 51, 50, 7896, 7863, 152, 104, 108, 7900, 7901, 7971, 25, 183, 7926, 7927, 4, 7984, 45, 185, 55, 7876, 6, 172, 7898, 175, 100, 7859, 57, 7948, 151, 167, 7897, 7882, 7962, 1278, 2065, 2064, 2063, 3363, 3365, 3364, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], maxlen=100), 'answered_correctly': deque([1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 1, 1, 1, 1, 0, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 1, 0, 1, 1, 1, 1, 0, 1, 1, 1, 1, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], maxlen=100), 'task_container_id': deque([1, 2, 0, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 40, 40, 41, 41, 41, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], maxlen=100), 'part_id': deque([5, 5, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 2, 3, 3, 3, 4, 4, 4, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], maxlen=100), 'padded': deque([False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True, True], maxlen=100)}</code></p>\n<p>cc <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>. </p>\n<p>Ty!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1074275,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-10T13:07:25.757000",
          "content": "<p>Hello mate,</p>\n<p>My input's shape is [batch_size, sequence_length, n_features] where those features are: the shifted answered_correctly, time features, question features, etc. Then I take every [:, :, feature] to create the embeddings. </p>\n<p>My output's shape is [batch_size, 1] (as well as the target), which means that for every sequence I have one output (0-1), that I then use to compute a binary logloss. </p>\n<p>I'm afraid that I can't increase my model size due to lack of computation power. My model is converging now, do you think that increasing the model size could improve the AUC from 0.763 to a relatively higher score?</p>\n<p>Also, I'm only using from 5M to 10M rows, depending on the experiment. And still some runs take several hours to complete.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1074496,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-10T18:10:42.197000",
          "content": "<p>Interesting, need to check why it's not the same case for me. Regarding that position encoding,  isn't it true that position encodings are independent of input length?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1074544,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-10T19:35:27.247000",
          "content": "<p>I've tried multiple options there: </p>\n<ul>\n<li>Learnable sequence dependant position encoding.</li>\n<li>Non learnable sequence dependant position encoding.</li>\n<li>Learnable sequence non dependant position encoding.</li>\n<li>Non learnable sequence non dependant position encoding.</li>\n</ul>\n<p>I've seen no important difference yet TBH, so I'm sitcking to non learnable non dependant since theyre cheaper and easier to set up. Just a bunch of sinusoidal vectors with shape (sequence_length, d_model).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1074576,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-10T20:11:05.287000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> In some papers I've read about KT and Transformer, some defined position encoding as a relative position (to a fixed sequence), so always 0, 1, 2, 3 … whatever the way you pick the sequence in the full time series. Sin/cos positional encoding (as in Attention is All you need) works for me, I've not tried direct 0, 1, 2, 3 …</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1074583,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-10T20:23:44.007000",
          "content": "<p>In a different discussion the current winner pointed out that the position encoding are actually learnable parameters (embedding vectors whose input is an integer and output is a vector of model dimension size). </p>\n<p>And what you said is true, in all those transformer based papers the position encoding is actually fixed - not increasing. Though, my intuition tells me that it shouldn't be as good as having an actual increasing position encoding for KT. Why?</p>\n<p>Imagine I have a sentence like <em>[..] and he said he was the […]</em>. You could actually complete that sentence with a high level of precision, even without knowing the beginning of the sentence. However, do you think a user will perform equally in the 1200th question than in the 70th even if their last sequence_length interactions were the same? </p>\n<p>My theory is that including that information should be relevant, and maybe decisive for those fighting for a 0.00x extra.</p>\n<p>Any thoughts?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1074602,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-10T21:00:21.573000",
          "content": "<p>True but you could also add such information in another category like:<br>\nNewbie = Interactions with 0-100<br>\nNovice = Interactions with 0-500<br>\nMedium = Interactions with 500-2000<br>\n…<br>\nExpert = 5000+</p>\n<p>I did not try it yet but I will.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1074647,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-10T22:35:56.640000",
          "content": "<p>Nice idea! Take into account that a user is first a newbie, then a novice… don’t leak information from the future!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1075080,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-11T11:28:23.277000",
          "content": "<p>I've seen in your notebook that you're using <code>task_container_id</code>, as it's increasing monotonically (except for some <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/189465\" target=\"_blank\">cases</a>) then it could also act as pseudo absolute sequence.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F0d6af8db417ea5f9393c11657ab345c7%2Ftask.png?generation=1605094256718118&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1075147,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-11T12:50:28.630000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> ,</p>\n<p>My current score on LB is a transformer, but I have difficulty to reproduce the results after adding more inference code for the validation. Something is really strange there, so my current LB might be kind of random. Hope I can find where the problem is. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1075152,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-11T12:53:52.103000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> What is your score on local validation?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1075198,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-11T13:34:05.177000",
          "content": "<p>Currently, I have only 3 CV measured, but it is for a smaller model, and they were run before the API update:</p>\n<pre><code>CV - 0.6651957631111145\nLB - 0.668\n\nCV - 0.6859979629516602\nLB - 0.690\n\nCV - 0.6827537417411804\nLB - 0.689\n\nCV - 0.679010808467865\nLB - 0.684\n</code></pre>\n<p>As I mentioned, I am still debugging the code because I am no longer to reproduce the same results as above. But the CV is tight to LB.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1075203,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-11T13:38:57.010000",
          "content": "<p>I have added task_container_id as embeddings even if in the papers they don't say anything about it because I found a certainly jump in score. I also tried to normalize it (dividing it by 9999) and feed it as continuous variable to the network but didn't seem to work well.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1076085,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-12T08:53:46.213000",
          "content": "<p>I saw your notebook, I am planning to release mine in PyTorch soon; Here's my question, your are doing a <strong>pooling</strong> o/p, that's why you are not seeing preds for every sequence element i feel as it <strong>reduces</strong> <code>(batch, steps, features)</code> (channel_last) to <code>(batch_size, features)</code>. Pl correct me if i misunderstood something.</p>\n<pre><code>    x = tf.keras.layers.GlobalAveragePooling1D()(x)\n    x = tf.keras.layers.Dropout(0.2)(x)    \n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1076107,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-12T09:12:56.603000",
          "content": "<p>That's it. The advantage is that I don't need lookahead mask. Maybe there's any drawback but I don't know.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1076179,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-12T09:58:43.537000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid/pytorch-demystifying-transformers\" target=\"_blank\">Here</a> it's as promised. Please let me know my mistakes!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1076623,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-12T17:42:09.390000",
          "content": "<p>I've tried to implement left padding mask (for users with less total interactions than sequence length - i.e. like in inference) to use <code>nn.Transformer</code> but I'm facing to some conditions that leads to NaN tensor (and then NaN loss) to due to both self attention mask and padding mask. Nice example is provided here:<br>\n<a href=\"https://discuss.pytorch.org/t/how-to-add-padding-mask-to-nn-transformerencoder-module/63390/7\" target=\"_blank\">https://discuss.pytorch.org/t/how-to-add-padding-mask-to-nn-transformerencoder-module/63390/7</a></p>\n<p>Root cause is mask conditions that lead to the following result in <code>nn.MultiheadAttention</code>:</p>\n<pre><code>x = torch.Tensor([[[float(\"-inf\"), float(\"-inf\"), float(\"-inf\")]]])\nsoftmax = torch.nn.Softmax(dim=-1)\nsoftmax(x)\n...\ntensor([[[   nan,    nan,    nan]]])\n</code></pre>\n<p>Some discussions about this issue:<br>\n<a href=\"https://github.com/pytorch/pytorch/issues/41508\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/41508</a><br>\n<a href=\"https://github.com/pytorch/pytorch/pull/42323\" target=\"_blank\">https://github.com/pytorch/pytorch/pull/42323</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1076656,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-12T18:06:49.403000",
          "content": "<p>You are masking everything in the sequence mate!</p>\n<pre><code>&gt;&gt;&gt; x = torch.Tensor([[[float(\"-inf\"), float(\"-inf\"), float(\"-inf\"), 0.5]]])\n&gt;&gt;&gt; torch.nn.Softmax(dim=-1)(x)\ntensor([[[0., 0., 0., 1.]]])\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1076668,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-12T18:19:20.970000",
          "content": "<p>Not in the input sequence (I've double checked) but within Attention matrix multiplication I've one (or more) rows with full <code>-inf</code>. Well, it's not a big problem for training as we can ignore users with small interactions but I'm wondering how it will behaves for new users in inference. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1076710,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-12T19:01:51.370000",
          "content": "<p>That means you are probably masking a full row, thus having a full minus infinite row.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1142976,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2021-01-07T17:43:09.117000",
      "content": "<p>Sorry, I've been quiet since a few couple of days but I've been very busy with my team to try to improve our models.<br>\nLast 5 submissions have been sent so now competition is almost completed. I would like to thanks <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> <a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> and all other contributors to this thread. It really helped us to make transformer(s) model(s) train better and work. I won't share any secret right now but I think top teams and you guys have found similar \"things\" that made transformer better and better.<br>\nI wish you the best for private LB and for year 2021!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1143552,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2021-01-08T00:22:58.553000",
          "content": "<p>Great Work Guy's! A competition to remember ❤️ and next time will work insanely harder to have that top ~1-2% finish! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1144673,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2021-01-08T15:50:12.887000",
      "content": "<p>I have shared my final solution here <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209793\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209793</a>. It has some tricks that weren't discussed in this thread, implemented during the last two weeks of competition. Grew me up from 0.794 to 0.800 in public LB, and I believe it could still have gotten a better score with more time to finetune.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1110427,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-12-12T18:16:20.857000",
      "content": "<p>Here is update after unrelenting work of 5 days 77.1 SAINT score…Now will try some other stuff </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1115292,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-16T07:00:42.037000",
          "content": "<p>Bravo! keep going Jaideep… see if some ensembles can take it further up by 1-2%age..</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1134417,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2021-01-01T09:30:34.067000",
          "content": "<p>Congratulations on finally solving the submission problem</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1134429,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2021-01-01T09:45:42.753000",
          "content": "<p>yea..I had to team up w office colleague and tried multiple experiments…Finally it boils down to a browser issue :)<br>\nbut I guess it is too late now. Dont have time to experiment anything</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1134750,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2021-01-01T14:46:09.757000",
          "content": "<p>Just simply enjoy the competition. I used to think that transformer can only be used in the field of NLP, growth of knowledge</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1109542,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2020-12-11T19:59:14.950000",
      "content": "<p>I didn't contribute much in this thread, rather I asked a lot of your approaches.</p>\n<p>I just published a training / validation in TensorFlow with TPU (also works with GPU), and it works also on Colab (minimal change required). No competition submission pipeline is provided - it is (a lot, really a lot) personal effort to this competition, and providing it will also be unfair to those who work hard - you definitely know it.</p>\n<p>In particular, I implemented the auto-regressive prediction - but I found it doesn't perform better than <code>pretend each question (to be predicted during inference) in a question bundle as a single question</code>. And it doesn't work with TPU (with GPU, it is OK), so I didn't continue with it. </p>\n<p>As documentation - I have to say it is far beyond enough. I still need to focus on the competition, hope you could understand. I will add more during the time.</p>\n<p>However, I am currently not able to get higher. With <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> latest sampling strategy, I tried to train a model with his size. It didn't work at the 1st try, but after playing a bit of lr, I get a CV which is as good as a larger model whose LB is about 0.778. I guess the gap between 0.778 and 0.781 could be the different model design or other factors.</p>\n<p>I am currently train a lager model with the latest sampling strategy - thanks for TPU and Colab, I have quite resource to do experiments. </p>\n<p>Here it is <a href=\"https://www.kaggle.com/yihdarshieh/tpu-track-knowledge-states-of-1m-students\" target=\"_blank\">TPU - Track knowledge states of 1M+ students</a>.</p>\n<p>Thanks for your shares of your approaches - especially to <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>, <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> and <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>, among others.</p>\n<p>Good luck!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1109600,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2020-12-11T21:34:38.613000",
          "content": "<blockquote>\n  <p>No competition submission pipeline is provided </p>\n</blockquote>\n<p>You mean inference code, right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1109604,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-11T21:49:57.237000",
          "content": "<p>Yes, no inference code  (there is a validation code, but it use the validation dataset stored in the tf record files, so the inference code for this part is not suitable for inference for the competition API setting)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1109612,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2020-12-11T21:58:18.213000",
          "content": "<p>I see. Well, I agree with that (it would be bad to mess up LB with forked-and-run notebooks). I'm trying to write inference code for the first notebook from this post (PyTorch one), but some reason it produces repeating nearly-zero values. Have you encountered such thing?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1109613,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-11T22:01:17.937000",
          "content": "<p>I don't use that notebook. I built my own code from the beginning - and I can say the inference code need quite efforts to make it right. You can ask the author to share some tips.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1109616,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2020-12-11T22:04:08.313000",
          "content": "<p>Hmm, I see. Thanks for the reply.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1109645,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-11T22:55:06.940000",
          "content": "<p>Nice contribution mate!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1115363,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-16T08:37:37.510000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>  that is great what was you LB  and CV with basic Sampling strategy . <br>\nHow diff is that from claverru .. i couldnt figure out much .</p>\n<p>my latest lb  is 77.3 vs CV of 74.81 .Trying to narrow down the gap.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119261,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-19T21:10:03.087000",
          "content": "<p>With that kernel, the best I can get is 0.781. For further improvement (if any), it won't be publish before the competition is finished.</p>\n<p>My sampling is random selection for each - but the nb of examples (each example is a sequence) from each user are different - but predetermined.</p>\n<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> sampling uses probability distribution for sampling from different users. So we still have different no. of sequences from different users, but that number is not predetermined, and in each epoch, some users might not been seen.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1120048,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-20T15:05:01.513000",
          "content": "<p>Ok , my best cv so far is 77.36 ,saint .will see how much lb does it gets . I suppose it should be less than 78 only ,higher cv don't get higher lb .</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1092406,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2020-11-26T19:26:03.690000",
      "content": "<p>Hello guys, I've been trying to follow this discussion but I've been a little busy lately. I have one remaining question that I've not seen here answered. </p>\n<p>What do you do with users with more interactions than your input length? Do you roll a window? I've tried that and noticed it was (logically) overfitting towards those interactions that appear in multiple rolls, giving me an absurd AUC (~85%). </p>\n<p>Imagine a window size of 3, having a user with:<br>\n[A, B, C, D, E, F] interactions.</p>\n<p>If I roll a window I obtain [[A, B, C], [B, C, D], [C, D, E], [D, E, F]]. As you can see, for example, the interaction C appears multiple times.</p>\n<p>I've also been looking for this in NLP literature but I can't find an answer. Would you share your approach? Or a pointer to a paper/article/post?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1092415,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-26T19:43:15.967000",
          "content": "<p>Not sure what other folk's are doing, but as of now, I simply split them into multiple sequences of fixed length.(total // SEQ_LEN will be the all data you get by that way + some residual). </p>\n<p>I am really interested in knowing why Transformer's work actually without knowing the user_ids. (e.g.)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1092422,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-26T19:54:37.507000",
          "content": "<p>I used to use random sample subsequence from users' history. But I found that potentially I overfit for users with much shorter sequences. For example, if a user has only 50 interaction. In each epoch, it will get the full history. While for a user having 1000 interaction, it only get one subsequence in a epoch.</p>\n<p>For your question, I would say Adiitya's approach is standard. You could probably have some overlapping thought. Like [A, B, C, D], [C, D, E, F] etc.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1092437,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-26T20:17:53.017000",
          "content": "<p>I see, I thought about that but didn't try it. If it is working for you I guess it's cool. If I get to set it up I will come to you with news. Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1092442,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-26T20:25:45.087000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> , The notion of user is based on the user's history. The user id is not essential, it is just a name. The real content is their history. The only problem of Transformer is that we can't use the full history , due to the timing and other resource constraint. I believe if we can use the full history (even not in this competition), we will get better results</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1092503,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-26T22:34:50.447000",
          "content": "<p>Indeed, if you could use a full user history, it would be even better than having an ID. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1092540,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-27T00:20:43.800000",
          "content": "<p>I trained on full history actually, an hour per epoch (~58-59mins) with 256 BS and a maximum of 5-7 epochs. (No apex as of now). But guess, there are bugs here and there, so :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1093191,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-27T14:27:48.153000",
          "content": "<p>1 hour per epoch? It's 10 minutes for me for 96M rows. Did you index your pandas dataframe by user_id? I'm using something like this:</p>\n<pre><code>def __getitem__(self, idx):\n    if torch.is_tensor(idx):\n        idx = idx.tolist()\n\n    if self.subset == \"test\":\n        # For test, we need one question only to answer (with related user)\n        row = self.df.iloc[idx:idx+1,:]\n    else:\n        # For train/valid we need series per user\n        user_id = self.dex[idx]\n        row = self.df.loc[user_id]\n\n    sample =  self.get_sample(row)\n    return sample\n</code></pre>\n<p>Without such index it's super slow even with multiple workers.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1093244,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-27T15:04:13.877000",
          "content": "<p>Ahh wow, that's pretty fast And I am not indexing the data-frame. Rather doing something incorrect it seems.</p>\n<p>I am not sure why it's slower for me; </p>\n<p>Here's my approach;<br>\nI prepared dataset in <a href=\"https://www.kaggle.com/adityaecdrid/pytorch-demystifying-transformers\" target=\"_blank\">this</a> way, except that's it's for all user data and those who have more than 100 SEQ's, i simply break them up. So my training <code>len(dataset)</code> is like 11_41_813. (with a BS of 256, that's like ~4.5K batches in total)</p>\n<pre><code># this is what i do when i prepare the dataset\ngrp.agg({\n    \"content_id\":list, \"answered_correctly\":list, \"task_container_id\":list,\n    \"part_id\":list, \"prior_question_elapsed_time\":list,\n</code></pre>\n<p>Actually I am detaching the loss etc from GPU to CPU, hence it's slow i believe, will remove that part and see it as well. (it's an anti-pattern but just for debugging etc)</p>\n<p>I will look into what you have suggested! Thanks for the tips.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1093254,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-27T15:11:55.760000",
              "content": "<p>To speed up you can also compute metric (add loss.item()) only every 50 batch iterations on training only. What does matter is the validation.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1093328,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-27T16:07:32.353000",
          "content": "<p>Hmm interesting pseudo-code; So if i understand it correctly, you are fetching random rows from the user whoever is at <strong>idx</strong>; And then you maintain the LAST_SEQ or something in <code>self.get_sample</code>? Thanks for your help and tips!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1094103,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-28T10:20:02.723000",
          "content": "<p>Wait what? 10 minutes for 100M rows per epoch? That just killed me. </p>",
          "votes": 3,
          "replies": [
            {
              "id": 1094155,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-28T11:10:55.883000",
              "content": "<p>14 minutes exactly, I'm running it on Colab/GPU (P100).<br>\nScreenshot below with in my current training with MSE loss:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Faaaf4526e1ad47e94f8eae2894f3f5df%2Ftrain_epoch.png?generation=1606561845720602&amp;alt=media\" alt=\"\"></p>\n<p>One epoch covers all 375k users but not all data for all users. It picks a sequence of 100 interactions (randomly) for each. So it's not 14min for 100M rows but 14min for around 37M rows. And all rows have <code>user_id</code> as index to speed up the rows retrievial once an user_id is picked. The dataframe is already sorted by user_id and timestamp, so the dataloader just needs to pad if needed and move data to GPU (4 workers). Also my train loop drops all useless move from GPU to CPU, it's quite important.</p>\n<p>Total RAM used during training: 8GB (all data prepared before and just loaded)<br>\nGPU RAM close to limit (16GB) with batch size = 256<br>\nAll tensors are <code>long</code> tensors</p>\n<p>I plan to try to run it on Kaggle kernel soon to see how it behaves.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 1094169,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-11-28T11:23:41.387000",
              "content": "<p>Beautiful! I am looking forward to your colab notebook link once we are done! And this explains the time_difference as well.</p>\n<blockquote>\n  <p>One epoch covers all 375k users but not all data for all users. It picks a sequence of 100 interactions (randomly) for each.</p>\n</blockquote>\n<p>What i am doing is all possible sequences for any user in any epoch capped at a SEQ_LEN as one row of training data from me. But your strategy is better i believe ❤️. Can't wait to go about doing this on V100's now!</p>\n<blockquote>\n  <p>Also my train loop drops all useless move from GPU to CPU, it's quite important.</p>\n</blockquote>\n<p>Very IMP!</p>\n<p>The next thing which i want to do is train an encoder first that can predict the next content_id and then add decoder etc to it if needed.</p>\n<p>Plus you are using MaskedBCE, only an encoder?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1094180,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-28T11:34:24.887000",
              "content": "<p>MaskedBCE is the name of my custom loss to workaround the padding mask issue I've with <code>nn.Transformer</code>. It allows to ignore padded parts.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1094356,
              "author_name": "Yih-Dar SHIEH",
              "author_url": "",
              "post_date": "2020-11-28T14:41:24.530000",
              "content": "<p>But I saw MSE on the screenshot?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 1094494,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-28T16:58:47.617000",
              "content": "<p>Yes, my MaskedBCE includes either BCE or MSE (so I should have named it differently).</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1094497,
              "author_name": "Yih-Dar SHIEH",
              "author_url": "",
              "post_date": "2020-11-28T17:02:50.440000",
              "content": "<p>Would you mind to share your CV/LB for MSE / BCE losses? I am busy to debug and haven't been able to try extra configuration yet …</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1094520,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-11-28T17:24:41.233000",
              "content": "<p>MSE loss should raise eyebrows..! Trying to think why it's for 😅</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1094537,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-28T17:42:53.180000",
              "content": "<p>Sure, for fold1: <br>\nBCE:  Loss=0.5627, AUC 0.7685, LB=0.775<br>\nMSE:  Loss=0.1907, AUC 0.7687, LB=0.768</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 1094632,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-11-28T19:17:06.310000",
              "content": "<p>Your results are a bit odd, usually i get a better score with mse than with BCE (between +0.5 and +1 pts in AUC) did you think about rescaling you labels (for exemple 0 to 0 and 1 to 10)? This might help your mse loss to perform better (for values between zero and 1 the mse loss is very low as square diminish the value between zero and one)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1094681,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-28T20:10:58.327000",
              "content": "<p>Nope, I will try now. Thanks.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1096094,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-30T08:01:07.643000",
              "content": "<p><a href=\"https://www.kaggle.com/rously\" target=\"_blank\">@rously</a> I've tried to scale up the labels to 0-10, moved loss to MSE, removed the sigmoid and scale down labels/predictions to 0-1 just before metrics but I get similar results. MSE loss: 19.16, AUC: 0.766</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1094147,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-28T11:04:56.293000",
          "content": "<p>Ya, I am also dead now 🤐🥺.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1094258,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-28T12:51:16.983000",
          "content": "<p>if you use the TPUs on collab you can even reduce the training time at about 1.5min per epoch, with the whole dataset seen at each epochs.</p>\n<p>Took me a bit of time to understand how to use properly TPU though<br>\n(you need to put your data on tfrecord format on a private google cloud storage and connect it to the collab env …)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1094278,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-28T13:07:14.517000",
          "content": "<p>Is that even worth it?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1094325,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-28T14:10:16.013000",
          "content": "<p>TFRecord isn't needed if you use PyTorch but it's tricky to get it working. If you are on Colab Pro and have sufficient RAM, it can be done but will take some time. Nonetheless 14-15 mins is fine for one epoch. <a href=\"https://www.kaggle.com/adityaecdrid/simple-xlmr-tpu-pytorch\" target=\"_blank\">ref</a>. </p>\n<p>Just a warning, It's not straightforward to do it and there are stability issues.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1094625,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-28T19:14:22.453000",
          "content": "<p>You can also bypass tfrecords on tensorflow if you can fit your whole dataset in ram, but my dataset processed takes more than 10go in total, so it is a bit tricky to use as is ^^</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1104665,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-07T07:07:56.327000",
          "content": "<p>'window size-1' dummy padding at beginning and end?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1091985,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-11-26T12:52:42.800000",
      "content": "<p>The best I can get so far CV=0.764 and LB=0.775<br>\nIt required more memory optimizations to make it works within 3 hours.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fbcf2f62ad526a9a1f64d31409d0e37c7%2Finference3.png?generation=1606395072707510&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1092032,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-26T13:42:46.757000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> , for LB, do you train on your validation dataset before submitting? Currently I can't pass 0.770, I am not sure if I should try to train on the validation dataset before submitting ….</p>",
          "votes": 2,
          "replies": [
            {
              "id": 1092299,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-26T17:20:27.993000",
              "content": "<p>No, just on train fold (1 fold currently). I'm going to try with another fold and ensemble both.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 1092041,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-11-26T13:50:03.303000",
          "content": "<p>my LB score can't pass 0.77 either</p>",
          "votes": 1,
          "replies": [
            {
              "id": 1109396,
              "author_name": "Jaideep",
              "author_url": "",
              "post_date": "2020-12-11T16:16:51.627000",
              "content": "<blockquote>\n  <p>my LB score can't pass 0.77 either<br>\n  <a href=\"https://www.kaggle.com/yangxiaoshuai\" target=\"_blank\">@yangxiaoshuai</a>  how are  you passing inputs.. <br>\n  i think there are two possible approaches. <br>\n  q[1:],qa[1:] then shifting embeddings<br>\n  or q[1:],qa[:-1] to take care of shifting of embeddings.</p>\n</blockquote>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1092078,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-26T14:24:06.773000",
          "content": "<p>I cannot beat .74 🥺; there are some bugs for sure and my CV is higher than LB when i use SAINT. <br>\nAny tips on this would be helpful?</p>\n<p>There are some bugs which i am aware of and trying to find a fix for them. But why do you want to make it work in 3 hours? We have 9 hours for GPU's submission.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 1092306,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-26T17:24:54.257000",
              "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> </p>\n<blockquote>\n  <p>But why do you want to make it work in 3 hours?</p>\n</blockquote>\n<p>Good question 👍 <br>\nIt's not a secret but I plan to ensemble it with other models so I need to make each as fast as possible. If SAINT+ takes 9 hours then I'm done with it.</p>\n<p>Currently I'm stuck with LGB models, I cannot train them anymore, too much features and not enough memory.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1092308,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-11-26T17:27:22.697000",
              "content": "<p>Ya, trying to squeeze out as much as possible so that i can have 2-3 models at-least;</p>\n<blockquote>\n  <p>Currently I'm stuck with LGB models, I cannot train them anymore, too much features and not enough memory.</p>\n</blockquote>\n<p>I am just curious, have you tried only running on 10M rows with SAINT's feats into your lgbm as well?</p>\n<p>Also, try to avoid \"object\" dtype, i saw my RAM exploding because if it; (I am sure you are aware of this, just leaving what i found)<br>\nTy!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1092318,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-26T17:33:31.487000",
              "content": "<blockquote>\n  <p>I am just curious, have you tried only running on 10M rows with SAINT's feats into your lgbm as well?</p>\n</blockquote>\n<p>Nope, not yet.</p>\n<blockquote>\n  <p>Also, try to avoid \"object\" dtype, i saw my RAM exploding because if it; (I am sure you are aware of this, just leaving what i found)</p>\n</blockquote>\n<p>Yes, it's not easy to optimize everything, for some structure I had to use built-in types and for other numpy to balance memory/compute performances.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1092357,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-11-26T18:20:49.713000",
              "content": "<p>Yep, it's more of a software challenge than ML :) (60-40%) (personal sentiments)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1093179,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-11-27T14:19:01.580000",
              "content": "<blockquote>\n  <p>Currently I'm stuck with LGB models, I cannot train them anymore, too much features and not enough memory.</p>\n</blockquote>\n<p>I cannot train on my lapi (16 gigs) for more than 8M rows now with 15 feats :(</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1092671,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2020-11-27T04:42:00.263000",
          "content": "<p>Similar here. Using 1kw training samples, my transformer model achieve 0.766 auc, while with 2kw training training samples, the LB AUC is 0.769. Haven't trained models using more data yet. But curious  about how to achieve AUC more than 0.79+ as the paper stated… And as for the inference time, I found it's very unstable. My submit time for transformer models change between 1hour and 3 hours(even for the same model, the submission time is changing). Similar things happen for LGB models submission…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1092963,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-27T10:37:53.010000",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> Inference time is also unstable for me, for the same model/inference code, it's between 2h15 to 3h. BTW: Which loss are you using?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1093331,
          "author_name": "HAO",
          "author_url": "",
          "post_date": "2020-11-27T16:09:36.173000",
          "content": "<p>Currently it's nn.CrossEntropyLoss, but I'm planning tuning it also the overall structure after my work with GBDT is done.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1093350,
          "author_name": "vvs",
          "author_url": "",
          "post_date": "2020-11-27T16:28:45.457000",
          "content": "<p>within 3 hours is so fast, use greedy decoding?<br>\nI'm struggling with inference time…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1074914,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2020-11-11T08:27:13.503000",
      "content": "<p>I have opened a Notebook with a light version of a Transformer Encoder in TensorFlow. Feel free to check it out to gather some ideas. <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">NOTEBOOK</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1096744,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-11-30T18:33:56.887000",
      "content": "<p>I've forked a SAKT <a href=\"https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing\" target=\"_blank\">kernel</a> which is super fast (train + test) within 3 hours. Thanks <a href=\"https://www.kaggle.com/wangsg\" target=\"_blank\">@wangsg</a> for the baseline.</p>\n<p>It shows a simple way to index data to speed up dataloading. It uses almost all data.<br>\nI've added random sequence picking + simple train/valid split + score fixes.</p>\n<p>SAKT is not SAINT as it uses an encoder only but it's a good place to start.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1096787,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-30T19:23:19.197000",
          "content": "<p>This is beautiful! Just curious, Have you plotted attention weights as well!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1100122,
          "author_name": "Gannon Reynolds",
          "author_url": "",
          "post_date": "2020-12-02T21:26:11.463000",
          "content": "<p>That SAKT kernel has been super useful, but it doesn't use close to nearly all the data. It uses one 100 max sequence from every user. But for some users that's less than 1 percent of their interactions. It ends up using only about a quarter of the potential data.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1085471,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-11-20T22:35:42.523000",
      "content": "<p>Another result with masked loss for padding (as I'm not able to have padding mask working in Transformer) and fixed total questions length (thanks <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> for the good catch!)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F71b24b4a24b4e6875a95c45c28cf7a29%2Finference2.png?generation=1605911666424581&amp;alt=media\" alt=\"\"> <br>\nI think we can do better.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1085626,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-21T03:16:41.407000",
          "content": "<p>Great Job Pal! Godspeed :)</p>\n<blockquote>\n  <p>masked loss for padding</p>\n</blockquote>\n<p>I am not sure I understood this part; Is it that you trained it just like we do it in NLP when using Bert etc with a MLM loss?</p>\n<p>Your input sequence could be something like <code>[CLS] last_5_tokens [SEP] remaining_tokens [PAD]…</code><br>\nand something similar to MLM loss on <code>remaining_tokens</code>? In that case, How are you passing other items e.g. elapsed_time etc and dealing with the fact that Bert restricts to 512 as max_seq_len?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1078159,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-11-14T13:03:24.867000",
      "content": "<p>Some additional results depending on train data size (sequence length = 100)</p>\n<ul>\n<li>33% of data: Local CV= 0.738</li>\n<li>50% of data: Local CV = 0.749</li>\n<li>90% of data: Local CV = 0.757</li>\n<li>95% of data: Local CV = 0.760 (update)</li>\n</ul>",
      "votes": 4,
      "replies": [
        {
          "id": 1079062,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-15T15:18:31.337000",
          "content": "<p>Hey mate, congrats on your LB update!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1079561,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-16T07:59:46.240000",
          "content": "<p>My current LB is with LGB not with transformer yet. I'm working on inference kernel with a trained transformer, I need to optimize it as it will exceed the 9h time limit. I should have first result this week.<br>\nOne difficulty is around the padding for new user.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1079590,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-16T09:13:11.703000",
          "content": "<p>Oh! Then could you tell us when it's done?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1079595,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-16T09:17:08.810000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1081047,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-16T18:47:42.317000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> Currently it does not fit within 9h. I need to try to optimize it.</p>\n<p><strong>Update</strong>:  I've finally reached acceptable inference time, I would like to share the pain points and some solutions:</p>\n<ul>\n<li>History: Keep last 100 interactions <strong>per user in memory</strong></li>\n<li>Build test sequence <strong>on-fly</strong> from current test_df and history for each user</li>\n<li>Forget pandas to concat or basic clean operations, move to plain <strong>list or numpy</strong>.</li>\n<li>Avoid workers &gt; 0 in Pytorch dataloader, batches are small due to API and for some reasons concurrent workers introduce large overhead.</li>\n</ul>\n<p>First test submission inference time: 5h</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1082265,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-17T18:38:54.173000",
          "content": "<p>After all optimizations, inference time = <strong>3h</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F6febdde32456300991e2b4ac63a07c41%2Finference.png?generation=1605638316021117&amp;alt=media\" alt=\"\"></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1082268,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-17T18:41:06.070000",
          "content": "<p>Wow <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> ❤️🎊🎉; You are gonna rock the LB soon! Congratulations! Now i can go back and start working again on the transformers! As it's possible to do it 😅</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1082276,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-17T18:48:05.783000",
          "content": "<p>Few doubts I have,</p>\n<ul>\n<li>History: Keep last 100 interactions per user in memory</li>\n</ul>\n<p>Did you use a deque for this ?</p>\n<ul>\n<li>Build test sequence on-fly from current test_df and history for each user</li>\n</ul>\n<p>Append operation on the deque would take care of it automatically i feel (but you have to use your own collate_fn)</p>\n<ul>\n<li>You are doing things that's mentioned in the paper more or less like triangular masking etc?</li>\n</ul>\n<p>Ty!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1082278,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-17T18:53:31.470000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> would you mind to share your model archeticture and model size? I have difficulty to go beyond 0.765 ….</p>\n<p>For example, do you use encoder-decoder? If so, how do you do inference when submitting?<br>\nBecause it seems we need to perform inference at several timestamps in each test batch, and it is kind time consuming, no?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1082293,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-17T19:15:14.757000",
          "content": "<p>Nice <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> mate, you really deserve. Also, I was wondering if any of you guys could help me out to understand something.</p>\n<p>I've been training with an only output per sequence, as you could have read before in this discussion. I'm trying to switch to a sequence output, having [A, B, C] with shifted answered_correctly as input and predicting answered_correctly for [A, B, C].</p>\n<p>My main concern is: what if you have the sequence [F, G, H]? Predicting the answered_correctly for F wouldn't make sense since in that sequence you dont have previous interactions.</p>\n<p>How do you train those interactions above the sequence length?</p>\n<p>EDIT: They don't say anything about this in the reference papers.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1082297,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-17T19:19:58.043000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> - not clear to me about your question. What the relationship between [A, B, C] and [F, G, H] …?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1082301,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-17T19:23:59.020000",
          "content": "<blockquote>\n  <p>My main concern is: what if you have the sequence [F, G, H]? Predicting the answered_correctly for F wouldn't make sense since in that sequence you dont have previous interactions.</p>\n</blockquote>\n<p>I guess, you are worrying about unseen users for which we have no history, right?<br>\nIn that case, I would default the preds to 0.67 for the first interaction. [keeping things simple, as i don't have a transformer working 😅]</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1082306,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-17T19:30:39.180000",
          "content": "<p>if the concern is the new user, than the model will learn the distribution of the answer correction for each question, independent of the users.</p>\n<p>Just like the prediction the first word in a corpus.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1082312,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-17T19:37:45.830000",
          "content": "<p>I'll try to explain again. Imagine you have 200 interactions for a user, and a window size/sequence length of 50. If you try to input [100:150] for example, the output for the first element in the sequence makes no sense, since in that specific sequence you don't have any prior interactions, but the user actually interacted 100 times before that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1082314,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-17T19:41:16.443000",
          "content": "<p>Have into account that you output the same number of elements than the input has. The last element of the output makes sense, but the rest of them don't (except for the first sequence length interactions).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1082322,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-17T19:48:48.623000",
          "content": "<p>This only applies for the training phase. In the inference, you just take whatever element in the output you have to. The question is: what is the utility of training N to N instead of N to 1?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1082330,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-17T19:55:40.287000",
          "content": "<p>I don't have a good answer for it. If we use window size, that's it. I think it as <code>a model use the previous fixed length of history to predict things</code>.</p>\n<p>If you want, a simple add-on will be use the absolute pos embedding --&gt; say the 1000th interaction. So at least you have something global in user history added into the model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1082339,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-17T20:09:43.127000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> No need to use dequeue, just use of built-in python list like:</p>\n<pre><code>my_list = [] # Start with empty list\nmy_list.extend([1,2,3]) # now list is [1,2,3]\nmy_list.extend([4,5,6,7,8,9,10,11]) # now list is [1,2,3,4,5,6,7,8,9,10,11]\nmy_list = my_list[-10:] # now list contains only the last 10 items: [2,3,4,5,6,7,8,9,10,11]\n</code></pre>\n<p>Maintain a list for each user and for each items (question_id, answer, elapsed time …) with the last 100 interactions and it will fit into memory and it will be much faster than <code>np.append(...)</code> or <code>pd.concat(...)</code> by an order of magnitude!</p>\n<p>I'm almost following the SAINT/SAINT+ papers. I've added PRIOR_QUESTION_HAD_EXPLANATION has a new feature and I'm using cos/sin positional encoding. My sequence length is 100. Embedding dimension is 256, so my batches are (BS, 100, 256) and output of my model is (BS, 100) followed by <code>nn.BCEWithLogitsLoss()</code> to compute loss and then <code>torch.sigmoid()</code> to get probabilities.</p>\n<p>Currently I'm not using padding mask at all, I'm padding with existing sequence from another user. For self attention, I'm using the triu mask.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1082348,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-17T20:17:34.923000",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> . Do you perform several inferences in each test bacth (because there are potentially multiple questions in each bundle )?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1082354,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-17T20:25:45.187000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> My model encoder + decoder, configuration is:</p>\n<pre><code>class raw_conf:\n\n    mtype = \"SAINT\"\n    backbone = \"transformer\" \n\n    pad_mode = \"random\" # i.e. pad with data from another user\n    seq_len = 100\n    embedding_dim = 256 # embed_dim must be divisible by num_heads\n    exercices_id_size = 13782\n    exercices_part_size = 7\n    response_size = 2 \n    elapsed_time_size = 300 \n    lag_time_size = 720 # It was 1440 in SAINT paper.\n    explanation_size = 2\n    position_encoding_enabled = True\n\n    # Model\n    nhead = 8\n    num_encoder_layers = 4\n    num_decoder_layers = 4\n    dim_feedforward = 2048\n    dropout = 0.1\n    activation = None\n</code></pre>\n<p>For inference, assuming an API batch with 16 users and 24 questions, I build - on-fly with user's history in memory - a tensor with (24, 100) and after prediction I get (24,1) by returning only the last answer (predition). The on-fly part is tricky (pick history + padding) and must be super fast to fit the 9h time limit.</p>\n<pre><code>def per_model_predict(model, X_test, features_cols=None):\n    test_dataset = RIIIDDataset(X_test, conf, None, subset=\"test\")\n    # Data loader must be super fast, the bottleneck is here.\n    test_loader = DataLoader(test_dataset, batch_size=conf.BATCH_SIZE, shuffle=False, num_workers=0, drop_last = False, pin_memory=False)\n    # Predict is done here, no bottleneck here.\n    _, y_prob = test_loop_fn(conf, test_loader, None, model, conf.L_DEVICE, verbose=False) # (BS, 100)\n    return y_prob[:, -1]\n</code></pre>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1082359,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-17T20:36:47.310000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> About:</p>\n<blockquote>\n  <p>If you try to input [100:150] for example, the output for the first element in the sequence makes no sense, since in that specific sequence you don't have any prior interactions, but the user actually interacted 100 times before that.</p>\n</blockquote>\n<p>You can ignore it in the loss by computing BCE only from 101 to 150.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1082361,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-17T20:41:44.790000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a></p>\n<blockquote>\n  <p>Do you perform several inferences in each test batch (because there are potentially multiple questions in each bundle )?</p>\n</blockquote>\n<p>Yes. If we've 3 questions in one task container like [Q1,Q2,Q3] for one user then it becomes a tensor of (3, 100) and then 3 independant predictions for this user. All the timestamps for Q1, Q2, Q3 are the same. I don't use the prediction from Q1 to predict Q2. I only use the history for this user each time.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1082381,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-17T21:09:06.680000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>, thank you very much, very clear explanation of the approach. Keep going!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1082384,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-11-17T21:11:55.900000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Why are you using sinusoidal position encoding? I'm noticing that my model struggles to converge when using learnable position embeddings. Did you have the same problem?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1082400,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-17T21:38:11.970000",
          "content": "<p>Sinuasoidal position encoding gives better performance on my training. With \"Noam\" LR warm up and LR always below 0.0008 I don't have convergence issue.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1082529,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-18T01:35:06.740000",
          "content": "<p>Thanks for the detailed responses;</p>\n<blockquote>\n  <p>I'm almost following the SAINT/SAINT+ papers. I've added PRIOR_QUESTION_HAD_EXPLANATION has a new feature and I'm using cos/sin positional encoding. My sequence length is 100. Embedding dimension is 256, so my batches are (BS, 100, 256) and output of my model is (BS, 100) followed by nn.BCEWithLogitsLoss() to compute loss and then torch.sigmoid() to get probabilities.</p>\n</blockquote>\n<p>Shouldn't the Transformer's decoder output's shape be (T, N, E) as per docs where T is the target sequence length, N is the batch size, E is the feature number (embedding_dims)?</p>\n<p>Is your target sequence length defined as 1 as on the output side, we don't need a sequence right? Sorry, but i am still very confused about this 🤐. Would appreciate is someone can help with the same.  If it's 1, then we don't need any mask's on decoder side, right? <br>\nTy! </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2F15f4aeda8965f6cce896f8d988f7bf16%2FScreenshot%202020-11-18%20at%207.04.17%20AM.png?generation=1605663277214065&amp;alt=media\" alt=\"Decoder_SAINT+\"></p>\n<blockquote>\n  <p>Sinuasoidal position encoding gives better performance on my training.</p>\n</blockquote>\n<p>This might be better actually as it's position agnostic and can easily extrapolate to sequences longer than our training sequences len.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1082800,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-11-18T08:33:28.613000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> About:</p>\n<blockquote>\n  <p>Shouldn't the Transformer's decoder output's shape be (T, N, E) </p>\n</blockquote>\n<p>Yes, it is so you have to make sure you're passing the correct input and getting the correct output. This is what I did with <code>transpose</code> before/after using <code>nn.Transformer</code>:</p>\n<pre><code>x_position_exercices = x_position_exercices.transpose(1,0) # (seq_len, BS, embedding_dim)\nx_position_responses = x_position_responses.transpose(1,0) # (seq_len, BS, embedding_dim)\nx_transformer = self.transformer(src=x_position_exercices, tgt=x_position_responses, src_mask=src_mask, tgt_mask=tgt_mask, memory_mask=mem_mask) # (seq_len, BS, embedding_dim)\nx_transformer = x_transformer.transpose(1,0) # (BS, seq_len, embedding_dim)\n</code></pre>\n<p>My output is (BS, 100), the model tries to predict each answer of the 100, not only the last one. I think predicting only the last one will also work (as it's only what we need for inference). I'm going to try it too.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1082953,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-18T12:30:47.747000",
          "content": "<p>Thanks a lot! It's quite helpful; Keep galloping on the LB :)</p>\n<p>Edit -&gt; You are using mem_mask as well, nice; Need to check that as well 😅</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1082968,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-18T12:58:17.367000",
              "content": "<p>And the more data you include, the more iterations you've per epoch and the less epochs you need. For example:</p>\n<pre><code>Fold 1 train users: 289764 valid users: 12452\nFold 1 train size: (86298794, 6) valid size: (863478, 6)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Ff1950fc7b58b4e1e56b64091a29cd8ce%2Ftrain_0.758.png?generation=1605704249682412&amp;alt=media\" alt=\"\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 1082969,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-11-18T13:00:43.143000",
              "content": "<p>Looking forward for a fork-able version of the same! Hehe; You can actually create more data this way by swapping content_ids in and out from different user attempts as it doesn't depends on the user. Not sure, maybe you can try to \"fake\" the data but it might impact the performance.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1082978,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-18T13:06:46.640000",
              "content": "<p>I'm not sure to publish my kernel now as it provides score in medals zone. Moreover, I'm using around 90% of data so It won't load in Kaggle.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 1082983,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-11-18T13:11:51.887000",
              "content": "<p>I completely agree, (was just kidding) you shouldn't do it. You have already provided enough information for everyone to implement their own now; GL!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1082994,
              "author_name": "Yih-Dar SHIEH",
              "author_url": "",
              "post_date": "2020-11-18T13:17:51.473000",
              "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>  <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> What is mem_mask? I never heard about it for transformer.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1082999,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-18T13:23:35.193000",
              "content": "<p>It's an attention mask for output of the decoder:<br>\n<a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html</a><br>\n<a href=\"https://discuss.pytorch.org/t/memory-mask-in-nn-transformer/55230\" target=\"_blank\">https://discuss.pytorch.org/t/memory-mask-in-nn-transformer/55230</a></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 1083387,
              "author_name": "Yih-Dar SHIEH",
              "author_url": "",
              "post_date": "2020-11-18T23:17:43.277000",
              "content": "<p>Thanks, I don't use pytorch, and i just call the mem mask as decoder to encoder attention mask😄</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1083372,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-18T22:41:23.803000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1083664,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-19T08:20:16.460000",
          "content": "<blockquote>\n  <p>exercices_id_size = 13782</p>\n</blockquote>\n<p>I am just trying to combine so many tips here and there for the NN's; <br>\nIt seems that you are training on lecture videos as well from the config's?</p>\n<p>Secondly, are you training it in an auto-regressive fashion? I don't think we are as we aren't generating anything here.</p>\n<p>From the Attention Is All You Need Paper,</p>\n<blockquote>\n  <p>At each step the model is auto-regressive[10], consuming the previously generated symbols as additional input when generating the next.</p>\n</blockquote>",
          "votes": 1,
          "replies": [
            {
              "id": 1083680,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-19T08:32:07.520000",
              "content": "<p>Yes, you're correct there are only 13523 questions and I planned to use lectures later (but I did not yet). Good catch! So it should be:</p>\n<p><code>exercices_id_size = 13523 # 13782</code></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 1083686,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-11-19T08:38:44.550000",
              "content": "<p>One more I have for you, there's a start token for everything on the decoder side, right? [Cs, Ps, ETs, LTs]</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1083694,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-19T08:43:33.757000",
              "content": "<p>Yes on all, I'm not sure of the best way to do it, so I did the following:</p>\n<pre><code># Add start token to correctness\nx_correctness = torch.roll(x_correctness, shifts=(0, 1, 0), dims=(0, 1, 0)) # Shift right the sequence\nx_correctness[:,0,:] = self.response_size # Start token\n</code></pre>\n<p>People with NL skills should know the best way, maybe directly embedding with padding index parameter=0.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 1083697,
              "author_name": "Yih-Dar SHIEH",
              "author_url": "",
              "post_date": "2020-11-19T08:54:57.330000",
              "content": "<p>Hi, do you add start token to the truncated sequence, or to the full history sequence (up to the current time) before the truncation? I mean, we might have [  Q100, Q101, ….] and [R100, R101] …., and if you add start token at this truncated sequence, we will have [START, R100, …]. Not sure if this is the best, but I don't think it doesn't hurt much.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1083781,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-11-19T11:28:44.163000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1085711,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-21T06:19:47.493000",
          "content": "<p><code>For inference, assuming an API batch with 16 users and 24 questions, I build - on-fly with user's history in memory - a tensor with (24, 100) and after prediction I get (24,1) by returning only the last answer (predition). The on-fly part is tricky (pick history + padding) and must be super fast to fit the 9h time limit.</code></p>\n<p>Can you share, how much it took for the sample example_test.csv? It's taking ~3 secs for all 4 iters of the sample test_set but it's not going to pass through that time limit of 9 hours as it has to be less than 0.55/iter but it's .75/iter for me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1085791,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-21T07:43:08.203000",
          "content": "<blockquote>\n  <p>For inference, assuming an API batch with 16 users and 24 questions, I build - on-fly with user's history in memory - a tensor with (24, 100) and after prediction I get (24,1) by returning only the last answer (predition). The on-fly part is tricky (pick history + padding) and must be super fast to fit the 9h time limit.</p>\n</blockquote>\n<p>I guess my init draft of inferencing with fixed total questions length is working (atleast on the dummy data),  <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> If i understand correctly, wouldn't this mean that you will be returning the pred for padded_token as well sometimes if the seq_len isn't equal to the max_value?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1085914,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-11-21T09:20:02.067000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> would you mind to share how you calculate lag time and how to use it. As mentioned in <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197528\" target=\"_blank\">this thread</a>, for the decoder, suppose we have Qn and R(n-1) and we want to predict R(n), we use the elapsed time for answering Q(n-1), that is the <code>priior_question_elapsed_time</code> given at R(n).</p>\n<p>For lag time, I think it would be similar: We need to use the lag time between the ending of answering Q(n-2) and the stating time for answering Q(n-1). But I have some trouble to calculate this using tensors. (otherwise I have to build this information directly in the datasets and convert them to tensors later …)</p>",
          "votes": 1,
          "replies": [
            {
              "id": 1085933,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-11-21T09:35:29.047000",
              "content": "<p>Start with something simple, timestamp(N) - timestamp(N-1), even if not the exact definition it' s a good estimation.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 1085959,
              "author_name": "Yih-Dar SHIEH",
              "author_url": "",
              "post_date": "2020-11-21T09:54:09.127000",
              "content": "<p>Yes, I might over-complicates a lot of things :) I will give it a try first. Thanks</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1085962,
              "author_name": "Aditya Soni",
              "author_url": "",
              "post_date": "2020-11-21T09:56:24.717000",
              "content": "<p>If we do this, we have to take care of some outliers as well i believe. difference b/w my last logout today and my next login time can bee quite large.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1085988,
              "author_name": "Yih-Dar SHIEH",
              "author_url": "",
              "post_date": "2020-11-21T10:11:13.800000",
              "content": "<p>they use a maximal lag time of 300 seconds and categorical features for it. So it shouldn't be a real problem for large lag time (will be truncated anyway)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1128442,
              "author_name": "Jaideep",
              "author_url": "",
              "post_date": "2020-12-27T12:51:21.317000",
              "content": "<blockquote>\n  <p>they use a maximal lag time of 300 seconds and categorical features for it. So it shouldn't be a real problem for large lag time (will be truncated anyway)</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> <br>\n1)to what max time lag time can be truncated is it 1400 minutes.. or 1400*60 sec?<br>\n2) by categorical embedding means each lag time is given a time bucket ? if it falls in some range </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1108534,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-10T19:04:38.120000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <br>\nhow do you pass memory mask… i currently do simply this,where future mask is simple np.triu<br>\n<code>att_mask = future_mask(x.size(0)).to(device)</code><br>\n<code>x1=self.transformer_decoder(response,att_output,tgt_mask=resp_attn,memory_mask=att_mask)</code></p>",
          "votes": 0,
          "replies": [
            {
              "id": 1108541,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-12-10T19:14:16.387000",
              "content": "<p>I'm using <code>nn.Transformer</code> directly. There is a parameter dedicated for mem_mask.<br>\n<a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1108675,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-10T22:24:29.047000",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> :</p>\n<p>Finally, there is minimal change in my code to having your sampling stragegy, something like</p>\n<pre><code>            ds = ds.flat_map(\n                lambda x: tf.data.Dataset.from_tensors(x).repeat(\n                    tf.cast(tf.random.uniform(shape=[]) &lt; self.training_sample_prob_table.lookup(x['user_id']), tf.int64)\n                )\n            )\n</code></pre>\n<p>Just repeat the (full) sequences of each user, depending on a probably pre-computed in a table like</p>\n<pre><code>        initializer = tf.lookup.KeyValueTensorInitializer(user_id_tensor, sample_prob_tensor)\n        training_sample_prob_table = tf.lookup.StaticHashTable(initializer, default_value=0.0)\n        self.training_sample_prob_table = training_sample_prob_table\n</code></pre>\n<p>I don't know yet how much it helps for me, because I am still not able to get 0.78 with your model size. I don't know what causes it:</p>\n<pre><code>- I use position embedding (learnable) for pos in [0, 1, ... window_size - 1]. I guess you use `sinusoidal` embedding?\n\n- You use softmax and I use sigmoid.\n\n- You use lag time, I don't\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1108896,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-11T05:50:29.493000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>  </p>\n<pre><code>src – the sequence to the encoder (required).\n\ntgt – the sequence to the decoder (required).\n\nsrc_mask – the additive mask for the src sequence (optional).\n\ntgt_mask – the additive mask for the tgt sequence (optional).\n\nmemory_mask – the additive mask for the encoder output (optional).\n</code></pre>\n<p>what i meant was for every mask  input we can pass this way ?</p>\n<pre><code>src_mask = future_mask(x.size(0)).to(device) #using triu\ntgt_mask=  future_mask(response.size(0)).to(device) #using triu\nmemory_mask =future_mask(x.size(0)).to(device) #using triu\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1134850,
      "author_name": "Gannon Reynolds",
      "author_url": "",
      "post_date": "2021-01-01T16:24:31.470000",
      "content": "<p>I'm in the mix with a SAINT implementation! 0.778 LB with just content_id so far. Super stoked to improve on that. </p>\n<p>I do have a question related to the paper though, if anyone cares to enlighten me. In the SAINT, paper they state that the decoder takes in a sequential input R of response embeddings with start token embedding S. As of now, my implementation doesn't have a unique starting token embedding. It's just padded with zeros. Is this much of an issue? Can anyone share how they're dealing with this start token embedding? I'm not sure I understand that part.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1134882,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2021-01-01T17:09:09.987000",
          "content": "<p>Congrats!</p>\n<p>I haven't heard the term 'stoked' since I left Northern California :-).</p>\n<p>Regarding the starter tokens on the decoder side—imo, <strong>none</strong> of them are necessary except, perhaps, on the 'correctness embedding'. I can't say more than that though =] we're getting near the finish line.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1134969,
          "author_name": "Darren Lahr",
          "author_url": "",
          "post_date": "2021-01-01T18:46:28.487000",
          "content": "<p>I incremented by 1 on the answered_correctly field and use (0 for padding, 1 for answered incorrectly, 2 for correct, 3 for start token)</p>\n<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> Good point around not needing the starter tokens, I had overlooked that point thanks :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1107907,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-12-10T04:30:34.847000",
      "content": "<p>Alright, so here's one of my concern which i am not able to help find answer to, Let's assume your batch looks like, (0 represents the padding token) (assume that next question you are going to attempt is the last non_zero value) (6,8,12,19,25,25)</p>\n<pre><code>tensor([[ 1,  2,  3,  4,  5,  6],\n        [ 7,  8,  0,  0,  0,  0],\n        [10, 11, 12,  0,  0,  0],\n        [14, 15, 16, 17, 18, 19],\n        [20, 21, 22, 23, 24, 25],\n        20, 21, 22, 23, 24, 25]])\n</code></pre>\n<p>So in this case, the src_mask i.e. the self_attention mask in the encoder side, that shouldn't simply be a upper-traingular matrix, right?</p>\n<p>Like if we give mask like the below, then that's incorrect, right?</p>\n<pre><code>tensor([[0., -inf, -inf, -inf, -inf, -inf],\n        [0., 0., -inf, -inf, -inf, -inf],\n        [0., 0., 0., -inf, -inf, -inf],\n        [0., 0., 0., 0., -inf, -inf],\n        [0., 0., 0., 0., 0., -inf],\n        [0., 0., 0., 0., 0., 0.]])\n</code></pre>\n<p>And it should actually be like this, right? (basically i would call it a length mask 😅)</p>\n<pre><code>tensor([[0., 0., 0., 0., 0., -inf],\n        [0., -inf, -inf, -inf, -inf, -inf],\n        [0., 0., -inf, -inf, -inf, -inf],\n        [0., 0., 0., 0., 0., -inf],\n        [0., 0., 0., 0., 0., -inf],\n        [0., 0., 0., 0., 0., -inf]])\n</code></pre>\n<p>Is my understanding apt?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1107933,
          "author_name": "Frank Pan",
          "author_url": "",
          "post_date": "2020-12-10T05:10:58.473000",
          "content": "<p>No, your understanding is not correct. The mask is not applied on this dimension, instead, it is applied to every attention head.<br>\nThe upper-triangular mask will make sure at each step we only look at the past.<br>\nIn you above example, every position needs a different mask. When you combine all the masks for all steps, you will get a upper-triangular matrix.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1107968,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-10T06:14:11.953000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> for the clarifications!  Then it will remain the same as tgt_mask as well? (look-ahead-mask from decoder point of view) as there also we need to prevent looking ahead at future tokens pretty much.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107983,
          "author_name": "Frank Pan",
          "author_url": "",
          "post_date": "2020-12-10T06:34:46.527000",
          "content": "<p>That is a different thing, although they are all called \"mask\". The target mask is only to filter out the steps we care about, i.e. excluding padding and lectures in this case.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107989,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-10T06:44:15.427000",
          "content": "<p>You don't use lecture information? (i.e. no attention to it?)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107993,
          "author_name": "Frank Pan",
          "author_url": "",
          "post_date": "2020-12-10T06:58:34.677000",
          "content": "<p>I use lecture information. What I mean here is that you might want to exclude lectures when calculating the loss and AUC. It won't be a very big problem if you don't, other than causing some distortion to the metrics.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1109058,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-11T09:39:55.617000",
          "content": "<p><a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a>  i dont use lectures and pass the mask this way  ,where future mask is np.triu.<br>\nI do this as i see in Architecture that every input to MultiAttn head is through Triu mask.</p>\n<pre><code>att_mask = future_mask(x.size(0)).to(device)\nresp_attn= future_mask(response.size(0)).to(device)\n        x1=self.transformer (src=x,tgt=response ,\n                             src_mask=att_mask,\n                               tgt_mask=resp_attn,\n                               memory_mask=att_mask)\n</code></pre>\n<p>Let me know if am following the paper correctly. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1109234,
          "author_name": "Frank Pan",
          "author_url": "",
          "post_date": "2020-12-11T12:59:44.987000",
          "content": "<p>Actually I just realized you guys are talking about the Transformer Module of Torch <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>. So what I said about tgt_mask above was not correct.</p>\n<p>To my understanding src_mask and tgt_mask should all be triu matrix. <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> your implementation looks good to me, as long as size(0) is the sequence length (not batch size) here. (Since the length of x and tgt is the same in this case, maybe you can just use att_mask for all three masks?)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1109303,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2020-12-11T14:17:29.713000",
          "content": "<p>I think src[tgt]_key_padding_mask can be used for individual sequence length masking in the transformer module of torch.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1117041,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-17T17:08:29.307000",
          "content": "<p>This was a great sub-thread conversation.</p>\n<p>Along the same lines, I am interested in how you all are handling <em>validation</em>. In my current scheme, I have a number of users who are only in train and not val, and a number of users who are in both. Since I don't have user_id embeddings, it makes the most sense that users in both should have their historical data available (fed to the transformer model) during val; but for validation sake, I don't want the validation loss/metric to include those historical values. Is anyone handling this using double masks, or how are you all dealing with that?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1119990,
          "author_name": "Frank Pan",
          "author_url": "",
          "post_date": "2020-12-20T14:18:18.647000",
          "content": "<p>I just use 95% users to train, 5% users for validation (only take the latest seqence length of history for validation). This is certainly not ideal as some of the targets in validation can only rely on none to very short history to predict, even if longer history exists. Therefore the local validation score will be lower than LB score.  But my submissions suggest higher local validation score always means higher LB score, so in practice this might be good enough.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1096342,
      "author_name": "TitiTest",
      "author_url": "",
      "post_date": "2020-11-30T12:29:36.033000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4570913%2Fb5c717170357fb281614b150c232e4a0%2Fxy2.png?generation=1606666420284171&amp;alt=media\" alt=\"\"><br>\nHi, this graph is AUC for given position in a 128 sequence length of a transformer model.<br>\nI think we can interpret the slope as the ability of the model to take into account the past …</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1096685,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-11-30T17:36:05.120000",
          "content": "<p>Is this for train data or validation?</p>\n<p>I have a similar trajectory as well.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2555700%2Fb25a23e4905be16922936b5f320e1b1d%2FScreenshot%202020-11-30%20223258.png?generation=1606757745828460&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1096702,
          "author_name": "TitiTest",
          "author_url": "",
          "post_date": "2020-11-30T17:50:39.300000",
          "content": "<p>Validation on about 2% of the dataset and you ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1096914,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-11-30T21:46:50.850000",
          "content": "<p>Validation, 3% of users. I run validation for complete users sequences (max 128).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1082028,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-11-17T14:24:40.857000",
      "content": "<p>I am struggling to make my lgbm inference properly for 4 days now, so forget about DL models inference here for now; (to save your time on more stronger features that can help your lgbm models);  Another reason is, one small mistake and the whole thing will blow, as it's quite unforgivable in nature</p>\n<p>Simple idea is to use the feature vectors from them as inputs to your lgbm to get started where you know you won't be breaking anything for sure! (as that should have some visible improvement on your lgbm model)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1138890,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2021-01-05T03:34:30.850000",
      "content": "<p>Not a great implementation, but as promised in my old kernel, here's the one which has <a href=\"https://www.kaggle.com/adityaecdrid/fork-of-saint-inference-ea970c\" target=\"_blank\">SAINT's Inference</a>. It's buggy somewhere (~.739 on LB )and I am not able to fix it on my own, so making it public, hoping it will help someone for sure!</p>\n<p>NB, there are bugs in it, so use it by analysing the same, the code's for reference only.</p>\n<p>Thanks for all your shares!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1133430,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2020-12-31T07:59:11.803000",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> Is there any chance you share a bit of the secret that boosts your LB from 0.798 to 0.806? That's very amazing, you are doing so great!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1133475,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-31T08:46:01.947000",
          "content": "<p>Here it's not easy to climb, the team is working hard. You also made solid progress! with transformer too I guess.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1133533,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-31T09:40:18.523000",
          "content": "<p>Lastest update 79.3 is saint plus. we are working  towards making it better<br>\n<a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>  did you use Lagtime , what is way u computed it  ts2-ts1 ?</p>",
          "votes": -2,
          "replies": [
            {
              "id": 1133577,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2020-12-31T10:44:40.370000",
              "content": "<p>Yes, timestamp(N) - timestamp(N-1). It's not exactly the same definition as in SAINT+ paper but an approximation</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1133583,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-31T10:54:34.727000",
          "content": "<p>Thanks, I'm just fixed a few bugs (2 to be precise) :) <br>\nNo change in the model architecture or pipeline.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1133664,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-31T12:17:00.807000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Yes, difficult … and yes, worked (and will work) hard, very hard …😂😆</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133753,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-31T13:53:33.353000",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> if possible, could you share a bit, other than the features used in the original SAINT / SAINT+, what are (some of) the extra things you have added to your model / features?</p>\n<p>I am surprised that you could get LB <code>0.792</code> with <code>d_model=128 with n_layer=2</code>. If you have no extra things in your model / features, I feel I must have done something wrong … (I also tried to debug, but the inputs for my training/validation has no difference to my submission pipeline … i.e. no bug found anymore)</p>\n<p>I hope I can hear a bit from you - at least give me some hope 😄</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133781,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-31T14:23:30.593000",
          "content": "<p>I reached 0.775 with pure saint iirc. I added the time features to get up to 0.788, fixed a few bugs and got 0.792. Increasing model size from there took me to 0.798 and fixing a few more bugs helped me reach 0.806. Now these bugs were present all along the way so not sure what to make of it.</p>\n<p>I think you should have 0.79 just by using SAINT+ with 128 d_model and  2 n_layers.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1133788,
          "author_name": "hrunic",
          "author_url": "",
          "post_date": "2020-12-31T14:26:17.910000",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> Do you use all the data for training? And how long does it take you to train one epoch?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1133800,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-31T14:38:28.557000",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> Thank you! One last question (at least for today), what's is your CV strategy, and what's your validation dataset size?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1133818,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-31T15:02:16.700000",
          "content": "<p>only use kaggle gpu?amazing</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133819,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-31T15:04:24.983000",
          "content": "<p>Just keeping a few users (3~4%) unseen for validation. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133822,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-31T15:05:48.417000",
          "content": "<p>Yeah, and TPUs.<br>\nLearned how to use TPUs 2 weeks ago. Got inspired by <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1133827,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-31T15:10:43.270000",
          "content": "<p>Looking forward to your proposal after the competition end.Can I know the parameters of your pure saint model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133830,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-31T15:13:21.390000",
          "content": "<p>d_model 128 and n_layers 2</p>\n<p>Edit: I'm using a modified version now with higher dimensions and layers.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133863,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-31T15:36:25.747000",
          "content": "<p>is it as heavy as mentioned in paper. I feel  512 dim, 6 /6  thats an  overkill given the resource every one has.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133866,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-31T15:37:17.490000",
          "content": "<p>I definitely helped a DL monster to grow further 😆</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133870,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-31T15:38:40.537000",
          "content": "<p>So you only validate on unseen users? I have keep users with some training history and some new users</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1133876,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-31T15:43:45.190000",
          "content": "<p>Hope my TPU notebooks get much more votes once you give your <code>thank you speech</code> after winning 😊</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1133917,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-31T16:35:41",
          "content": "<p><a href=\"https://www.kaggle.com/jjaideepvalani\" target=\"_blank\">@jjaideepvalani</a></p>\n<p>It's currently at 512 with 4 layers but I think it overfits a bit. I'll try with 256. Thing is the time difference in training with 256 vs 512 is not more than ~60 secs per epoch so I go with 512 while training.</p>\n<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> Amen to that :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1133932,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2020-12-31T16:59:57.437000",
          "content": "<p>My goal of Saint model is to achieve the effect of the paper, I am just a novice, and cost too much time in lgb, I can not achieve a high score。</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1134939,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-01-01T18:05:32.597000",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> will we have neg lag times ,apart from the ones that correspond to timestep -0 ?<br>\nif yes around how many ?</p>",
          "votes": -3,
          "replies": []
        },
        {
          "id": 1134956,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2021-01-01T18:21:11.023000",
          "content": "<p>lag time all above 0,</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1134965,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2021-01-01T18:36:55.250000",
          "content": "<p>I can't comment on the time feature for now</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1142307,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2021-01-07T09:59:32.260000",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>  congrats for  continuing to be in gold hunt.<br>\nLooks like after hard efforts we might miss by few miles :)</p>\n<p>Any thing one can think around to go beyond .806 .  :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1142313,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2021-01-07T10:02:55.740000",
          "content": "<p>Better data usage </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142320,
          "author_name": "Dean",
          "author_url": "",
          "post_date": "2021-01-07T10:08:48.490000",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> Your team have great work and have good rank.<br>\nIf you don't mind. How about your single model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142821,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-01-07T15:58:53.963000",
          "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a>  around same as ensemble only.. close to what  bro abdur  had got earlier few days ago. <br>\nBtw any way to fasten up the inference. I put optimized merges only, but i think too much time goes in status updates. using np.append.</p>\n<p>I regret that i lost my resources so couldnt contribute to experiments but was trying to compensate best by doing analysis and coming up with new features , improving scores</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1103856,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2020-12-06T11:19:10.600000",
      "content": "<p>Any of you still working this way?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1103878,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-06T11:50:29.150000",
          "content": "<p>Okay, I think I'm sharing a full description of my current best model since I have taken some ideas from this discussion:</p>\n<ul>\n<li>CV: 0.7610, LB: 0.767</li>\n<li>Model<ul>\n<li>2 Encoder Layers, 2 Decoder Layers</li>\n<li>Model dimension: 128</li>\n<li>Number of heads: 4</li>\n<li>Sequence Length: 96</li>\n<li>Padding mask and look-ahead mask in both encoder and decoder.</li>\n<li>Padding-masked loss.</li></ul></li>\n<li>Features (differences with SAINT+):<ul>\n<li>No absolute position encoding. Instead relative position in the MHA layers.</li>\n<li>Using task_container_id embeddings which I add to the decoder.</li>\n<li>I use lag (scaled 0-1), elapsed_time (scaled 0-1) and prior_question_had_explanation as continous features (they all together in a dense layer that transform to model dimension size, activated with tanh). It is added to the decoder.</li></ul></li>\n<li>Training procedure:<ul>\n<li>Batch size: 512</li>\n<li>Optimizer: Adam (beta_1=0.9, beta_2=0.999, epsilon=1e-8) with a LR scheduler as stated in SAINT+ paper.</li>\n<li>Loss: Categorical crossentropy (2 output neurons activated with softmax).</li>\n<li>AUC computed taking only into account last non-padded element in every sequence.</li></ul></li>\n<li>Input flow:<ul>\n<li>80 M rows minus lectures (311567 users).</li>\n<li>For training, random sample a sequence per user per epoch, having a linear increasing probability (the last sequence in a user has double the chances to be selected than the first one).</li>\n<li>For validation I use every possible sequence for 4% of the users. </li></ul></li>\n</ul>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 1104350,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-06T20:51:51.400000",
          "content": "<p>For me num_layers = 2 has been more effective, the model diverges (AUC gets stuck around 0.63) if I increase this even to 3.<br>\nSimilar is the case with d_model, if I increase this from 128, the model again diverges. </p>\n<p>CV: 0.775 (No LB for this)</p>\n<p>Maybe it's an issue with my implementation.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1104406,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-06T22:47:13.907000",
          "content": "<p>Mine still converges when bigger but no needed. With my current size is actually overfitting in a few epochs, and training keeps improving over time (hence showing the model can still learn). I still don't understand how can we obtain the metris reported by SAINT+ neither why they use such a big model (4 encoder layers, 4 decoder layers, 512 model dimension).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1106063,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-08T13:26:40.173000",
          "content": "<p>Still on it but fighting with memory/time out issues since a few days. Fighting for both training and inference!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1106077,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-08T13:50:18.117000",
          "content": "<blockquote>\n  <p>AUC computed taking only into account last non-padded element in every sequence</p>\n</blockquote>\n<p>Hmm, that's interesting. Have you tried computing it for all non-padded elements?</p>\n<blockquote>\n  <p>Loss: Categorical crossentropy (2 output neurons activated with softmax).</p>\n</blockquote>\n<p>In paper, they used a sigmoid at the end. Plus, if we use ignore_index in the loss for the padded element, that's should bring in the same effevt, right? (wrt NN.CrossEntropy, torch).</p>\n<p>Also, I use 4 encoder/decoder with 8 Attention heads and hidden dim as 256/512. 2048 is too high for me.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1106230,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-08T16:09:49.697000",
          "content": "<p>I have tried, yes, but I found it to be less accurate.</p>\n<p>I think it should be practically the same, I read some guys pointing out that 2-softmax gives better results than 1-sigmoid but can't confirm.</p>\n<p>That model you describe is too big I think, do you really need it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1107298,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-09T15:15:54.613000",
          "content": "<blockquote>\n  <p>That model you describe is too big I think, do you really need it?</p>\n</blockquote>\n<p>Well, i don't need it because a 2 layer model also scores the same almost for me. Was just trying to mimic the paper, that's it. Great to see you up on LB 🎊🎊</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119389,
          "author_name": "alexxu",
          "author_url": "",
          "post_date": "2020-12-20T02:21:08.490000",
          "content": "<p>Hi claverru, I trained a four-layer SAINT+, with final LB 0.78. But I find it's quite hard to blend this model with LGBM. It always leads to Submission Scoring Error. I suppose it is due to running over time.  Do you meet with any similar problems?  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1100139,
      "author_name": "Gannon Reynolds",
      "author_url": "",
      "post_date": "2020-12-02T21:56:04.503000",
      "content": "<p>MPWARE, without divulging too much information, can you tell me if you're adding engineered features to your SAINT model? I'm particularly interested in engineered continuous features. I'm amazed at how well SAKT does with just the content_id as input. It does better than a tabular model I've been working on with several engineered features. I've just been wondering how many of these engineered features are being captured indirectly by the sequences, and how useful it would be to add them to my SAKT model. Cheers </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1100145,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-02T22:05:03.633000",
          "content": "<p><a href=\"https://www.kaggle.com/gannonreynolds\" target=\"_blank\">@gannonreynolds</a> Zero engineered features in my SAINT implementation. I've tried to add <code>tags1</code> and <code>task_container_id</code>  related to questions but it did not help. </p>\n<blockquote>\n  <p>I'm amazed at how well SAKT does with just the content_id as input</p>\n</blockquote>\n<p>Yes, that's the magic of Transformer. It learns hidden patterns from sequences (it seems that Transformers could overperform some CNN models for computer vision). It's promising. <br>\nCurrently I'm trying TabNet with and without engineered features but I'm not able to beat any of my other models.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1100169,
          "author_name": "Gannon Reynolds",
          "author_url": "",
          "post_date": "2020-12-02T22:23:08.390000",
          "content": "<p>Wow that's really impressive, thanks for the intel. I had become pretty uninspired just running into a wall with the performance of my tabular model, but this thread has re-sparked my enthusiasm. Exciting stuff</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1086735,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-11-22T02:15:46.330000",
      "content": "<p></p>\n<p>Had a first successful submission at .708;</p>\n<p>Edit -&gt; Current best is at .740 ~2.5-3 hours inference.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1141340,
      "author_name": "Rodolphe Lampe",
      "author_url": "",
      "post_date": "2021-01-06T16:35:58.273000",
      "content": "<p>This competition is so stressful and frustrating with these technical constraints. I hope my last submission will work …</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F371132%2F657af4d0a68945a185d57f1198dbe893%2FScreenshot_2021-01-06%20TensorBoard(1).png?generation=1609950875720205&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1141365,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2021-01-06T16:51:36.053000",
          "content": "<p>Mine keeps giving submission errors after 4 hours GPU time and 4 hours of submission. If anyone has a clue do let me know. It is some silly mistake but I am doing so many things at last min that I just am gave up on this error</p>\n<p>calculate lagtime list for the test_df received<br>\nnote that max_timestamp_u_dict has the last timestamp for each user</p>\n<p>prev_test_df = test_df.copy()<br>\nquestion_len=len( test_df[test_df['content_type_id'] == 0])<br>\nlagtime_list = np.zeros(question_len, dtype=np.float32)<br>\ni=0</p>\n<p>for j, (user_id,content_type_id,timestamp,content_id) in enumerate(zip(test_df['user_id'].values,test_df['content_type_id'].values,test_df['timestamp'].values, test_df['content_id'].values)):<br>\nif(content_type_id==0):<br>\nif user_id in max_timestamp_u_dict['max_time_stamp'].keys():<br>\nlagtime_list[i]=int((timestamp - max_timestamp_u_dict['max_time_stamp'][user_id]) /(1000360010))<br>\nmax_timestamp_u_dict['max_time_stamp'][user_id] = timestamp<br>\nelse:<br>\nlagtime_list[i] = int(lagtime_mean)<br>\nmax_timestamp_u_dict['max_time_stamp'].update({user_id:timestamp})<br>\ni=i+1</p>\n<p>Now lagtime is in a list. We just have to assign it when processing the next test_df as below:<br>\nprev_test_df = prev_test_df[prev_test_df.content_type_id == False].reset_index(drop=True)<br>\nprev_test_df[\"lagtime\"]=lagtime_list<br>\nprev_test_df['lagtime'].fillna(int(lagtime_mean), inplace=True)<br>\nprev_test_df.lagtime=prev_test_df.lagtime.astype('int')</p>\n<p>dump the state<br>\nprev_group = prev_test_df[['user_id', 'content_id', 'answered_correctly', 'lagtime']].groupby('user_id').apply(lambda r: (r['content_id'].values,r['answered_correctly'].values,r['lagtime'].values))</p>\n<p>Somehow this gives scoring error and my GPU time is out<br>\nAny hints, clues appreciated</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1141460,
          "author_name": "Gannon Reynolds",
          "author_url": "",
          "post_date": "2021-01-06T18:12:25.363000",
          "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> Have you tried running it against the CV script from <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a> ? That was useful for me because you can test your inference against data without having to submit it. So if it errors out, you know why. Or if it doesn't error out, in my experience, the error probably has something to do with how new users are handled.</p>\n<p>Also, if you wrap your code blocks in <code></code> for comments, it makes it much easier to read, for anyone trying to help.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1141480,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2021-01-06T18:21:45.917000",
          "content": "<p><a href=\"https://www.kaggle.com/gannonreynolds\" target=\"_blank\">@gannonreynolds</a> Thanks for the update. I havent taken a look at that. Will do so tomorrow if time permits.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1141517,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2021-01-06T18:35:33.597000",
          "content": "<p>you may need TPU if you have used 40hours gpu</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1140053,
      "author_name": "qiaqia",
      "author_url": "",
      "post_date": "2021-01-05T19:17:52.897000",
      "content": "<p>without time features 0.778, however it's too late and add time feature reduce score.didnt find the reason yet</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1138465,
      "author_name": "Gannon Reynolds",
      "author_url": "",
      "post_date": "2021-01-04T17:49:22.060000",
      "content": "<p>I'm curious about the usefulness of the exercise categories used in the SAINT paper. I would imagine the embeddings of the content ids would do a better job at finding the relationships between questions better than a human deciding what the tags should be. Am I off base in thinking that? Has anyone compared models run with the provided tags and question details vs without?</p>\n<p>I understand if no one wants to discuss SAINT implementations this close to the end of the comp. I'm just curious and have very limited GPU time left to test it myself.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1138501,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2021-01-04T18:22:30.133000",
          "content": "<p>The category is just the TOEIC part. Sure, the question_id embeddings are more useful, but the TOEIC part (category) also provides signal. Think of it as a type of question clustering. Like \"Reading\", or \"Listening\", or \"Comprehension\", etc.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F2ee06e85eca1af9fdb28c38be8b44eeb%2FScreen%20Shot%202021-01-04%20at%2012.20.58%20PM.png?generation=1609784547132431&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1138521,
          "author_name": "Gannon Reynolds",
          "author_url": "",
          "post_date": "2021-01-04T18:43:07.927000",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> Thanks for the response. I imagine it doesn't hurt to include it, if you have the time and computing resources to do so. I've just been wondering how useful it is. If a question only has one part (category), the relevant information about that category should be learned by the embedding of that question id right? I would imagine the embedding learns the part of the question implicitly, and with more nuance, since it's being represented by a higher dimensional tensor. </p>\n<p>I'm still pretty new to the idea of embeddings, so I'm trying to learn what exactly their capabilities and limitations are.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1138530,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2021-01-04T18:49:14",
          "content": "<p>At the end of the day, DL is just finding statistical correlations within the data to the point of over-fitting and then early stopping before validation performance gets any worse. It can also be proven that given enough hidden units, a single layer MLP can approximate any function =]. But sometimes the 'art' is about finding helpful ways to assist or even coax the net into behaving nicely. Adding things like part embeddings, as a form a clustering, help the net learn better to associate questions within the same part (perhaps the user sucks at the skills necessary for some part and so user-responses belonging to it suffer), even though as you mention, an exercise embedding at the content_id level can/should/does do this as well. Best to try w/ and w/o the part embedding and you'll see it'll help improve validation score before over-fitting occurs =].</p>\n<p>Easy way to think about it—imagine you have only 10 / 100M responses to a particular question. With a 128 or 256 content_id embedding, these few samples might not be sufficient for the net to properly learn to associate those questions with the appropriate part, without our assistance. In our dataset, we have a number of questions that are even just asked a single time!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1138538,
          "author_name": "Gannon Reynolds",
          "author_url": "",
          "post_date": "2021-01-04T18:55:52.577000",
          "content": "<p>That's some really interesting food for thought. Thanks <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1133918,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-12-31T16:35:58.110000",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> does  LR scheduler type make a difference ,</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1133939,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-31T17:09:29.180000",
          "content": "<p>I've only used the Noam LR so can't compare.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1128432,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-27T12:41:03.187000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1123628,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-23T11:36:32.517000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1122162,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-22T08:36:59.323000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1122185,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-22T08:52:41.020000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1122194,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-22T08:56:52.120000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1122689,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-22T16:15:35.640000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 1122711,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-12-22T16:39:32.227000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1122747,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-22T17:17:29.293000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1122753,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-22T17:23:50.847000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1122792,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-22T17:53:10.610000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1122872,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-22T18:50:47.967000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1123175,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-23T01:54:39.550000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1123251,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-23T04:42:22.387000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1123285,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-23T05:20:22.527000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1123778,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-23T13:48:07.677000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1124578,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-24T04:12:57.107000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135374,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-02T07:50:01.863000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1119945,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-20T13:26:58.133000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1097973,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-01T11:28:15.390000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1098062,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-01T12:22:50.817000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1098082,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-01T12:30:27.560000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1098166,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-01T13:30:47.450000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1098238,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-01T14:09:16.583000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1098575,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-01T17:54:25.457000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1098639,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-01T18:45:34.897000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1101879,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-04T11:11:49.110000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1097346,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-01T03:03:09.067000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1096962,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-30T22:48:46.227000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1094094,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-28T10:11:56.640000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1094101,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-28T10:17:04.413000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1093402,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-27T17:23:36.967000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1100422,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-03T04:43:34.783000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1100457,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-03T05:12:36.203000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1083824,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-19T12:29:34.437000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1083870,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-19T13:30:14.420000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1083880,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-19T13:41:51.013000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 1083885,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-11-19T13:51:26.177000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1085595,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-11-21T02:33:39.110000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 1085865,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-11-21T08:49:38.607000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 1086071,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-11-21T10:42:39",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 1086078,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-11-21T10:49:19.803000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1086080,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-11-21T10:52:02.230000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1086085,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-11-21T10:56:59.260000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1086147,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-11-21T12:36:58.803000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1087072,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-11-22T10:08:15.260000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1108903,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-11T06:04:18.300000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1142042,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-07T05:23:00.030000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1087159,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-22T12:19:50.060000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1084568,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-20T06:55:38.887000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1095395,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-29T15:08:07.370000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1124810,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-24T08:08:16.417000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1071126": "Hi all,\n\nIs someone able to reach [SAINT](https://arxiv.org/abs/2002.07033) benchmark on RIIID dataset?\nThey claim AUC around 0.78-0.79\n\nThey use Exercice ID (content_id), Exercice category (could be question part), elapsed time, lag time  + (had explanation flag optionally) and answers.\n\n**Current status** (sum up of all comments below) on **2021/01/03**:\n* Pytorch `nn.Transformer` based implementation (not published yet)\n* Sequences from around 334k users in training, 14k in validation\n* My best local validation score: Fold1= 0.7903, LB=0.795\n* Other best score from @claverru LB=0.794\n* ~~Padding mask issue for users with total interaction lower than sequence length~~\n* Inference time (private notebook) = Between 2h15min and 3h\n\n\n2 public notebooks implemented:\n* Tensorflow: https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\n* Pytorch: https://www.kaggle.com/adityaecdrid/pytorch-demystifying-transformers\n\nOther public notebooks related:\n* SAKT (CV 0.745/LB 0.752): https://www.kaggle.com/mpware/sakt-fork\n* SAKT (LB 0.751): https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing\n* SAKT (LB 0.541) https://www.kaggle.com/leadbest/sakt-self-attentive-knowledge-tracing-submitter\n \n",
    "1107283": "Hey guys, I finally got it. Still not trained with all the data, I will when I have more RAM (it is suposed to be coming). Single model very close to SAINT+:\n- [Epochs: 24] loss: 0.5324 - custom_auc: 0.7811 - val_loss: 0.5316 - val_custom_auc: 0.7768\n- LB: 0.781\n\nAsk me anything about it. Probably the most important keypoint is the sequence sampling strategy. It is what really boosted my model.",
    "1081936": "I've tried several variant of saint, and only have a CV score to provide, no LB yet:\nVariants : \n- Saint++ implemented with the same parameters as in the papers and the same features : CV Roc AUC ~0.757\n- Decoder architecture trained with both a next question prediction loss and a next correct prediction loss : CV Roc AUC ~0.762\n- Encoder architecture (bidirectional) pretrained with a masked question modeling loss then finetuned with a sequence pair classification loss (seq 1 = question to predict, seq 2 = user history) : CV ROC AUC : ~0.77\n- Replace the pure attention layers of transformers by LSTM + Attention : no performance improvement and instability during training.\n\nFor each of these architectures I chose a sequence length of 128, a depth of model of 512 and a ffn representation size of 1024. Optimizer was adam with lr 3e-5 (more caused gradient explosion) and bs either 32 or 64.\nNumber of layers per block was either 4 or 8\n\nConvergence is often decided during the first few epochs, but increase slitghly when continuing training (i've trained the encoder on MQM loss for about 8h on a RTX2070)\n\n\nGiven the very low differences between the different architectures performances, i'd say that the architecture is not really impactfull for the modelisation (providing that you have a good one).\n\nSo for next iteration i might focus more on how to build embeddings for the different sequences (right now i use on embedding layer for each input sequence and i do a summation)\n\nIf you have any ideas of other variants to try or new features to include i'm open for discussion :)\n\n\nPS : For the ones trying to replicate the encoder results, the MQM loss need to be quickstart by first making a directionnal encoder, training for a few epochs with a next question prediction loss and then adding the look ahead masks and swithing to MQM loss (it fails to start when only on MQM loss)\n\n",
    "1112659": "PyTorch implementation of SAINT - https://github.com/arshadshk/SAINT-pytorch\n\nLet me know if any corrections or suggestions are there.\n(Also have a look at SAKT-PyTorch https://github.com/arshadshk/SAKT-pytorch )",
    "1128932": "Gotta tell you guys, my last update brought me from 784 to 794. Only one detail got me there, related with timestamp (and I think it is still improvable). For those who haven't worked around this feature hard enough, I really recommend it.",
    "1071786": "I've read both SAINT and SAINT+ papers but some questions remain.\n\n**Input sequences** should be something like:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fce0c6347798d2f33707b51b623e4a714%2Fsequences.png?generation=1604752912346221&alt=media)\n\n**Start token(s):**\n*The decoder takes O and another sequential input Re = [S,Re1,··· ,Rek−1] of response embeddings with the start token embedding S.*\n\nSomething like this should work for start token:\n```\n...\nresponse_size = 2\nself.response_embedding = nn.Embedding(response_size + 1, embedding_dim) # +1 to include start token\n...\nx_correctness = self.response_embedding(response_sequence)\n...\n# Add start token to correctness\nx_correctness = torch.roll(x_correctness, shifts=(0, 1, 0), dims=(0, 1, 0)) # Shift right the sequence\nx_correctness[:,0,:] = self.response_size # Start token\n```\nx_correctness is embedding of response sequence (0,1,1,1,0 ...), response_size=2, so with token it will be (2,0,1,1,1,0 ...)\n\n\n**Self attention mask:**\n```\n# If a BoolTensor is provided, the positions with the value of True will be ignored while the position with the value of False will be unchanged.\n# tensor([[False,  True,  True,  True],\n#         [False, False,  True,  True],\n#         [False, False, False,  True],\n#         [False, False, False, False]])    \ndef generate_mask(self, size, diagonal=1):        \n    return torch.triu(torch.ones(size, size)==1, diagonal=diagonal)\n```\n\n**Learning rate:**\nWarmup from 2000 to 4000 iterations works.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F9d3dc4354e1cb1ba653cfb8adba9e84a%2Flr.png?generation=1605009427277426&alt=media)\n",
    "1142976": "Sorry, I've been quiet since a few couple of days but I've been very busy with my team to try to improve our models.\nLast 5 submissions have been sent so now competition is almost completed. I would like to thanks @adityaecdrid @yihdarshieh @claverru @abdurrafae @jaideepvalani and all other contributors to this thread. It really helped us to make transformer(s) model(s) train better and work. I won't share any secret right now but I think top teams and you guys have found similar \"things\" that made transformer better and better.\nI wish you the best for private LB and for year 2021!",
    "1144673": "I have shared my final solution here [https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209793](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209793). It has some tricks that weren't discussed in this thread, implemented during the last two weeks of competition. Grew me up from 0.794 to 0.800 in public LB, and I believe it could still have gotten a better score with more time to finetune.",
    "1110427": "Here is update after unrelenting work of 5 days 77.1 SAINT score...Now will try some other stuff ",
    "1109542": "I didn't contribute much in this thread, rather I asked a lot of your approaches.\n\nI just published a training / validation in TensorFlow with TPU (also works with GPU), and it works also on Colab (minimal change required). No competition submission pipeline is provided - it is (a lot, really a lot) personal effort to this competition, and providing it will also be unfair to those who work hard - you definitely know it.\n\nIn particular, I implemented the auto-regressive prediction - but I found it doesn't perform better than `pretend each question (to be predicted during inference) in a question bundle as a single question`. And it doesn't work with TPU (with GPU, it is OK), so I didn't continue with it. \n\nAs documentation - I have to say it is far beyond enough. I still need to focus on the competition, hope you could understand. I will add more during the time.\n\nHowever, I am currently not able to get higher. With @claverru latest sampling strategy, I tried to train a model with his size. It didn't work at the 1st try, but after playing a bit of lr, I get a CV which is as good as a larger model whose LB is about 0.778. I guess the gap between 0.778 and 0.781 could be the different model design or other factors.\n\nI am currently train a lager model with the latest sampling strategy - thanks for TPU and Colab, I have quite resource to do experiments. \n\nHere it is [TPU - Track knowledge states of 1M+ students](https://www.kaggle.com/yihdarshieh/tpu-track-knowledge-states-of-1m-students).\n\nThanks for your shares of your approaches - especially to @claverru, @adityaecdrid and @mpware, among others.\n\nGood luck!",
    "1092406": "Hello guys, I've been trying to follow this discussion but I've been a little busy lately. I have one remaining question that I've not seen here answered. \n\nWhat do you do with users with more interactions than your input length? Do you roll a window? I've tried that and noticed it was (logically) overfitting towards those interactions that appear in multiple rolls, giving me an absurd AUC (~85%). \n\nImagine a window size of 3, having a user with:\n[A, B, C, D, E, F] interactions.\n\nIf I roll a window I obtain [[A, B, C], [B, C, D], [C, D, E], [D, E, F]]. As you can see, for example, the interaction C appears multiple times.\n\nI've also been looking for this in NLP literature but I can't find an answer. Would you share your approach? Or a pointer to a paper/article/post?",
    "1091985": "The best I can get so far CV=0.764 and LB=0.775\nIt required more memory optimizations to make it works within 3 hours.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fbcf2f62ad526a9a1f64d31409d0e37c7%2Finference3.png?generation=1606395072707510&alt=media)",
    "1074914": "I have opened a Notebook with a light version of a Transformer Encoder in TensorFlow. Feel free to check it out to gather some ideas. [NOTEBOOK](https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public)",
    "1096744": "I've forked a SAKT [kernel](https://www.kaggle.com/wangsg/a-self-attentive-model-for-knowledge-tracing) which is super fast (train + test) within 3 hours. Thanks @wangsg for the baseline.\n\nIt shows a simple way to index data to speed up dataloading. It uses almost all data.\nI've added random sequence picking + simple train/valid split + score fixes.\n\nSAKT is not SAINT as it uses an encoder only but it's a good place to start.",
    "1085471": "Another result with masked loss for padding (as I'm not able to have padding mask working in Transformer) and fixed total questions length (thanks @adityaecdrid for the good catch!)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F71b24b4a24b4e6875a95c45c28cf7a29%2Finference2.png?generation=1605911666424581&alt=media) \nI think we can do better.",
    "1078159": "Some additional results depending on train data size (sequence length = 100)\n- 33% of data: Local CV= 0.738\n- 50% of data: Local CV = 0.749\n- 90% of data: Local CV = 0.757\n- 95% of data: Local CV = 0.760 (update)",
    "1134850": "I'm in the mix with a SAINT implementation! 0.778 LB with just content_id so far. Super stoked to improve on that. \n\nI do have a question related to the paper though, if anyone cares to enlighten me. In the SAINT, paper they state that the decoder takes in a sequential input R of response embeddings with start token embedding S. As of now, my implementation doesn't have a unique starting token embedding. It's just padded with zeros. Is this much of an issue? Can anyone share how they're dealing with this start token embedding? I'm not sure I understand that part.",
    "1107907": "Alright, so here's one of my concern which i am not able to help find answer to, Let's assume your batch looks like, (0 represents the padding token) (assume that next question you are going to attempt is the last non_zero value) (6,8,12,19,25,25)\n\n```\ntensor([[ 1,  2,  3,  4,  5,  6],\n        [ 7,  8,  0,  0,  0,  0],\n        [10, 11, 12,  0,  0,  0],\n        [14, 15, 16, 17, 18, 19],\n        [20, 21, 22, 23, 24, 25],\n        20, 21, 22, 23, 24, 25]])\n```\nSo in this case, the src_mask i.e. the self_attention mask in the encoder side, that shouldn't simply be a upper-traingular matrix, right?\n\nLike if we give mask like the below, then that's incorrect, right?\n```\ntensor([[0., -inf, -inf, -inf, -inf, -inf],\n        [0., 0., -inf, -inf, -inf, -inf],\n        [0., 0., 0., -inf, -inf, -inf],\n        [0., 0., 0., 0., -inf, -inf],\n        [0., 0., 0., 0., 0., -inf],\n        [0., 0., 0., 0., 0., 0.]])\n```\n\nAnd it should actually be like this, right? (basically i would call it a length mask 😅)\n\n```\ntensor([[0., 0., 0., 0., 0., -inf],\n        [0., -inf, -inf, -inf, -inf, -inf],\n        [0., 0., -inf, -inf, -inf, -inf],\n        [0., 0., 0., 0., 0., -inf],\n        [0., 0., 0., 0., 0., -inf],\n        [0., 0., 0., 0., 0., -inf]])\n```\n\nIs my understanding apt?",
    "1096342": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4570913%2Fb5c717170357fb281614b150c232e4a0%2Fxy2.png?generation=1606666420284171&alt=media)\nHi, this graph is AUC for given position in a 128 sequence length of a transformer model.\nI think we can interpret the slope as the ability of the model to take into account the past ...\n ",
    "1082028": "I am struggling to make my lgbm inference properly for 4 days now, so forget about DL models inference here for now; (to save your time on more stronger features that can help your lgbm models);  Another reason is, one small mistake and the whole thing will blow, as it's quite unforgivable in nature\n\nSimple idea is to use the feature vectors from them as inputs to your lgbm to get started where you know you won't be breaking anything for sure! (as that should have some visible improvement on your lgbm model)",
    "1138890": "Not a great implementation, but as promised in my old kernel, here's the one which has [SAINT's Inference](https://www.kaggle.com/adityaecdrid/fork-of-saint-inference-ea970c). It's buggy somewhere (~.739 on LB )and I am not able to fix it on my own, so making it public, hoping it will help someone for sure!\n\nNB, there are bugs in it, so use it by analysing the same, the code's for reference only.\n\nThanks for all your shares!",
    "1133430": "@abdurrafae Is there any chance you share a bit of the secret that boosts your LB from 0.798 to 0.806? That's very amazing, you are doing so great!",
    "1142307": "@mpware  congrats for  continuing to be in gold hunt.\nLooks like after hard efforts we might miss by few miles :)\n\nAny thing one can think around to go beyond .806 .  :)\n",
    "1103856": "Any of you still working this way?",
    "1100139": "MPWARE, without divulging too much information, can you tell me if you're adding engineered features to your SAINT model? I'm particularly interested in engineered continuous features. I'm amazed at how well SAKT does with just the content_id as input. It does better than a tabular model I've been working on with several engineered features. I've just been wondering how many of these engineered features are being captured indirectly by the sequences, and how useful it would be to add them to my SAKT model. Cheers ",
    "1086735": "~~So, My inference design finally works ~2.5-3 hours but the performance isn't quite expected. (~.67) 🥺. So i was looking for tips as to how can one debug the same except printing tensors 😅? Ty!\n~~\n\nHad a first successful submission at .708;\n\nEdit -> Current best is at .740 ~2.5-3 hours inference.",
    "1141340": "This competition is so stressful and frustrating with these technical constraints. I hope my last submission will work ...\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F371132%2F657af4d0a68945a185d57f1198dbe893%2FScreenshot_2021-01-06%20TensorBoard(1).png?generation=1609950875720205&alt=media)",
    "1140053": "without time features 0.778, however it's too late and add time feature reduce score.didnt find the reason yet",
    "1138465": "I'm curious about the usefulness of the exercise categories used in the SAINT paper. I would imagine the embeddings of the content ids would do a better job at finding the relationships between questions better than a human deciding what the tags should be. Am I off base in thinking that? Has anyone compared models run with the provided tags and question details vs without?\n\nI understand if no one wants to discuss SAINT implementations this close to the end of the comp. I'm just curious and have very limited GPU time left to test it myself.",
    "1133918": "@abdurrafae does  LR scheduler type make a difference ,",
    "1128432": "My Saint Benchmark\nCV 0.789 LB 0.783\n",
    "1123628": "what a cool work!",
    "1122162": "Dose anyone know how to convert prior_elasped_time to continuous embedding as input of Transformer?",
    "1119945": "@mpware , your team's current score LB 0.797, is it a score obtained from combining LGBM with your transformer's model (LB=0.783)?",
    "1097973": "So i was checking my work, it seems that in SAINT they apply masking for every Multi-Head Attention layers in decoder step? Is my understanding correct?",
    "1097346": "If you don't mind, can anyone reveal how high the score can be achieved with a single nn model?👀👀👀",
    "1096962": "Hi @mpware thanks for this information and the time and effort put into it.",
    "1094094": "I have one extra concern in addition to time: how much memory is required for full data training in Transformer?",
    "1093402": "I am not sure where there's a bug in my pipeline for me but it seems but content_id's sequence can be quite similar at times... (check user_id's 204790744 and 189703047)\n\nAttached is what I see,\n\n```\n{'user_id': 189703047, 'content_id': [7900, 7876, 175, 1278, 2065, 2064, 2063, 3364, 3365, 3363, 2948, 2946, 2947, 2595, 2594, 2593, 4492, 4120, 4696, 6116, 6173, 6370##, 6911, 6910, 6909, 6908, 7219, 7218, 7216, 7217]}\n\n{'user_id': 204790744, 'content_id': [7900, 7876, 175, 1278, 2064, 2063, 2065, 3365, 3364, 3363, 2946, 2947, 2948, 2595, 2594, 2593, 4492, 4120, 4696, 6116, 6173, 6370\", 6879, 6880, 6878, 6877, 7219, 7216, 7218, 7217, 4237, 5675, 6276, 9063, 4740, 5154, 6439, 9163, 5328, 4537]}\n```",
    "1083824": "Can you share @mpware as to what's your accuracy stats? For me, it's pretty bad and AUC is around .67 for SAINT+ on unseen users. [WIP] Ty!",
    "1142042": "",
    "1087159": "",
    "1084568": "",
    "1095395": "thank you for sharing!",
    "1124810": "Thanks for sharing and updating! "
  }
}