{
  "id": 200715,
  "title": "Overfitting with the Transformers",
  "url": "/competitions/riiid-test-answer-prediction/discussion/200715",
  "author_name": "Rodolphe Lampe",
  "post_date": "2020-12-01T14:41:58.525000",
  "votes": 22,
  "comment_count": 39,
  "views": 0,
  "content": "<p>I'd like to discuss about SAINT and the Transformers in general :</p>\n<ul>\n<li><p>I saw a few notebooks about them and it seems like people are reimplementing the various layers of the Transformers but it seems to me that SAINT uses exactly the implementation made in Pytorch using <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html</a> You need to pay attention to the masks using the pytorch function <a href=\"https://pytorch.org/docs/stable/_modules/torch/nn/modules/transformer.html#Transformer.generate_square_subsequent_mask\" target=\"_blank\">https://pytorch.org/docs/stable/_modules/torch/nn/modules/transformer.html#Transformer.generate_square_subsequent_mask</a> for the decoder inputs, encoder inputs and memory inputs. About the various sequence lengths, I had problems using the parameters *_key_padding_mask so, instead I didn't use them but the loss knows what is a padding and what's not so it can ignore the padded part of a sequence.</p></li>\n<li><p>As we are encoding the content_id and the position (following SAINT), it seems that it creates overfitting (at least for us) and it makes sense to me as the content_id + position is plenty of information for a big capacity model to memorize. The Transformers in NLP may not encounter that problem because the NLP datasets are so huge but, for Riiid, even if the dataset is big, it might be a problem. Did you encounter overfitting ? Our train loss keeps improving while the validation is plateauing :</p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F371132%2Fc7c3cfa449863901c65bb678bf80b300%2FScreenshot_2020-12-01%20TensorBoard(1).png?generation=1606833595078574&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1098275,
      "postDate": "2020-12-01T14:41:58.527Z",
      "content": "<p>I'd like to discuss about SAINT and the Transformers in general :</p>\n<ul>\n<li><p>I saw a few notebooks about them and it seems like people are reimplementing the various layers of the Transformers but it seems to me that SAINT uses exactly the implementation made in Pytorch using <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html</a> You need to pay attention to the masks using the pytorch function <a href=\"https://pytorch.org/docs/stable/_modules/torch/nn/modules/transformer.html#Transformer.generate_square_subsequent_mask\" target=\"_blank\">https://pytorch.org/docs/stable/_modules/torch/nn/modules/transformer.html#Transformer.generate_square_subsequent_mask</a> for the decoder inputs, encoder inputs and memory inputs. About the various sequence lengths, I had problems using the parameters *_key_padding_mask so, instead I didn't use them but the loss knows what is a padding and what's not so it can ignore the padded part of a sequence.</p></li>\n<li><p>As we are encoding the content_id and the position (following SAINT), it seems that it creates overfitting (at least for us) and it makes sense to me as the content_id + position is plenty of information for a big capacity model to memorize. The Transformers in NLP may not encounter that problem because the NLP datasets are so huge but, for Riiid, even if the dataset is big, it might be a problem. Did you encounter overfitting ? Our train loss keeps improving while the validation is plateauing :</p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F371132%2Fc7c3cfa449863901c65bb678bf80b300%2FScreenshot_2020-12-01%20TensorBoard(1).png?generation=1606833595078574&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I'd like to discuss about SAINT and the Transformers in general :\n\n- I saw a few notebooks about them and it seems like people are reimplementing the various layers of the Transformers but it seems to me that SAINT uses exactly the implementation made in Pytorch using https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html You need to pay attention to the masks using the pytorch function https://pytorch.org/docs/stable/_modules/torch/nn/modules/transformer.html#Transformer.generate_square_subsequent_mask for the decoder inputs, encoder inputs and memory inputs. About the various sequence lengths, I had problems using the parameters *_key_padding_mask so, instead I didn't use them but the loss knows what is a padding and what's not so it can ignore the padded part of a sequence.\n\n- As we are encoding the content_id and the position (following SAINT), it seems that it creates overfitting (at least for us) and it makes sense to me as the content_id + position is plenty of information for a big capacity model to memorize. The Transformers in NLP may not encounter that problem because the NLP datasets are so huge but, for Riiid, even if the dataset is big, it might be a problem. Did you encounter overfitting ? Our train loss keeps improving while the validation is plateauing :\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F371132%2Fc7c3cfa449863901c65bb678bf80b300%2FScreenshot_2020-12-01%20TensorBoard(1).png?generation=1606833595078574&alt=media)\n\n",
      "votes": 22
    },
    {
      "id": 1098610,
      "postDate": "2020-12-01T18:21:33.177Z",
      "content": "<p>I have also noticed a quick overfitting. Depending on my model's size and amount of data, my validation AUC starts decreasing around epoch 5-15 while my training score keeps improving. </p>\n<p>Dropping absolute position encoding and adding relative position encoding to the MultiHeadAttention layers improves the performance for me. I haven't seen any official implementation but it's not hard to code it following the paper <a href=\"https://arxiv.org/abs/1803.02155\" target=\"_blank\">https://arxiv.org/abs/1803.02155</a>. Let me know if you try this and works for you.</p>\n<p>Adding lag didn't improve my model. </p>\n<p>I'm still not sure about adding task_container_id information, which I tried too.</p>\n<ul>\n<li>As a discrete feature converted to embeddings it should helps the model to learn certain alumns' patterns. Though for the latest ids maybe it is overfitting towards the few users that reached that state.</li>\n<li>As a continuous feature it should help the model to know the alumn's <em>expertise</em> but I dont have any experimental trace yet. </li>\n</ul>",
      "rawMarkdown": "I have also noticed a quick overfitting. Depending on my model's size and amount of data, my validation AUC starts decreasing around epoch 5-15 while my training score keeps improving. \n\nDropping absolute position encoding and adding relative position encoding to the MultiHeadAttention layers improves the performance for me. I haven't seen any official implementation but it's not hard to code it following the paper [https://arxiv.org/abs/1803.02155](https://arxiv.org/abs/1803.02155). Let me know if you try this and works for you.\n\nAdding lag didn't improve my model. \n\nI'm still not sure about adding task_container_id information, which I tried too.\n- As a discrete feature converted to embeddings it should helps the model to learn certain alumns' patterns. Though for the latest ids maybe it is overfitting towards the few users that reached that state.\n- As a continuous feature it should help the model to know the alumn's _expertise_ but I dont have any experimental trace yet. ",
      "votes": 5,
      "replies": [
        {
          "id": 1098629,
          "postDate": "2020-12-01T18:36:57.280Z",
          "content": "<p>I would not expect <code>task_container_id</code> (extra) helpful -- it is kind of positional information. </p>\n<p>While we use window size, and if we assign position from <code>0</code> to <code>window_size - 1</code>, isn't it (kind) of relative position? Probably I need to read the paper you gave.</p>",
          "rawMarkdown": "I would not expect `task_container_id` (extra) helpful -- it is kind of positional information. \n\nWhile we use window size, and if we assign position from `0` to `window_size - 1`, isn't it (kind) of relative position? Probably I need to read the paper you gave."
        },
        {
          "id": 1098648,
          "postDate": "2020-12-01T18:53:55.440Z",
          "content": "<p>I pretty recommend it since one of the autors is Vaswani.</p>",
          "rawMarkdown": "I pretty recommend it since one of the autors is Vaswani.",
          "votes": 1
        },
        {
          "id": 1098654,
          "postDate": "2020-12-01T18:59:01.033Z",
          "content": "<p>Thanks! Would you mind to share how much it increase your score after using relative pos?</p>",
          "rawMarkdown": "Thanks! Would you mind to share how much it increase your score after using relative pos?",
          "votes": 1
        },
        {
          "id": 1098752,
          "postDate": "2020-12-01T20:22:00.860Z",
          "content": "<p>I didn't record it, but something around +0.05 % in AUC.</p>",
          "rawMarkdown": "I didn't record it, but something around +0.05 % in AUC.",
          "votes": 1
        },
        {
          "id": 1098767,
          "postDate": "2020-12-01T20:41:02.873Z",
          "content": "<p>Using <code>task_container_id</code> might have benefit over just usual pos info. due to it can learn which questions form a bundle.</p>",
          "rawMarkdown": "Using `task_container_id` might have benefit over just usual pos info. due to it can learn which questions form a bundle.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1098440,
      "postDate": "2020-12-01T16:41:15.697Z",
      "content": "<p>It depends on how much data and LR policy. It's around 24 epochs for my training.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F5b19aad581563723bfef451ab75b3516%2Ftrain_0.764.png?generation=1606840794474610&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "It depends on how much data and LR policy. It's around 24 epochs for my training.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F5b19aad581563723bfef451ab75b3516%2Ftrain_0.764.png?generation=1606840794474610&alt=media)",
      "votes": 5,
      "replies": [
        {
          "id": 1098458,
          "postDate": "2020-12-01T16:49:35.133Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> What's the maximal value of your lr? I use Noam LR with maximal lr between 1e-3 and 5e-4, but get best CV at 5-7 epoch :(</p>",
          "rawMarkdown": "@mpware What's the maximal value of your lr? I use Noam LR with maximal lr between 1e-3 and 5e-4, but get best CV at 5-7 epoch :("
        },
        {
          "id": 1098466,
          "postDate": "2020-12-01T16:52:24.887Z",
          "content": "<p>MaxLR is around 0.0006</p>",
          "rawMarkdown": "MaxLR is around 0.0006",
          "votes": 2
        }
      ]
    },
    {
      "id": 1099084,
      "postDate": "2020-12-02T04:35:26.337Z",
      "content": "<p>Well there's relatively a new kid in the town, descending from the author of SAKT, <a href=\"https://dl.acm.org/doi/pdf/10.1145/3340531.3411994\" target=\"_blank\">RKT : Relation-Aware Self-Attention for Knowledge Tracing</a>. <strong>Very very well written paper</strong> once more and has the potential to beat SAINT[+]…</p>",
      "rawMarkdown": "Well there's relatively a new kid in the town, descending from the author of SAKT, [RKT : Relation-Aware Self-Attention for Knowledge Tracing](https://dl.acm.org/doi/pdf/10.1145/3340531.3411994). **Very very well written paper** once more and has the potential to beat SAINT[+]...",
      "votes": 4,
      "replies": [
        {
          "id": 1099217,
          "postDate": "2020-12-02T07:26:08.467Z",
          "content": "<p>Thanks for sharing, seems very interesting !</p>",
          "rawMarkdown": "Thanks for sharing, seems very interesting !",
          "votes": 2
        },
        {
          "id": 1099450,
          "postDate": "2020-12-02T11:24:54.750Z",
          "content": "<p>it's indeed much interesting but I am not sure how to build that contingency matrix. Would appreciate thoughts on that. (We can have lot of pairs here, right? Or Maybe we can split them into sub matrices based on \"part_id\")</p>",
          "rawMarkdown": "it's indeed much interesting but I am not sure how to build that contingency matrix. Would appreciate thoughts on that. (We can have lot of pairs here, right? Or Maybe we can split them into sub matrices based on \"part_id\")"
        },
        {
          "id": 1104768,
          "postDate": "2020-12-07T07:55:42.747Z",
          "content": "<p>code seems to have been taken offline..</p>",
          "rawMarkdown": "code seems to have been taken offline.."
        },
        {
          "id": 1112823,
          "postDate": "2020-12-14T23:21:03.127Z",
          "content": "<p>Both the code and dataset are still available… =) however, without the original competition question text, it is not possible for us to compute the exercise relation matrix, which is calculated using text embedding similarity.</p>",
          "rawMarkdown": "Both the code and dataset are still available... =) however, without the original competition question text, it is not possible for us to compute the exercise relation matrix, which is calculated using text embedding similarity.",
          "votes": 1
        },
        {
          "id": 1112939,
          "postDate": "2020-12-15T03:11:38.863Z",
          "content": "<blockquote>\n  <p>code seems to have been taken offline..</p>\n</blockquote>\n<p>Just check git log's and go to the previous commit :)</p>\n<blockquote>\n  <p>Both the code and dataset are still available</p>\n</blockquote>\n<p>Exactly!</p>",
          "rawMarkdown": ">code seems to have been taken offline..\n\nJust check git log's and go to the previous commit :)\n\n>Both the code and dataset are still available\n\nExactly!"
        }
      ]
    },
    {
      "id": 1118262,
      "postDate": "2020-12-18T22:15:44.107Z",
      "content": "<p>To give you news, I found an explanation about the quick overfitting. It's quite interesting and the fix is very useful to get rid of this overfitting and improve the score quite a lot. I'll be happy to share it at the end of the competition.</p>",
      "rawMarkdown": "To give you news, I found an explanation about the quick overfitting. It's quite interesting and the fix is very useful to get rid of this overfitting and improve the score quite a lot. I'll be happy to share it at the end of the competition.",
      "votes": 1,
      "replies": [
        {
          "id": 1118389,
          "postDate": "2020-12-19T03:01:36.393Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        },
        {
          "id": 1118552,
          "postDate": "2020-12-19T07:15:26.010Z",
          "content": "<p>Did you use dropouts? :D</p>",
          "rawMarkdown": "Did you use dropouts? :D"
        },
        {
          "id": 1118949,
          "postDate": "2020-12-19T14:59:44.063Z",
          "content": "<p>Actually it's not hard or tricky but all notebooks on Saint/SAKT are doing something \"wrong\" and correcting it helps a lot. It's not about the model but more about how you use the data. I'll write something about it, maybe in a new post.</p>",
          "rawMarkdown": "Actually it's not hard or tricky but all notebooks on Saint/SAKT are doing something \"wrong\" and correcting it helps a lot. It's not about the model but more about how you use the data. I'll write something about it, maybe in a new post.",
          "votes": 2
        },
        {
          "id": 1119048,
          "postDate": "2020-12-19T17:10:27.943Z",
          "content": "<p>Thanks for the comment! I am going to check it once more and it can hopefully help me reach ~.78 then with SAKT alone hopefully 😅</p>",
          "rawMarkdown": "Thanks for the comment! I am going to check it once more and it can hopefully help me reach ~.78 then with SAKT alone hopefully 😅"
        },
        {
          "id": 1119243,
          "postDate": "2020-12-19T20:53:26.737Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205368\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205368</a></p>",
          "rawMarkdown": "https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205368",
          "votes": 1
        }
      ]
    },
    {
      "id": 1098722,
      "postDate": "2020-12-01T19:46:58.363Z",
      "content": "<p>I've seen that although my loss for validation stops improving earlier, the AUC for validation keeps pace with training metrics for a long time. I think it's probably cause my model params are quite limited d_model as 128 and only 2 num_layers on SAINT architecture. </p>\n<p>Blue - Training, Orange - Validation</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2555700%2F63120b5c599ded149c221204574ee92a%2FAUC.png?generation=1606852048798788&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2555700%2F012b327bf0cec3a2dac3afce32b5c166%2FLos.png?generation=1606851880755866&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I've seen that although my loss for validation stops improving earlier, the AUC for validation keeps pace with training metrics for a long time. I think it's probably cause my model params are quite limited d_model as 128 and only 2 num_layers on SAINT architecture. \n\nBlue - Training, Orange - Validation\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2555700%2F63120b5c599ded149c221204574ee92a%2FAUC.png?generation=1606852048798788&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2555700%2F012b327bf0cec3a2dac3afce32b5c166%2FLos.png?generation=1606851880755866&alt=media)\n",
      "votes": 1
    },
    {
      "id": 1101442,
      "postDate": "2020-12-03T22:46:16.540Z",
      "content": "<p>I tried to reduce my learning rate (Noam schedule) - Before: min/max lr = 5e-5 and  5e-4. After: min/max lr = 1e-5 and 5e-5. Previously, I got the best CV at epoch 6, and now I get a higher CV at epoch 30 (I haven't checked at which epoch I got the highest CV with this smaller lr yet.)</p>\n<p>For CV, I got an increase of 0.003, this is the biggest jump recently. The submission is still running, I will update the results tomorrow.</p>\n<p>However, if you have enough resource and time, you might try yourself. The idea is that to make the model not converge so quickly - Let it see more examples to decide to converge.</p>\n<p>Hope this helps. (But not sure if it worth the time/resource to get this extra 0.003)</p>",
      "rawMarkdown": "I tried to reduce my learning rate (Noam schedule) - Before: min/max lr = 5e-5 and  5e-4. After: min/max lr = 1e-5 and 5e-5. Previously, I got the best CV at epoch 6, and now I get a higher CV at epoch 30 (I haven't checked at which epoch I got the highest CV with this smaller lr yet.)\n\nFor CV, I got an increase of 0.003, this is the biggest jump recently. The submission is still running, I will update the results tomorrow.\n\nHowever, if you have enough resource and time, you might try yourself. The idea is that to make the model not converge so quickly - Let it see more examples to decide to converge.\n\nHope this helps. (But not sure if it worth the time/resource to get this extra 0.003)",
      "votes": 2,
      "replies": [
        {
          "id": 1101608,
          "postDate": "2020-12-04T04:20:02.770Z",
          "content": "<blockquote>\n  <p>For CV, I got an increase of 0.003, this is the biggest jump recently. The submission is still running, I will update the results tomorrow.</p>\n</blockquote>\n<p>🎉🎉</p>",
          "rawMarkdown": ">For CV, I got an increase of 0.003, this is the biggest jump recently. The submission is still running, I will update the results tomorrow.\n\n🎉🎉"
        },
        {
          "id": 1101921,
          "postDate": "2020-12-04T12:01:59.593Z",
          "content": "<p>For LB, I got 0.002 increase. And epoch 30 is the epoch with the highest CV.</p>",
          "rawMarkdown": "For LB, I got 0.002 increase. And epoch 30 is the epoch with the highest CV.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1099215,
      "postDate": "2020-12-02T07:23:51.610Z",
      "content": "<p>It might be due to my own implementation as I tried a few things like concatenating content_id embeddings and positional embeddings instead of summing them (I will go back to the normal behavior). But I tried something : I erased all the content_id_embeddings (setting 0 everywhere) to understand if it might be the cause of overfitting. The blue line is the experiment while the grey one uses the content_id_embeddings. We see much less overfitting. It seems to me that the content_id_embeddings (with the positional embeddings) are way too much used by the transformers to memorize the data.<br>\nI have 4 layers and d_model = 64. I'll try to decrease the number of layers to 2. My lr is 1e-3</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F371132%2F7c0c44d6f3755a20635c53066e9e4319%2FScreenshot_2020-12-02%20TensorBoard.png?generation=1606893817690516&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "It might be due to my own implementation as I tried a few things like concatenating content_id embeddings and positional embeddings instead of summing them (I will go back to the normal behavior). But I tried something : I erased all the content_id_embeddings (setting 0 everywhere) to understand if it might be the cause of overfitting. The blue line is the experiment while the grey one uses the content_id_embeddings. We see much less overfitting. It seems to me that the content_id_embeddings (with the positional embeddings) are way too much used by the transformers to memorize the data.\nI have 4 layers and d_model = 64. I'll try to decrease the number of layers to 2. My lr is 1e-3\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F371132%2F7c0c44d6f3755a20635c53066e9e4319%2FScreenshot_2020-12-02%20TensorBoard.png?generation=1606893817690516&alt=media)",
      "votes": 2,
      "replies": [
        {
          "id": 1118422,
          "postDate": "2020-12-19T04:07:35.163Z",
          "content": "<p>May I ask how big is your model? </p>",
          "rawMarkdown": "May I ask how big is your model? "
        }
      ]
    },
    {
      "id": 1098344,
      "postDate": "2020-12-01T15:33:25.673Z",
      "content": "<p>I get the best validation / LB scores at about 5-7 epochs. Further training doesn't help anymore.</p>",
      "rawMarkdown": "I get the best validation / LB scores at about 5-7 epochs. Further training doesn't help anymore.",
      "votes": 2
    },
    {
      "id": 1112807,
      "postDate": "2020-12-14T22:52:52.923Z",
      "content": "<blockquote>\n  <p>but it seems to me that SAINT uses exactly the implementation made in Pytorch using</p>\n</blockquote>\n<p>Are you sure about this? Looking at the transformerencoder layer and transformerdecoder layer in pytorch, what is described in the original and saint+ papers for the model architecture is <strong>not</strong> the verbatim pytorch implementation. I've found at least three differences. The most obvious of which being the feed forward dimension—which saint doesn't project at all, and that the attention is all you need paper squeezes to d_model//8 (their explanation was to keep the number of parameters consistent with an 8headed mha vs regular singular attending). Even there the pytorch default ffn dim is 2048, which is crazy since a lot of language models have 768 dim, so an unsuspecting user just plugging in the module would actually be expanding the dimensionality.</p>",
      "rawMarkdown": "> but it seems to me that SAINT uses exactly the implementation made in Pytorch using\n\nAre you sure about this? Looking at the transformerencoder layer and transformerdecoder layer in pytorch, what is described in the original and saint+ papers for the model architecture is **not** the verbatim pytorch implementation. I've found at least three differences. The most obvious of which being the feed forward dimension—which saint doesn't project at all, and that the attention is all you need paper squeezes to d_model//8 (their explanation was to keep the number of parameters consistent with an 8headed mha vs regular singular attending). Even there the pytorch default ffn dim is 2048, which is crazy since a lot of language models have 768 dim, so an unsuspecting user just plugging in the module would actually be expanding the dimensionality.",
      "replies": [
        {
          "id": 1112942,
          "postDate": "2020-12-15T03:16:12.010Z",
          "content": "<blockquote>\n  <p>The most obvious of which being the feed forward dimension—which saint doesn't project at all</p>\n</blockquote>\n<p>Not sure i got this :( SAINT does project Q,K,V in MHA. Plus, <code>A concatenation of h attention heads is multiplied by W O to aggregate the outputs of different attention heads</code></p>\n<p>They do split into multiple MHA's, right?</p>\n<p>The rest other points, yep, there's a subtle difference in them but hey, it shouldn't spoil the score so much on LB when you make a sub with a model that has ~.767 on both train/val even if you ignore these diffs…</p>",
          "rawMarkdown": ">The most obvious of which being the feed forward dimension—which saint doesn't project at all\n\nNot sure i got this :( SAINT does project Q,K,V in MHA. Plus, `A concatenation of h attention heads is multiplied by W O to aggregate the outputs of different attention heads`\n\nThey do split into multiple MHA's, right?\n\nThe rest other points, yep, there's a subtle difference in them but hey, it shouldn't spoil the score so much on LB when you make a sub with a model that has ~.767 on both train/val even if you ignore these diffs..."
        },
        {
          "id": 1112946,
          "postDate": "2020-12-15T03:24:47.763Z",
          "content": "<p>I'll recheck, but I was certain in the saint paper's FFN, they kept the feed forward dimension == d_model. Also, in Saint, they start with layernorm, whereas in pytorch, layernorm comes after.</p>",
          "rawMarkdown": "I'll recheck, but I was certain in the saint paper's FFN, they kept the feed forward dimension == d_model. Also, in Saint, they start with layernorm, whereas in pytorch, layernorm comes after."
        },
        {
          "id": 1118256,
          "postDate": "2020-12-18T22:12:50.667Z",
          "content": "<p>The hyperparameter can be adjusted to have ff_dimension == d_model. About the layer norm, yes it's a difference but it doesn't change much the algorithm as the pytorch implementation has a layer norm at the end (and saint moved it at the beginning) and so when you compose layers it doesn't change much (only at the very beginning and the very end), there is also a layernorm over the memory.<br>\nI just tested both way to do it and the results are pretty much the same.<br>\nI hope I didn't miss something else.</p>",
          "rawMarkdown": "The hyperparameter can be adjusted to have ff_dimension == d_model. About the layer norm, yes it's a difference but it doesn't change much the algorithm as the pytorch implementation has a layer norm at the end (and saint moved it at the beginning) and so when you compose layers it doesn't change much (only at the very beginning and the very end), there is also a layernorm over the memory.\nI just tested both way to do it and the results are pretty much the same.\nI hope I didn't miss something else.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1098547,
      "postDate": "2020-12-01T17:38:44.577Z",
      "content": "<blockquote>\n  <p>Our train loss keeps improving while the validation is plateauing :</p>\n</blockquote>\n<p>I don't get it, why is this a problem? Isn't this a normal train-validation loss behaviour?</p>",
      "rawMarkdown": "> Our train loss keeps improving while the validation is plateauing :\n\nI don't get it, why is this a problem? Isn't this a normal train-validation loss behaviour?",
      "replies": [
        {
          "id": 1098619,
          "postDate": "2020-12-01T18:33:11.583Z",
          "content": "<p>The big problem is that we can't reproduce SAINT / SAINT+ scores (at least, not from the public discussion I have seen on this forum). And the quick overfitting might be a factor that prevent us getting SAINT / SAINT+ scores.</p>",
          "rawMarkdown": "The big problem is that we can't reproduce SAINT / SAINT+ scores (at least, not from the public discussion I have seen on this forum). And the quick overfitting might be a factor that prevent us getting SAINT / SAINT+ scores."
        },
        {
          "id": 1098640,
          "postDate": "2020-12-01T18:48:07.900Z",
          "content": "<p>My model also starts overfitting after 5 epochs. Maybe that's because there is repetitiveness inside one epoch.</p>",
          "rawMarkdown": "My model also starts overfitting after 5 epochs. Maybe that's because there is repetitiveness inside one epoch."
        },
        {
          "id": 1098815,
          "postDate": "2020-12-01T21:36:53.270Z",
          "content": "<p><a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a> , do you use transformer-like model like SAINT(+)? If so, really impressive your score, because a lot of us can't even pass 0.78.</p>",
          "rawMarkdown": "@bacicnikola , do you use transformer-like model like SAINT(+)? If so, really impressive your score, because a lot of us can't even pass 0.78."
        },
        {
          "id": 1098877,
          "postDate": "2020-12-01T23:17:26.163Z",
          "content": "<p>A transformer, but I'm not sure how much it resembles SAINT(+) at this point.<br>\nThank you, but who knows, maybe I'm overfitting :)</p>",
          "rawMarkdown": "A transformer, but I'm not sure how much it resembles SAINT(+) at this point.\nThank you, but who knows, maybe I'm overfitting :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 1098541,
      "postDate": "2020-12-01T17:36:55.533Z",
      "content": "<p>For <code>*_key_padding_mask</code>, we can do it in this way, you will be padding your content_id with a fix PAD token, so where ever it's not that, you have to make that as False, wherever it's PAD, it will be True.</p>\n<p>Ref:</p>\n<blockquote>\n  <p>[src/tgt/memory]_key_padding_mask provides specified elements in the key to be ignored by the attention. If a ByteTensor is provided, the non-zero positions will be ignored while the zero positions will be unchanged. If a BoolTensor is provided, the positions with the value of <code>True</code> will be ignored while the position with the value of <code>False</code> will be unchanged.</p>\n</blockquote>",
      "rawMarkdown": "For `*_key_padding_mask`, we can do it in this way, you will be padding your content_id with a fix PAD token, so where ever it's not that, you have to make that as False, wherever it's PAD, it will be True.\n\nRef:\n>  [src/tgt/memory]_key_padding_mask provides specified elements in the key to be ignored by the attention. If a ByteTensor is provided, the non-zero positions will be ignored while the zero positions will be unchanged. If a BoolTensor is provided, the positions with the value of ``True`` will be ignored while the position with the value of ``False`` will be unchanged."
    },
    {
      "id": 1098288,
      "postDate": "2020-12-01T14:49:38.927Z",
      "content": "<p>Yep, I have this behaviour right after 8-9 epochs (d_model as 256) almost always. It's prone to overfit easily. (My Val on 2.5 M rows flatten around .760)</p>",
      "rawMarkdown": "Yep, I have this behaviour right after 8-9 epochs (d_model as 256) almost always. It's prone to overfit easily. (My Val on 2.5 M rows flatten around .760)"
    }
  ],
  "comments": [
    {
      "id": 1098610,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2020-12-01T18:21:33.177000",
      "content": "<p>I have also noticed a quick overfitting. Depending on my model's size and amount of data, my validation AUC starts decreasing around epoch 5-15 while my training score keeps improving. </p>\n<p>Dropping absolute position encoding and adding relative position encoding to the MultiHeadAttention layers improves the performance for me. I haven't seen any official implementation but it's not hard to code it following the paper <a href=\"https://arxiv.org/abs/1803.02155\" target=\"_blank\">https://arxiv.org/abs/1803.02155</a>. Let me know if you try this and works for you.</p>\n<p>Adding lag didn't improve my model. </p>\n<p>I'm still not sure about adding task_container_id information, which I tried too.</p>\n<ul>\n<li>As a discrete feature converted to embeddings it should helps the model to learn certain alumns' patterns. Though for the latest ids maybe it is overfitting towards the few users that reached that state.</li>\n<li>As a continuous feature it should help the model to know the alumn's <em>expertise</em> but I dont have any experimental trace yet. </li>\n</ul>",
      "votes": 5,
      "replies": [
        {
          "id": 1098629,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-01T18:36:57.280000",
          "content": "<p>I would not expect <code>task_container_id</code> (extra) helpful -- it is kind of positional information. </p>\n<p>While we use window size, and if we assign position from <code>0</code> to <code>window_size - 1</code>, isn't it (kind) of relative position? Probably I need to read the paper you gave.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1098648,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-01T18:53:55.440000",
          "content": "<p>I pretty recommend it since one of the autors is Vaswani.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1098654,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-01T18:59:01.033000",
          "content": "<p>Thanks! Would you mind to share how much it increase your score after using relative pos?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1098752,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-01T20:22:00.860000",
          "content": "<p>I didn't record it, but something around +0.05 % in AUC.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1098767,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-01T20:41:02.873000",
          "content": "<p>Using <code>task_container_id</code> might have benefit over just usual pos info. due to it can learn which questions form a bundle.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1098440,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-12-01T16:41:15.697000",
      "content": "<p>It depends on how much data and LR policy. It's around 24 epochs for my training.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F5b19aad581563723bfef451ab75b3516%2Ftrain_0.764.png?generation=1606840794474610&amp;alt=media\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 1098458,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-01T16:49:35.133000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> What's the maximal value of your lr? I use Noam LR with maximal lr between 1e-3 and 5e-4, but get best CV at 5-7 epoch :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1098466,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-12-01T16:52:24.887000",
          "content": "<p>MaxLR is around 0.0006</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1099084,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-12-02T04:35:26.337000",
      "content": "<p>Well there's relatively a new kid in the town, descending from the author of SAKT, <a href=\"https://dl.acm.org/doi/pdf/10.1145/3340531.3411994\" target=\"_blank\">RKT : Relation-Aware Self-Attention for Knowledge Tracing</a>. <strong>Very very well written paper</strong> once more and has the potential to beat SAINT[+]…</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1099217,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2020-12-02T07:26:08.467000",
          "content": "<p>Thanks for sharing, seems very interesting !</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1099450,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-02T11:24:54.750000",
          "content": "<p>it's indeed much interesting but I am not sure how to build that contingency matrix. Would appreciate thoughts on that. (We can have lot of pairs here, right? Or Maybe we can split them into sub matrices based on \"part_id\")</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1104768,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-07T07:55:42.747000",
          "content": "<p>code seems to have been taken offline..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1112823,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-14T23:21:03.127000",
          "content": "<p>Both the code and dataset are still available… =) however, without the original competition question text, it is not possible for us to compute the exercise relation matrix, which is calculated using text embedding similarity.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1112939,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-15T03:11:38.863000",
          "content": "<blockquote>\n  <p>code seems to have been taken offline..</p>\n</blockquote>\n<p>Just check git log's and go to the previous commit :)</p>\n<blockquote>\n  <p>Both the code and dataset are still available</p>\n</blockquote>\n<p>Exactly!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1118262,
      "author_name": "Rodolphe Lampe",
      "author_url": "",
      "post_date": "2020-12-18T22:15:44.107000",
      "content": "<p>To give you news, I found an explanation about the quick overfitting. It's quite interesting and the fix is very useful to get rid of this overfitting and improve the score quite a lot. I'll be happy to share it at the end of the competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1118389,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-19T03:01:36.393000",
          "content": "",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1118552,
          "author_name": "AbdurRafae",
          "author_url": "",
          "post_date": "2020-12-19T07:15:26.010000",
          "content": "<p>Did you use dropouts? :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1118949,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2020-12-19T14:59:44.063000",
          "content": "<p>Actually it's not hard or tricky but all notebooks on Saint/SAKT are doing something \"wrong\" and correcting it helps a lot. It's not about the model but more about how you use the data. I'll write something about it, maybe in a new post.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1119048,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-19T17:10:27.943000",
          "content": "<p>Thanks for the comment! I am going to check it once more and it can hopefully help me reach ~.78 then with SAKT alone hopefully 😅</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1119243,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2020-12-19T20:53:26.737000",
          "content": "<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205368\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/205368</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1098722,
      "author_name": "AbdurRafae",
      "author_url": "",
      "post_date": "2020-12-01T19:46:58.363000",
      "content": "<p>I've seen that although my loss for validation stops improving earlier, the AUC for validation keeps pace with training metrics for a long time. I think it's probably cause my model params are quite limited d_model as 128 and only 2 num_layers on SAINT architecture. </p>\n<p>Blue - Training, Orange - Validation</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2555700%2F63120b5c599ded149c221204574ee92a%2FAUC.png?generation=1606852048798788&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2555700%2F012b327bf0cec3a2dac3afce32b5c166%2FLos.png?generation=1606851880755866&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1101442,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2020-12-03T22:46:16.540000",
      "content": "<p>I tried to reduce my learning rate (Noam schedule) - Before: min/max lr = 5e-5 and  5e-4. After: min/max lr = 1e-5 and 5e-5. Previously, I got the best CV at epoch 6, and now I get a higher CV at epoch 30 (I haven't checked at which epoch I got the highest CV with this smaller lr yet.)</p>\n<p>For CV, I got an increase of 0.003, this is the biggest jump recently. The submission is still running, I will update the results tomorrow.</p>\n<p>However, if you have enough resource and time, you might try yourself. The idea is that to make the model not converge so quickly - Let it see more examples to decide to converge.</p>\n<p>Hope this helps. (But not sure if it worth the time/resource to get this extra 0.003)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1101608,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-04T04:20:02.770000",
          "content": "<blockquote>\n  <p>For CV, I got an increase of 0.003, this is the biggest jump recently. The submission is still running, I will update the results tomorrow.</p>\n</blockquote>\n<p>🎉🎉</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1101921,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-04T12:01:59.593000",
          "content": "<p>For LB, I got 0.002 increase. And epoch 30 is the epoch with the highest CV.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1099215,
      "author_name": "Rodolphe Lampe",
      "author_url": "",
      "post_date": "2020-12-02T07:23:51.610000",
      "content": "<p>It might be due to my own implementation as I tried a few things like concatenating content_id embeddings and positional embeddings instead of summing them (I will go back to the normal behavior). But I tried something : I erased all the content_id_embeddings (setting 0 everywhere) to understand if it might be the cause of overfitting. The blue line is the experiment while the grey one uses the content_id_embeddings. We see much less overfitting. It seems to me that the content_id_embeddings (with the positional embeddings) are way too much used by the transformers to memorize the data.<br>\nI have 4 layers and d_model = 64. I'll try to decrease the number of layers to 2. My lr is 1e-3</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F371132%2F7c0c44d6f3755a20635c53066e9e4319%2FScreenshot_2020-12-02%20TensorBoard.png?generation=1606893817690516&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1118422,
          "author_name": "Shuhao Cao",
          "author_url": "",
          "post_date": "2020-12-19T04:07:35.163000",
          "content": "<p>May I ask how big is your model? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1098344,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2020-12-01T15:33:25.673000",
      "content": "<p>I get the best validation / LB scores at about 5-7 epochs. Further training doesn't help anymore.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1112807,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-12-14T22:52:52.923000",
      "content": "<blockquote>\n  <p>but it seems to me that SAINT uses exactly the implementation made in Pytorch using</p>\n</blockquote>\n<p>Are you sure about this? Looking at the transformerencoder layer and transformerdecoder layer in pytorch, what is described in the original and saint+ papers for the model architecture is <strong>not</strong> the verbatim pytorch implementation. I've found at least three differences. The most obvious of which being the feed forward dimension—which saint doesn't project at all, and that the attention is all you need paper squeezes to d_model//8 (their explanation was to keep the number of parameters consistent with an 8headed mha vs regular singular attending). Even there the pytorch default ffn dim is 2048, which is crazy since a lot of language models have 768 dim, so an unsuspecting user just plugging in the module would actually be expanding the dimensionality.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1112942,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-15T03:16:12.010000",
          "content": "<blockquote>\n  <p>The most obvious of which being the feed forward dimension—which saint doesn't project at all</p>\n</blockquote>\n<p>Not sure i got this :( SAINT does project Q,K,V in MHA. Plus, <code>A concatenation of h attention heads is multiplied by W O to aggregate the outputs of different attention heads</code></p>\n<p>They do split into multiple MHA's, right?</p>\n<p>The rest other points, yep, there's a subtle difference in them but hey, it shouldn't spoil the score so much on LB when you make a sub with a model that has ~.767 on both train/val even if you ignore these diffs…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1112946,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-15T03:24:47.763000",
          "content": "<p>I'll recheck, but I was certain in the saint paper's FFN, they kept the feed forward dimension == d_model. Also, in Saint, they start with layernorm, whereas in pytorch, layernorm comes after.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1118256,
          "author_name": "Rodolphe Lampe",
          "author_url": "",
          "post_date": "2020-12-18T22:12:50.667000",
          "content": "<p>The hyperparameter can be adjusted to have ff_dimension == d_model. About the layer norm, yes it's a difference but it doesn't change much the algorithm as the pytorch implementation has a layer norm at the end (and saint moved it at the beginning) and so when you compose layers it doesn't change much (only at the very beginning and the very end), there is also a layernorm over the memory.<br>\nI just tested both way to do it and the results are pretty much the same.<br>\nI hope I didn't miss something else.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1098547,
      "author_name": "Nikola Bacic",
      "author_url": "",
      "post_date": "2020-12-01T17:38:44.577000",
      "content": "<blockquote>\n  <p>Our train loss keeps improving while the validation is plateauing :</p>\n</blockquote>\n<p>I don't get it, why is this a problem? Isn't this a normal train-validation loss behaviour?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1098619,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-01T18:33:11.583000",
          "content": "<p>The big problem is that we can't reproduce SAINT / SAINT+ scores (at least, not from the public discussion I have seen on this forum). And the quick overfitting might be a factor that prevent us getting SAINT / SAINT+ scores.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1098640,
          "author_name": "Nikola Bacic",
          "author_url": "",
          "post_date": "2020-12-01T18:48:07.900000",
          "content": "<p>My model also starts overfitting after 5 epochs. Maybe that's because there is repetitiveness inside one epoch.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1098815,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-12-01T21:36:53.270000",
          "content": "<p><a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a> , do you use transformer-like model like SAINT(+)? If so, really impressive your score, because a lot of us can't even pass 0.78.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1098877,
          "author_name": "Nikola Bacic",
          "author_url": "",
          "post_date": "2020-12-01T23:17:26.163000",
          "content": "<p>A transformer, but I'm not sure how much it resembles SAINT(+) at this point.<br>\nThank you, but who knows, maybe I'm overfitting :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1098541,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-12-01T17:36:55.533000",
      "content": "<p>For <code>*_key_padding_mask</code>, we can do it in this way, you will be padding your content_id with a fix PAD token, so where ever it's not that, you have to make that as False, wherever it's PAD, it will be True.</p>\n<p>Ref:</p>\n<blockquote>\n  <p>[src/tgt/memory]_key_padding_mask provides specified elements in the key to be ignored by the attention. If a ByteTensor is provided, the non-zero positions will be ignored while the zero positions will be unchanged. If a BoolTensor is provided, the positions with the value of <code>True</code> will be ignored while the position with the value of <code>False</code> will be unchanged.</p>\n</blockquote>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1098288,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-12-01T14:49:38.927000",
      "content": "<p>Yep, I have this behaviour right after 8-9 epochs (d_model as 256) almost always. It's prone to overfit easily. (My Val on 2.5 M rows flatten around .760)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1098275": "I'd like to discuss about SAINT and the Transformers in general :\n\n- I saw a few notebooks about them and it seems like people are reimplementing the various layers of the Transformers but it seems to me that SAINT uses exactly the implementation made in Pytorch using https://pytorch.org/docs/stable/generated/torch.nn.Transformer.html You need to pay attention to the masks using the pytorch function https://pytorch.org/docs/stable/_modules/torch/nn/modules/transformer.html#Transformer.generate_square_subsequent_mask for the decoder inputs, encoder inputs and memory inputs. About the various sequence lengths, I had problems using the parameters *_key_padding_mask so, instead I didn't use them but the loss knows what is a padding and what's not so it can ignore the padded part of a sequence.\n\n- As we are encoding the content_id and the position (following SAINT), it seems that it creates overfitting (at least for us) and it makes sense to me as the content_id + position is plenty of information for a big capacity model to memorize. The Transformers in NLP may not encounter that problem because the NLP datasets are so huge but, for Riiid, even if the dataset is big, it might be a problem. Did you encounter overfitting ? Our train loss keeps improving while the validation is plateauing :\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F371132%2Fc7c3cfa449863901c65bb678bf80b300%2FScreenshot_2020-12-01%20TensorBoard(1).png?generation=1606833595078574&alt=media)\n\n",
    "1098610": "I have also noticed a quick overfitting. Depending on my model's size and amount of data, my validation AUC starts decreasing around epoch 5-15 while my training score keeps improving. \n\nDropping absolute position encoding and adding relative position encoding to the MultiHeadAttention layers improves the performance for me. I haven't seen any official implementation but it's not hard to code it following the paper [https://arxiv.org/abs/1803.02155](https://arxiv.org/abs/1803.02155). Let me know if you try this and works for you.\n\nAdding lag didn't improve my model. \n\nI'm still not sure about adding task_container_id information, which I tried too.\n- As a discrete feature converted to embeddings it should helps the model to learn certain alumns' patterns. Though for the latest ids maybe it is overfitting towards the few users that reached that state.\n- As a continuous feature it should help the model to know the alumn's _expertise_ but I dont have any experimental trace yet. ",
    "1098440": "It depends on how much data and LR policy. It's around 24 epochs for my training.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2F5b19aad581563723bfef451ab75b3516%2Ftrain_0.764.png?generation=1606840794474610&alt=media)",
    "1099084": "Well there's relatively a new kid in the town, descending from the author of SAKT, [RKT : Relation-Aware Self-Attention for Knowledge Tracing](https://dl.acm.org/doi/pdf/10.1145/3340531.3411994). **Very very well written paper** once more and has the potential to beat SAINT[+]...",
    "1118262": "To give you news, I found an explanation about the quick overfitting. It's quite interesting and the fix is very useful to get rid of this overfitting and improve the score quite a lot. I'll be happy to share it at the end of the competition.",
    "1098722": "I've seen that although my loss for validation stops improving earlier, the AUC for validation keeps pace with training metrics for a long time. I think it's probably cause my model params are quite limited d_model as 128 and only 2 num_layers on SAINT architecture. \n\nBlue - Training, Orange - Validation\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2555700%2F63120b5c599ded149c221204574ee92a%2FAUC.png?generation=1606852048798788&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2555700%2F012b327bf0cec3a2dac3afce32b5c166%2FLos.png?generation=1606851880755866&alt=media)\n",
    "1101442": "I tried to reduce my learning rate (Noam schedule) - Before: min/max lr = 5e-5 and  5e-4. After: min/max lr = 1e-5 and 5e-5. Previously, I got the best CV at epoch 6, and now I get a higher CV at epoch 30 (I haven't checked at which epoch I got the highest CV with this smaller lr yet.)\n\nFor CV, I got an increase of 0.003, this is the biggest jump recently. The submission is still running, I will update the results tomorrow.\n\nHowever, if you have enough resource and time, you might try yourself. The idea is that to make the model not converge so quickly - Let it see more examples to decide to converge.\n\nHope this helps. (But not sure if it worth the time/resource to get this extra 0.003)",
    "1099215": "It might be due to my own implementation as I tried a few things like concatenating content_id embeddings and positional embeddings instead of summing them (I will go back to the normal behavior). But I tried something : I erased all the content_id_embeddings (setting 0 everywhere) to understand if it might be the cause of overfitting. The blue line is the experiment while the grey one uses the content_id_embeddings. We see much less overfitting. It seems to me that the content_id_embeddings (with the positional embeddings) are way too much used by the transformers to memorize the data.\nI have 4 layers and d_model = 64. I'll try to decrease the number of layers to 2. My lr is 1e-3\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F371132%2F7c0c44d6f3755a20635c53066e9e4319%2FScreenshot_2020-12-02%20TensorBoard.png?generation=1606893817690516&alt=media)",
    "1098344": "I get the best validation / LB scores at about 5-7 epochs. Further training doesn't help anymore.",
    "1112807": "> but it seems to me that SAINT uses exactly the implementation made in Pytorch using\n\nAre you sure about this? Looking at the transformerencoder layer and transformerdecoder layer in pytorch, what is described in the original and saint+ papers for the model architecture is **not** the verbatim pytorch implementation. I've found at least three differences. The most obvious of which being the feed forward dimension—which saint doesn't project at all, and that the attention is all you need paper squeezes to d_model//8 (their explanation was to keep the number of parameters consistent with an 8headed mha vs regular singular attending). Even there the pytorch default ffn dim is 2048, which is crazy since a lot of language models have 768 dim, so an unsuspecting user just plugging in the module would actually be expanding the dimensionality.",
    "1098547": "> Our train loss keeps improving while the validation is plateauing :\n\nI don't get it, why is this a problem? Isn't this a normal train-validation loss behaviour?",
    "1098541": "For `*_key_padding_mask`, we can do it in this way, you will be padding your content_id with a fix PAD token, so where ever it's not that, you have to make that as False, wherever it's PAD, it will be True.\n\nRef:\n>  [src/tgt/memory]_key_padding_mask provides specified elements in the key to be ignored by the attention. If a ByteTensor is provided, the non-zero positions will be ignored while the zero positions will be unchanged. If a BoolTensor is provided, the positions with the value of ``True`` will be ignored while the position with the value of ``False`` will be unchanged.",
    "1098288": "Yep, I have this behaviour right after 8-9 epochs (d_model as 256) almost always. It's prone to overfit easily. (My Val on 2.5 M rows flatten around .760)"
  }
}