{
  "id": 195949,
  "title": "About the position encoding in SAINT papers",
  "url": "/competitions/riiid-test-answer-prediction/discussion/195949",
  "author_name": "",
  "post_date": "2020-11-08T12:52:33.636513700Z",
  "votes": 16,
  "comment_count": 40,
  "views": 0,
  "content": "<p>After getting stuck with my current approach, I've been paying attention to both papers shared by our host (<a href=\"https://arxiv.org/abs/2010.12042\" target=\"_blank\">https://arxiv.org/abs/2010.12042</a> and <a href=\"https://arxiv.org/abs/2002.07033\" target=\"_blank\">https://arxiv.org/abs/2002.07033</a>).</p>\n<p>I'm trying to develop a light adaptation to later be able to work around it. </p>\n<p>My main concern at the moment is about the position encoding, which seems confusing to me, and maybe anyone could clarify.</p>\n<p>How do we set up the maximum position? If I understand correctly, in the paper Attention is All You Need, we sample a sinusoidal-like vector with the maximum position, and then take from it only the window for the current chunk. </p>\n<p>For instance, we can set a maximum position as 1024, but then we only take indices [0:256] from the vector for the first 256 tokens. Did I understand properly?</p>\n<p>If yes, then how do we set the maximum position for our users' histories? In theory we could have a larger history in the test set than in the training set.</p>\n<p>Thank you in advance fellas.</p>",
  "messages": [
    {
      "id": "1072564",
      "postDate": "11/08/2020 12:52:33",
      "content": "<p>After getting stuck with my current approach, I've been paying attention to both papers shared by our host (<a href=\"https://arxiv.org/abs/2010.12042\" target=\"_blank\">https://arxiv.org/abs/2010.12042</a> and <a href=\"https://arxiv.org/abs/2002.07033\" target=\"_blank\">https://arxiv.org/abs/2002.07033</a>).</p>\n<p>I'm trying to develop a light adaptation to later be able to work around it. </p>\n<p>My main concern at the moment is about the position encoding, which seems confusing to me, and maybe anyone could clarify.</p>\n<p>How do we set up the maximum position? If I understand correctly, in the paper Attention is All You Need, we sample a sinusoidal-like vector with the maximum position, and then take from it only the window for the current chunk. </p>\n<p>For instance, we can set a maximum position as 1024, but then we only take indices [0:256] from the vector for the first 256 tokens. Did I understand properly?</p>\n<p>If yes, then how do we set the maximum position for our users' histories? In theory we could have a larger history in the test set than in the training set.</p>\n<p>Thank you in advance fellas.</p>",
      "rawMarkdown": "After getting stuck with my current approach, I've been paying attention to both papers shared by our host ([https://arxiv.org/abs/2010.12042](https://arxiv.org/abs/2010.12042) and [https://arxiv.org/abs/2002.07033](https://arxiv.org/abs/2002.07033)).\n\nI'm trying to develop a light adaptation to later be able to work around it. \n\nMy main concern at the moment is about the position encoding, which seems confusing to me, and maybe anyone could clarify.\n\nHow do we set up the maximum position? If I understand correctly, in the paper Attention is All You Need, we sample a sinusoidal-like vector with the maximum position, and then take from it only the window for the current chunk. \n\nFor instance, we can set a maximum position as 1024, but then we only take indices [0:256] from the vector for the first 256 tokens. Did I understand properly?\n\nIf yes, then how do we set the maximum position for our users' histories? In theory we could have a larger history in the test set than in the training set.\n\nThank you in advance fellas.",
      "votes": null
    },
    {
      "id": "1072602",
      "postDate": "11/08/2020 13:50:04",
      "content": "<p>Usually you cap the sequence length up to <code>max_len</code> (sometimes referred as window size). <br>\nYou do this for the test set as well.<br>\nThen the maximum possible  is f.e. <code>max_len+1</code> (it doesn't necessarily have to be the value 1024)</p>",
      "rawMarkdown": "Usually you cap the sequence length up to `max_len` (sometimes referred as window size). \nYou do this for the test set as well.\nThen the maximum possible  is f.e. `max_len+1` (it doesn't necessarily have to be the value 1024)",
      "votes": null
    },
    {
      "id": "1072618",
      "postDate": "11/08/2020 14:00:54",
      "content": "<p>Yeah, you cap the sequence. But you cannot cap the position encoding in this competition. You would be saying that the first interaction is in the same place than, say, the 512th at some point.</p>",
      "rawMarkdown": "Yeah, you cap the sequence. But you cannot cap the position encoding in this competition. You would be saying that the first interaction is in the same place than, say, the 512th at some point.",
      "votes": null
    },
    {
      "id": "1072624",
      "postDate": "11/08/2020 14:07:49",
      "content": "<p>Look at the diagram I quickly sketched. I have a window size of 100 but to be able to know the exact ordinal interaction I need to have this information in the position encoding.</p>",
      "rawMarkdown": "Look at the diagram I quickly sketched. I have a window size of 100 but to be able to know the exact ordinal interaction I need to have this information in the position encoding.",
      "votes": null
    },
    {
      "id": "1072628",
      "postDate": "11/08/2020 14:17:21",
      "content": "<p>Yes, we do have that problem here. It's basically like <code>position_no % max_seq</code>. And the embedding for the same layer would go from 0-99 i feel.</p>",
      "rawMarkdown": "Yes, we do have that problem here. It's basically like `position_no % max_seq`. And the embedding for the same layer would go from 0-99 i feel.",
      "votes": null
    },
    {
      "id": "1072652",
      "postDate": "11/08/2020 14:53:10",
      "content": "<p>I agree, there will be some loss of info with window resizing. Currently I only use <code>position_no % max_seq</code> to at least incorporate the relative positions within the window.</p>",
      "rawMarkdown": "I agree, there will be some loss of info with window resizing. Currently I only use ` position_no % max_seq` to at least incorporate the relative positions within the window.",
      "votes": null
    },
    {
      "id": "1072660",
      "postDate": "11/08/2020 15:04:07",
      "content": "<p>Currently, I have a maximal seq len as 512, and the window size is 128. For each truncated history sequence of length window size (i.e 128), my position is always from 0 to 127. </p>\n<p>This is not ideal, because it losses the actual interaction position (a user's interaction history might be very long).</p>",
      "rawMarkdown": "Currently, I have a maximal seq len as 512, and the window size is 128. For each truncated history sequence of length window size (i.e 128), my position is always from 0 to 127. \n\nThis is not ideal, because it losses the actual interaction position (a user's interaction history might be very long).",
      "votes": null
    },
    {
      "id": "1072672",
      "postDate": "11/08/2020 15:19:03",
      "content": "<p>What I have done (don't know if it is the best solution neither if it is even correct) is:</p>\n<ul>\n<li>Generating position encodings with my d_model and a maximum position of 40K (the larger user history I've seen is around 16K).</li>\n<li>Create a column with <code>np.arange(len(user))</code> and gather those indices from the position encodings above. In tensorflow it is <code>p = tf.gather(position_encoding, indices, axis=0)</code>, where <ul>\n<li><code>shape(position_encoding) = (40K, d_model)</code></li>\n<li><code>shape(indices) = (batch_size, window_size)</code></li>\n<li><code>shape(p) = (batch_size, window_size, d_model)</code></li></ul></li>\n</ul>",
      "rawMarkdown": "What I have done (don't know if it is the best solution neither if it is even correct) is:\n- Generating position encodings with my d_model and a maximum position of 40K (the larger user history I've seen is around 16K).\n- Create a column with `np.arange(len(user))` and gather those indices from the position encodings above. In tensorflow it is `p = tf.gather(position_encoding, indices, axis=0)`, where \n  - `shape(position_encoding) = (40K, d_model)`\n  - `shape(indices) = (batch_size, window_size)`\n  - `shape(p) = (batch_size, window_size, d_model)`",
      "votes": null
    },
    {
      "id": "1072685",
      "postDate": "11/08/2020 15:33:37",
      "content": "<p>It's fine, the only cost is about that you are using a larger embedding matrix for positions - and it <em>might</em> affect the  memory usage - especially if you use GPU and you want to do ensembling. It doesn't affect the inference timing because tf.gather should be quite fast and after it, the actually computation is based on <code>p</code> only, which is relatively small.</p>",
      "rawMarkdown": "It's fine, the only cost is about that you are using a larger embedding matrix for positions - and it *might* affect the  memory usage - especially if you use GPU and you want to do ensembling. It doesn't affect the inference timing because tf.gather should be quite fast and after it, the actually computation is based on `p` only, which is relatively small.",
      "votes": null
    },
    {
      "id": "1072697",
      "postDate": "11/08/2020 15:42:45",
      "content": "<p>Well, for a model dimension of 64, I'd have a (40 000, 64) tensor, which is around 82M bits in case theyre float 32, giving as a result ~10.3 MBytes. If my numbers are correct, I think we can deal with it. Though maybe 40K is too much, given that the most prolific alumn has 16k interactions, maybe with a length of 20K we are safe.</p>",
      "rawMarkdown": "Well, for a model dimension of 64, I'd have a (40 000, 64) tensor, which is around 82M bits in case theyre float 32, giving as a result ~10.3 MBytes. If my numbers are correct, I think we can deal with it. Though maybe 40K is too much, given that the most prolific alumn has 16k interactions, maybe with a length of 20K we are safe.",
      "votes": null
    },
    {
      "id": "1072701",
      "postDate": "11/08/2020 15:46:34",
      "content": "<p>Yep, if you stick to dim 64, it won't be a problem. Is dim 64 the one gives your current score? If so, I am surprised :)</p>",
      "rawMarkdown": "Yep, if you stick to dim 64, it won't be a problem. Is dim 64 the one gives your current score? If so, I am surprised :)",
      "votes": null
    },
    {
      "id": "1072710",
      "postDate": "11/08/2020 15:54:35",
      "content": "<p>My current score has a model dimension of 20 I think. I've changed it so many times.</p>\n<p>EDIT: Solution -&gt; Score</p>",
      "rawMarkdown": "My current score has a model dimension of 20 I think. I've changed it so many times.\n\nEDIT: Solution -> Score",
      "votes": null
    },
    {
      "id": "1072719",
      "postDate": "11/08/2020 16:06:51",
      "content": "<p>The problem is that I have to squeeze my model to fit in the GPU quota. My humble 1060 can't deal with a Transformer based model until there's any open source lineal attention or I get any 3080 xD. And even squeezing the model I smash the quota just during the weekend. </p>",
      "rawMarkdown": "The problem is that I have to squeeze my model to fit in the GPU quota. My humble 1060 can't deal with a Transformer based model until there's any open source lineal attention or I get any 3080 xD. And even squeezing the model I smash the quota just during the weekend.",
      "votes": null
    },
    {
      "id": "1072731",
      "postDate": "11/08/2020 16:19:24",
      "content": "<p>try google colab (free or pro version)</p>",
      "rawMarkdown": "try google colab (free or pro version)",
      "votes": null
    },
    {
      "id": "1072737",
      "postDate": "11/08/2020 16:22:54",
      "content": "<p>Hum, yeah but how do you import your data? The comfortable way I know is by connecting your Google Drive and it's terribly slow.</p>",
      "rawMarkdown": "Hum, yeah but how do you import your data? The comfortable way I know is by connecting your Google Drive and it's terribly slow.",
      "votes": null
    },
    {
      "id": "1072752",
      "postDate": "11/08/2020 16:37:47",
      "content": "<p>you can put the date to google bucket, and download it from your colab notebook - i think. I haven't done this for this competition yet, but I have done it before for another competition.</p>",
      "rawMarkdown": "you can put the date to google bucket, and download it from your colab notebook - i think. I haven't done this for this competition yet, but I have done it before for another competition.",
      "votes": null
    },
    {
      "id": "1072756",
      "postDate": "11/08/2020 16:45:55",
      "content": "<p>You can try this, there's an extension named as curl-wget, install it on Chrome. And then go to kaggle's data download section for the comp and simply hit the download all and then cancel it. Grab the curl link generated from that extension and copy/paste that to colab. Or, Use kaggle API Or, push it to S3.</p>\n<p>NB I/O should be a bottle neck nonetheless. (imo, not 100% sure)</p>",
      "rawMarkdown": "You can try this, there's an extension named as curl-wget, install it on Chrome. And then go to kaggle's data download section for the comp and simply hit the download all and then cancel it. Grab the curl link generated from that extension and copy/paste that to colab. Or, Use kaggle API Or, push it to S3.\n\nNB I/O should be a bottle neck nonetheless. (imo, not 100% sure)",
      "votes": null
    },
    {
      "id": "1072983",
      "postDate": "11/09/2020 01:15:26",
      "content": "<p>Bro, IMHO, positional encodings help the attention mechanism figure out the <em>relative</em> position of inputs. Even if you encode the 1000th interaction of the user as pos 1000, since the window size is 100 and the network only has access to the last 100 interactions, its better(for consistency) to pos encode only the window size positions(i.e. 100). Because even if the network needs to attend to inputs before the window, it cant.</p>",
      "rawMarkdown": "Bro, IMHO, positional encodings help the attention mechanism figure out the *relative* position of inputs. Even if you encode the 1000th interaction of the user as pos 1000, since the window size is 100 and the network only has access to the last 100 interactions, its better(for consistency) to pos encode only the window size positions(i.e. 100). Because even if the network needs to attend to inputs before the window, it cant.",
      "votes": null
    },
    {
      "id": "1073110",
      "postDate": "11/09/2020 06:20:08",
      "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>   Can you elaborate a bit on the difference of maximal sequence len and window size?</p>",
      "rawMarkdown": "yihdarshieh   Can you elaborate a bit on the difference of maximal sequence len and window size?",
      "votes": null
    },
    {
      "id": "1073226",
      "postDate": "11/09/2020 10:07:05",
      "content": "<p>Hum, those are the things that I tried to discuss here in this post. So, are you sure about that? Then there's literally no difference between the first and the 1000th interaction, and intuitively the alumn should have improved. </p>",
      "rawMarkdown": "Hum, those are the things that I tried to discuss here in this post. So, are you sure about that? Then there's literally no difference between the first and the 1000th interaction, and intuitively the alumn should have improved.",
      "votes": null
    },
    {
      "id": "1073231",
      "postDate": "11/09/2020 10:17:12",
      "content": "<p>In the paper <a href=\"https://arxiv.org/pdf/2002.07033.pdf\" target=\"_blank\">Towards an Appropriate Query, Key, and ValueComputation for Knowledge Tracing</a> they say:</p>\n<p><em>Position: The position (1st, 2nd, …) of an exercise or a\nresponse in the input sequence is represented as a position\nembedding vector. The position embeddings are shared\nacross the exercise sequence and the response sequence.</em></p>\n<p>The truth is that it isn't clear to me yet 😑.</p>",
      "rawMarkdown": "In the paper [Towards an Appropriate Query, Key, and ValueComputation for Knowledge Tracing](https://arxiv.org/pdf/2002.07033.pdf) they say:\n\n*Position: The position (1st, 2nd, ...) of an exercise or a\nresponse in the input sequence is represented as a position\nembedding vector. The position embeddings are shared\nacross the exercise sequence and the response sequence.*\n\nThe truth is that it isn't clear to me yet 😑.",
      "votes": null
    },
    {
      "id": "1073256",
      "postDate": "11/09/2020 10:59:57",
      "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> <br>\nIn the SAINT paper, they dont use the sinusoidal embeddings. They make pos embeddings a learnable parameter as well. Although I am not sure if it makes the model any better. But that is the reason for this line \"The position embeddings are shared across the exercise sequence and the response sequence.\". If they were the sinusoidal embeddings, they wouldnt have any parameters to share (they are constant). Here, think of position embedding as an encoding from a variable X of cardinality = max_seq_len , mapping it to X_hat which is a vector of dimensionality d_model. Does this help?</p>",
      "rawMarkdown": "claverru \nIn the SAINT paper, they dont use the sinusoidal embeddings. They make pos embeddings a learnable parameter as well. Although I am not sure if it makes the model any better. But that is the reason for this line \"The position embeddings are shared across the exercise sequence and the response sequence.\". If they were the sinusoidal embeddings, they wouldnt have any parameters to share (they are constant). Here, think of position embedding as an encoding from a variable X of cardinality = max_seq_len , mapping it to X_hat which is a vector of dimensionality d_model. Does this help?",
      "votes": null
    },
    {
      "id": "1073264",
      "postDate": "11/09/2020 11:06:46",
      "content": "<p>To add to this, I think a benefit of using learned position embeddings is that you can actually encode, say position 5000, if you actually want to encode the session count of the user instead of just the local position of the interaction. That might have some benefit to it.</p>",
      "rawMarkdown": "To add to this, I think a benefit of using learned position embeddings is that you can actually encode, say position 5000, if you actually want to encode the session count of the user instead of just the local position of the interaction. That might have some benefit to it.",
      "votes": null
    },
    {
      "id": "1073267",
      "postDate": "11/09/2020 11:09:42",
      "content": "<p>Yeah, a lot. I'm re-reading it and you're right. It seems like a learnable parameter. </p>",
      "rawMarkdown": "Yeah, a lot. I'm re-reading it and you're right. It seems like a learnable parameter.",
      "votes": null
    },
    {
      "id": "1073342",
      "postDate": "11/09/2020 12:53:55",
      "content": "<p>Nice to know that! 👍</p>",
      "rawMarkdown": "Nice to know that! 👍",
      "votes": null
    },
    {
      "id": "1074634",
      "postDate": "11/10/2020 22:10:30",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> They could be the same. I just fixed a maximal seq len 512 for all experimentation. And try to used different window size (64, 128, 256 etc ) &lt;= 512.</p>",
      "rawMarkdown": "abdurrafae They could be the same. I just fixed a maximal seq len 512 for all experimentation. And try to used different window size (64, 128, 256 etc ) <= 512.",
      "votes": null
    },
    {
      "id": "1074709",
      "postDate": "11/11/2020 00:52:34",
      "content": "<p>Oh, that makes sense</p>",
      "rawMarkdown": "Oh, that makes sense",
      "votes": null
    },
    {
      "id": "1074915",
      "postDate": "11/11/2020 08:27:52",
      "content": "<p>I have opened a Notebook with a light version of a Transformer Encoder in TensorFlow. Feel free to check it out to gather some ideas. <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">NOTEBOOK</a></p>",
      "rawMarkdown": "I have opened a Notebook with a light version of a Transformer Encoder in TensorFlow. Feel free to check it out to gather some ideas. [NOTEBOOK](https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public)",
      "votes": null
    },
    {
      "id": "1075540",
      "postDate": "11/11/2020 19:12:37",
      "content": "<p>This did the trick to use Colab easily:</p>\n<pre><code>!pip install -q kaggle\n\n!mkdir ~/.kaggle\n!cp kaggle.json ~/.kaggle/\n!chmod 600 ~/.kaggle/kaggle.json \n!kaggle competitions download -c riiid-test-answer-prediction\n</code></pre>",
      "rawMarkdown": "This did the trick to use Colab easily:\n```\n!pip install -q kaggle\n\n!mkdir ~/.kaggle\n!cp kaggle.json ~/.kaggle/\n!chmod 600 ~/.kaggle/kaggle.json \n!kaggle competitions download -c riiid-test-answer-prediction\n```",
      "votes": null
    },
    {
      "id": "1075559",
      "postDate": "11/11/2020 19:26:12",
      "content": "<p>You won't be able to use the package <code>riiideducation</code> since colab use python 3.6. But you can work with train / validation, good enough.</p>",
      "rawMarkdown": "You won't be able to use the package `riiideducation` since colab use python 3.6. But you can work with train / validation, good enough.",
      "votes": null
    },
    {
      "id": "1075571",
      "postDate": "11/11/2020 19:32:09",
      "content": "<p>I can use it to train models which is my main bottleneck :)</p>",
      "rawMarkdown": "I can use it to train models which is my main bottleneck :)",
      "votes": null
    },
    {
      "id": "1075575",
      "postDate": "11/11/2020 19:38:22",
      "content": "<p>BTW, I saw that in another thread, you mentioned you stuck at 0.6x. But you LB is 0.763 ??</p>",
      "rawMarkdown": "BTW, I saw that in another thread, you mentioned you stuck at 0.6x. But you LB is 0.763 ??",
      "votes": null
    },
    {
      "id": "1075604",
      "postDate": "11/11/2020 20:06:18",
      "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> , just a tip if helpful - I use Google Bucket to save the checkpoint I obtained from Colab. Don't use path like './' to save models or their checkpoints on colab - otherwise don't forget to send it to a permanent place once the training is done.</p>",
      "rawMarkdown": "claverru , just a tip if helpful - I use Google Bucket to save the checkpoint I obtained from Colab. Don't use path like './' to save models or their checkpoints on colab - otherwise don't forget to send it to a permanent place once the training is done.",
      "votes": null
    },
    {
      "id": "1075611",
      "postDate": "11/11/2020 20:13:40",
      "content": "<p>I think I meant 60th in the LB. That've changed a bit since then he. Let's see if I can comeback though.</p>",
      "rawMarkdown": "I think I meant 60th in the LB. That've changed a bit since then he. Let's see if I can comeback though.",
      "votes": null
    },
    {
      "id": "1087404",
      "postDate": "11/22/2020 17:18:34",
      "content": "<p>I was just pondering about this -: Is it possible to somehow decay attention weights over time? Idea is very similar to how we learn things and then we forget them as well. Plus too long questions back in the past don't help much either if you are going to attempt a new one. </p>\n<p>So what it's supposed to ensure is that if a user faces a new question the very old experiences shouldn't be helpful/relevant.</p>\n<p>Thoughts?</p>",
      "rawMarkdown": "I was just pondering about this -: Is it possible to somehow decay attention weights over time? Idea is very similar to how we learn things and then we forget them as well. Plus too long questions back in the past don't help much either if you are going to attempt a new one. \n\nSo what it's supposed to ensure is that if a user faces a new question the very old experiences shouldn't be helpful/relevant.\n\nThoughts?",
      "votes": null
    },
    {
      "id": "1087418",
      "postDate": "11/22/2020 17:33:10",
      "content": "<p>Do you mean over training time?</p>",
      "rawMarkdown": "Do you mean over training time?",
      "votes": null
    },
    {
      "id": "1087420",
      "postDate": "11/22/2020 17:36:09",
      "content": "<p>Anyways it weird since they are activated by softmax. One the other hand, one thing is the attention an alumn pays to previous questions and another one is the attention an element in a sequence pays to another element in the sequence.</p>",
      "rawMarkdown": "Anyways it weird since they are activated by softmax. One the other hand, one thing is the attention an alumn pays to previous questions and another one is the attention an element in a sequence pays to another element in the sequence.",
      "votes": null
    },
    {
      "id": "1087751",
      "postDate": "11/23/2020 04:04:32",
      "content": "<blockquote>\n  <p>Do you mean over training time?</p>\n</blockquote>\n<p>Yes. But i guess maybe i am just over thinking. First should achieve good results and then think about all this.</p>",
      "rawMarkdown": ">Do you mean over training time?\n\nYes. But i guess maybe i am just over thinking. First should achieve good results and then think about all this.",
      "votes": null
    },
    {
      "id": "1088072",
      "postDate": "11/23/2020 09:43:38",
      "content": "<p>Don't they mention lag time as a temporal feature in Saint plus?</p>",
      "rawMarkdown": "Don't they mention lag time as a temporal feature in Saint plus?",
      "votes": null
    },
    {
      "id": "1111721",
      "postDate": "12/14/2020 00:49:08",
      "content": "<p><a href=\"https://www.kaggle.com/shivanandmn/riiid-sakt\" target=\"_blank\">https://www.kaggle.com/shivanandmn/riiid-sakt</a> . I am also tried to implement SAINT, but I am getting an embedding related error, would anybody help me. I am a  <strong>newbie</strong> </p>",
      "rawMarkdown": "https://www.kaggle.com/shivanandmn/riiid-sakt . I am also tried to implement SAINT, but I am getting an embedding related error, would anybody help me. I am a  **newbie**",
      "votes": null
    },
    {
      "id": "1116164",
      "postDate": "12/16/2020 23:09:29",
      "content": "<p>24 days later, these same thoughts are entering into my mind.</p>\n<p>MHA will build a bunch of attention heads for us. There's no reason not to try a separate path with a \"static\" or time-based attention head. The way it'd work is that for static head, the attention matrix could be based on something like a number of questions. So questions that are more than n questions away get decayed by log function, only decaying in the past direction. The advantage of this is that we can broadcast across the entire batch. With a time-based attention head, a separate max_len x max_len attention needs to be built for each sample in the batch, where the decay is based on timestamp. This time, where 0 = ts of the question at the position, and the further away historical we move from that, the more decayed.</p>",
      "rawMarkdown": "24 days later, these same thoughts are entering into my mind.\n\nMHA will build a bunch of attention heads for us. There's no reason not to try a separate path with a \"static\" or time-based attention head. The way it'd work is that for static head, the attention matrix could be based on something like a number of questions. So questions that are more than n questions away get decayed by log function, only decaying in the past direction. The advantage of this is that we can broadcast across the entire batch. With a time-based attention head, a separate max_len x max_len attention needs to be built for each sample in the batch, where the decay is based on timestamp. This time, where 0 = ts of the question at the position, and the further away historical we move from that, the more decayed.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1072602,
      "author_name": "rafiko1",
      "author_url": "",
      "post_date": "11/08/2020 13:50:04",
      "content": "<p>Usually you cap the sequence length up to <code>max_len</code> (sometimes referred as window size). <br>\nYou do this for the test set as well.<br>\nThen the maximum possible  is f.e. <code>max_len+1</code> (it doesn't necessarily have to be the value 1024)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1072618,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/08/2020 14:00:54",
          "content": "<p>Yeah, you cap the sequence. But you cannot cap the position encoding in this competition. You would be saying that the first interaction is in the same place than, say, the 512th at some point.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072624,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/08/2020 14:07:49",
          "content": "<p>Look at the diagram I quickly sketched. I have a window size of 100 but to be able to know the exact ordinal interaction I need to have this information in the position encoding.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072628,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "11/08/2020 14:17:21",
          "content": "<p>Yes, we do have that problem here. It's basically like <code>position_no % max_seq</code>. And the embedding for the same layer would go from 0-99 i feel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072652,
          "author_name": "rafiko1",
          "author_url": "",
          "post_date": "11/08/2020 14:53:10",
          "content": "<p>I agree, there will be some loss of info with window resizing. Currently I only use <code>position_no % max_seq</code> to at least incorporate the relative positions within the window.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072660,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "11/08/2020 15:04:07",
          "content": "<p>Currently, I have a maximal seq len as 512, and the window size is 128. For each truncated history sequence of length window size (i.e 128), my position is always from 0 to 127. </p>\n<p>This is not ideal, because it losses the actual interaction position (a user's interaction history might be very long).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1073110,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "11/09/2020 06:20:08",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>   Can you elaborate a bit on the difference of maximal sequence len and window size?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1074634,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "11/10/2020 22:10:30",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> They could be the same. I just fixed a maximal seq len 512 for all experimentation. And try to used different window size (64, 128, 256 etc ) &lt;= 512.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1074709,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "11/11/2020 00:52:34",
          "content": "<p>Oh, that makes sense</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1072672,
      "author_name": "claverru",
      "author_url": "",
      "post_date": "11/08/2020 15:19:03",
      "content": "<p>What I have done (don't know if it is the best solution neither if it is even correct) is:</p>\n<ul>\n<li>Generating position encodings with my d_model and a maximum position of 40K (the larger user history I've seen is around 16K).</li>\n<li>Create a column with <code>np.arange(len(user))</code> and gather those indices from the position encodings above. In tensorflow it is <code>p = tf.gather(position_encoding, indices, axis=0)</code>, where <ul>\n<li><code>shape(position_encoding) = (40K, d_model)</code></li>\n<li><code>shape(indices) = (batch_size, window_size)</code></li>\n<li><code>shape(p) = (batch_size, window_size, d_model)</code></li></ul></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1072685,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "11/08/2020 15:33:37",
          "content": "<p>It's fine, the only cost is about that you are using a larger embedding matrix for positions - and it <em>might</em> affect the  memory usage - especially if you use GPU and you want to do ensembling. It doesn't affect the inference timing because tf.gather should be quite fast and after it, the actually computation is based on <code>p</code> only, which is relatively small.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072697,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/08/2020 15:42:45",
          "content": "<p>Well, for a model dimension of 64, I'd have a (40 000, 64) tensor, which is around 82M bits in case theyre float 32, giving as a result ~10.3 MBytes. If my numbers are correct, I think we can deal with it. Though maybe 40K is too much, given that the most prolific alumn has 16k interactions, maybe with a length of 20K we are safe.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072701,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "11/08/2020 15:46:34",
          "content": "<p>Yep, if you stick to dim 64, it won't be a problem. Is dim 64 the one gives your current score? If so, I am surprised :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072710,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/08/2020 15:54:35",
          "content": "<p>My current score has a model dimension of 20 I think. I've changed it so many times.</p>\n<p>EDIT: Solution -&gt; Score</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072719,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/08/2020 16:06:51",
          "content": "<p>The problem is that I have to squeeze my model to fit in the GPU quota. My humble 1060 can't deal with a Transformer based model until there's any open source lineal attention or I get any 3080 xD. And even squeezing the model I smash the quota just during the weekend. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072731,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "11/08/2020 16:19:24",
          "content": "<p>try google colab (free or pro version)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072737,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/08/2020 16:22:54",
          "content": "<p>Hum, yeah but how do you import your data? The comfortable way I know is by connecting your Google Drive and it's terribly slow.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072752,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "11/08/2020 16:37:47",
          "content": "<p>you can put the date to google bucket, and download it from your colab notebook - i think. I haven't done this for this competition yet, but I have done it before for another competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072756,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "11/08/2020 16:45:55",
          "content": "<p>You can try this, there's an extension named as curl-wget, install it on Chrome. And then go to kaggle's data download section for the comp and simply hit the download all and then cancel it. Grab the curl link generated from that extension and copy/paste that to colab. Or, Use kaggle API Or, push it to S3.</p>\n<p>NB I/O should be a bottle neck nonetheless. (imo, not 100% sure)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1072983,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "11/09/2020 01:15:26",
          "content": "<p>Bro, IMHO, positional encodings help the attention mechanism figure out the <em>relative</em> position of inputs. Even if you encode the 1000th interaction of the user as pos 1000, since the window size is 100 and the network only has access to the last 100 interactions, its better(for consistency) to pos encode only the window size positions(i.e. 100). Because even if the network needs to attend to inputs before the window, it cant.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1073226,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/09/2020 10:07:05",
          "content": "<p>Hum, those are the things that I tried to discuss here in this post. So, are you sure about that? Then there's literally no difference between the first and the 1000th interaction, and intuitively the alumn should have improved. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1073231,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/09/2020 10:17:12",
          "content": "<p>In the paper <a href=\"https://arxiv.org/pdf/2002.07033.pdf\" target=\"_blank\">Towards an Appropriate Query, Key, and ValueComputation for Knowledge Tracing</a> they say:</p>\n<p><em>Position: The position (1st, 2nd, …) of an exercise or a\nresponse in the input sequence is represented as a position\nembedding vector. The position embeddings are shared\nacross the exercise sequence and the response sequence.</em></p>\n<p>The truth is that it isn't clear to me yet 😑.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1073256,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "11/09/2020 10:59:57",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> <br>\nIn the SAINT paper, they dont use the sinusoidal embeddings. They make pos embeddings a learnable parameter as well. Although I am not sure if it makes the model any better. But that is the reason for this line \"The position embeddings are shared across the exercise sequence and the response sequence.\". If they were the sinusoidal embeddings, they wouldnt have any parameters to share (they are constant). Here, think of position embedding as an encoding from a variable X of cardinality = max_seq_len , mapping it to X_hat which is a vector of dimensionality d_model. Does this help?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1073264,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "11/09/2020 11:06:46",
          "content": "<p>To add to this, I think a benefit of using learned position embeddings is that you can actually encode, say position 5000, if you actually want to encode the session count of the user instead of just the local position of the interaction. That might have some benefit to it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1073267,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/09/2020 11:09:42",
          "content": "<p>Yeah, a lot. I'm re-reading it and you're right. It seems like a learnable parameter. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1073342,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "11/09/2020 12:53:55",
          "content": "<p>Nice to know that! 👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075540,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/11/2020 19:12:37",
          "content": "<p>This did the trick to use Colab easily:</p>\n<pre><code>!pip install -q kaggle\n\n!mkdir ~/.kaggle\n!cp kaggle.json ~/.kaggle/\n!chmod 600 ~/.kaggle/kaggle.json \n!kaggle competitions download -c riiid-test-answer-prediction\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075559,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "11/11/2020 19:26:12",
          "content": "<p>You won't be able to use the package <code>riiideducation</code> since colab use python 3.6. But you can work with train / validation, good enough.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075571,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/11/2020 19:32:09",
          "content": "<p>I can use it to train models which is my main bottleneck :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075575,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "11/11/2020 19:38:22",
          "content": "<p>BTW, I saw that in another thread, you mentioned you stuck at 0.6x. But you LB is 0.763 ??</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075604,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "11/11/2020 20:06:18",
          "content": "<p><a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> , just a tip if helpful - I use Google Bucket to save the checkpoint I obtained from Colab. Don't use path like './' to save models or their checkpoints on colab - otherwise don't forget to send it to a permanent place once the training is done.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075611,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/11/2020 20:13:40",
          "content": "<p>I think I meant 60th in the LB. That've changed a bit since then he. Let's see if I can comeback though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1074915,
      "author_name": "claverru",
      "author_url": "",
      "post_date": "11/11/2020 08:27:52",
      "content": "<p>I have opened a Notebook with a light version of a Transformer Encoder in TensorFlow. Feel free to check it out to gather some ideas. <a href=\"https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public\" target=\"_blank\">NOTEBOOK</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1087404,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "11/22/2020 17:18:34",
      "content": "<p>I was just pondering about this -: Is it possible to somehow decay attention weights over time? Idea is very similar to how we learn things and then we forget them as well. Plus too long questions back in the past don't help much either if you are going to attempt a new one. </p>\n<p>So what it's supposed to ensure is that if a user faces a new question the very old experiences shouldn't be helpful/relevant.</p>\n<p>Thoughts?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1087418,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/22/2020 17:33:10",
          "content": "<p>Do you mean over training time?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1087420,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "11/22/2020 17:36:09",
          "content": "<p>Anyways it weird since they are activated by softmax. One the other hand, one thing is the attention an alumn pays to previous questions and another one is the attention an element in a sequence pays to another element in the sequence.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1087751,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "11/23/2020 04:04:32",
          "content": "<blockquote>\n  <p>Do you mean over training time?</p>\n</blockquote>\n<p>Yes. But i guess maybe i am just over thinking. First should achieve good results and then think about all this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1088072,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "11/23/2020 09:43:38",
          "content": "<p>Don't they mention lag time as a temporal feature in Saint plus?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1116164,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/16/2020 23:09:29",
          "content": "<p>24 days later, these same thoughts are entering into my mind.</p>\n<p>MHA will build a bunch of attention heads for us. There's no reason not to try a separate path with a \"static\" or time-based attention head. The way it'd work is that for static head, the attention matrix could be based on something like a number of questions. So questions that are more than n questions away get decayed by log function, only decaying in the past direction. The advantage of this is that we can broadcast across the entire batch. With a time-based attention head, a separate max_len x max_len attention needs to be built for each sample in the batch, where the decay is based on timestamp. This time, where 0 = ts of the question at the position, and the further away historical we move from that, the more decayed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1111721,
      "author_name": "shivanandmn",
      "author_url": "",
      "post_date": "12/14/2020 00:49:08",
      "content": "<p><a href=\"https://www.kaggle.com/shivanandmn/riiid-sakt\" target=\"_blank\">https://www.kaggle.com/shivanandmn/riiid-sakt</a> . I am also tried to implement SAINT, but I am getting an embedding related error, would anybody help me. I am a  <strong>newbie</strong> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1072564": "After getting stuck with my current approach, I've been paying attention to both papers shared by our host ([https://arxiv.org/abs/2010.12042](https://arxiv.org/abs/2010.12042) and [https://arxiv.org/abs/2002.07033](https://arxiv.org/abs/2002.07033)).\n\nI'm trying to develop a light adaptation to later be able to work around it. \n\nMy main concern at the moment is about the position encoding, which seems confusing to me, and maybe anyone could clarify.\n\nHow do we set up the maximum position? If I understand correctly, in the paper Attention is All You Need, we sample a sinusoidal-like vector with the maximum position, and then take from it only the window for the current chunk. \n\nFor instance, we can set a maximum position as 1024, but then we only take indices [0:256] from the vector for the first 256 tokens. Did I understand properly?\n\nIf yes, then how do we set the maximum position for our users' histories? In theory we could have a larger history in the test set than in the training set.\n\nThank you in advance fellas.",
    "1072602": "Usually you cap the sequence length up to `max_len` (sometimes referred as window size). \nYou do this for the test set as well.\nThen the maximum possible  is f.e. `max_len+1` (it doesn't necessarily have to be the value 1024)",
    "1072618": "Yeah, you cap the sequence. But you cannot cap the position encoding in this competition. You would be saying that the first interaction is in the same place than, say, the 512th at some point.",
    "1072624": "Look at the diagram I quickly sketched. I have a window size of 100 but to be able to know the exact ordinal interaction I need to have this information in the position encoding.",
    "1072628": "Yes, we do have that problem here. It's basically like `position_no % max_seq`. And the embedding for the same layer would go from 0-99 i feel.",
    "1072652": "I agree, there will be some loss of info with window resizing. Currently I only use ` position_no % max_seq` to at least incorporate the relative positions within the window.",
    "1072660": "Currently, I have a maximal seq len as 512, and the window size is 128. For each truncated history sequence of length window size (i.e 128), my position is always from 0 to 127. \n\nThis is not ideal, because it losses the actual interaction position (a user's interaction history might be very long).",
    "1072672": "What I have done (don't know if it is the best solution neither if it is even correct) is:\n- Generating position encodings with my d_model and a maximum position of 40K (the larger user history I've seen is around 16K).\n- Create a column with `np.arange(len(user))` and gather those indices from the position encodings above. In tensorflow it is `p = tf.gather(position_encoding, indices, axis=0)`, where \n  - `shape(position_encoding) = (40K, d_model)`\n  - `shape(indices) = (batch_size, window_size)`\n  - `shape(p) = (batch_size, window_size, d_model)`",
    "1072685": "It's fine, the only cost is about that you are using a larger embedding matrix for positions - and it *might* affect the  memory usage - especially if you use GPU and you want to do ensembling. It doesn't affect the inference timing because tf.gather should be quite fast and after it, the actually computation is based on `p` only, which is relatively small.",
    "1072697": "Well, for a model dimension of 64, I'd have a (40 000, 64) tensor, which is around 82M bits in case theyre float 32, giving as a result ~10.3 MBytes. If my numbers are correct, I think we can deal with it. Though maybe 40K is too much, given that the most prolific alumn has 16k interactions, maybe with a length of 20K we are safe.",
    "1072701": "Yep, if you stick to dim 64, it won't be a problem. Is dim 64 the one gives your current score? If so, I am surprised :)",
    "1072710": "My current score has a model dimension of 20 I think. I've changed it so many times.\n\nEDIT: Solution -> Score",
    "1072719": "The problem is that I have to squeeze my model to fit in the GPU quota. My humble 1060 can't deal with a Transformer based model until there's any open source lineal attention or I get any 3080 xD. And even squeezing the model I smash the quota just during the weekend.",
    "1072731": "try google colab (free or pro version)",
    "1072737": "Hum, yeah but how do you import your data? The comfortable way I know is by connecting your Google Drive and it's terribly slow.",
    "1072752": "you can put the date to google bucket, and download it from your colab notebook - i think. I haven't done this for this competition yet, but I have done it before for another competition.",
    "1072756": "You can try this, there's an extension named as curl-wget, install it on Chrome. And then go to kaggle's data download section for the comp and simply hit the download all and then cancel it. Grab the curl link generated from that extension and copy/paste that to colab. Or, Use kaggle API Or, push it to S3.\n\nNB I/O should be a bottle neck nonetheless. (imo, not 100% sure)",
    "1072983": "Bro, IMHO, positional encodings help the attention mechanism figure out the *relative* position of inputs. Even if you encode the 1000th interaction of the user as pos 1000, since the window size is 100 and the network only has access to the last 100 interactions, its better(for consistency) to pos encode only the window size positions(i.e. 100). Because even if the network needs to attend to inputs before the window, it cant.",
    "1073110": "yihdarshieh   Can you elaborate a bit on the difference of maximal sequence len and window size?",
    "1073226": "Hum, those are the things that I tried to discuss here in this post. So, are you sure about that? Then there's literally no difference between the first and the 1000th interaction, and intuitively the alumn should have improved.",
    "1073231": "In the paper [Towards an Appropriate Query, Key, and ValueComputation for Knowledge Tracing](https://arxiv.org/pdf/2002.07033.pdf) they say:\n\n*Position: The position (1st, 2nd, ...) of an exercise or a\nresponse in the input sequence is represented as a position\nembedding vector. The position embeddings are shared\nacross the exercise sequence and the response sequence.*\n\nThe truth is that it isn't clear to me yet 😑.",
    "1073256": "claverru \nIn the SAINT paper, they dont use the sinusoidal embeddings. They make pos embeddings a learnable parameter as well. Although I am not sure if it makes the model any better. But that is the reason for this line \"The position embeddings are shared across the exercise sequence and the response sequence.\". If they were the sinusoidal embeddings, they wouldnt have any parameters to share (they are constant). Here, think of position embedding as an encoding from a variable X of cardinality = max_seq_len , mapping it to X_hat which is a vector of dimensionality d_model. Does this help?",
    "1073264": "To add to this, I think a benefit of using learned position embeddings is that you can actually encode, say position 5000, if you actually want to encode the session count of the user instead of just the local position of the interaction. That might have some benefit to it.",
    "1073267": "Yeah, a lot. I'm re-reading it and you're right. It seems like a learnable parameter.",
    "1073342": "Nice to know that! 👍",
    "1074634": "abdurrafae They could be the same. I just fixed a maximal seq len 512 for all experimentation. And try to used different window size (64, 128, 256 etc ) <= 512.",
    "1074709": "Oh, that makes sense",
    "1074915": "I have opened a Notebook with a light version of a Transformer Encoder in TensorFlow. Feel free to check it out to gather some ideas. [NOTEBOOK](https://www.kaggle.com/claverru/demystifying-transformers-let-s-make-it-public)",
    "1075540": "This did the trick to use Colab easily:\n```\n!pip install -q kaggle\n\n!mkdir ~/.kaggle\n!cp kaggle.json ~/.kaggle/\n!chmod 600 ~/.kaggle/kaggle.json \n!kaggle competitions download -c riiid-test-answer-prediction\n```",
    "1075559": "You won't be able to use the package `riiideducation` since colab use python 3.6. But you can work with train / validation, good enough.",
    "1075571": "I can use it to train models which is my main bottleneck :)",
    "1075575": "BTW, I saw that in another thread, you mentioned you stuck at 0.6x. But you LB is 0.763 ??",
    "1075604": "claverru , just a tip if helpful - I use Google Bucket to save the checkpoint I obtained from Colab. Don't use path like './' to save models or their checkpoints on colab - otherwise don't forget to send it to a permanent place once the training is done.",
    "1075611": "I think I meant 60th in the LB. That've changed a bit since then he. Let's see if I can comeback though.",
    "1087404": "I was just pondering about this -: Is it possible to somehow decay attention weights over time? Idea is very similar to how we learn things and then we forget them as well. Plus too long questions back in the past don't help much either if you are going to attempt a new one. \n\nSo what it's supposed to ensure is that if a user faces a new question the very old experiences shouldn't be helpful/relevant.\n\nThoughts?",
    "1087418": "Do you mean over training time?",
    "1087420": "Anyways it weird since they are activated by softmax. One the other hand, one thing is the attention an alumn pays to previous questions and another one is the attention an element in a sequence pays to another element in the sequence.",
    "1087751": ">Do you mean over training time?\n\nYes. But i guess maybe i am just over thinking. First should achieve good results and then think about all this.",
    "1088072": "Don't they mention lag time as a temporal feature in Saint plus?",
    "1111721": "https://www.kaggle.com/shivanandmn/riiid-sakt . I am also tried to implement SAINT, but I am getting an embedding related error, would anybody help me. I am a  **newbie**",
    "1116164": "24 days later, these same thoughts are entering into my mind.\n\nMHA will build a bunch of attention heads for us. There's no reason not to try a separate path with a \"static\" or time-based attention head. The way it'd work is that for static head, the attention matrix could be based on something like a number of questions. So questions that are more than n questions away get decayed by log function, only decaying in the past direction. The advantage of this is that we can broadcast across the entire batch. With a time-based attention head, a separate max_len x max_len attention needs to be built for each sample in the batch, where the decay is based on timestamp. This time, where 0 = ts of the question at the position, and the further away historical we move from that, the more decayed."
  },
  "source": "meta"
}