{
  "id": 206185,
  "title": "Race to the finish line - SAKT/SSAKT and SAINT/SAINT+",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206185",
  "author_name": "Allohvk",
  "post_date": "2020-12-23T14:48:27.796000",
  "votes": 27,
  "comment_count": 13,
  "views": 0,
  "content": "<p>The number of teams has now crossed 3000 and my best wishes to all the teams!!!</p>\n<p>Going by the discussions and public kernels, these are the two models which seem to be rage. Saint is in fact tested on the EdNet database and so we can definitely rely on most of the interesting observations made in the paper. In one of the discussion threads on Saint, there seemed to be a bit of confusion on the design and hence I thought I will quickly bring out the key differences between these two models in case it benefits anyone. </p>\n<p>KEY DIFFERENCES: </p>\n<ul>\n<li><p>SAINT input is ridiculously simple. There is just the Exercise and Exercise category + of course the position on the encoder end. Thats it. The category is mostly the ‘part’. Its decoder side is even more crazy. Just contains the response and position. But look at the amazing results. Just by using the exercise(+part) and only the response, the transformer architecture is able to throw up such wonderful scores. </p></li>\n<li><p>Now let us come to the areas of confusion. (1) There is no interaction-embedding used in SAINT. Interaction = exercise+response as one entity. They do mention interactions in the paper but this is only for comparing SAINT benchmarks versus other models. SAINT itself uses ONLY exercises in the encoder as feed and NOT interactions. One of the helpful reference SAINT model implementation written in Kaggle uses ‘interactions’ on the encoder end. Perhaps it is just a semantics issue and the input is meant to be the exercise? I am not sure but if you want to replicate SAINT, you may want to keep this point in mind if you are forking that code<br>\n(2) I saw in one of the discussions that SAINT was using elapsed time, response time etc. Upon going thru’ the paper I was surprised to see that these were NOT being used at all. See Fig 3 in the original SAINT paper. They were being used only to show in the ablations that adding elapsed time and timestamp made NO difference to the score. This is kind of non-intuitive. However I later realised my mistake when I saw the SAINT+ paper, where they did add temporal information the decoder and this did make a marginal improvement to their scores. </p></li>\n<li><p>SAKT on the other hand uses interactions on the encoder end. On the decoder end it uses only exercises. This is similar to most other models and this is why I kind of like SAINT. They have a clean segregation of question and response and do not mix them up. This also results in better Key, Query, Value combinations:<br>\nSAINT: Key=value=encoder output = Exercise; Query = Response (a lil hard to grasp intuitively at least initially)<br>\nSAKT: Key=value= encoder output =Interactions; Query = Exercise (very intuitive and hence adopted by many models)<br>\nOf course for self-attention at both ends, the same K,Q,V are used respectively<br>\nWhy is the segregation of exercise (encoder) and response (decoder) important? They nicely show the difference in attention weights between the encoder and the decoder block in their Fig 7. If you look at it, you will immediately notice that the encoder weights are sparse and the decoder weights are quite rich. So possibly they mean to say that relationships are better captured this way (if segregated) rather than if we use interactions (exercise+responses) at the encoder end.</p></li>\n<li><p>SAKT has only one attention block that uses exercise embeddings as queries and interaction embeddings as keys/values. It is not really a transformer architecture in that sense. It does not use self-attention to embed the exercises, responses or the interactions. In fact the authors report a decrease in AUC when the SAKT attention block is stacked multiple times. SSAKT solves this issue by applying self-attention on exercises before supplying them as queries. The outputs of the exercise self-attention block and the exercise-interaction attention block enters the corresponding following blocks as inputs for their attention layers. So the reference architecture being used in Kaggle is more based on SSAKT rather than the original SAKT. <br>\nEDIT - THIS IS INCORRECT. The Kaggle kernels seem to follow the original SAKT paper to the hilt and don't have self attention.</p></li>\n<li><p>On the other hand, SAINT’s architecture is more easily stackable and provides better performance. They have 4 blocks - each recognising more and more complex relations compared to its previous block. Notice the attention weights in Fig 9 where they show the differences between attention weights after the first block and the 4th block. The last block is definitely richer and more thorough</p></li>\n</ul>\n<p>Lastly Saint had 10% dropout and Saint+ had 0% dropout..Seems counter intuitive to me. Overfitting?<br>\n<br><br>\nWhat else could make a big difference to this competition?</p>\n<ul>\n<li>Reductions (taking mean of questions in a bundle or even in a session) - probably not; local attention (probably yes); intelligent sample selection(hell yes); clever use of temporal information(yes but possibly limited?); clustering (Yes - definitely exercise, but students can be challenging); clever ensembles (not a 50-50) - yes; architecture improvements (possibly not..we may not see too much variation from Santa..but I would really love it if someone brings in convolutions or graphs into the game); brute power (possibly No); powerful features (possibly not in the transformer models but yes in others and ensembles); passing attention weights to the next step (possibly yes but only if the next exercise is in the same cluster); decaying attention weights(even a crude formula) - probably yes; leveraging tags and lecture information (Intuitively yes but who knows)..</li>\n</ul>\n<p>Unfortunately it is all guesswork because while the competition has entered its last phase and just when I finally seem to have gotten time from work to spend a few hours, I noticed that my submissions are not getting recognised and have raised a support ticket with Kaggle..No response so far… and I am twiddling my thumbs. If any of you have any suggestions to fix the issue do let me know… I just forked <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> API Detailed Introduction kernel. Did not change anything. Did a 'Run and commit All'. It runs fine and creates the submission file. I then submit the file, I get the message that Kaggle is going to run my notebook privately and score me. Waited for few hours to days….but my score is not getting reflected in scoreboard. No error message also under submissions. Tried multiple times. Each time, I re-run, I get the same message - \"You have 5 submissions remaining today\" which is also surprising since I remember reading somewhere that erroneous submissions are also counted under daily score. Any suggestions?</p>\n<p>My series on RiiiD (in case you liked this discussion):</p>\n<p>From Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481</a></p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203184\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203184</a> - Hidden features and possible architectures</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185</a> - Some additional clarifications on SAKT/SAINT</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584</a> - A small discussion on position embeddings for those interested.</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719</a> - On lectures, the art of forgetting and why I retired hurt</p>",
  "messages": [
    {
      "id": 1123864,
      "postDate": "2020-12-23T14:48:27.797Z",
      "content": "<p>The number of teams has now crossed 3000 and my best wishes to all the teams!!!</p>\n<p>Going by the discussions and public kernels, these are the two models which seem to be rage. Saint is in fact tested on the EdNet database and so we can definitely rely on most of the interesting observations made in the paper. In one of the discussion threads on Saint, there seemed to be a bit of confusion on the design and hence I thought I will quickly bring out the key differences between these two models in case it benefits anyone. </p>\n<p>KEY DIFFERENCES: </p>\n<ul>\n<li><p>SAINT input is ridiculously simple. There is just the Exercise and Exercise category + of course the position on the encoder end. Thats it. The category is mostly the ‘part’. Its decoder side is even more crazy. Just contains the response and position. But look at the amazing results. Just by using the exercise(+part) and only the response, the transformer architecture is able to throw up such wonderful scores. </p></li>\n<li><p>Now let us come to the areas of confusion. (1) There is no interaction-embedding used in SAINT. Interaction = exercise+response as one entity. They do mention interactions in the paper but this is only for comparing SAINT benchmarks versus other models. SAINT itself uses ONLY exercises in the encoder as feed and NOT interactions. One of the helpful reference SAINT model implementation written in Kaggle uses ‘interactions’ on the encoder end. Perhaps it is just a semantics issue and the input is meant to be the exercise? I am not sure but if you want to replicate SAINT, you may want to keep this point in mind if you are forking that code<br>\n(2) I saw in one of the discussions that SAINT was using elapsed time, response time etc. Upon going thru’ the paper I was surprised to see that these were NOT being used at all. See Fig 3 in the original SAINT paper. They were being used only to show in the ablations that adding elapsed time and timestamp made NO difference to the score. This is kind of non-intuitive. However I later realised my mistake when I saw the SAINT+ paper, where they did add temporal information the decoder and this did make a marginal improvement to their scores. </p></li>\n<li><p>SAKT on the other hand uses interactions on the encoder end. On the decoder end it uses only exercises. This is similar to most other models and this is why I kind of like SAINT. They have a clean segregation of question and response and do not mix them up. This also results in better Key, Query, Value combinations:<br>\nSAINT: Key=value=encoder output = Exercise; Query = Response (a lil hard to grasp intuitively at least initially)<br>\nSAKT: Key=value= encoder output =Interactions; Query = Exercise (very intuitive and hence adopted by many models)<br>\nOf course for self-attention at both ends, the same K,Q,V are used respectively<br>\nWhy is the segregation of exercise (encoder) and response (decoder) important? They nicely show the difference in attention weights between the encoder and the decoder block in their Fig 7. If you look at it, you will immediately notice that the encoder weights are sparse and the decoder weights are quite rich. So possibly they mean to say that relationships are better captured this way (if segregated) rather than if we use interactions (exercise+responses) at the encoder end.</p></li>\n<li><p>SAKT has only one attention block that uses exercise embeddings as queries and interaction embeddings as keys/values. It is not really a transformer architecture in that sense. It does not use self-attention to embed the exercises, responses or the interactions. In fact the authors report a decrease in AUC when the SAKT attention block is stacked multiple times. SSAKT solves this issue by applying self-attention on exercises before supplying them as queries. The outputs of the exercise self-attention block and the exercise-interaction attention block enters the corresponding following blocks as inputs for their attention layers. So the reference architecture being used in Kaggle is more based on SSAKT rather than the original SAKT. <br>\nEDIT - THIS IS INCORRECT. The Kaggle kernels seem to follow the original SAKT paper to the hilt and don't have self attention.</p></li>\n<li><p>On the other hand, SAINT’s architecture is more easily stackable and provides better performance. They have 4 blocks - each recognising more and more complex relations compared to its previous block. Notice the attention weights in Fig 9 where they show the differences between attention weights after the first block and the 4th block. The last block is definitely richer and more thorough</p></li>\n</ul>\n<p>Lastly Saint had 10% dropout and Saint+ had 0% dropout..Seems counter intuitive to me. Overfitting?<br>\n<br><br>\nWhat else could make a big difference to this competition?</p>\n<ul>\n<li>Reductions (taking mean of questions in a bundle or even in a session) - probably not; local attention (probably yes); intelligent sample selection(hell yes); clever use of temporal information(yes but possibly limited?); clustering (Yes - definitely exercise, but students can be challenging); clever ensembles (not a 50-50) - yes; architecture improvements (possibly not..we may not see too much variation from Santa..but I would really love it if someone brings in convolutions or graphs into the game); brute power (possibly No); powerful features (possibly not in the transformer models but yes in others and ensembles); passing attention weights to the next step (possibly yes but only if the next exercise is in the same cluster); decaying attention weights(even a crude formula) - probably yes; leveraging tags and lecture information (Intuitively yes but who knows)..</li>\n</ul>\n<p>Unfortunately it is all guesswork because while the competition has entered its last phase and just when I finally seem to have gotten time from work to spend a few hours, I noticed that my submissions are not getting recognised and have raised a support ticket with Kaggle..No response so far… and I am twiddling my thumbs. If any of you have any suggestions to fix the issue do let me know… I just forked <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> API Detailed Introduction kernel. Did not change anything. Did a 'Run and commit All'. It runs fine and creates the submission file. I then submit the file, I get the message that Kaggle is going to run my notebook privately and score me. Waited for few hours to days….but my score is not getting reflected in scoreboard. No error message also under submissions. Tried multiple times. Each time, I re-run, I get the same message - \"You have 5 submissions remaining today\" which is also surprising since I remember reading somewhere that erroneous submissions are also counted under daily score. Any suggestions?</p>\n<p>My series on RiiiD (in case you liked this discussion):</p>\n<p>From Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481</a></p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203184\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203184</a> - Hidden features and possible architectures</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185</a> - Some additional clarifications on SAKT/SAINT</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584</a> - A small discussion on position embeddings for those interested.</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719</a> - On lectures, the art of forgetting and why I retired hurt</p>",
      "rawMarkdown": "The number of teams has now crossed 3000 and my best wishes to all the teams!!!\n\nGoing by the discussions and public kernels, these are the two models which seem to be rage. Saint is in fact tested on the EdNet database and so we can definitely rely on most of the interesting observations made in the paper. In one of the discussion threads on Saint, there seemed to be a bit of confusion on the design and hence I thought I will quickly bring out the key differences between these two models in case it benefits anyone. \n\nKEY DIFFERENCES: \n- SAINT input is ridiculously simple. There is just the Exercise and Exercise category + of course the position on the encoder end. Thats it. The category is mostly the ‘part’. Its decoder side is even more crazy. Just contains the response and position. But look at the amazing results. Just by using the exercise(+part) and only the response, the transformer architecture is able to throw up such wonderful scores. \n- Now let us come to the areas of confusion. (1) There is no interaction-embedding used in SAINT. Interaction = exercise+response as one entity. They do mention interactions in the paper but this is only for comparing SAINT benchmarks versus other models. SAINT itself uses ONLY exercises in the encoder as feed and NOT interactions. One of the helpful reference SAINT model implementation written in Kaggle uses ‘interactions’ on the encoder end. Perhaps it is just a semantics issue and the input is meant to be the exercise? I am not sure but if you want to replicate SAINT, you may want to keep this point in mind if you are forking that code\n(2) I saw in one of the discussions that SAINT was using elapsed time, response time etc. Upon going thru’ the paper I was surprised to see that these were NOT being used at all. See Fig 3 in the original SAINT paper. They were being used only to show in the ablations that adding elapsed time and timestamp made NO difference to the score. This is kind of non-intuitive. However I later realised my mistake when I saw the SAINT+ paper, where they did add temporal information the decoder and this did make a marginal improvement to their scores. \n\n- SAKT on the other hand uses interactions on the encoder end. On the decoder end it uses only exercises. This is similar to most other models and this is why I kind of like SAINT. They have a clean segregation of question and response and do not mix them up. This also results in better Key, Query, Value combinations:\nSAINT: Key=value=encoder output = Exercise; Query = Response (a lil hard to grasp intuitively at least initially)\nSAKT: Key=value= encoder output =Interactions; Query = Exercise (very intuitive and hence adopted by many models)\nOf course for self-attention at both ends, the same K,Q,V are used respectively\nWhy is the segregation of exercise (encoder) and response (decoder) important? They nicely show the difference in attention weights between the encoder and the decoder block in their Fig 7. If you look at it, you will immediately notice that the encoder weights are sparse and the decoder weights are quite rich. So possibly they mean to say that relationships are better captured this way (if segregated) rather than if we use interactions (exercise+responses) at the encoder end.\n\n- SAKT has only one attention block that uses exercise embeddings as queries and interaction embeddings as keys/values. It is not really a transformer architecture in that sense. It does not use self-attention to embed the exercises, responses or the interactions. In fact the authors report a decrease in AUC when the SAKT attention block is stacked multiple times. SSAKT solves this issue by applying self-attention on exercises before supplying them as queries. The outputs of the exercise self-attention block and the exercise-interaction attention block enters the corresponding following blocks as inputs for their attention layers. So the reference architecture being used in Kaggle is more based on SSAKT rather than the original SAKT. \nEDIT - THIS IS INCORRECT. The Kaggle kernels seem to follow the original SAKT paper to the hilt and don't have self attention.\n\n- On the other hand, SAINT’s architecture is more easily stackable and provides better performance. They have 4 blocks - each recognising more and more complex relations compared to its previous block. Notice the attention weights in Fig 9 where they show the differences between attention weights after the first block and the 4th block. The last block is definitely richer and more thorough\n\nLastly Saint had 10% dropout and Saint+ had 0% dropout..Seems counter intuitive to me. Overfitting?\n<br>\nWhat else could make a big difference to this competition?\n- Reductions (taking mean of questions in a bundle or even in a session) - probably not; local attention (probably yes); intelligent sample selection(hell yes); clever use of temporal information(yes but possibly limited?); clustering (Yes - definitely exercise, but students can be challenging); clever ensembles (not a 50-50) - yes; architecture improvements (possibly not..we may not see too much variation from Santa..but I would really love it if someone brings in convolutions or graphs into the game); brute power (possibly No); powerful features (possibly not in the transformer models but yes in others and ensembles); passing attention weights to the next step (possibly yes but only if the next exercise is in the same cluster); decaying attention weights(even a crude formula) - probably yes; leveraging tags and lecture information (Intuitively yes but who knows)..\n\nUnfortunately it is all guesswork because while the competition has entered its last phase and just when I finally seem to have gotten time from work to spend a few hours, I noticed that my submissions are not getting recognised and have raised a support ticket with Kaggle..No response so far… and I am twiddling my thumbs. If any of you have any suggestions to fix the issue do let me know... I just forked @sohier API Detailed Introduction kernel. Did not change anything. Did a 'Run and commit All'. It runs fine and creates the submission file. I then submit the file, I get the message that Kaggle is going to run my notebook privately and score me. Waited for few hours to days….but my score is not getting reflected in scoreboard. No error message also under submissions. Tried multiple times. Each time, I re-run, I get the same message - \"You have 5 submissions remaining today\" which is also surprising since I remember reading somewhere that erroneous submissions are also counted under daily score. Any suggestions?\n\nMy series on RiiiD (in case you liked this discussion):\n\nFrom Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203184 - Hidden features and possible architectures\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185 - Some additional clarifications on SAKT/SAINT\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584 - A small discussion on position embeddings for those interested.\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719 - On lectures, the art of forgetting and why I retired hurt",
      "votes": 27
    },
    {
      "id": 1124094,
      "postDate": "2020-12-23T17:10:47.307Z",
      "content": "<p>What is SSAKT? I found SAKT paper, but could not find SSAKT.</p>",
      "rawMarkdown": "What is SSAKT? I found SAKT paper, but could not find SSAKT.",
      "votes": 1,
      "replies": [
        {
          "id": 1124095,
          "postDate": "2020-12-23T17:12:01.713Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1124115,
          "postDate": "2020-12-23T17:24:10.413Z",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> my bad..there is no official such paper. SSAKT is the term used for stacked SAKT in the SAINT paper. They did it because as part of their comparison studies they wanted to show how their architecture stacks up easily versus competition. You can refer to Saint paper for SSAKT details.. sorry for teh confusion </p>",
          "rawMarkdown": "@mamasinkgs my bad..there is no official such paper. SSAKT is the term used for stacked SAKT in the SAINT paper. They did it because as part of their comparison studies they wanted to show how their architecture stacks up easily versus competition. You can refer to Saint paper for SSAKT details.. sorry for teh confusion ",
          "votes": 1
        },
        {
          "id": 1124130,
          "postDate": "2020-12-23T17:34:56.963Z",
          "content": "<p>Thank you, I found SSAKT in SAINT paper! I will give it a try :)</p>",
          "rawMarkdown": "Thank you, I found SSAKT in SAINT paper! I will give it a try :)",
          "votes": 2
        },
        {
          "id": 1126496,
          "postDate": "2020-12-25T16:52:44.837Z",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> - nope..it doesent perform well against Saint even if stacked…BTW great to see u back on the top of the leaderboard  :)</p>",
          "rawMarkdown": "@mamasinkgs - nope..it doesent perform well against Saint even if stacked...BTW great to see u back on the top of the leaderboard  :)"
        }
      ]
    },
    {
      "id": 1124009,
      "postDate": "2020-12-23T16:26:22.723Z",
      "content": "<p>Unlike us, SAINT also has an absolute timestamp. While their encoding of it was naive in my opinion (a unique latent vector is assigned for every hour of the year, 8766 total), I believe there would have been a lot of room  for further bolstering these models with time-series like features if only they would have provided us absolute timestamps as well rather than relative ones. Perhaps after the competition I'll look around to see if I can get my hands on the EdNet dataset and test it out.</p>\n<p>I hope you're able to resolve the submission scoring issues. In my experience, there's no dedicated team that'll get you through that—you'll just have to tinker around with your sub and seek communal assistance. Comment out every line and then work forward from there.</p>",
      "rawMarkdown": "Unlike us, SAINT also has an absolute timestamp. While their encoding of it was naive in my opinion (a unique latent vector is assigned for every hour of the year, 8766 total), I believe there would have been a lot of room  for further bolstering these models with time-series like features if only they would have provided us absolute timestamps as well rather than relative ones. Perhaps after the competition I'll look around to see if I can get my hands on the EdNet dataset and test it out.\n\nI hope you're able to resolve the submission scoring issues. In my experience, there's no dedicated team that'll get you through that—you'll just have to tinker around with your sub and seek communal assistance. Comment out every line and then work forward from there.",
      "votes": 2,
      "replies": [
        {
          "id": 1124084,
          "postDate": "2020-12-23T17:04:25.783Z",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> for your suggestions! Will do a clean debug tomorrow. Worst case I will just create a new account and see (I suspect is it real Kaggle bug not program bug as barebones submission also not working)..</p>\n<p>Regarding Saint, you are correct.. I suspect there could have been room for improvement. When we talk of Time-series we usually talk of a fixed frequency..(say daily stocks etc). In this case of Riiid, the frequency is not fixed. Positonal encoding can only bring in the order of sequence. A good feature related to timestamp HAS to improve the score…hence was a bit surprised to see their paper..though they did some time related improvements in Saint+</p>\n<p>BTW we probably could work around with these relative timestamps to produce some absolute ones..We just give a arbitrary start date to each user…and then do our calculations based on that absolute date..Not sure if it will work though</p>",
          "rawMarkdown": "thanks @authman for your suggestions! Will do a clean debug tomorrow. Worst case I will just create a new account and see (I suspect is it real Kaggle bug not program bug as barebones submission also not working)..\n\nRegarding Saint, you are correct.. I suspect there could have been room for improvement. When we talk of Time-series we usually talk of a fixed frequency..(say daily stocks etc). In this case of Riiid, the frequency is not fixed. Positonal encoding can only bring in the order of sequence. A good feature related to timestamp HAS to improve the score...hence was a bit surprised to see their paper..though they did some time related improvements in Saint+\n\nBTW we probably could work around with these relative timestamps to produce some absolute ones..We just give a arbitrary start date to each user...and then do our calculations based on that absolute date..Not sure if it will work though\n\n"
        },
        {
          "id": 1124109,
          "postDate": "2020-12-23T17:20:14.557Z",
          "content": "<p>Don't create a new account!</p>\n<p>At least not before getting a response from a Kaggle admin. These types of simultaneous code-rerun competitions are notorious for providing little to no debugging information on submission issues, which causes many people grief and I hope Kaggle finds a better way of balancing helping devs debug vs leaking information about private set. But on a new account, even though your intentions are good, the site's TOS will likely mark you as a cheater and you'll end up having any submission made dropped from the LB.</p>\n<blockquote>\n  <p>BTW we probably could work around with these relative timestamps to produce some absolute ones..We just give a arbitrary start date to each user…and then do our calculations based on that absolute date..Not sure if it will work though</p>\n</blockquote>\n<p>I tried this yesterday and can confirm there is no signal. There's definitely a pattern, but no signal. I tried with both continuous and discretized values binned into hours for 24h in a day. If you look at the graphs, given the initial timestamp, people interact a lot closer to that time. This isn't just an artifact of new accounts either. If you limit to just accounts with, e.g. 10-100 interactions, 100-300 interactions, or 300+ interactions, the pattern remains constant. Unfortunately, so too does the mean accuracy ratio as well as standard deviations. The only thing that changes is the count, which while interesting, doesn't help us out.</p>",
          "rawMarkdown": "Don't create a new account!\n\nAt least not before getting a response from a Kaggle admin. These types of simultaneous code-rerun competitions are notorious for providing little to no debugging information on submission issues, which causes many people grief and I hope Kaggle finds a better way of balancing helping devs debug vs leaking information about private set. But on a new account, even though your intentions are good, the site's TOS will likely mark you as a cheater and you'll end up having any submission made dropped from the LB.\n\n> BTW we probably could work around with these relative timestamps to produce some absolute ones..We just give a arbitrary start date to each user…and then do our calculations based on that absolute date..Not sure if it will work though\n\nI tried this yesterday and can confirm there is no signal. There's definitely a pattern, but no signal. I tried with both continuous and discretized values binned into hours for 24h in a day. If you look at the graphs, given the initial timestamp, people interact a lot closer to that time. This isn't just an artifact of new accounts either. If you limit to just accounts with, e.g. 10-100 interactions, 100-300 interactions, or 300+ interactions, the pattern remains constant. Unfortunately, so too does the mean accuracy ratio as well as standard deviations. The only thing that changes is the count, which while interesting, doesn't help us out.",
          "votes": 2
        },
        {
          "id": 1124121,
          "postDate": "2020-12-23T17:28:00.257Z",
          "content": "<p>The shape of the image remains the same even when you remove the first few transactions a user makes to account for the majority of users that have very few interactions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F32b9895fedd229b6f99dd2eb47ccabd8%2FScreen%20Shot%202020-12-23%20at%2011.28.20%20AM.png?generation=1608744513383742&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "The shape of the image remains the same even when you remove the first few transactions a user makes to account for the majority of users that have very few interactions.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F32b9895fedd229b6f99dd2eb47ccabd8%2FScreen%20Shot%202020-12-23%20at%2011.28.20%20AM.png?generation=1608744513383742&alt=media)",
          "votes": 1
        },
        {
          "id": 1124131,
          "postDate": "2020-12-23T17:35:16.617Z",
          "content": "<p>hmm..this is interesting. wonder what we are missing here..</p>",
          "rawMarkdown": "hmm..this is interesting. wonder what we are missing here.."
        },
        {
          "id": 1125753,
          "postDate": "2020-12-25T03:11:40.180Z",
          "content": "<p>Do NOT create new account, because then you might get banned for multiple account usage!</p>",
          "rawMarkdown": "Do NOT create new account, because then you might get banned for multiple account usage!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1125724,
      "postDate": "2020-12-25T02:34:46.723Z",
      "content": "<p>Are there TensorFlow implementation of these models?</p>",
      "rawMarkdown": "Are there TensorFlow implementation of these models?",
      "replies": [
        {
          "id": 1125936,
          "postDate": "2020-12-25T07:28:42.343Z",
          "content": "<p>I saw a couple of Keras codes in the public Kaggle kernels. Possibly they need to be optimised. Pytorch is new to me (well, so was Keras 2 months back but that is diff story) but considering the nice high scoring kernels posted in Pytorch, I decided to switch. I am not repenting my decision at all..Torch rocks. It is goodbye keras from now on for me :) </p>",
          "rawMarkdown": "I saw a couple of Keras codes in the public Kaggle kernels. Possibly they need to be optimised. Pytorch is new to me (well, so was Keras 2 months back but that is diff story) but considering the nice high scoring kernels posted in Pytorch, I decided to switch. I am not repenting my decision at all..Torch rocks. It is goodbye keras from now on for me :) "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1124094,
      "author_name": "mamas",
      "author_url": "",
      "post_date": "2020-12-23T17:10:47.307000",
      "content": "<p>What is SSAKT? I found SAKT paper, but could not find SSAKT.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1124095,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-23T17:12:01.713000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1124115,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-23T17:24:10.413000",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> my bad..there is no official such paper. SSAKT is the term used for stacked SAKT in the SAINT paper. They did it because as part of their comparison studies they wanted to show how their architecture stacks up easily versus competition. You can refer to Saint paper for SSAKT details.. sorry for teh confusion </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1124130,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2020-12-23T17:34:56.963000",
          "content": "<p>Thank you, I found SSAKT in SAINT paper! I will give it a try :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1126496,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-25T16:52:44.837000",
          "content": "<p><a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> - nope..it doesent perform well against Saint even if stacked…BTW great to see u back on the top of the leaderboard  :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1124009,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-12-23T16:26:22.723000",
      "content": "<p>Unlike us, SAINT also has an absolute timestamp. While their encoding of it was naive in my opinion (a unique latent vector is assigned for every hour of the year, 8766 total), I believe there would have been a lot of room  for further bolstering these models with time-series like features if only they would have provided us absolute timestamps as well rather than relative ones. Perhaps after the competition I'll look around to see if I can get my hands on the EdNet dataset and test it out.</p>\n<p>I hope you're able to resolve the submission scoring issues. In my experience, there's no dedicated team that'll get you through that—you'll just have to tinker around with your sub and seek communal assistance. Comment out every line and then work forward from there.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1124084,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-23T17:04:25.783000",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> for your suggestions! Will do a clean debug tomorrow. Worst case I will just create a new account and see (I suspect is it real Kaggle bug not program bug as barebones submission also not working)..</p>\n<p>Regarding Saint, you are correct.. I suspect there could have been room for improvement. When we talk of Time-series we usually talk of a fixed frequency..(say daily stocks etc). In this case of Riiid, the frequency is not fixed. Positonal encoding can only bring in the order of sequence. A good feature related to timestamp HAS to improve the score…hence was a bit surprised to see their paper..though they did some time related improvements in Saint+</p>\n<p>BTW we probably could work around with these relative timestamps to produce some absolute ones..We just give a arbitrary start date to each user…and then do our calculations based on that absolute date..Not sure if it will work though</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1124109,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-23T17:20:14.557000",
          "content": "<p>Don't create a new account!</p>\n<p>At least not before getting a response from a Kaggle admin. These types of simultaneous code-rerun competitions are notorious for providing little to no debugging information on submission issues, which causes many people grief and I hope Kaggle finds a better way of balancing helping devs debug vs leaking information about private set. But on a new account, even though your intentions are good, the site's TOS will likely mark you as a cheater and you'll end up having any submission made dropped from the LB.</p>\n<blockquote>\n  <p>BTW we probably could work around with these relative timestamps to produce some absolute ones..We just give a arbitrary start date to each user…and then do our calculations based on that absolute date..Not sure if it will work though</p>\n</blockquote>\n<p>I tried this yesterday and can confirm there is no signal. There's definitely a pattern, but no signal. I tried with both continuous and discretized values binned into hours for 24h in a day. If you look at the graphs, given the initial timestamp, people interact a lot closer to that time. This isn't just an artifact of new accounts either. If you limit to just accounts with, e.g. 10-100 interactions, 100-300 interactions, or 300+ interactions, the pattern remains constant. Unfortunately, so too does the mean accuracy ratio as well as standard deviations. The only thing that changes is the count, which while interesting, doesn't help us out.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1124121,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-12-23T17:28:00.257000",
          "content": "<p>The shape of the image remains the same even when you remove the first few transactions a user makes to account for the majority of users that have very few interactions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F933480%2F32b9895fedd229b6f99dd2eb47ccabd8%2FScreen%20Shot%202020-12-23%20at%2011.28.20%20AM.png?generation=1608744513383742&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1124131,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-23T17:35:16.617000",
          "content": "<p>hmm..this is interesting. wonder what we are missing here..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1125753,
          "author_name": "AmorfEvo",
          "author_url": "",
          "post_date": "2020-12-25T03:11:40.180000",
          "content": "<p>Do NOT create new account, because then you might get banned for multiple account usage!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1125724,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2020-12-25T02:34:46.723000",
      "content": "<p>Are there TensorFlow implementation of these models?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1125936,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-25T07:28:42.343000",
          "content": "<p>I saw a couple of Keras codes in the public Kaggle kernels. Possibly they need to be optimised. Pytorch is new to me (well, so was Keras 2 months back but that is diff story) but considering the nice high scoring kernels posted in Pytorch, I decided to switch. I am not repenting my decision at all..Torch rocks. It is goodbye keras from now on for me :) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1123864": "The number of teams has now crossed 3000 and my best wishes to all the teams!!!\n\nGoing by the discussions and public kernels, these are the two models which seem to be rage. Saint is in fact tested on the EdNet database and so we can definitely rely on most of the interesting observations made in the paper. In one of the discussion threads on Saint, there seemed to be a bit of confusion on the design and hence I thought I will quickly bring out the key differences between these two models in case it benefits anyone. \n\nKEY DIFFERENCES: \n- SAINT input is ridiculously simple. There is just the Exercise and Exercise category + of course the position on the encoder end. Thats it. The category is mostly the ‘part’. Its decoder side is even more crazy. Just contains the response and position. But look at the amazing results. Just by using the exercise(+part) and only the response, the transformer architecture is able to throw up such wonderful scores. \n- Now let us come to the areas of confusion. (1) There is no interaction-embedding used in SAINT. Interaction = exercise+response as one entity. They do mention interactions in the paper but this is only for comparing SAINT benchmarks versus other models. SAINT itself uses ONLY exercises in the encoder as feed and NOT interactions. One of the helpful reference SAINT model implementation written in Kaggle uses ‘interactions’ on the encoder end. Perhaps it is just a semantics issue and the input is meant to be the exercise? I am not sure but if you want to replicate SAINT, you may want to keep this point in mind if you are forking that code\n(2) I saw in one of the discussions that SAINT was using elapsed time, response time etc. Upon going thru’ the paper I was surprised to see that these were NOT being used at all. See Fig 3 in the original SAINT paper. They were being used only to show in the ablations that adding elapsed time and timestamp made NO difference to the score. This is kind of non-intuitive. However I later realised my mistake when I saw the SAINT+ paper, where they did add temporal information the decoder and this did make a marginal improvement to their scores. \n\n- SAKT on the other hand uses interactions on the encoder end. On the decoder end it uses only exercises. This is similar to most other models and this is why I kind of like SAINT. They have a clean segregation of question and response and do not mix them up. This also results in better Key, Query, Value combinations:\nSAINT: Key=value=encoder output = Exercise; Query = Response (a lil hard to grasp intuitively at least initially)\nSAKT: Key=value= encoder output =Interactions; Query = Exercise (very intuitive and hence adopted by many models)\nOf course for self-attention at both ends, the same K,Q,V are used respectively\nWhy is the segregation of exercise (encoder) and response (decoder) important? They nicely show the difference in attention weights between the encoder and the decoder block in their Fig 7. If you look at it, you will immediately notice that the encoder weights are sparse and the decoder weights are quite rich. So possibly they mean to say that relationships are better captured this way (if segregated) rather than if we use interactions (exercise+responses) at the encoder end.\n\n- SAKT has only one attention block that uses exercise embeddings as queries and interaction embeddings as keys/values. It is not really a transformer architecture in that sense. It does not use self-attention to embed the exercises, responses or the interactions. In fact the authors report a decrease in AUC when the SAKT attention block is stacked multiple times. SSAKT solves this issue by applying self-attention on exercises before supplying them as queries. The outputs of the exercise self-attention block and the exercise-interaction attention block enters the corresponding following blocks as inputs for their attention layers. So the reference architecture being used in Kaggle is more based on SSAKT rather than the original SAKT. \nEDIT - THIS IS INCORRECT. The Kaggle kernels seem to follow the original SAKT paper to the hilt and don't have self attention.\n\n- On the other hand, SAINT’s architecture is more easily stackable and provides better performance. They have 4 blocks - each recognising more and more complex relations compared to its previous block. Notice the attention weights in Fig 9 where they show the differences between attention weights after the first block and the 4th block. The last block is definitely richer and more thorough\n\nLastly Saint had 10% dropout and Saint+ had 0% dropout..Seems counter intuitive to me. Overfitting?\n<br>\nWhat else could make a big difference to this competition?\n- Reductions (taking mean of questions in a bundle or even in a session) - probably not; local attention (probably yes); intelligent sample selection(hell yes); clever use of temporal information(yes but possibly limited?); clustering (Yes - definitely exercise, but students can be challenging); clever ensembles (not a 50-50) - yes; architecture improvements (possibly not..we may not see too much variation from Santa..but I would really love it if someone brings in convolutions or graphs into the game); brute power (possibly No); powerful features (possibly not in the transformer models but yes in others and ensembles); passing attention weights to the next step (possibly yes but only if the next exercise is in the same cluster); decaying attention weights(even a crude formula) - probably yes; leveraging tags and lecture information (Intuitively yes but who knows)..\n\nUnfortunately it is all guesswork because while the competition has entered its last phase and just when I finally seem to have gotten time from work to spend a few hours, I noticed that my submissions are not getting recognised and have raised a support ticket with Kaggle..No response so far… and I am twiddling my thumbs. If any of you have any suggestions to fix the issue do let me know... I just forked @sohier API Detailed Introduction kernel. Did not change anything. Did a 'Run and commit All'. It runs fine and creates the submission file. I then submit the file, I get the message that Kaggle is going to run my notebook privately and score me. Waited for few hours to days….but my score is not getting reflected in scoreboard. No error message also under submissions. Tried multiple times. Each time, I re-run, I get the same message - \"You have 5 submissions remaining today\" which is also surprising since I remember reading somewhere that erroneous submissions are also counted under daily score. Any suggestions?\n\nMy series on RiiiD (in case you liked this discussion):\n\nFrom Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203184 - Hidden features and possible architectures\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185 - Some additional clarifications on SAKT/SAINT\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584 - A small discussion on position embeddings for those interested.\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719 - On lectures, the art of forgetting and why I retired hurt",
    "1124094": "What is SSAKT? I found SAKT paper, but could not find SSAKT.",
    "1124009": "Unlike us, SAINT also has an absolute timestamp. While their encoding of it was naive in my opinion (a unique latent vector is assigned for every hour of the year, 8766 total), I believe there would have been a lot of room  for further bolstering these models with time-series like features if only they would have provided us absolute timestamps as well rather than relative ones. Perhaps after the competition I'll look around to see if I can get my hands on the EdNet dataset and test it out.\n\nI hope you're able to resolve the submission scoring issues. In my experience, there's no dedicated team that'll get you through that—you'll just have to tinker around with your sub and seek communal assistance. Comment out every line and then work forward from there.",
    "1125724": "Are there TensorFlow implementation of these models?"
  }
}