{
  "id": 189398,
  "title": "[placeholder] transformer for knowledge tracing",
  "url": "/competitions/riiid-test-answer-prediction/discussion/189398",
  "author_name": "hengck23",
  "post_date": "2020-10-07T13:27:33.633000",
  "votes": 53,
  "comment_count": 27,
  "views": 0,
  "content": "<p>this post will be updated as i go along. it is like a mini-diary of my work.</p>\n<p>… to be updated ….</p>\n<p>for a start: <a href=\"https://paperswithcode.com/task/knowledge-tracing\" target=\"_blank\">https://paperswithcode.com/task/knowledge-tracing</a></p>\n<p>definition: <br>\nKnowledge tracing—where a machine models the knowledge of a student as they interact with coursework, see [1]</p>\n<p>[1] 'Deep Knowledge Tracing'- Chris Piech, nips 2015</p>",
  "messages": [
    {
      "id": 1040971,
      "postDate": "2020-10-07T13:27:33.633Z",
      "content": "<p>this post will be updated as i go along. it is like a mini-diary of my work.</p>\n<p>… to be updated ….</p>\n<p>for a start: <a href=\"https://paperswithcode.com/task/knowledge-tracing\" target=\"_blank\">https://paperswithcode.com/task/knowledge-tracing</a></p>\n<p>definition: <br>\nKnowledge tracing—where a machine models the knowledge of a student as they interact with coursework, see [1]</p>\n<p>[1] 'Deep Knowledge Tracing'- Chris Piech, nips 2015</p>",
      "rawMarkdown": "this post will be updated as i go along. it is like a mini-diary of my work.\n\n... to be updated ....\n\nfor a start: https://paperswithcode.com/task/knowledge-tracing\n\ndefinition: \nKnowledge tracing—where a machine models the knowledge of a student as they interact with coursework, see [1]\n\n[1] 'Deep Knowledge Tracing'- Chris Piech, nips 2015",
      "votes": 53
    },
    {
      "id": 1041887,
      "postDate": "2020-10-08T00:43:14.067Z",
      "content": "<p>after a night of paper reading and youtube surfacing for quick tutorials, I think a basic model looks like this<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7a22a6edb1adfd2a59f2d124bc773745%2FSelection_166.png?generation=1602117792236510&amp;alt=media\" alt=\"\"></p>\n<p>using this model, my first step is to make a naive model, i.e. prediction using average values </p>",
      "rawMarkdown": "after a night of paper reading and youtube surfacing for quick tutorials, I think a basic model looks like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7a22a6edb1adfd2a59f2d124bc773745%2FSelection_166.png?generation=1602117792236510&alt=media)\n\nusing this model, my first step is to make a naive model, i.e. prediction using average values ",
      "votes": 5,
      "replies": [
        {
          "id": 1042092,
          "postDate": "2020-10-08T03:59:59.757Z",
          "content": "<p>it is a seq-to-seq problem</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3d2d55d0babd7309ea6693bdfa4a703b%2FSelection_173.png?generation=1602129597919384&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "it is a seq-to-seq problem\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3d2d55d0babd7309ea6693bdfa4a703b%2FSelection_173.png?generation=1602129597919384&alt=media)",
          "votes": 6
        },
        {
          "id": 1042157,
          "postDate": "2020-10-08T05:16:58.820Z",
          "content": "<p>Probably we could do some bayesian step at the end</p>",
          "rawMarkdown": "Probably we could do some bayesian step at the end"
        },
        {
          "id": 1043455,
          "postDate": "2020-10-09T02:16:29.300Z",
          "content": "<p>Have you read the SAINT paper yet? </p>\n<ul>\n<li><a href=\"https://arxiv.org/pdf/2002.07033.pdf\" target=\"_blank\">https://arxiv.org/pdf/2002.07033.pdf</a></li>\n</ul>",
          "rawMarkdown": "Have you read the SAINT paper yet? \n\n* https://arxiv.org/pdf/2002.07033.pdf",
          "votes": 1
        },
        {
          "id": 1043907,
          "postDate": "2020-10-09T10:21:38.277Z",
          "content": "<p><a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> <br>\nthanks for the paper. yes this is one of the paper in my list. if kagglers has paper that they wish to implement, they can suggest here.</p>\n<p>it may be easier if a group of kagglers develop some baseline methods together and we can cross check each other bug.</p>\n<p>i also looking for papers that actually do recommendation but can be converted to knowledge tracking.</p>",
          "rawMarkdown": "@tuckerarrants \nthanks for the paper. yes this is one of the paper in my list. if kagglers has paper that they wish to implement, they can suggest here.\n\nit may be easier if a group of kagglers develop some baseline methods together and we can cross check each other bug.\n\ni also looking for papers that actually do recommendation but can be converted to knowledge tracking.",
          "votes": 1
        },
        {
          "id": 1046054,
          "postDate": "2020-10-11T10:08:23.377Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23/notebookcdb764afc6?scriptVersionId=44475365\" target=\"_blank\">https://www.kaggle.com/hengck23/notebookcdb764afc6?scriptVersionId=44475365</a><br>\naverage prediction baseline : LB 0.741</p>",
          "rawMarkdown": "https://www.kaggle.com/hengck23/notebookcdb764afc6?scriptVersionId=44475365\naverage prediction baseline : LB 0.741",
          "votes": 6
        },
        {
          "id": 1046191,
          "postDate": "2020-10-11T12:25:33.277Z",
          "content": "<p>I completed the static average approach. The next step is <strong>online average prediction</strong> which is another baseline before the time-series approach.</p>",
          "rawMarkdown": "I completed the static average approach. The next step is **online average prediction** which is another baseline before the time-series approach.",
          "votes": 1
        },
        {
          "id": 1046330,
          "postDate": "2020-10-11T14:53:37.190Z",
          "content": "<p>Just curious Heng, how do you plan to decide the window_size? Plus in test_data using API, we don't get the pupils data in one-shot, it comes in breaks at any random group_num.</p>\n<pre><code>def moving_average(series, n):\n    \"\"\"\n        Calculate average of last n observations\n    \"\"\"\n    return np.average(series[-n:])\n\nmoving_average(date_sales, 24) # prediction for the last observed day (past 24 hours)\n</code></pre>\n<p>Something like this. (<a href=\"https://www.kaggle.com/adityaecdrid/my-first-time-series-comp-added-prophet\" target=\"_blank\">ref</a>)</p>",
          "rawMarkdown": "Just curious Heng, how do you plan to decide the window_size? Plus in test_data using API, we don't get the pupils data in one-shot, it comes in breaks at any random group_num.\n\n```\ndef moving_average(series, n):\n    \"\"\"\n        Calculate average of last n observations\n    \"\"\"\n    return np.average(series[-n:])\n\nmoving_average(date_sales, 24) # prediction for the last observed day (past 24 hours)\n```\n\nSomething like this. ([ref](https://www.kaggle.com/adityaecdrid/my-first-time-series-comp-added-prophet))",
          "votes": 1
        },
        {
          "id": 1046407,
          "postDate": "2020-10-11T16:15:41.953Z",
          "content": "<p>But in inference you dont know the question sequence because with the submission API you have to preeict a batch before getting the next one, how would you adress that problem?</p>",
          "rawMarkdown": "But in inference you dont know the question sequence because with the submission API you have to preeict a batch before getting the next one, how would you adress that problem?",
          "votes": 1
        },
        {
          "id": 1046471,
          "postDate": "2020-10-11T17:27:31.853Z",
          "content": "<p>If you have the same user in train and test, you could just concatenate their new sequences to the old ones you’ve already seen in train as the API delivers it to you. If there are new users, you start a new sequence for them and predict with what you know of the questions and other metadata. </p>",
          "rawMarkdown": "If you have the same user in train and test, you could just concatenate their new sequences to the old ones you’ve already seen in train as the API delivers it to you. If there are new users, you start a new sequence for them and predict with what you know of the questions and other metadata. ",
          "votes": 2
        },
        {
          "id": 1046481,
          "postDate": "2020-10-11T17:45:35.757Z",
          "content": "<p>Not enough memory to store all users seqs, besides the issue that is not ensured that test seqs came after train seqs</p>",
          "rawMarkdown": "Not enough memory to store all users seqs, besides the issue that is not ensured that test seqs came after train seqs"
        },
        {
          "id": 1046490,
          "postDate": "2020-10-11T17:53:08.140Z",
          "content": "<p>i suppose these baselines will not be very high score, but it gives you an idea how the later rnn, transformer would perform.</p>\n<p>i am thinking of  just implementing a moving average counter like this:</p>\n<p><a href=\"https://www.kaggle.com/kneroma/riid-user-and-content-mean-predictor\" target=\"_blank\">https://www.kaggle.com/kneroma/riid-user-and-content-mean-predictor</a></p>",
          "rawMarkdown": "i suppose these baselines will not be very high score, but it gives you an idea how the later rnn, transformer would perform.\n\ni am thinking of  just implementing a moving average counter like this:\n\nhttps://www.kaggle.com/kneroma/riid-user-and-content-mean-predictor\n\n ",
          "votes": 2
        },
        {
          "id": 1050537,
          "postDate": "2020-10-15T14:15:29.993Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> New to time series, could you explain what is relation between baseline models correlate with seq2seq rnn, transformer models ?.. </p>\n<p>Do you mean is moving average counter features give an idea of will these features work well in seq2seq models or not ?</p>",
          "rawMarkdown": "@hengck23 New to time series, could you explain what is relation between baseline models correlate with seq2seq rnn, transformer models ?.. \n\nDo you mean is moving average counter features give an idea of will these features work well in seq2seq models or not ?"
        },
        {
          "id": 1054239,
          "postDate": "2020-10-19T19:23:13.263Z",
          "content": "<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> In this context, a simple seq2seq model would be using a sequence of questions that a particular student saw (chronologically) as the input sequence and using the sequence of answers to those questions as the output sequence. This way, we train a model to 'translate' between questions and answers.</p>",
          "rawMarkdown": "@seshurajup In this context, a simple seq2seq model would be using a sequence of questions that a particular student saw (chronologically) as the input sequence and using the sequence of answers to those questions as the output sequence. This way, we train a model to 'translate' between questions and answers."
        },
        {
          "id": 1054307,
          "postDate": "2020-10-19T20:08:40.097Z",
          "content": "<p>On private test you only can solve one question at a time, so its partial seq to seq, your seq is the past questions &amp; answers and the output must be 0/1 answered correctly to that question</p>",
          "rawMarkdown": "On private test you only can solve one question at a time, so its partial seq to seq, your seq is the past questions & answers and the output must be 0/1 answered correctly to that question"
        },
        {
          "id": 1063002,
          "postDate": "2020-10-28T12:18:25.333Z",
          "content": "<p>here is the update on the online average baseline: LB 0.748<br>\n(the code is not efficient and takes 6hr to run)</p>\n<p><a href=\"https://www.kaggle.com/hengck23/notebookf100a7833c\" target=\"_blank\">https://www.kaggle.com/hengck23/notebookf100a7833c</a></p>",
          "rawMarkdown": "here is the update on the online average baseline: LB 0.748\n(the code is not efficient and takes 6hr to run)\n\nhttps://www.kaggle.com/hengck23/notebookf100a7833c",
          "votes": 2
        },
        {
          "id": 1066893,
          "postDate": "2020-11-02T07:28:42.293Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  sorry it would be a very basic question as i  am quite new to this space so trying to build up the learning..<br>\nhow do we interpret users responses to Questions with time based  features in terms of embeddings. <br>\nExample see below excerpt from SAINT paper</p>\n<p>Excercise and response embeddings what info would these embeddings have..</p>\n<pre><code>In this paper, we propose a novel Transformer-based model for\nknowledge tracing, SAINT: Separated Self-AttentIve Neural\nKnowledge Tracing. SAINT has an encoder-decoder structure where the exercise and response embedding sequences\nseparately enter, respectively, the encoder and the decoder\n</code></pre>",
          "rawMarkdown": "@hengck23  sorry it would be a very basic question as i  am quite new to this space so trying to build up the learning..\nhow do we interpret users responses to Questions with time based  features in terms of embeddings. \nExample see below excerpt from SAINT paper\n\nExcercise and response embeddings what info would these embeddings have..\n```\nIn this paper, we propose a novel Transformer-based model for\nknowledge tracing, SAINT: Separated Self-AttentIve Neural\nKnowledge Tracing. SAINT has an encoder-decoder structure where the exercise and response embedding sequences\nseparately enter, respectively, the encoder and the decoder\n```\n\n",
          "votes": 1
        },
        {
          "id": 1097750,
          "postDate": "2020-12-01T08:10:19.530Z",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> I am not aware which exact paper you are referring to. But in general, I reviewed a bunch of papers from there and most use Attention in some way or the other. I don't think there would be a big difference between the various models (based on the 2-3 that I saw), nonetheless let me dig further and revert.</p>\n<p>To answer your questions:</p>\n<ul>\n<li>Embeddings are used to represent each question and QnA pair. They take a bunch of params like the exerciseID, category, position, elapsed time, timestamp etc and create the embeddings during training </li>\n<li>This competition becomes a form of supervised sequence learning task - given student’s past exercise interactions X = (x1, x2, . . . , xt), predict her next interaction xt+1. x = (e,r), where e is the exercise and r is the correctness of the student’s answer. This data is available at every past timestep 't'</li>\n<li>Because knowledge learning is a sequential process (for e.g. you learn equations of motion first..then probably Newton's laws, then gravity etc..where things flow in a sequence..likewise peformance on a particular exercise is dependent on her performance on the past exercises esp. ones related to that exercise) so RNN models can be used to nicely encode this representation..</li>\n<li>With RNN comes the problem of long sequence representation</li>\n<li>This is solved by attention. This esp. helps to focus on the scores of 'specific' past exercises during prediction</li>\n<li>Some papers have used attention with RNN to augment the scores and some have removed RNNs altogether and used attention in all its glory (transformer version - the encoder part related to self attention) </li>\n<li>There are minor differences between the papers on how the keys, queries and values need to be computed otherwise these represent the standard transformer encoder model introduced in 'Attention is all u need'</li>\n</ul>\n<p>Hope this intro helps.</p>",
          "rawMarkdown": "@jaideepvalani I am not aware which exact paper you are referring to. But in general, I reviewed a bunch of papers from there and most use Attention in some way or the other. I don't think there would be a big difference between the various models (based on the 2-3 that I saw), nonetheless let me dig further and revert.\n\nTo answer your questions:\n- Embeddings are used to represent each question and QnA pair. They take a bunch of params like the exerciseID, category, position, elapsed time, timestamp etc and create the embeddings during training \n- This competition becomes a form of supervised sequence learning task - given student’s past exercise interactions X = (x1, x2, . . . , xt), predict her next interaction xt+1. x = (e,r), where e is the exercise and r is the correctness of the student’s answer. This data is available at every past timestep 't'\n- Because knowledge learning is a sequential process (for e.g. you learn equations of motion first..then probably Newton's laws, then gravity etc..where things flow in a sequence..likewise peformance on a particular exercise is dependent on her performance on the past exercises esp. ones related to that exercise) so RNN models can be used to nicely encode this representation..\n- With RNN comes the problem of long sequence representation\n- This is solved by attention. This esp. helps to focus on the scores of 'specific' past exercises during prediction\n- Some papers have used attention with RNN to augment the scores and some have removed RNNs altogether and used attention in all its glory (transformer version - the encoder part related to self attention) \n- There are minor differences between the papers on how the keys, queries and values need to be computed otherwise these represent the standard transformer encoder model introduced in 'Attention is all u need'\n\nHope this intro helps.",
          "votes": 2
        },
        {
          "id": 1102936,
          "postDate": "2020-12-05T13:48:39.133Z",
          "content": "<p>I have made an attempt to analyze existing literature and shared the summary: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481</a></p>",
          "rawMarkdown": "I have made an attempt to analyze existing literature and shared the summary: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481",
          "votes": 2
        },
        {
          "id": 1103048,
          "postDate": "2020-12-05T15:38:49.977Z",
          "content": "<p><a href=\"https://www.kaggle.com/allovhk\" target=\"_blank\">@allovhk</a>   what is the difference between excercise embeddings and interaction embeddings <br>\ni get confused when  i check SAKT Model  in public kernel.  <br>\n    <code>\ntarget_id = q[1:] # Ignore first item 1 to 99\n     label = qa[1:] # Ignore first item 1 to 99\n      # x = np.zeros(self.max_seq-1, dtype=int)\n     x = q[:-1].copy() # 0 to 98\n     print(x.shape)\n     x += (qa[:-1] == 1) * self.n_skill # y = et + rt x E\n     print('after',x.shape,target_id.shape)\n</code><br>\nhere i see that shifted exercises attempted are passed as exercise embeddings and t-1  are passed as interaction embeddings.</p>",
          "rawMarkdown": "   @allovhk   what is the difference between excercise embeddings and interaction embeddings \ni get confused when  i check SAKT Model  in public kernel.  \n\n      ```\n  target_id = q[1:] # Ignore first item 1 to 99\n        label = qa[1:] # Ignore first item 1 to 99\n\n        # x = np.zeros(self.max_seq-1, dtype=int)\n        x = q[:-1].copy() # 0 to 98\n        print(x.shape)\n        x += (qa[:-1] == 1) * self.n_skill # y = et + rt x E\n        print('after',x.shape,target_id.shape)\n```\nhere i see that shifted exercises attempted are passed as exercise embeddings and t-1  are passed as interaction embeddings."
        },
        {
          "id": 1103328,
          "postDate": "2020-12-05T20:21:12.093Z",
          "content": "<p>Have anyone tried answer prediction (A, B, C, D) approach? Is it better than just correctness prediction? My intuition tells me that the question embeddings already have that information without making it explicit.</p>",
          "rawMarkdown": "Have anyone tried answer prediction (A, B, C, D) approach? Is it better than just correctness prediction? My intuition tells me that the question embeddings already have that information without making it explicit."
        }
      ]
    },
    {
      "id": 1063235,
      "postDate": "2020-10-28T16:44:51.017Z",
      "content": "<p>I have a doubt here, from the paper,</p>\n<blockquote>\n  <p>the encoder takes a sequence of exercise embeddings 𝐸𝑒 =[𝐸𝑒,𝐸𝑒,…,𝐸𝑒]  E𝑇<br>\n  as input and pass an output sequence 𝑂𝑒𝑛𝑐 =[𝑂𝑒𝑛𝑐,𝑂𝑒𝑛𝑐,…,𝑂𝑒𝑛𝑐] to the decoder. The decoder<br>\n  additionally takes a shifted response embedding sequence 𝑅𝑒 = [𝑆𝑒,𝑅1,𝑅2,…,𝑅t-1 ]as input (till T-1)<br>\n  first element is a start token embedding, to produce the final output sequence 𝑐ˆ = [𝑐1ˆ , 𝑐2ˆ , . . . , 𝑐Tˆ ].</p>\n</blockquote>\n<p>So that means that we have to keep shifting the input data as well and keep it at one timestamp ahead than what the decoder can access?</p>\n<p>e.g.<br>\nwhile training, we will keep a loop that will keep on shifting the content_ids let's say like this<br>\n4481,  9780,  <br>\n4481,  9780,  3632,  <br>\n4481,  9780,  3632,  6397,  <br>\n4481,  9780,  3632,  6397,  5327,</p>\n<p>and at decoder side we have to generate something like this? <br>\n4481,  <br>\n4481,  9780, <br>\n4481,  9780,  3632,<br>\n4481,  9780,  3632,  6397</p>\n<p>Do we have to create a mask for target or how do we generate something like this efficiently? But even if we generate a mask, it's not going to take in a batch, right?</p>",
      "rawMarkdown": "I have a doubt here, from the paper,\n>the encoder takes a sequence of exercise embeddings 𝐸𝑒 =[𝐸𝑒,𝐸𝑒,...,𝐸𝑒]  E𝑇\nas input and pass an output sequence 𝑂𝑒𝑛𝑐 =[𝑂𝑒𝑛𝑐,𝑂𝑒𝑛𝑐,...,𝑂𝑒𝑛𝑐] to the decoder. The decoder\nadditionally takes a shifted response embedding sequence 𝑅𝑒 = [𝑆𝑒,𝑅1,𝑅2,...,𝑅t-1 ]as input (till T-1)\nfirst element is a start token embedding, to produce the final output sequence 𝑐ˆ = [𝑐1ˆ , 𝑐2ˆ , . . . , 𝑐Tˆ ].\n\nSo that means that we have to keep shifting the input data as well and keep it at one timestamp ahead than what the decoder can access?\n\ne.g.\nwhile training, we will keep a loop that will keep on shifting the content_ids let's say like this\n4481,  9780,  \n4481,  9780,  3632,  \n4481,  9780,  3632,  6397,  \n4481,  9780,  3632,  6397,  5327,\n\nand at decoder side we have to generate something like this? \n4481,  \n4481,  9780, \n4481,  9780,  3632,\n4481,  9780,  3632,  6397\n\nDo we have to create a mask for target or how do we generate something like this efficiently? But even if we generate a mask, it's not going to take in a batch, right?",
      "votes": 2,
      "replies": [
        {
          "id": 1063261,
          "postDate": "2020-10-28T17:07:19.753Z",
          "content": "<p>We only need to predict if the current question is correct or not so you just can put both inputs into encoder and let it make all the work, you dont need the decoder because you dont need to reconstruct any sequence</p>",
          "rawMarkdown": "We only need to predict if the current question is correct or not so you just can put both inputs into encoder and let it make all the work, you dont need the decoder because you dont need to reconstruct any sequence",
          "votes": 2
        },
        {
          "id": 1063271,
          "postDate": "2020-10-28T17:21:52.660Z",
          "content": "<p>Hmm, sorry but i didn't understand what do you mean by \"both inputs\"? Plus on the paper, they do use decoder section as well with temporal feats and shifted embedding sequences in saint+ and only responses sequences in saint.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Faec31bc4c285fe94dd34543159389fb7%2FScreenshot%202020-10-28%20at%2010.51.12%20PM.png?generation=1603905698133797&amp;alt=media\" alt=\"model_arch_saint+\"></p>",
          "rawMarkdown": "Hmm, sorry but i didn't understand what do you mean by \"both inputs\"? Plus on the paper, they do use decoder section as well with temporal feats and shifted embedding sequences in saint+ and only responses sequences in saint.\n\n![model_arch_saint+](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Faec31bc4c285fe94dd34543159389fb7%2FScreenshot%202020-10-28%20at%2010.51.12%20PM.png?generation=1603905698133797&alt=media)"
        },
        {
          "id": 1063288,
          "postDate": "2020-10-28T17:43:26.553Z",
          "content": "<p>All inputs into encoder. On the figure the output is a sequence but we dont need that we only need current question ….</p>",
          "rawMarkdown": "All inputs into encoder. On the figure the output is a sequence but we dont need that we only need current question ....",
          "votes": 2
        },
        {
          "id": 1063296,
          "postDate": "2020-10-28T17:57:43.170Z",
          "content": "<p>Ahh thanks a lot! I see now what you meant!</p>",
          "rawMarkdown": "Ahh thanks a lot! I see now what you meant!"
        }
      ]
    },
    {
      "id": 1065266,
      "postDate": "2020-10-31T05:20:56.963Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1041887,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-10-08T00:43:14.067000",
      "content": "<p>after a night of paper reading and youtube surfacing for quick tutorials, I think a basic model looks like this<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7a22a6edb1adfd2a59f2d124bc773745%2FSelection_166.png?generation=1602117792236510&amp;alt=media\" alt=\"\"></p>\n<p>using this model, my first step is to make a naive model, i.e. prediction using average values </p>",
      "votes": 5,
      "replies": [
        {
          "id": 1042092,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-10-08T03:59:59.757000",
          "content": "<p>it is a seq-to-seq problem</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3d2d55d0babd7309ea6693bdfa4a703b%2FSelection_173.png?generation=1602129597919384&amp;alt=media\" alt=\"\"></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1042157,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2020-10-08T05:16:58.820000",
          "content": "<p>Probably we could do some bayesian step at the end</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1043455,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2020-10-09T02:16:29.300000",
          "content": "<p>Have you read the SAINT paper yet? </p>\n<ul>\n<li><a href=\"https://arxiv.org/pdf/2002.07033.pdf\" target=\"_blank\">https://arxiv.org/pdf/2002.07033.pdf</a></li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1043907,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-10-09T10:21:38.277000",
          "content": "<p><a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> <br>\nthanks for the paper. yes this is one of the paper in my list. if kagglers has paper that they wish to implement, they can suggest here.</p>\n<p>it may be easier if a group of kagglers develop some baseline methods together and we can cross check each other bug.</p>\n<p>i also looking for papers that actually do recommendation but can be converted to knowledge tracking.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1046054,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-10-11T10:08:23.377000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23/notebookcdb764afc6?scriptVersionId=44475365\" target=\"_blank\">https://www.kaggle.com/hengck23/notebookcdb764afc6?scriptVersionId=44475365</a><br>\naverage prediction baseline : LB 0.741</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1046191,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-10-11T12:25:33.277000",
          "content": "<p>I completed the static average approach. The next step is <strong>online average prediction</strong> which is another baseline before the time-series approach.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1046330,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-11T14:53:37.190000",
          "content": "<p>Just curious Heng, how do you plan to decide the window_size? Plus in test_data using API, we don't get the pupils data in one-shot, it comes in breaks at any random group_num.</p>\n<pre><code>def moving_average(series, n):\n    \"\"\"\n        Calculate average of last n observations\n    \"\"\"\n    return np.average(series[-n:])\n\nmoving_average(date_sales, 24) # prediction for the last observed day (past 24 hours)\n</code></pre>\n<p>Something like this. (<a href=\"https://www.kaggle.com/adityaecdrid/my-first-time-series-comp-added-prophet\" target=\"_blank\">ref</a>)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1046407,
          "author_name": "EnricRovira",
          "author_url": "",
          "post_date": "2020-10-11T16:15:41.953000",
          "content": "<p>But in inference you dont know the question sequence because with the submission API you have to preeict a batch before getting the next one, how would you adress that problem?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1046471,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2020-10-11T17:27:31.853000",
          "content": "<p>If you have the same user in train and test, you could just concatenate their new sequences to the old ones you’ve already seen in train as the API delivers it to you. If there are new users, you start a new sequence for them and predict with what you know of the questions and other metadata. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1046481,
          "author_name": "EnricRovira",
          "author_url": "",
          "post_date": "2020-10-11T17:45:35.757000",
          "content": "<p>Not enough memory to store all users seqs, besides the issue that is not ensured that test seqs came after train seqs</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1046490,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-10-11T17:53:08.140000",
          "content": "<p>i suppose these baselines will not be very high score, but it gives you an idea how the later rnn, transformer would perform.</p>\n<p>i am thinking of  just implementing a moving average counter like this:</p>\n<p><a href=\"https://www.kaggle.com/kneroma/riid-user-and-content-mean-predictor\" target=\"_blank\">https://www.kaggle.com/kneroma/riid-user-and-content-mean-predictor</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1050537,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2020-10-15T14:15:29.993000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> New to time series, could you explain what is relation between baseline models correlate with seq2seq rnn, transformer models ?.. </p>\n<p>Do you mean is moving average counter features give an idea of will these features work well in seq2seq models or not ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1054239,
          "author_name": "Tucker Arrants",
          "author_url": "",
          "post_date": "2020-10-19T19:23:13.263000",
          "content": "<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> In this context, a simple seq2seq model would be using a sequence of questions that a particular student saw (chronologically) as the input sequence and using the sequence of answers to those questions as the output sequence. This way, we train a model to 'translate' between questions and answers.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1054307,
          "author_name": "EnricRovira",
          "author_url": "",
          "post_date": "2020-10-19T20:08:40.097000",
          "content": "<p>On private test you only can solve one question at a time, so its partial seq to seq, your seq is the past questions &amp; answers and the output must be 0/1 answered correctly to that question</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1063002,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-10-28T12:18:25.333000",
          "content": "<p>here is the update on the online average baseline: LB 0.748<br>\n(the code is not efficient and takes 6hr to run)</p>\n<p><a href=\"https://www.kaggle.com/hengck23/notebookf100a7833c\" target=\"_blank\">https://www.kaggle.com/hengck23/notebookf100a7833c</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1066893,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-11-02T07:28:42.293000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  sorry it would be a very basic question as i  am quite new to this space so trying to build up the learning..<br>\nhow do we interpret users responses to Questions with time based  features in terms of embeddings. <br>\nExample see below excerpt from SAINT paper</p>\n<p>Excercise and response embeddings what info would these embeddings have..</p>\n<pre><code>In this paper, we propose a novel Transformer-based model for\nknowledge tracing, SAINT: Separated Self-AttentIve Neural\nKnowledge Tracing. SAINT has an encoder-decoder structure where the exercise and response embedding sequences\nseparately enter, respectively, the encoder and the decoder\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1097750,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-01T08:10:19.530000",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> I am not aware which exact paper you are referring to. But in general, I reviewed a bunch of papers from there and most use Attention in some way or the other. I don't think there would be a big difference between the various models (based on the 2-3 that I saw), nonetheless let me dig further and revert.</p>\n<p>To answer your questions:</p>\n<ul>\n<li>Embeddings are used to represent each question and QnA pair. They take a bunch of params like the exerciseID, category, position, elapsed time, timestamp etc and create the embeddings during training </li>\n<li>This competition becomes a form of supervised sequence learning task - given student’s past exercise interactions X = (x1, x2, . . . , xt), predict her next interaction xt+1. x = (e,r), where e is the exercise and r is the correctness of the student’s answer. This data is available at every past timestep 't'</li>\n<li>Because knowledge learning is a sequential process (for e.g. you learn equations of motion first..then probably Newton's laws, then gravity etc..where things flow in a sequence..likewise peformance on a particular exercise is dependent on her performance on the past exercises esp. ones related to that exercise) so RNN models can be used to nicely encode this representation..</li>\n<li>With RNN comes the problem of long sequence representation</li>\n<li>This is solved by attention. This esp. helps to focus on the scores of 'specific' past exercises during prediction</li>\n<li>Some papers have used attention with RNN to augment the scores and some have removed RNNs altogether and used attention in all its glory (transformer version - the encoder part related to self attention) </li>\n<li>There are minor differences between the papers on how the keys, queries and values need to be computed otherwise these represent the standard transformer encoder model introduced in 'Attention is all u need'</li>\n</ul>\n<p>Hope this intro helps.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1102936,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-05T13:48:39.133000",
          "content": "<p>I have made an attempt to analyze existing literature and shared the summary: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1103048,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-05T15:38:49.977000",
          "content": "<p><a href=\"https://www.kaggle.com/allovhk\" target=\"_blank\">@allovhk</a>   what is the difference between excercise embeddings and interaction embeddings <br>\ni get confused when  i check SAKT Model  in public kernel.  <br>\n    <code>\ntarget_id = q[1:] # Ignore first item 1 to 99\n     label = qa[1:] # Ignore first item 1 to 99\n      # x = np.zeros(self.max_seq-1, dtype=int)\n     x = q[:-1].copy() # 0 to 98\n     print(x.shape)\n     x += (qa[:-1] == 1) * self.n_skill # y = et + rt x E\n     print('after',x.shape,target_id.shape)\n</code><br>\nhere i see that shifted exercises attempted are passed as exercise embeddings and t-1  are passed as interaction embeddings.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1103328,
          "author_name": "Claudio Verdú Ruiz",
          "author_url": "",
          "post_date": "2020-12-05T20:21:12.093000",
          "content": "<p>Have anyone tried answer prediction (A, B, C, D) approach? Is it better than just correctness prediction? My intuition tells me that the question embeddings already have that information without making it explicit.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1063235,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-10-28T16:44:51.017000",
      "content": "<p>I have a doubt here, from the paper,</p>\n<blockquote>\n  <p>the encoder takes a sequence of exercise embeddings 𝐸𝑒 =[𝐸𝑒,𝐸𝑒,…,𝐸𝑒]  E𝑇<br>\n  as input and pass an output sequence 𝑂𝑒𝑛𝑐 =[𝑂𝑒𝑛𝑐,𝑂𝑒𝑛𝑐,…,𝑂𝑒𝑛𝑐] to the decoder. The decoder<br>\n  additionally takes a shifted response embedding sequence 𝑅𝑒 = [𝑆𝑒,𝑅1,𝑅2,…,𝑅t-1 ]as input (till T-1)<br>\n  first element is a start token embedding, to produce the final output sequence 𝑐ˆ = [𝑐1ˆ , 𝑐2ˆ , . . . , 𝑐Tˆ ].</p>\n</blockquote>\n<p>So that means that we have to keep shifting the input data as well and keep it at one timestamp ahead than what the decoder can access?</p>\n<p>e.g.<br>\nwhile training, we will keep a loop that will keep on shifting the content_ids let's say like this<br>\n4481,  9780,  <br>\n4481,  9780,  3632,  <br>\n4481,  9780,  3632,  6397,  <br>\n4481,  9780,  3632,  6397,  5327,</p>\n<p>and at decoder side we have to generate something like this? <br>\n4481,  <br>\n4481,  9780, <br>\n4481,  9780,  3632,<br>\n4481,  9780,  3632,  6397</p>\n<p>Do we have to create a mask for target or how do we generate something like this efficiently? But even if we generate a mask, it's not going to take in a batch, right?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1063261,
          "author_name": "EnricRovira",
          "author_url": "",
          "post_date": "2020-10-28T17:07:19.753000",
          "content": "<p>We only need to predict if the current question is correct or not so you just can put both inputs into encoder and let it make all the work, you dont need the decoder because you dont need to reconstruct any sequence</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1063271,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-28T17:21:52.660000",
          "content": "<p>Hmm, sorry but i didn't understand what do you mean by \"both inputs\"? Plus on the paper, they do use decoder section as well with temporal feats and shifted embedding sequences in saint+ and only responses sequences in saint.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F835774%2Faec31bc4c285fe94dd34543159389fb7%2FScreenshot%202020-10-28%20at%2010.51.12%20PM.png?generation=1603905698133797&amp;alt=media\" alt=\"model_arch_saint+\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1063288,
          "author_name": "EnricRovira",
          "author_url": "",
          "post_date": "2020-10-28T17:43:26.553000",
          "content": "<p>All inputs into encoder. On the figure the output is a sequence but we dont need that we only need current question ….</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1063296,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-28T17:57:43.170000",
          "content": "<p>Ahh thanks a lot! I see now what you meant!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1065266,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-31T05:20:56.963000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1040971": "this post will be updated as i go along. it is like a mini-diary of my work.\n\n... to be updated ....\n\nfor a start: https://paperswithcode.com/task/knowledge-tracing\n\ndefinition: \nKnowledge tracing—where a machine models the knowledge of a student as they interact with coursework, see [1]\n\n[1] 'Deep Knowledge Tracing'- Chris Piech, nips 2015",
    "1041887": "after a night of paper reading and youtube surfacing for quick tutorials, I think a basic model looks like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7a22a6edb1adfd2a59f2d124bc773745%2FSelection_166.png?generation=1602117792236510&alt=media)\n\nusing this model, my first step is to make a naive model, i.e. prediction using average values ",
    "1063235": "I have a doubt here, from the paper,\n>the encoder takes a sequence of exercise embeddings 𝐸𝑒 =[𝐸𝑒,𝐸𝑒,...,𝐸𝑒]  E𝑇\nas input and pass an output sequence 𝑂𝑒𝑛𝑐 =[𝑂𝑒𝑛𝑐,𝑂𝑒𝑛𝑐,...,𝑂𝑒𝑛𝑐] to the decoder. The decoder\nadditionally takes a shifted response embedding sequence 𝑅𝑒 = [𝑆𝑒,𝑅1,𝑅2,...,𝑅t-1 ]as input (till T-1)\nfirst element is a start token embedding, to produce the final output sequence 𝑐ˆ = [𝑐1ˆ , 𝑐2ˆ , . . . , 𝑐Tˆ ].\n\nSo that means that we have to keep shifting the input data as well and keep it at one timestamp ahead than what the decoder can access?\n\ne.g.\nwhile training, we will keep a loop that will keep on shifting the content_ids let's say like this\n4481,  9780,  \n4481,  9780,  3632,  \n4481,  9780,  3632,  6397,  \n4481,  9780,  3632,  6397,  5327,\n\nand at decoder side we have to generate something like this? \n4481,  \n4481,  9780, \n4481,  9780,  3632,\n4481,  9780,  3632,  6397\n\nDo we have to create a mask for target or how do we generate something like this efficiently? But even if we generate a mask, it's not going to take in a batch, right?",
    "1065266": ""
  }
}