{
  "id": 201481,
  "title": "From Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time",
  "url": "/competitions/riiid-test-answer-prediction/discussion/201481",
  "author_name": "Allohvk",
  "post_date": "2020-12-05T08:19:38.323000",
  "votes": 99,
  "comment_count": 13,
  "views": 0,
  "content": "<ul>\n<li>A general intro to the domain</li>\n<li>Quick summary of key models in plain English</li>\n<li>My interpretation of the salient approaches in top 5 papers on this domain, few code links &amp; a good video I found</li>\n<li>Few concluding remarks</li>\n</ul>\n<p>Target Audience - Newcomers like me who have no clue on the domain &amp; want a 15-min 'starter-kit' before going on to tackle the competition</p>\n<p><br><br></p>\n<h2>GENERAL INTRO TO THE DOMAIN</h2>\n<p>Knowledge Tracing (KT) is the art of modeling the knowledge of a student as she interacts with her course work. It is a core part of Intelligent Tutoring Systems (ITS) which aim to provide a personalized learning experience to students. The data for Knowledge Tracing (KT) comes from 'interactions' which happen when the students perform 'exercise's, and this allows the prediction of their performance on future exercises. Interactions contains the student 'response's. These are binary values indicating whether the student answered the exercise correctly or not and the student's knowledge is updated accordingly. In certain models, exercises are tagged to 'Knowledge-concepts' or 'Skill-tags' by domain experts and these additional tags help the models perform better. </p>\n<p>So KT is all about:<br>\nStudent interactions (exams) -&gt; Knowledge model(skill proficiency) --&gt; Prediction (on new exams)</p>\n<p>In general, most models right from Bayesian ones in early nineties have performed reasonably well and helped pedagogy evolve. The explosion of data in recent times naturally benefited the deep net models and though traditional models made a valiant comeback in between, most of the papers in the past 2 years have all been about deep models. The Covid situation has pushed even more students towards online education which has resulted in further data explosion and 2021 could be the year of mining all this data. Like how 2018 belonged to NLP, 2021 could well belong to Knowledge Tracing (KT) and we can definitely expect more such competitions. Co-incidentally, almost all SOTA models in the past 2 years use at least some part of the 'Attention is all you need' architecture by Vaswani et al. which changed the field of NLP in 2018.</p>\n<p><br><br></p>\n<h2>QUICK SUMMARY OF KEY MODELS IN PLAIN ENGLISH</h2>\n<p>Bayesian Knowledge Tracing (BKT) models knowledge state as a set of binaries - each representing the understanding or non-understanding of a single KC (Knowledge Concept). So it is essentially a set of tuples. Hidden Markov Models are used to update these, as the student answers exercises. The model assumes the learner never forgets and every new question has a fixed probability of helping the student understand the KC. BKT models were proposed as early as 1972 but refined in 95 by Corbett &amp; Anderson. It held sway for a good 2 decades with minor changes along the way (incorporating the possibility of guess-work, silly-mistakes etc). It was challenged by DL methods in 2015 and made a brilliant comeback by incorporating features like forgetting etc and defeated the benchmarks set by DL models. Nonetheless with the explosion of data and the proliferation of 'attention' mechanism, this model seems to have receded to the background in the past 2 years. A pity.</p>\n<p>Unlike BKT which assumed that the knowledge state at any time-step depends only on the previous step state, RNNs were able to capture the complexity and diversity of data over time effectively and rode the next wave. Also, with deep nets - domain expertise no longer became a mandatory requirement for modeling (though it is still a huge added advantage). So any Tom, Dick or Allohvk could now use a large dataset, a GPU and a few lines of Keras code and make an attempt to model student knowledge. Deep knowledge tracing (DKT) systems use an LSTM which at each time-step, takes the 'interaction' tuple as input. The tuple is simply encoded using a one-hot vector. The output is a vector of length equal to the number of questions, each element representing the predicted probability that the student would correctly answer the same. When it came out, the results of DKT were significantly better than BKT, though it was later found to that there were some data preprocessing issues (an inadvertent mistake) which resulted in higher scores which tempered some of the jubilation.</p>\n<p>Nonetheless, this success set the tone for coming waves. DKT was followed by Dynamic Key-Value Memory Networks (DKVMN) to improve DKT’s structure. These are also LSTM based but recognize the fact that the hidden state in the LSTM network has limited power and augment this with memory matrices to (better) store information. DKVMN maps higher level skill-tags to exercises and can observe which skill tags are needed for which exercise, find out which areas the student needs to be trained on to improve her/his performance and get better scores or tell which skill-tag was lacking in the student's output due to which the scores were lower. They need to be used in combination with domain specific annotations to fully exploit the power. The ability to predict which specific skill tag a student needs to be trained upon to improve the upcoming (exam/competition) scores seems to be a dream for students/parents/teachers and wannabe Kaggle GMs alike :) </p>\n<p>Most original implementations of KT used one-hot encoding mechanisms to encode interactions. With the explosion of data, some moved onto word embeddings while others insisted that embeddings are an overkill as in this domain, the distance between different samples have no major correlation unlike in NLP. They use OHE but reduce the dimensionality leveraging certain techniques. I believe this is a minority voice now as most 2018-2020 papers that I saw were all about embeddings. Secondly, in many papers exercises have been grouped into KCs and all exercises in one KC are treated as the same exercise. Basically all exercises covering the same concept are treated as a single exercise. So Q = C and Qt = Ct. This is done to keep things simple.</p>\n<p>The last couple of years have followed a predictable path with Attention being introduced followed by self-attention and transformer-like architectures. DKVMN took a back seat to transformer-like architectures which is again a pity as I feel DKVMN's scope had not fully been exploited. </p>\n<p>Exercise-Enhanced Recurrent Neural Network (EERNN) use a Bi-LSTM network to obtain the text embedding of each question, and concatenates the embedding with that of the corresponding student-interaction tuple. The concatenated embeddings now contain the text feature of the questions and are fed into a LSTM network. EERNN also uses attention mechanism to aggregate all the hidden states of the past LSTM units. </p>\n<p>The self-attentive knowledge tracing (SAKT) method introduced self-attention to KT and is inspired by various pieces of the transformer architecture from Vaswani's paper. The query is an exercise embedding vector, and key and value are interaction embedding vectors. The code is at: <a href=\"https://github.com/shalini1194/SAKT/tree/master/2019-EDM\" target=\"_blank\">https://github.com/shalini1194/SAKT/tree/master/2019-EDM</a>. </p>\n<p>To put it very briefly, SAKT identifies relevant Knowledge Concepts (KCs) from the past interactions and then predicts student’s performance based on her performance on those KCs by 'attend'ing to them. Note that they use the exercises themselves as KCs and the semantics can get a bit confusing due to this. There is an embedding layer for the exercises and an embedding layer for the interactions. When predicting the response to an exercise, the model uses the exercise embedding as a query and the interaction embeddings as keys and values and builds a set of attention weights which is then used for predictions. Note that at every time stamp t+1, the current exercise et+1's embedding is used to query the all past interactions (xt) which act as both the key and value. The other way to look at it would be to consider the model with inputs x1,x2,…,xt−1 and shift the exercise sequence one position ahead, e2,e3,…, et and the output being the correctness of the response r2,r3,…,rt. </p>\n<p>This is a simple enough model and the code link shared above can be taken as a starting point to the competition.</p>\n<p>When it came out, SAKT's results were exceptionally good. However, another paper did point out a small un-intentional bug in SAKT (see <a href=\"https://arxiv.org/pdf/2008.01169.pdf\" target=\"_blank\">https://arxiv.org/pdf/2008.01169.pdf</a>) which may have inflated the AUC a bit. Nonetheless, this paper was the first to usher in the age of transformers in KT universe and deserves praise. They use 'self-attention' and do away with RNNs altogether. For encoding temporal info, they have 'position-encoding' as in the transformer paper to maintain sequence knowledge, residual connections to propagate the lower layer features to the higher layers, the feed forward layer to bring in an element of non-linearity and a normalizing layer to help stabilize and accelerate…all these were inspired by the original transformer paper.  The most important was the use of multi-head attention. Multiple heads help in capturing the attention weights in different subspaces. It was bit hard for me to grasp this intuitively - why MHA works here? The whole idea of MHA in the original 'attention' paper was to capture different types of relationship between words. Possibly here, it may be capturing relations across different intervals of time? Anyway, as per the SAKT paper, the authors clearly demonstrate that performance drops when single attention heads were used. </p>\n<p>Though this is not a discussion on 'attention', it may be pertinent to define it in 1 line because nearly all the SOTA models use Attention in one form or the other. When presented with a question for which we want a prediction, the 'attention' model identifies relevant past interactions – it 'attends' to past interactions – giving higher weightage to interactions that matter while drowning out interactions that don't matter and then predicts future performance from these interactions. Which interactions to drown out and which to amplify is determined by a simple feed-forward network. The concept of 'self-attention' takes it one step further where every step in the sequence attends to every other step in the same sequence and attempts to build relationships within the sequence (before doing the inter-sequence Attention as mentioned above). So self-attention could be used to better embed interactions and exercises individually before a final attention layer where they connect to each other. Vaswani et al. also decided to use separate embeddings in the form of key, query and values instead of using just one embedding as in the original Bahdanau Attention paper in 2014. </p>\n<p>For a more detailed explanation of 'Attention' and an introduction to the Key, Query, Value concepts, please refer to: Craft your own Attention layer in 6 lines- The essence of Attention across all its intoxicating flavors ( <a href=\"https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e\" target=\"_blank\">https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e</a> ). The transformer explained in just 1 line is - a bunch of components like self-attention coupled with a few neat hacks like multi-head attention (repeat the Key, Query, Value  generation multiple times with different initial weights so as to more richly capture all relationships), position encoding (we are doing away with RNN so retain some way of encoding positions), residuals (to propagate the lower layer features to the higher layers), the feed forward layer(to bring in non-linearity) and a normalizing layers(stabilize and speeden). These blocks are there at both the encoder and decoder end and of course there is a connection between them. Almost all SOTA papers use some or all components of this architecture. All models require a masking mechanism that prevents the current position from 'attend'ing to subsequent positions. </p>\n<p><br><br></p>\n<h2>MY INTERPRETATION OF THE SALIENT FEATURES OF TOP 5 SOTA MODELS IN THIS DOMAIN</h2>\n<p>Disclaimer - I evaluated the top dozen odd papers with public access which were returned by a Google search. I am sure there are many other good papers as well</p>\n<p>Most of the below papers have come out in the last year or so around the same time. This has 2 interesting ramifications. They don't benchmark against one another. Also some of the claimed 'innovations' made have been replicated in another paper. Of course this is perfectly explainable since papers take a long time to get peer-reviewed and get published so there could be multiple groups working on same 'innovations' at a time.</p>\n<p>We will discuss them very briefly in NO PARTICULAR ORDER. I will only point out the salient items. I will continue to use the terminology - interactions, exercises, responses and Knowledge Concepts (KC) across papers even though the authors may have used different terminology. In some places, I may have inadvertently used the term 'question' instead of 'exercise'.</p>\n<p><br></p>\n<h2>Deep Knowledge Tracing with Transformers: </h2>\n<p><a href=\"http://link-springer-com-443.webvpn.fjmu.edu.cn/chapter/10.1007%2F978-3-030-52240-7_46\" target=\"_blank\">http://link-springer-com-443.webvpn.fjmu.edu.cn/chapter/10.1007%2F978-3-030-52240-7_46</a>. <br>\nThis is a 2020 paper. They introduce 2 specific innovations. Instead of directly encoding the questions, they define a separate set of weights which relate Knowledge Concepts(KC) to questions and then apply that to transform and come up with final representation of the exercises. Secondly they introduce a time decay function - they allow the attention weights to decay - this basically brings in the element of 'forgetting'. </p>\n<p>The model works like this - Take the interaction and pass it to an interaction-embedding layer. The output is the interaction-embedding (big surprise :) !) which is passed to the transformer block. We will come to the transformer block layer but first let us see what this interaction embedding layer is doing. Here, every KC is represented by a vector. A weight matrix is learnt during raining and this that associates the exercise in the interaction to every KC available. A weighted sum of this association(weight*KC vector) is taken and will represent the interaction. 1 point came to my mind - The interaction consists of the student response in addition to the exercise. Would it not make more sense to just see how the 'exercise' relates to KC vectors instead of seeing how the 'interaction' relates to the KC vectors?</p>\n<p>Let us now check out the 'transformer block'. This is not really a complete transformer as such but includes a few key components of the original transformer architecture - a self-attention layer, a FeedForward to bring in non-linearity and normalization layers. The authors talk of residual connections but I couldn't see this in their image. Anyway this transformer block takes the 'interaction embedding' and creates a Key, query, value for each. Based on that it calculates the attention weights and then the 'interaction-context' for each time-step which is then used for prediction. </p>\n<p>However there seem to be few missing info here - It looks like a simplified arch is shown and not the final one. For one thing, I feel almost certain that a multi-head attention would have been used…possibly positional encodings as well. The actual conf paper hyperlink does not work. I also couldn't find the source code for this paper. They also partially explain the model. In the block diagram provided, I could see a similar interaction-embedding layer and a transformer block which takes only the exercises as input (aah.. to some extent, my doubt above is cleared…) So looking at this image, I am guessing that the same process of embedding is carried out for the 'exercise's also (they really should title it exercise-embedding layer and not interaction-embedding). This exercise-embedding is now fed to another transformer block ALONG WITH THE 'interaction-context' to get the final context - one for each time-step. </p>\n<p>Finally a dense layer is used for predictions. </p>\n<p>So to summarize, salient features would be:</p>\n<ul>\n<li>A learnable weight matrix to map exercises to KCs - now we start out with better embeddings as we have used 'Attention' while preparing the initial embeddings itself (note this does not seem to be self-attention)</li>\n<li>A time-decay function for attention weights to model forgetting</li>\n<li>An encoder-decoder type of architecture where interaction-embeddings of t-1 steps and exercise-embeddings of t steps interact via attention mechanism (using Q, K, V concepts from the transformer paper) to generate a final context which is then used for predictions via a Dense layer</li>\n</ul>\n<p><br></p>\n<h2>Let us now look at SAINT - Separated Self-AttentIve Neural Knowledge Tracing: </h2>\n<p><a href=\"https://arxiv.org/pdf/2002.07033.pdf\" target=\"_blank\">https://arxiv.org/pdf/2002.07033.pdf</a>. This is Gold standard and probably came out before the above paper as they clearly mention that they are the first to introduce encoder-decoder transformer-like structure. They also introduce innovations in discovering the right key, query, values. </p>\n<p>The encoder takes the sequence of exercise embeddings as queries, keys and values, and produces output through a repeated self-attention mechanism. The decoder takes response embeddings (shifted right and prefixed with a start-token) as queries, keys and values then alternately applies self-attention and attention layers to the encoder output, thereby somewhat mimicking an actual transformer. This clean separation of input allows them to stack attention layers multiple times. In a nutshell, SAINT encoder takes the exercise embeddings and feeds the processed output to the decoder. The decoder meanwhile is ready with a first self-attention layer whose input is the response embeddings (along with some meta info) and which attends to itself (key, query, value all the same) and prepares a response embedding. Now the next attention block kicks in where the 'query' is the output of the first decoder layer. The keys and values are the encoder output. Finally, a prediction layer, consisting of a linear transformation layer followed by a sigmoid operation, is applied to the output of last layer so that the final decoder output is a series of probability values.</p>\n<p>They have additional meta data associated with questions and responses - things like response time, type of exercise etc - all of which is part of the respective embedding. Apart from this, they use multi-head attention, FF and residuals just as in the original transformer architecture.</p>\n<p>Few salient points:</p>\n<ul>\n<li>SOTA model - Has set the benchmark and seems to be quite thorough</li>\n<li>Deeper than other models - Stacking self-attention blocks improves performance unlike in other models like SAKT where perf deteriorates</li>\n<li>First to introduce an encoder-decode type of arch. Encoder feed is the exercises. Decoder feeds are the responses and the encoder output</li>\n<li>Key point - They do not generate interaction-embeddings but instead generate separate exercise-embeddings and response-embeddings. These are the encoder/decoder respectively. They do generate interaction-embeddings in their paper just to show in comparison studies that it does not add much value</li>\n<li>Use of meta tags along with question and responses</li>\n<li>Neat Experiments with masking and different key, query, values combinations before arriving at their final model</li>\n<li>Important: It must be noted how they incorporate temporal features only on the decoder end and show that this decision achieves the best AUC compared to incorporating them in the encoder input OR both the encoder/decoder</li>\n<li>I didn't see an explicit positional encoding. May not make a big difference</li>\n<li>I didn't see any 'forgetting' feature - This is interesting. Some sort of time-decay could have been applied to the weights. Ideally this should improve the scores and I am fairly sure this would have been experimented with before the paper was published. So there could be a good reason to give this a skip. Could the model be using some of the existing features and produce attention weights that have factored in some sort of time-decay already? Or is it just that 'forgetting' is at best a tricky area to tackle and in reality is far more complicated that what is typically modeled? </li>\n</ul>\n<p>The best performing model of SAINT has 4 layers and a latent space dimension of 512. I couldn't find associated source code. </p>\n<p><br></p>\n<h2>Deep Knowledge Tracing with Convolutions: </h2>\n<p><a href=\"https://arxiv.org/pdf/2008.01169.pdf\" target=\"_blank\">https://arxiv.org/pdf/2008.01169.pdf</a>.<br>\nThis is called the Convolutional Knowledge Tracing (CKT) model and was the first to introduce a convolutional network in the knowledge tracing. In addition to modeling the long-term effect of the entire question-answer sequence, CKT also strengthens the short-term effect of recent questions using 3D convolutions, thereby more effectively modeling the forgetting curve in the learning process. This is important because many studies have shown that in as little as 20 minutes after learning, only 58% of memory is retained, and this number drops to 25.4% in 6 days. They reshape the interaction-embeddings into a matrix and then use convolutions to extract important spatial patterns.</p>\n<p>Note that they use RNNs. The LSTM network outputs a hidden representation at every step t, which represents the hidden state of the processed interactions in the sequence so far. They now reshape each of the previous k embeddings (k is the window - so it takes embeddings from step t − k + 1 to t) into a matrix and then stack the k matrices to form a 3D tensor. They now use 3D convolutional network to learn from the tensor, which eventually outputs a hidden representation with length equal to that of the hidden representation generated by the LSTM network. They 'fuse' the two hidden representations and this fused representation is transformed to predict the student’s response.</p>\n<p>This is a refreshingly different approach. This model seems slightly older than the others nonetheless I liked its idea of using convolutions to capture spatial patterns. The way they fuse the 2 representations is also interesting. How does this way of adding temporal information compare to self-attention? Can self-attention be augmented further by convolutions? So instead of fusing the LSTM state with the convolution representation does it make sense to fuse it with the embeddings generated from self-attention (and of course do away with LSTMs altogether)? I guess it may not give better results because self-attention combined with MHA (and the whole lot of paraphernalia from the transformer - residuals, FF, norm layers, positional encoding etc) do completely account for temporal information as well as bring out deep representations…but is there any other way convolutions can benefit KT? This is a 100K dollar question</p>\n<p><br></p>\n<h2>Relation-aware self-attention model for Knowledge Tracing</h2>\n<p>One of the author of SAKT has come out with a new paper a couple of months back and introduces - RKT - Relation-aware self-attention model for Knowledge Tracing. This strengthens relationship between exercises using their textual content AS WELL AS student interaction data AND the forget behavior information (by modeling an exponentially decaying kernel function). Initially I did not understand why they needed student interaction data to capture relationship between exercises. I guess this is needed because apart from the textual similarity, many exercises need common skills - if two exercises are related, then the performance on one affects the other. Putting this in the reverse way, performance in interactions can identify relationships between exercises. We can say that if two exercises are textually very close to each other, then it makes sense to check student performance in interactions for these exercises and use it to further strengthen the relationship between those 2 exercises. So two exercise texts which are very similar and student interactions yield the same response for both (say both are correct) could be given higher weightage over two exercises that have similar text but student interactions show different responses (one correct, one wrong). In that sense, student interactions can be taken into account when building relationships between exercises. They use Phi coefficients (a measure of association for two binary variables - in this case the student responses) to build the exercise relationship and then add the textual similarity (provided it is within a certain distance) to derive the relation coefficients else make it zero. It seemed to be a neat innovative approach to me. Initially when reading the paper you may feel they seem to be prioritizing interaction performance over textual similarity but that is not the case. Relationships are only built if the questions are textually near to each other else the weight is zero!</p>\n<p>This exercise-relation information is represented by a set of relation coefficients. As before, they use the self-attention mechanism to learn the attention weights corresponding to the previous interaction for predicting whether a student will provide correct answer to the next exercise. These weights are now augmented by the relation coefficients to provide even better predictions. </p>\n<p>Another interesting point to note would be - They use the textual content of exercises to create simple embeddings for each word (not BERT). They use word2vec and use TEX tokens to transform equations in Maths exams. The exercise embedding is then a weighted combination of the embedding of all words present in the text of the exercise leveraging Smooth Inverse Frequency (SIF). SIF downgrades unimportant words and keeps relevant info only. The distance between exercises is then computed. I think they use cosine similarity.</p>\n<p>Lastly let us look at the ablations in their study. It is mighty interesting - particularly the one related to exercise relationship building. They compared model performances using a) only textual similarities b) only student interaction performance and c) using only KC tags (so all exercises in a same KC tag are related (weight=1) else they are not (weight=0) with d - the combination of a and b. The observations are as follows - 'c' performs the worst! This is expected as exercises can span KCs. 'b' seems to perform better than 'a' while 'd' is the best. Let us analyze 'b' again. The authors feel that even if textual content of two exercises are not similar, the association of knowledge involved in solving the two exercises could be high. Maybe I am totally wrong, but my understanding is this - With 'b' we are just clustering exercises into groups based on complexity. On one spectrum are exercises which are consistently solved correctly by most students (or is it just 1 student?), the other end of the spectrum consists of exercises that nobody could solve and the rest fall somewhere in between. While this may help in prediction, will this help in pedagogy? However a combination of 'a' and 'b' definitely makes sense and this is what the authors have chosen to go with and this is what gets the best results in the study. All in all, a fascinating paper.</p>\n<p><br></p>\n<h2>Context-Aware Attentive Knowledge Tracing</h2>\n<p><a href=\"https://arxiv.org/pdf/2007.12324.pdf\" target=\"_blank\">https://arxiv.org/pdf/2007.12324.pdf</a>.<br>\nAttentive knowledge tracing (AKT) brings in the following innovations: </p>\n<ul>\n<li>Instead of using raw question and response embeddings, they use context-aware representations of past questions and responses by taking a learner’s practice history into account using a modified version of attention. These modified representations reflect each learner’s actual comprehension of the question and the knowledge they actually acquire, given their personal response history. </li>\n<li>They modified the scaled dot product attention and have the attention weights decay exponentially based on the context-aware relative distance measure</li>\n<li>Use the Rasch model to bring out similarity &amp; differences in exercises within the same KC. Thus questions labeled as covering the same concept should not be treated as the same as they could have important individual differences</li>\n</ul>\n<p>There are four components: two self-attentive encoders, one for exercises and one for knowledge acquisition (I believe in plain words this is the interaction-embedding layer), a single attention-based knowledge retriever, and a feed-forward response prediction model.</p>\n<ul>\n<li>The two self-attentive encoders learn context-aware representations of the exercises and interactions. The first is the exercise encoder, which produces modified, contextualized representations of each exercise, given the sequence of exercises the learner has previously practiced on. The context-aware embedding of each exercise depends on both itself and the past exercises. The second is the knowledge encoder, which produces modified, contextualized representations of the knowledge the learner acquired while responding to past questions. This is the interaction-embedding </li>\n<li>The knowledge retriever, which retrieves knowledge acquired in the past that is relevant to the current question using an attention mechanism</li>\n<li>The response prediction model predicts the learner’s response to the current question using the retrieved knowledge</li>\n</ul>\n<p>Both encoders employ the self-attention mechanism. The knowledge retriever, on the other hand, uses the embedding of the current exercise as query, the keys are all the past exercises embeddings and values would be the interaction-embeddings. Notice how keys and values are different unlike in normal attention scenarios. For e.g. SAKT uses exercise embeddings as queries and interaction-embeddings for key as well as values. In AKT, exercise embeddings are used as queries and keys and values are the interaction embeddings. This is an interesting change and they claim that this method is more effective.</p>\n<p>The other interesting bit is the modifications made to the scaled dot product attention that is used in all other models. They say that scaled dot-product attention mechanism is not going to model memories decay effectively nor strengthen recent interactions. They add a multiplicative exponential decay term to the attention scores and a learnable decay rate parameter. So now the attention weights for the current question on a past question depends not only on the similarity between the corresponding query and key, but also on the relative number of time steps between them.</p>\n<p>The source code is available at: <a href=\"https://github.com/arghosh/AKT\" target=\"_blank\">https://github.com/arghosh/AKT</a></p>\n<p><br><br>\nAdditional notes: I later discovered a successor SAINT+ released a couple of months back which further pushes the benchmark by incorporating two additional temporal feature embeddings into the response embeddings: elapsed time, the time taken for a student to answer, and lag time, the time interval between adjacent learning activities. The lag time is used to build in 'forget'fulness (yay!). They show that the elapsed time is better represented as a continuous embedding instead of categorical which is somewhat counter-intuitive. Ideally 3-4 categories based on Z-score should have sufficed but of course this is just my opinion. Saint+ performs slightly better than Saint and is the new benchmark!</p>\n<p>Additional links:</p>\n<ul>\n<li><a href=\"https://github.com/thosgt/edm_main_algorithms:\" target=\"_blank\">https://github.com/thosgt/edm_main_algorithms:</a> Independent repository containing some models like DKT, SAKT etc</li>\n<li>Similar to above - contains a bunch of models - <a href=\"https://github.com/seewoo5/KT\" target=\"_blank\">https://github.com/seewoo5/KT</a></li>\n<li>A good introductory video: <a href=\"https://www.youtube.com/watch?v=CzRmRZNpB1Y\" target=\"_blank\">https://www.youtube.com/watch?v=CzRmRZNpB1Y</a></li>\n</ul>\n<p><br><br></p>\n<h2>GENERAL OBSERVATIONS:</h2>\n<p>My domain knowledge has accrued over a grand total of (the past) 36 hours… but perhaps viewing this domain as a rank outsider has its own advantages. My random thoughts (note - this is about the domain not the competition):</p>\n<ol>\n<li>1. While the current focus is on personalized recommendations, could there be plenty to exploit by generalization? Basically clustering of exercises or clustering of KCs, student profiles etc? The results could further aid personalization. But they have other implications. We could have signals that can be leveraged by tutors instead of students. If a set of questions are consistently answererd incorrectly, this could be a signal that the KCs related to those may need to be augmented with better course contents? </li>\n<li>While self-attention and transformer-like architectures are all the rage currently, can we leverage recommendation systems? Basically by using k-nearest on student profiles, we can predict how a student(/s) can perform by just looking at a similar profile from a past student who has competed the full course. Then use elements of recommendation systems to design the best course-path for this student?</li>\n<li>Can we bring in elements of reinforcement learning?</li>\n<li>Can we model the art of forgetting better? We forget a lot in the first 24-48 hours. After that the decay slows down. Beyond a certain point, knowledge does not decay! Also teh decay rate could vary across KCs. Maybe we could 'learn' the decay rate of cluster of KCs?</li>\n<li>Can we generate relevant KCs from exercises using DL methods? Some kind of topic modeling? Could this help? In general, question tagging can be further explored. This has nothing to do with individual student responses and is more generic</li>\n<li>Relevance of other cues could matter much more. For e.g. engagement of the student during the course is a much better predictor..tracking of eye movements etc?</li>\n<li>Performance is so much dependent on external factors. Brilliant students who have done well in the past  may start deteriorating if they have personal issues (death in the family, divorce, drugs, bad company, abuse etc)..The opposite too could happen. The shift does not happen overnight. Can we capture patterns to identify these shifts before they fully materialize? Interventions like counseling could help immensely if done at the right time</li>\n<li>How do we spot the super-performers? Most of the work is aimed at improving the scores of the majority. Super-performers may need a different kind of mentoring and customization</li>\n<li>How do we identify hidden talents - This is different from super-performers who are all-rounders. Hidden talents may be exceptionally good at only a certain type of exercises or subjects but average otherwise and typically do not stand out in any list. However they have the potential to become true geniuses in the area of their interest if encouraged</li>\n<li>Could graph based models be exploited better?</li>\n<li>How can we better leverage the attention weights? Luong et al. show that there is a great benefit in passing along the attention weights to the next timestep so that subsequent steps have an idea of what worked in the past. Of course that was in the field of NLP. Would it help here? We have seen in one of the papers above how they relate exercises to one another. In that case, does it makes sense to share past attention weight history at least among similar exercises in future timesteps?</li>\n</ol>\n<p>This article is a result of a quick weekend binge of reading various papers on the subject and my hurried interpretation of them. Please excuse typos/other lapses if any. I will review and correct as we go along.</p>\n<p>Note - I have not yet explored the actual Riiid dataset as of now because I wanted to have an unbiased look at existing literature and well, to be frank, my interests lie more in the application of technology rather than the technology  or the competition itself. Next weekend, I will try to create a more specific EDA and analysis on Riiid, but that may not be needed as I already see many great kernels…</p>\n<p>This has gone beyond what I set out to write but the papers were pretty exciting and I couldn't resist sharing my interpretation and thoughts on them. We are also about a month away from the deadline and I hope the couple of code sources I have pointed out will augment the existing Kaggle kernels in a positive way and not cause any type of shakeups.</p>\n<p>All the very best to all participants!</p>\n<p>Remaining series:</p>\n<p>From Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481</a></p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203184\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203184</a> - Hidden features and possible architectures</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185</a> - Some additional clarifications on SAKT/SAINT</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584</a> - A small discussion on position embeddings for those interested.</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719</a> - On lectures, the art of forgetting and why I retired hurt</p>",
  "messages": [
    {
      "id": 1102712,
      "postDate": "2020-12-05T08:19:38.323Z",
      "content": "<ul>\n<li>A general intro to the domain</li>\n<li>Quick summary of key models in plain English</li>\n<li>My interpretation of the salient approaches in top 5 papers on this domain, few code links &amp; a good video I found</li>\n<li>Few concluding remarks</li>\n</ul>\n<p>Target Audience - Newcomers like me who have no clue on the domain &amp; want a 15-min 'starter-kit' before going on to tackle the competition</p>\n<p><br><br></p>\n<h2>GENERAL INTRO TO THE DOMAIN</h2>\n<p>Knowledge Tracing (KT) is the art of modeling the knowledge of a student as she interacts with her course work. It is a core part of Intelligent Tutoring Systems (ITS) which aim to provide a personalized learning experience to students. The data for Knowledge Tracing (KT) comes from 'interactions' which happen when the students perform 'exercise's, and this allows the prediction of their performance on future exercises. Interactions contains the student 'response's. These are binary values indicating whether the student answered the exercise correctly or not and the student's knowledge is updated accordingly. In certain models, exercises are tagged to 'Knowledge-concepts' or 'Skill-tags' by domain experts and these additional tags help the models perform better. </p>\n<p>So KT is all about:<br>\nStudent interactions (exams) -&gt; Knowledge model(skill proficiency) --&gt; Prediction (on new exams)</p>\n<p>In general, most models right from Bayesian ones in early nineties have performed reasonably well and helped pedagogy evolve. The explosion of data in recent times naturally benefited the deep net models and though traditional models made a valiant comeback in between, most of the papers in the past 2 years have all been about deep models. The Covid situation has pushed even more students towards online education which has resulted in further data explosion and 2021 could be the year of mining all this data. Like how 2018 belonged to NLP, 2021 could well belong to Knowledge Tracing (KT) and we can definitely expect more such competitions. Co-incidentally, almost all SOTA models in the past 2 years use at least some part of the 'Attention is all you need' architecture by Vaswani et al. which changed the field of NLP in 2018.</p>\n<p><br><br></p>\n<h2>QUICK SUMMARY OF KEY MODELS IN PLAIN ENGLISH</h2>\n<p>Bayesian Knowledge Tracing (BKT) models knowledge state as a set of binaries - each representing the understanding or non-understanding of a single KC (Knowledge Concept). So it is essentially a set of tuples. Hidden Markov Models are used to update these, as the student answers exercises. The model assumes the learner never forgets and every new question has a fixed probability of helping the student understand the KC. BKT models were proposed as early as 1972 but refined in 95 by Corbett &amp; Anderson. It held sway for a good 2 decades with minor changes along the way (incorporating the possibility of guess-work, silly-mistakes etc). It was challenged by DL methods in 2015 and made a brilliant comeback by incorporating features like forgetting etc and defeated the benchmarks set by DL models. Nonetheless with the explosion of data and the proliferation of 'attention' mechanism, this model seems to have receded to the background in the past 2 years. A pity.</p>\n<p>Unlike BKT which assumed that the knowledge state at any time-step depends only on the previous step state, RNNs were able to capture the complexity and diversity of data over time effectively and rode the next wave. Also, with deep nets - domain expertise no longer became a mandatory requirement for modeling (though it is still a huge added advantage). So any Tom, Dick or Allohvk could now use a large dataset, a GPU and a few lines of Keras code and make an attempt to model student knowledge. Deep knowledge tracing (DKT) systems use an LSTM which at each time-step, takes the 'interaction' tuple as input. The tuple is simply encoded using a one-hot vector. The output is a vector of length equal to the number of questions, each element representing the predicted probability that the student would correctly answer the same. When it came out, the results of DKT were significantly better than BKT, though it was later found to that there were some data preprocessing issues (an inadvertent mistake) which resulted in higher scores which tempered some of the jubilation.</p>\n<p>Nonetheless, this success set the tone for coming waves. DKT was followed by Dynamic Key-Value Memory Networks (DKVMN) to improve DKT’s structure. These are also LSTM based but recognize the fact that the hidden state in the LSTM network has limited power and augment this with memory matrices to (better) store information. DKVMN maps higher level skill-tags to exercises and can observe which skill tags are needed for which exercise, find out which areas the student needs to be trained on to improve her/his performance and get better scores or tell which skill-tag was lacking in the student's output due to which the scores were lower. They need to be used in combination with domain specific annotations to fully exploit the power. The ability to predict which specific skill tag a student needs to be trained upon to improve the upcoming (exam/competition) scores seems to be a dream for students/parents/teachers and wannabe Kaggle GMs alike :) </p>\n<p>Most original implementations of KT used one-hot encoding mechanisms to encode interactions. With the explosion of data, some moved onto word embeddings while others insisted that embeddings are an overkill as in this domain, the distance between different samples have no major correlation unlike in NLP. They use OHE but reduce the dimensionality leveraging certain techniques. I believe this is a minority voice now as most 2018-2020 papers that I saw were all about embeddings. Secondly, in many papers exercises have been grouped into KCs and all exercises in one KC are treated as the same exercise. Basically all exercises covering the same concept are treated as a single exercise. So Q = C and Qt = Ct. This is done to keep things simple.</p>\n<p>The last couple of years have followed a predictable path with Attention being introduced followed by self-attention and transformer-like architectures. DKVMN took a back seat to transformer-like architectures which is again a pity as I feel DKVMN's scope had not fully been exploited. </p>\n<p>Exercise-Enhanced Recurrent Neural Network (EERNN) use a Bi-LSTM network to obtain the text embedding of each question, and concatenates the embedding with that of the corresponding student-interaction tuple. The concatenated embeddings now contain the text feature of the questions and are fed into a LSTM network. EERNN also uses attention mechanism to aggregate all the hidden states of the past LSTM units. </p>\n<p>The self-attentive knowledge tracing (SAKT) method introduced self-attention to KT and is inspired by various pieces of the transformer architecture from Vaswani's paper. The query is an exercise embedding vector, and key and value are interaction embedding vectors. The code is at: <a href=\"https://github.com/shalini1194/SAKT/tree/master/2019-EDM\" target=\"_blank\">https://github.com/shalini1194/SAKT/tree/master/2019-EDM</a>. </p>\n<p>To put it very briefly, SAKT identifies relevant Knowledge Concepts (KCs) from the past interactions and then predicts student’s performance based on her performance on those KCs by 'attend'ing to them. Note that they use the exercises themselves as KCs and the semantics can get a bit confusing due to this. There is an embedding layer for the exercises and an embedding layer for the interactions. When predicting the response to an exercise, the model uses the exercise embedding as a query and the interaction embeddings as keys and values and builds a set of attention weights which is then used for predictions. Note that at every time stamp t+1, the current exercise et+1's embedding is used to query the all past interactions (xt) which act as both the key and value. The other way to look at it would be to consider the model with inputs x1,x2,…,xt−1 and shift the exercise sequence one position ahead, e2,e3,…, et and the output being the correctness of the response r2,r3,…,rt. </p>\n<p>This is a simple enough model and the code link shared above can be taken as a starting point to the competition.</p>\n<p>When it came out, SAKT's results were exceptionally good. However, another paper did point out a small un-intentional bug in SAKT (see <a href=\"https://arxiv.org/pdf/2008.01169.pdf\" target=\"_blank\">https://arxiv.org/pdf/2008.01169.pdf</a>) which may have inflated the AUC a bit. Nonetheless, this paper was the first to usher in the age of transformers in KT universe and deserves praise. They use 'self-attention' and do away with RNNs altogether. For encoding temporal info, they have 'position-encoding' as in the transformer paper to maintain sequence knowledge, residual connections to propagate the lower layer features to the higher layers, the feed forward layer to bring in an element of non-linearity and a normalizing layer to help stabilize and accelerate…all these were inspired by the original transformer paper.  The most important was the use of multi-head attention. Multiple heads help in capturing the attention weights in different subspaces. It was bit hard for me to grasp this intuitively - why MHA works here? The whole idea of MHA in the original 'attention' paper was to capture different types of relationship between words. Possibly here, it may be capturing relations across different intervals of time? Anyway, as per the SAKT paper, the authors clearly demonstrate that performance drops when single attention heads were used. </p>\n<p>Though this is not a discussion on 'attention', it may be pertinent to define it in 1 line because nearly all the SOTA models use Attention in one form or the other. When presented with a question for which we want a prediction, the 'attention' model identifies relevant past interactions – it 'attends' to past interactions – giving higher weightage to interactions that matter while drowning out interactions that don't matter and then predicts future performance from these interactions. Which interactions to drown out and which to amplify is determined by a simple feed-forward network. The concept of 'self-attention' takes it one step further where every step in the sequence attends to every other step in the same sequence and attempts to build relationships within the sequence (before doing the inter-sequence Attention as mentioned above). So self-attention could be used to better embed interactions and exercises individually before a final attention layer where they connect to each other. Vaswani et al. also decided to use separate embeddings in the form of key, query and values instead of using just one embedding as in the original Bahdanau Attention paper in 2014. </p>\n<p>For a more detailed explanation of 'Attention' and an introduction to the Key, Query, Value concepts, please refer to: Craft your own Attention layer in 6 lines- The essence of Attention across all its intoxicating flavors ( <a href=\"https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e\" target=\"_blank\">https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e</a> ). The transformer explained in just 1 line is - a bunch of components like self-attention coupled with a few neat hacks like multi-head attention (repeat the Key, Query, Value  generation multiple times with different initial weights so as to more richly capture all relationships), position encoding (we are doing away with RNN so retain some way of encoding positions), residuals (to propagate the lower layer features to the higher layers), the feed forward layer(to bring in non-linearity) and a normalizing layers(stabilize and speeden). These blocks are there at both the encoder and decoder end and of course there is a connection between them. Almost all SOTA papers use some or all components of this architecture. All models require a masking mechanism that prevents the current position from 'attend'ing to subsequent positions. </p>\n<p><br><br></p>\n<h2>MY INTERPRETATION OF THE SALIENT FEATURES OF TOP 5 SOTA MODELS IN THIS DOMAIN</h2>\n<p>Disclaimer - I evaluated the top dozen odd papers with public access which were returned by a Google search. I am sure there are many other good papers as well</p>\n<p>Most of the below papers have come out in the last year or so around the same time. This has 2 interesting ramifications. They don't benchmark against one another. Also some of the claimed 'innovations' made have been replicated in another paper. Of course this is perfectly explainable since papers take a long time to get peer-reviewed and get published so there could be multiple groups working on same 'innovations' at a time.</p>\n<p>We will discuss them very briefly in NO PARTICULAR ORDER. I will only point out the salient items. I will continue to use the terminology - interactions, exercises, responses and Knowledge Concepts (KC) across papers even though the authors may have used different terminology. In some places, I may have inadvertently used the term 'question' instead of 'exercise'.</p>\n<p><br></p>\n<h2>Deep Knowledge Tracing with Transformers: </h2>\n<p><a href=\"http://link-springer-com-443.webvpn.fjmu.edu.cn/chapter/10.1007%2F978-3-030-52240-7_46\" target=\"_blank\">http://link-springer-com-443.webvpn.fjmu.edu.cn/chapter/10.1007%2F978-3-030-52240-7_46</a>. <br>\nThis is a 2020 paper. They introduce 2 specific innovations. Instead of directly encoding the questions, they define a separate set of weights which relate Knowledge Concepts(KC) to questions and then apply that to transform and come up with final representation of the exercises. Secondly they introduce a time decay function - they allow the attention weights to decay - this basically brings in the element of 'forgetting'. </p>\n<p>The model works like this - Take the interaction and pass it to an interaction-embedding layer. The output is the interaction-embedding (big surprise :) !) which is passed to the transformer block. We will come to the transformer block layer but first let us see what this interaction embedding layer is doing. Here, every KC is represented by a vector. A weight matrix is learnt during raining and this that associates the exercise in the interaction to every KC available. A weighted sum of this association(weight*KC vector) is taken and will represent the interaction. 1 point came to my mind - The interaction consists of the student response in addition to the exercise. Would it not make more sense to just see how the 'exercise' relates to KC vectors instead of seeing how the 'interaction' relates to the KC vectors?</p>\n<p>Let us now check out the 'transformer block'. This is not really a complete transformer as such but includes a few key components of the original transformer architecture - a self-attention layer, a FeedForward to bring in non-linearity and normalization layers. The authors talk of residual connections but I couldn't see this in their image. Anyway this transformer block takes the 'interaction embedding' and creates a Key, query, value for each. Based on that it calculates the attention weights and then the 'interaction-context' for each time-step which is then used for prediction. </p>\n<p>However there seem to be few missing info here - It looks like a simplified arch is shown and not the final one. For one thing, I feel almost certain that a multi-head attention would have been used…possibly positional encodings as well. The actual conf paper hyperlink does not work. I also couldn't find the source code for this paper. They also partially explain the model. In the block diagram provided, I could see a similar interaction-embedding layer and a transformer block which takes only the exercises as input (aah.. to some extent, my doubt above is cleared…) So looking at this image, I am guessing that the same process of embedding is carried out for the 'exercise's also (they really should title it exercise-embedding layer and not interaction-embedding). This exercise-embedding is now fed to another transformer block ALONG WITH THE 'interaction-context' to get the final context - one for each time-step. </p>\n<p>Finally a dense layer is used for predictions. </p>\n<p>So to summarize, salient features would be:</p>\n<ul>\n<li>A learnable weight matrix to map exercises to KCs - now we start out with better embeddings as we have used 'Attention' while preparing the initial embeddings itself (note this does not seem to be self-attention)</li>\n<li>A time-decay function for attention weights to model forgetting</li>\n<li>An encoder-decoder type of architecture where interaction-embeddings of t-1 steps and exercise-embeddings of t steps interact via attention mechanism (using Q, K, V concepts from the transformer paper) to generate a final context which is then used for predictions via a Dense layer</li>\n</ul>\n<p><br></p>\n<h2>Let us now look at SAINT - Separated Self-AttentIve Neural Knowledge Tracing: </h2>\n<p><a href=\"https://arxiv.org/pdf/2002.07033.pdf\" target=\"_blank\">https://arxiv.org/pdf/2002.07033.pdf</a>. This is Gold standard and probably came out before the above paper as they clearly mention that they are the first to introduce encoder-decoder transformer-like structure. They also introduce innovations in discovering the right key, query, values. </p>\n<p>The encoder takes the sequence of exercise embeddings as queries, keys and values, and produces output through a repeated self-attention mechanism. The decoder takes response embeddings (shifted right and prefixed with a start-token) as queries, keys and values then alternately applies self-attention and attention layers to the encoder output, thereby somewhat mimicking an actual transformer. This clean separation of input allows them to stack attention layers multiple times. In a nutshell, SAINT encoder takes the exercise embeddings and feeds the processed output to the decoder. The decoder meanwhile is ready with a first self-attention layer whose input is the response embeddings (along with some meta info) and which attends to itself (key, query, value all the same) and prepares a response embedding. Now the next attention block kicks in where the 'query' is the output of the first decoder layer. The keys and values are the encoder output. Finally, a prediction layer, consisting of a linear transformation layer followed by a sigmoid operation, is applied to the output of last layer so that the final decoder output is a series of probability values.</p>\n<p>They have additional meta data associated with questions and responses - things like response time, type of exercise etc - all of which is part of the respective embedding. Apart from this, they use multi-head attention, FF and residuals just as in the original transformer architecture.</p>\n<p>Few salient points:</p>\n<ul>\n<li>SOTA model - Has set the benchmark and seems to be quite thorough</li>\n<li>Deeper than other models - Stacking self-attention blocks improves performance unlike in other models like SAKT where perf deteriorates</li>\n<li>First to introduce an encoder-decode type of arch. Encoder feed is the exercises. Decoder feeds are the responses and the encoder output</li>\n<li>Key point - They do not generate interaction-embeddings but instead generate separate exercise-embeddings and response-embeddings. These are the encoder/decoder respectively. They do generate interaction-embeddings in their paper just to show in comparison studies that it does not add much value</li>\n<li>Use of meta tags along with question and responses</li>\n<li>Neat Experiments with masking and different key, query, values combinations before arriving at their final model</li>\n<li>Important: It must be noted how they incorporate temporal features only on the decoder end and show that this decision achieves the best AUC compared to incorporating them in the encoder input OR both the encoder/decoder</li>\n<li>I didn't see an explicit positional encoding. May not make a big difference</li>\n<li>I didn't see any 'forgetting' feature - This is interesting. Some sort of time-decay could have been applied to the weights. Ideally this should improve the scores and I am fairly sure this would have been experimented with before the paper was published. So there could be a good reason to give this a skip. Could the model be using some of the existing features and produce attention weights that have factored in some sort of time-decay already? Or is it just that 'forgetting' is at best a tricky area to tackle and in reality is far more complicated that what is typically modeled? </li>\n</ul>\n<p>The best performing model of SAINT has 4 layers and a latent space dimension of 512. I couldn't find associated source code. </p>\n<p><br></p>\n<h2>Deep Knowledge Tracing with Convolutions: </h2>\n<p><a href=\"https://arxiv.org/pdf/2008.01169.pdf\" target=\"_blank\">https://arxiv.org/pdf/2008.01169.pdf</a>.<br>\nThis is called the Convolutional Knowledge Tracing (CKT) model and was the first to introduce a convolutional network in the knowledge tracing. In addition to modeling the long-term effect of the entire question-answer sequence, CKT also strengthens the short-term effect of recent questions using 3D convolutions, thereby more effectively modeling the forgetting curve in the learning process. This is important because many studies have shown that in as little as 20 minutes after learning, only 58% of memory is retained, and this number drops to 25.4% in 6 days. They reshape the interaction-embeddings into a matrix and then use convolutions to extract important spatial patterns.</p>\n<p>Note that they use RNNs. The LSTM network outputs a hidden representation at every step t, which represents the hidden state of the processed interactions in the sequence so far. They now reshape each of the previous k embeddings (k is the window - so it takes embeddings from step t − k + 1 to t) into a matrix and then stack the k matrices to form a 3D tensor. They now use 3D convolutional network to learn from the tensor, which eventually outputs a hidden representation with length equal to that of the hidden representation generated by the LSTM network. They 'fuse' the two hidden representations and this fused representation is transformed to predict the student’s response.</p>\n<p>This is a refreshingly different approach. This model seems slightly older than the others nonetheless I liked its idea of using convolutions to capture spatial patterns. The way they fuse the 2 representations is also interesting. How does this way of adding temporal information compare to self-attention? Can self-attention be augmented further by convolutions? So instead of fusing the LSTM state with the convolution representation does it make sense to fuse it with the embeddings generated from self-attention (and of course do away with LSTMs altogether)? I guess it may not give better results because self-attention combined with MHA (and the whole lot of paraphernalia from the transformer - residuals, FF, norm layers, positional encoding etc) do completely account for temporal information as well as bring out deep representations…but is there any other way convolutions can benefit KT? This is a 100K dollar question</p>\n<p><br></p>\n<h2>Relation-aware self-attention model for Knowledge Tracing</h2>\n<p>One of the author of SAKT has come out with a new paper a couple of months back and introduces - RKT - Relation-aware self-attention model for Knowledge Tracing. This strengthens relationship between exercises using their textual content AS WELL AS student interaction data AND the forget behavior information (by modeling an exponentially decaying kernel function). Initially I did not understand why they needed student interaction data to capture relationship between exercises. I guess this is needed because apart from the textual similarity, many exercises need common skills - if two exercises are related, then the performance on one affects the other. Putting this in the reverse way, performance in interactions can identify relationships between exercises. We can say that if two exercises are textually very close to each other, then it makes sense to check student performance in interactions for these exercises and use it to further strengthen the relationship between those 2 exercises. So two exercise texts which are very similar and student interactions yield the same response for both (say both are correct) could be given higher weightage over two exercises that have similar text but student interactions show different responses (one correct, one wrong). In that sense, student interactions can be taken into account when building relationships between exercises. They use Phi coefficients (a measure of association for two binary variables - in this case the student responses) to build the exercise relationship and then add the textual similarity (provided it is within a certain distance) to derive the relation coefficients else make it zero. It seemed to be a neat innovative approach to me. Initially when reading the paper you may feel they seem to be prioritizing interaction performance over textual similarity but that is not the case. Relationships are only built if the questions are textually near to each other else the weight is zero!</p>\n<p>This exercise-relation information is represented by a set of relation coefficients. As before, they use the self-attention mechanism to learn the attention weights corresponding to the previous interaction for predicting whether a student will provide correct answer to the next exercise. These weights are now augmented by the relation coefficients to provide even better predictions. </p>\n<p>Another interesting point to note would be - They use the textual content of exercises to create simple embeddings for each word (not BERT). They use word2vec and use TEX tokens to transform equations in Maths exams. The exercise embedding is then a weighted combination of the embedding of all words present in the text of the exercise leveraging Smooth Inverse Frequency (SIF). SIF downgrades unimportant words and keeps relevant info only. The distance between exercises is then computed. I think they use cosine similarity.</p>\n<p>Lastly let us look at the ablations in their study. It is mighty interesting - particularly the one related to exercise relationship building. They compared model performances using a) only textual similarities b) only student interaction performance and c) using only KC tags (so all exercises in a same KC tag are related (weight=1) else they are not (weight=0) with d - the combination of a and b. The observations are as follows - 'c' performs the worst! This is expected as exercises can span KCs. 'b' seems to perform better than 'a' while 'd' is the best. Let us analyze 'b' again. The authors feel that even if textual content of two exercises are not similar, the association of knowledge involved in solving the two exercises could be high. Maybe I am totally wrong, but my understanding is this - With 'b' we are just clustering exercises into groups based on complexity. On one spectrum are exercises which are consistently solved correctly by most students (or is it just 1 student?), the other end of the spectrum consists of exercises that nobody could solve and the rest fall somewhere in between. While this may help in prediction, will this help in pedagogy? However a combination of 'a' and 'b' definitely makes sense and this is what the authors have chosen to go with and this is what gets the best results in the study. All in all, a fascinating paper.</p>\n<p><br></p>\n<h2>Context-Aware Attentive Knowledge Tracing</h2>\n<p><a href=\"https://arxiv.org/pdf/2007.12324.pdf\" target=\"_blank\">https://arxiv.org/pdf/2007.12324.pdf</a>.<br>\nAttentive knowledge tracing (AKT) brings in the following innovations: </p>\n<ul>\n<li>Instead of using raw question and response embeddings, they use context-aware representations of past questions and responses by taking a learner’s practice history into account using a modified version of attention. These modified representations reflect each learner’s actual comprehension of the question and the knowledge they actually acquire, given their personal response history. </li>\n<li>They modified the scaled dot product attention and have the attention weights decay exponentially based on the context-aware relative distance measure</li>\n<li>Use the Rasch model to bring out similarity &amp; differences in exercises within the same KC. Thus questions labeled as covering the same concept should not be treated as the same as they could have important individual differences</li>\n</ul>\n<p>There are four components: two self-attentive encoders, one for exercises and one for knowledge acquisition (I believe in plain words this is the interaction-embedding layer), a single attention-based knowledge retriever, and a feed-forward response prediction model.</p>\n<ul>\n<li>The two self-attentive encoders learn context-aware representations of the exercises and interactions. The first is the exercise encoder, which produces modified, contextualized representations of each exercise, given the sequence of exercises the learner has previously practiced on. The context-aware embedding of each exercise depends on both itself and the past exercises. The second is the knowledge encoder, which produces modified, contextualized representations of the knowledge the learner acquired while responding to past questions. This is the interaction-embedding </li>\n<li>The knowledge retriever, which retrieves knowledge acquired in the past that is relevant to the current question using an attention mechanism</li>\n<li>The response prediction model predicts the learner’s response to the current question using the retrieved knowledge</li>\n</ul>\n<p>Both encoders employ the self-attention mechanism. The knowledge retriever, on the other hand, uses the embedding of the current exercise as query, the keys are all the past exercises embeddings and values would be the interaction-embeddings. Notice how keys and values are different unlike in normal attention scenarios. For e.g. SAKT uses exercise embeddings as queries and interaction-embeddings for key as well as values. In AKT, exercise embeddings are used as queries and keys and values are the interaction embeddings. This is an interesting change and they claim that this method is more effective.</p>\n<p>The other interesting bit is the modifications made to the scaled dot product attention that is used in all other models. They say that scaled dot-product attention mechanism is not going to model memories decay effectively nor strengthen recent interactions. They add a multiplicative exponential decay term to the attention scores and a learnable decay rate parameter. So now the attention weights for the current question on a past question depends not only on the similarity between the corresponding query and key, but also on the relative number of time steps between them.</p>\n<p>The source code is available at: <a href=\"https://github.com/arghosh/AKT\" target=\"_blank\">https://github.com/arghosh/AKT</a></p>\n<p><br><br>\nAdditional notes: I later discovered a successor SAINT+ released a couple of months back which further pushes the benchmark by incorporating two additional temporal feature embeddings into the response embeddings: elapsed time, the time taken for a student to answer, and lag time, the time interval between adjacent learning activities. The lag time is used to build in 'forget'fulness (yay!). They show that the elapsed time is better represented as a continuous embedding instead of categorical which is somewhat counter-intuitive. Ideally 3-4 categories based on Z-score should have sufficed but of course this is just my opinion. Saint+ performs slightly better than Saint and is the new benchmark!</p>\n<p>Additional links:</p>\n<ul>\n<li><a href=\"https://github.com/thosgt/edm_main_algorithms:\" target=\"_blank\">https://github.com/thosgt/edm_main_algorithms:</a> Independent repository containing some models like DKT, SAKT etc</li>\n<li>Similar to above - contains a bunch of models - <a href=\"https://github.com/seewoo5/KT\" target=\"_blank\">https://github.com/seewoo5/KT</a></li>\n<li>A good introductory video: <a href=\"https://www.youtube.com/watch?v=CzRmRZNpB1Y\" target=\"_blank\">https://www.youtube.com/watch?v=CzRmRZNpB1Y</a></li>\n</ul>\n<p><br><br></p>\n<h2>GENERAL OBSERVATIONS:</h2>\n<p>My domain knowledge has accrued over a grand total of (the past) 36 hours… but perhaps viewing this domain as a rank outsider has its own advantages. My random thoughts (note - this is about the domain not the competition):</p>\n<ol>\n<li>1. While the current focus is on personalized recommendations, could there be plenty to exploit by generalization? Basically clustering of exercises or clustering of KCs, student profiles etc? The results could further aid personalization. But they have other implications. We could have signals that can be leveraged by tutors instead of students. If a set of questions are consistently answererd incorrectly, this could be a signal that the KCs related to those may need to be augmented with better course contents? </li>\n<li>While self-attention and transformer-like architectures are all the rage currently, can we leverage recommendation systems? Basically by using k-nearest on student profiles, we can predict how a student(/s) can perform by just looking at a similar profile from a past student who has competed the full course. Then use elements of recommendation systems to design the best course-path for this student?</li>\n<li>Can we bring in elements of reinforcement learning?</li>\n<li>Can we model the art of forgetting better? We forget a lot in the first 24-48 hours. After that the decay slows down. Beyond a certain point, knowledge does not decay! Also teh decay rate could vary across KCs. Maybe we could 'learn' the decay rate of cluster of KCs?</li>\n<li>Can we generate relevant KCs from exercises using DL methods? Some kind of topic modeling? Could this help? In general, question tagging can be further explored. This has nothing to do with individual student responses and is more generic</li>\n<li>Relevance of other cues could matter much more. For e.g. engagement of the student during the course is a much better predictor..tracking of eye movements etc?</li>\n<li>Performance is so much dependent on external factors. Brilliant students who have done well in the past  may start deteriorating if they have personal issues (death in the family, divorce, drugs, bad company, abuse etc)..The opposite too could happen. The shift does not happen overnight. Can we capture patterns to identify these shifts before they fully materialize? Interventions like counseling could help immensely if done at the right time</li>\n<li>How do we spot the super-performers? Most of the work is aimed at improving the scores of the majority. Super-performers may need a different kind of mentoring and customization</li>\n<li>How do we identify hidden talents - This is different from super-performers who are all-rounders. Hidden talents may be exceptionally good at only a certain type of exercises or subjects but average otherwise and typically do not stand out in any list. However they have the potential to become true geniuses in the area of their interest if encouraged</li>\n<li>Could graph based models be exploited better?</li>\n<li>How can we better leverage the attention weights? Luong et al. show that there is a great benefit in passing along the attention weights to the next timestep so that subsequent steps have an idea of what worked in the past. Of course that was in the field of NLP. Would it help here? We have seen in one of the papers above how they relate exercises to one another. In that case, does it makes sense to share past attention weight history at least among similar exercises in future timesteps?</li>\n</ol>\n<p>This article is a result of a quick weekend binge of reading various papers on the subject and my hurried interpretation of them. Please excuse typos/other lapses if any. I will review and correct as we go along.</p>\n<p>Note - I have not yet explored the actual Riiid dataset as of now because I wanted to have an unbiased look at existing literature and well, to be frank, my interests lie more in the application of technology rather than the technology  or the competition itself. Next weekend, I will try to create a more specific EDA and analysis on Riiid, but that may not be needed as I already see many great kernels…</p>\n<p>This has gone beyond what I set out to write but the papers were pretty exciting and I couldn't resist sharing my interpretation and thoughts on them. We are also about a month away from the deadline and I hope the couple of code sources I have pointed out will augment the existing Kaggle kernels in a positive way and not cause any type of shakeups.</p>\n<p>All the very best to all participants!</p>\n<p>Remaining series:</p>\n<p>From Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481</a></p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203184\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203184</a> - Hidden features and possible architectures</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185</a> - Some additional clarifications on SAKT/SAINT</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584</a> - A small discussion on position embeddings for those interested.</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719</a> - On lectures, the art of forgetting and why I retired hurt</p>",
      "rawMarkdown": "- A general intro to the domain\n- Quick summary of key models in plain English\n- My interpretation of the salient approaches in top 5 papers on this domain, few code links & a good video I found\n- Few concluding remarks\n\nTarget Audience - Newcomers like me who have no clue on the domain & want a 15-min 'starter-kit' before going on to tackle the competition\n\n<br><br>\n\nGENERAL INTRO TO THE DOMAIN\n-----------------------------------\nKnowledge Tracing (KT) is the art of modeling the knowledge of a student as she interacts with her course work. It is a core part of Intelligent Tutoring Systems (ITS) which aim to provide a personalized learning experience to students. The data for Knowledge Tracing (KT) comes from 'interactions' which happen when the students perform 'exercise's, and this allows the prediction of their performance on future exercises. Interactions contains the student 'response's. These are binary values indicating whether the student answered the exercise correctly or not and the student's knowledge is updated accordingly. In certain models, exercises are tagged to 'Knowledge-concepts' or 'Skill-tags' by domain experts and these additional tags help the models perform better. \n\nSo KT is all about:\nStudent interactions (exams) -> Knowledge model(skill proficiency) --> Prediction (on new exams)\n\nIn general, most models right from Bayesian ones in early nineties have performed reasonably well and helped pedagogy evolve. The explosion of data in recent times naturally benefited the deep net models and though traditional models made a valiant comeback in between, most of the papers in the past 2 years have all been about deep models. The Covid situation has pushed even more students towards online education which has resulted in further data explosion and 2021 could be the year of mining all this data. Like how 2018 belonged to NLP, 2021 could well belong to Knowledge Tracing (KT) and we can definitely expect more such competitions. Co-incidentally, almost all SOTA models in the past 2 years use at least some part of the 'Attention is all you need' architecture by Vaswani et al. which changed the field of NLP in 2018.\n\n\n<br><br>\nQUICK SUMMARY OF KEY MODELS IN PLAIN ENGLISH\n----------------------------------------------------------\nBayesian Knowledge Tracing (BKT) models knowledge state as a set of binaries - each representing the understanding or non-understanding of a single KC (Knowledge Concept). So it is essentially a set of tuples. Hidden Markov Models are used to update these, as the student answers exercises. The model assumes the learner never forgets and every new question has a fixed probability of helping the student understand the KC. BKT models were proposed as early as 1972 but refined in 95 by Corbett & Anderson. It held sway for a good 2 decades with minor changes along the way (incorporating the possibility of guess-work, silly-mistakes etc). It was challenged by DL methods in 2015 and made a brilliant comeback by incorporating features like forgetting etc and defeated the benchmarks set by DL models. Nonetheless with the explosion of data and the proliferation of 'attention' mechanism, this model seems to have receded to the background in the past 2 years. A pity.\n\nUnlike BKT which assumed that the knowledge state at any time-step depends only on the previous step state, RNNs were able to capture the complexity and diversity of data over time effectively and rode the next wave. Also, with deep nets - domain expertise no longer became a mandatory requirement for modeling (though it is still a huge added advantage). So any Tom, Dick or Allohvk could now use a large dataset, a GPU and a few lines of Keras code and make an attempt to model student knowledge. Deep knowledge tracing (DKT) systems use an LSTM which at each time-step, takes the 'interaction' tuple as input. The tuple is simply encoded using a one-hot vector. The output is a vector of length equal to the number of questions, each element representing the predicted probability that the student would correctly answer the same. When it came out, the results of DKT were significantly better than BKT, though it was later found to that there were some data preprocessing issues (an inadvertent mistake) which resulted in higher scores which tempered some of the jubilation.\n\nNonetheless, this success set the tone for coming waves. DKT was followed by Dynamic Key-Value Memory Networks (DKVMN) to improve DKT’s structure. These are also LSTM based but recognize the fact that the hidden state in the LSTM network has limited power and augment this with memory matrices to (better) store information. DKVMN maps higher level skill-tags to exercises and can observe which skill tags are needed for which exercise, find out which areas the student needs to be trained on to improve her/his performance and get better scores or tell which skill-tag was lacking in the student's output due to which the scores were lower. They need to be used in combination with domain specific annotations to fully exploit the power. The ability to predict which specific skill tag a student needs to be trained upon to improve the upcoming (exam/competition) scores seems to be a dream for students/parents/teachers and wannabe Kaggle GMs alike :) \n\nMost original implementations of KT used one-hot encoding mechanisms to encode interactions. With the explosion of data, some moved onto word embeddings while others insisted that embeddings are an overkill as in this domain, the distance between different samples have no major correlation unlike in NLP. They use OHE but reduce the dimensionality leveraging certain techniques. I believe this is a minority voice now as most 2018-2020 papers that I saw were all about embeddings. Secondly, in many papers exercises have been grouped into KCs and all exercises in one KC are treated as the same exercise. Basically all exercises covering the same concept are treated as a single exercise. So Q = C and Qt = Ct. This is done to keep things simple.\n\nThe last couple of years have followed a predictable path with Attention being introduced followed by self-attention and transformer-like architectures. DKVMN took a back seat to transformer-like architectures which is again a pity as I feel DKVMN's scope had not fully been exploited. \n\nExercise-Enhanced Recurrent Neural Network (EERNN) use a Bi-LSTM network to obtain the text embedding of each question, and concatenates the embedding with that of the corresponding student-interaction tuple. The concatenated embeddings now contain the text feature of the questions and are fed into a LSTM network. EERNN also uses attention mechanism to aggregate all the hidden states of the past LSTM units. \n\nThe self-attentive knowledge tracing (SAKT) method introduced self-attention to KT and is inspired by various pieces of the transformer architecture from Vaswani's paper. The query is an exercise embedding vector, and key and value are interaction embedding vectors. The code is at: https://github.com/shalini1194/SAKT/tree/master/2019-EDM. \n\nTo put it very briefly, SAKT identifies relevant Knowledge Concepts (KCs) from the past interactions and then predicts student’s performance based on her performance on those KCs by 'attend'ing to them. Note that they use the exercises themselves as KCs and the semantics can get a bit confusing due to this. There is an embedding layer for the exercises and an embedding layer for the interactions. When predicting the response to an exercise, the model uses the exercise embedding as a query and the interaction embeddings as keys and values and builds a set of attention weights which is then used for predictions. Note that at every time stamp t+1, the current exercise et+1's embedding is used to query the all past interactions (xt) which act as both the key and value. The other way to look at it would be to consider the model with inputs x1,x2,...,xt−1 and shift the exercise sequence one position ahead, e2,e3,..., et and the output being the correctness of the response r2,r3,...,rt. \n\nThis is a simple enough model and the code link shared above can be taken as a starting point to the competition.\n\nWhen it came out, SAKT's results were exceptionally good. However, another paper did point out a small un-intentional bug in SAKT (see https://arxiv.org/pdf/2008.01169.pdf) which may have inflated the AUC a bit. Nonetheless, this paper was the first to usher in the age of transformers in KT universe and deserves praise. They use 'self-attention' and do away with RNNs altogether. For encoding temporal info, they have 'position-encoding' as in the transformer paper to maintain sequence knowledge, residual connections to propagate the lower layer features to the higher layers, the feed forward layer to bring in an element of non-linearity and a normalizing layer to help stabilize and accelerate...all these were inspired by the original transformer paper.  The most important was the use of multi-head attention. Multiple heads help in capturing the attention weights in different subspaces. It was bit hard for me to grasp this intuitively - why MHA works here? The whole idea of MHA in the original 'attention' paper was to capture different types of relationship between words. Possibly here, it may be capturing relations across different intervals of time? Anyway, as per the SAKT paper, the authors clearly demonstrate that performance drops when single attention heads were used. \n\nThough this is not a discussion on 'attention', it may be pertinent to define it in 1 line because nearly all the SOTA models use Attention in one form or the other. When presented with a question for which we want a prediction, the 'attention' model identifies relevant past interactions – it 'attends' to past interactions – giving higher weightage to interactions that matter while drowning out interactions that don't matter and then predicts future performance from these interactions. Which interactions to drown out and which to amplify is determined by a simple feed-forward network. The concept of 'self-attention' takes it one step further where every step in the sequence attends to every other step in the same sequence and attempts to build relationships within the sequence (before doing the inter-sequence Attention as mentioned above). So self-attention could be used to better embed interactions and exercises individually before a final attention layer where they connect to each other. Vaswani et al. also decided to use separate embeddings in the form of key, query and values instead of using just one embedding as in the original Bahdanau Attention paper in 2014. \n\nFor a more detailed explanation of 'Attention' and an introduction to the Key, Query, Value concepts, please refer to: Craft your own Attention layer in 6 lines- The essence of Attention across all its intoxicating flavors ( https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e ). The transformer explained in just 1 line is - a bunch of components like self-attention coupled with a few neat hacks like multi-head attention (repeat the Key, Query, Value  generation multiple times with different initial weights so as to more richly capture all relationships), position encoding (we are doing away with RNN so retain some way of encoding positions), residuals (to propagate the lower layer features to the higher layers), the feed forward layer(to bring in non-linearity) and a normalizing layers(stabilize and speeden). These blocks are there at both the encoder and decoder end and of course there is a connection between them. Almost all SOTA papers use some or all components of this architecture. All models require a masking mechanism that prevents the current position from 'attend'ing to subsequent positions. \n\n\n\n\n<br><br>\nMY INTERPRETATION OF THE SALIENT FEATURES OF TOP 5 SOTA MODELS IN THIS DOMAIN\n------------------------------------------------------------------------------\nDisclaimer - I evaluated the top dozen odd papers with public access which were returned by a Google search. I am sure there are many other good papers as well\n\nMost of the below papers have come out in the last year or so around the same time. This has 2 interesting ramifications. They don't benchmark against one another. Also some of the claimed 'innovations' made have been replicated in another paper. Of course this is perfectly explainable since papers take a long time to get peer-reviewed and get published so there could be multiple groups working on same 'innovations' at a time.\n\nWe will discuss them very briefly in NO PARTICULAR ORDER. I will only point out the salient items. I will continue to use the terminology - interactions, exercises, responses and Knowledge Concepts (KC) across papers even though the authors may have used different terminology. In some places, I may have inadvertently used the term 'question' instead of 'exercise'.\n\n<br>\nDeep Knowledge Tracing with Transformers: \n--------------\nhttp://link-springer-com-443.webvpn.fjmu.edu.cn/chapter/10.1007%2F978-3-030-52240-7_46. \nThis is a 2020 paper. They introduce 2 specific innovations. Instead of directly encoding the questions, they define a separate set of weights which relate Knowledge Concepts(KC) to questions and then apply that to transform and come up with final representation of the exercises. Secondly they introduce a time decay function - they allow the attention weights to decay - this basically brings in the element of 'forgetting'. \n\nThe model works like this - Take the interaction and pass it to an interaction-embedding layer. The output is the interaction-embedding (big surprise :) !) which is passed to the transformer block. We will come to the transformer block layer but first let us see what this interaction embedding layer is doing. Here, every KC is represented by a vector. A weight matrix is learnt during raining and this that associates the exercise in the interaction to every KC available. A weighted sum of this association(weight*KC vector) is taken and will represent the interaction. 1 point came to my mind - The interaction consists of the student response in addition to the exercise. Would it not make more sense to just see how the 'exercise' relates to KC vectors instead of seeing how the 'interaction' relates to the KC vectors?\n\nLet us now check out the 'transformer block'. This is not really a complete transformer as such but includes a few key components of the original transformer architecture - a self-attention layer, a FeedForward to bring in non-linearity and normalization layers. The authors talk of residual connections but I couldn't see this in their image. Anyway this transformer block takes the 'interaction embedding' and creates a Key, query, value for each. Based on that it calculates the attention weights and then the 'interaction-context' for each time-step which is then used for prediction. \n\nHowever there seem to be few missing info here - It looks like a simplified arch is shown and not the final one. For one thing, I feel almost certain that a multi-head attention would have been used...possibly positional encodings as well. The actual conf paper hyperlink does not work. I also couldn't find the source code for this paper. They also partially explain the model. In the block diagram provided, I could see a similar interaction-embedding layer and a transformer block which takes only the exercises as input (aah.. to some extent, my doubt above is cleared...) So looking at this image, I am guessing that the same process of embedding is carried out for the 'exercise's also (they really should title it exercise-embedding layer and not interaction-embedding). This exercise-embedding is now fed to another transformer block ALONG WITH THE 'interaction-context' to get the final context - one for each time-step. \n\nFinally a dense layer is used for predictions. \n\nSo to summarize, salient features would be:\n- A learnable weight matrix to map exercises to KCs - now we start out with better embeddings as we have used 'Attention' while preparing the initial embeddings itself (note this does not seem to be self-attention)\n- A time-decay function for attention weights to model forgetting\n- An encoder-decoder type of architecture where interaction-embeddings of t-1 steps and exercise-embeddings of t steps interact via attention mechanism (using Q, K, V concepts from the transformer paper) to generate a final context which is then used for predictions via a Dense layer\n\n\n<br>\n\nLet us now look at SAINT - Separated Self-AttentIve Neural Knowledge Tracing: \n--------------\nhttps://arxiv.org/pdf/2002.07033.pdf. This is Gold standard and probably came out before the above paper as they clearly mention that they are the first to introduce encoder-decoder transformer-like structure. They also introduce innovations in discovering the right key, query, values. \n\nThe encoder takes the sequence of exercise embeddings as queries, keys and values, and produces output through a repeated self-attention mechanism. The decoder takes response embeddings (shifted right and prefixed with a start-token) as queries, keys and values then alternately applies self-attention and attention layers to the encoder output, thereby somewhat mimicking an actual transformer. This clean separation of input allows them to stack attention layers multiple times. In a nutshell, SAINT encoder takes the exercise embeddings and feeds the processed output to the decoder. The decoder meanwhile is ready with a first self-attention layer whose input is the response embeddings (along with some meta info) and which attends to itself (key, query, value all the same) and prepares a response embedding. Now the next attention block kicks in where the 'query' is the output of the first decoder layer. The keys and values are the encoder output. Finally, a prediction layer, consisting of a linear transformation layer followed by a sigmoid operation, is applied to the output of last layer so that the final decoder output is a series of probability values.\n\nThey have additional meta data associated with questions and responses - things like response time, type of exercise etc - all of which is part of the respective embedding. Apart from this, they use multi-head attention, FF and residuals just as in the original transformer architecture.\n\nFew salient points:\n- SOTA model - Has set the benchmark and seems to be quite thorough\n- Deeper than other models - Stacking self-attention blocks improves performance unlike in other models like SAKT where perf deteriorates\n- First to introduce an encoder-decode type of arch. Encoder feed is the exercises. Decoder feeds are the responses and the encoder output\n- Key point - They do not generate interaction-embeddings but instead generate separate exercise-embeddings and response-embeddings. These are the encoder/decoder respectively. They do generate interaction-embeddings in their paper just to show in comparison studies that it does not add much value\n- Use of meta tags along with question and responses\n- Neat Experiments with masking and different key, query, values combinations before arriving at their final model\n- Important: It must be noted how they incorporate temporal features only on the decoder end and show that this decision achieves the best AUC compared to incorporating them in the encoder input OR both the encoder/decoder\n- I didn't see an explicit positional encoding. May not make a big difference\n- I didn't see any 'forgetting' feature - This is interesting. Some sort of time-decay could have been applied to the weights. Ideally this should improve the scores and I am fairly sure this would have been experimented with before the paper was published. So there could be a good reason to give this a skip. Could the model be using some of the existing features and produce attention weights that have factored in some sort of time-decay already? Or is it just that 'forgetting' is at best a tricky area to tackle and in reality is far more complicated that what is typically modeled? \n\nThe best performing model of SAINT has 4 layers and a latent space dimension of 512. I couldn't find associated source code. \n\n\n\n<br>\nDeep Knowledge Tracing with Convolutions: \n--------------\nhttps://arxiv.org/pdf/2008.01169.pdf.\nThis is called the Convolutional Knowledge Tracing (CKT) model and was the first to introduce a convolutional network in the knowledge tracing. In addition to modeling the long-term effect of the entire question-answer sequence, CKT also strengthens the short-term effect of recent questions using 3D convolutions, thereby more effectively modeling the forgetting curve in the learning process. This is important because many studies have shown that in as little as 20 minutes after learning, only 58% of memory is retained, and this number drops to 25.4% in 6 days. They reshape the interaction-embeddings into a matrix and then use convolutions to extract important spatial patterns.\n\nNote that they use RNNs. The LSTM network outputs a hidden representation at every step t, which represents the hidden state of the processed interactions in the sequence so far. They now reshape each of the previous k embeddings (k is the window - so it takes embeddings from step t − k + 1 to t) into a matrix and then stack the k matrices to form a 3D tensor. They now use 3D convolutional network to learn from the tensor, which eventually outputs a hidden representation with length equal to that of the hidden representation generated by the LSTM network. They 'fuse' the two hidden representations and this fused representation is transformed to predict the student’s response.\n\nThis is a refreshingly different approach. This model seems slightly older than the others nonetheless I liked its idea of using convolutions to capture spatial patterns. The way they fuse the 2 representations is also interesting. How does this way of adding temporal information compare to self-attention? Can self-attention be augmented further by convolutions? So instead of fusing the LSTM state with the convolution representation does it make sense to fuse it with the embeddings generated from self-attention (and of course do away with LSTMs altogether)? I guess it may not give better results because self-attention combined with MHA (and the whole lot of paraphernalia from the transformer - residuals, FF, norm layers, positional encoding etc) do completely account for temporal information as well as bring out deep representations...but is there any other way convolutions can benefit KT? This is a 100K dollar question\n\n\n\n<br>\nRelation-aware self-attention model for Knowledge Tracing\n--------------\nOne of the author of SAKT has come out with a new paper a couple of months back and introduces - RKT - Relation-aware self-attention model for Knowledge Tracing. This strengthens relationship between exercises using their textual content AS WELL AS student interaction data AND the forget behavior information (by modeling an exponentially decaying kernel function). Initially I did not understand why they needed student interaction data to capture relationship between exercises. I guess this is needed because apart from the textual similarity, many exercises need common skills - if two exercises are related, then the performance on one affects the other. Putting this in the reverse way, performance in interactions can identify relationships between exercises. We can say that if two exercises are textually very close to each other, then it makes sense to check student performance in interactions for these exercises and use it to further strengthen the relationship between those 2 exercises. So two exercise texts which are very similar and student interactions yield the same response for both (say both are correct) could be given higher weightage over two exercises that have similar text but student interactions show different responses (one correct, one wrong). In that sense, student interactions can be taken into account when building relationships between exercises. They use Phi coefficients (a measure of association for two binary variables - in this case the student responses) to build the exercise relationship and then add the textual similarity (provided it is within a certain distance) to derive the relation coefficients else make it zero. It seemed to be a neat innovative approach to me. Initially when reading the paper you may feel they seem to be prioritizing interaction performance over textual similarity but that is not the case. Relationships are only built if the questions are textually near to each other else the weight is zero!\n\nThis exercise-relation information is represented by a set of relation coefficients. As before, they use the self-attention mechanism to learn the attention weights corresponding to the previous interaction for predicting whether a student will provide correct answer to the next exercise. These weights are now augmented by the relation coefficients to provide even better predictions. \n\nAnother interesting point to note would be - They use the textual content of exercises to create simple embeddings for each word (not BERT). They use word2vec and use TEX tokens to transform equations in Maths exams. The exercise embedding is then a weighted combination of the embedding of all words present in the text of the exercise leveraging Smooth Inverse Frequency (SIF). SIF downgrades unimportant words and keeps relevant info only. The distance between exercises is then computed. I think they use cosine similarity.\n\nLastly let us look at the ablations in their study. It is mighty interesting - particularly the one related to exercise relationship building. They compared model performances using a) only textual similarities b) only student interaction performance and c) using only KC tags (so all exercises in a same KC tag are related (weight=1) else they are not (weight=0) with d - the combination of a and b. The observations are as follows - 'c' performs the worst! This is expected as exercises can span KCs. 'b' seems to perform better than 'a' while 'd' is the best. Let us analyze 'b' again. The authors feel that even if textual content of two exercises are not similar, the association of knowledge involved in solving the two exercises could be high. Maybe I am totally wrong, but my understanding is this - With 'b' we are just clustering exercises into groups based on complexity. On one spectrum are exercises which are consistently solved correctly by most students (or is it just 1 student?), the other end of the spectrum consists of exercises that nobody could solve and the rest fall somewhere in between. While this may help in prediction, will this help in pedagogy? However a combination of 'a' and 'b' definitely makes sense and this is what the authors have chosen to go with and this is what gets the best results in the study. All in all, a fascinating paper.\n\n\n\n<br>\nContext-Aware Attentive Knowledge Tracing\n--------------\nhttps://arxiv.org/pdf/2007.12324.pdf.\nAttentive knowledge tracing (AKT) brings in the following innovations: \n- Instead of using raw question and response embeddings, they use context-aware representations of past questions and responses by taking a learner’s practice history into account using a modified version of attention. These modified representations reflect each learner’s actual comprehension of the question and the knowledge they actually acquire, given their personal response history. \n- They modified the scaled dot product attention and have the attention weights decay exponentially based on the context-aware relative distance measure\n- Use the Rasch model to bring out similarity & differences in exercises within the same KC. Thus questions labeled as covering the same concept should not be treated as the same as they could have important individual differences\n\nThere are four components: two self-attentive encoders, one for exercises and one for knowledge acquisition (I believe in plain words this is the interaction-embedding layer), a single attention-based knowledge retriever, and a feed-forward response prediction model.\n- The two self-attentive encoders learn context-aware representations of the exercises and interactions. The first is the exercise encoder, which produces modified, contextualized representations of each exercise, given the sequence of exercises the learner has previously practiced on. The context-aware embedding of each exercise depends on both itself and the past exercises. The second is the knowledge encoder, which produces modified, contextualized representations of the knowledge the learner acquired while responding to past questions. This is the interaction-embedding \n- The knowledge retriever, which retrieves knowledge acquired in the past that is relevant to the current question using an attention mechanism\n- The response prediction model predicts the learner’s response to the current question using the retrieved knowledge\n\nBoth encoders employ the self-attention mechanism. The knowledge retriever, on the other hand, uses the embedding of the current exercise as query, the keys are all the past exercises embeddings and values would be the interaction-embeddings. Notice how keys and values are different unlike in normal attention scenarios. For e.g. SAKT uses exercise embeddings as queries and interaction-embeddings for key as well as values. In AKT, exercise embeddings are used as queries and keys and values are the interaction embeddings. This is an interesting change and they claim that this method is more effective.\n\nThe other interesting bit is the modifications made to the scaled dot product attention that is used in all other models. They say that scaled dot-product attention mechanism is not going to model memories decay effectively nor strengthen recent interactions. They add a multiplicative exponential decay term to the attention scores and a learnable decay rate parameter. So now the attention weights for the current question on a past question depends not only on the similarity between the corresponding query and key, but also on the relative number of time steps between them.\n\nThe source code is available at: https://github.com/arghosh/AKT\n\n\n<br>\nAdditional notes: I later discovered a successor SAINT+ released a couple of months back which further pushes the benchmark by incorporating two additional temporal feature embeddings into the response embeddings: elapsed time, the time taken for a student to answer, and lag time, the time interval between adjacent learning activities. The lag time is used to build in 'forget'fulness (yay!). They show that the elapsed time is better represented as a continuous embedding instead of categorical which is somewhat counter-intuitive. Ideally 3-4 categories based on Z-score should have sufficed but of course this is just my opinion. Saint+ performs slightly better than Saint and is the new benchmark!\n\nAdditional links:\n- https://github.com/thosgt/edm_main_algorithms: Independent repository containing some models like DKT, SAKT etc\n- Similar to above - contains a bunch of models - https://github.com/seewoo5/KT\n- A good introductory video: https://www.youtube.com/watch?v=CzRmRZNpB1Y\n\n\n\n<br><br>\n\nGENERAL OBSERVATIONS:\n-----------------------\nMy domain knowledge has accrued over a grand total of (the past) 36 hours... but perhaps viewing this domain as a rank outsider has its own advantages. My random thoughts (note - this is about the domain not the competition):\n1. 1. While the current focus is on personalized recommendations, could there be plenty to exploit by generalization? Basically clustering of exercises or clustering of KCs, student profiles etc? The results could further aid personalization. But they have other implications. We could have signals that can be leveraged by tutors instead of students. If a set of questions are consistently answererd incorrectly, this could be a signal that the KCs related to those may need to be augmented with better course contents? \n2. While self-attention and transformer-like architectures are all the rage currently, can we leverage recommendation systems? Basically by using k-nearest on student profiles, we can predict how a student(/s) can perform by just looking at a similar profile from a past student who has competed the full course. Then use elements of recommendation systems to design the best course-path for this student?\n3. Can we bring in elements of reinforcement learning?\n4. Can we model the art of forgetting better? We forget a lot in the first 24-48 hours. After that the decay slows down. Beyond a certain point, knowledge does not decay! Also teh decay rate could vary across KCs. Maybe we could 'learn' the decay rate of cluster of KCs?\n5. Can we generate relevant KCs from exercises using DL methods? Some kind of topic modeling? Could this help? In general, question tagging can be further explored. This has nothing to do with individual student responses and is more generic\n6. Relevance of other cues could matter much more. For e.g. engagement of the student during the course is a much better predictor..tracking of eye movements etc?\n7. Performance is so much dependent on external factors. Brilliant students who have done well in the past  may start deteriorating if they have personal issues (death in the family, divorce, drugs, bad company, abuse etc)..The opposite too could happen. The shift does not happen overnight. Can we capture patterns to identify these shifts before they fully materialize? Interventions like counseling could help immensely if done at the right time\n8. How do we spot the super-performers? Most of the work is aimed at improving the scores of the majority. Super-performers may need a different kind of mentoring and customization\n9. How do we identify hidden talents - This is different from super-performers who are all-rounders. Hidden talents may be exceptionally good at only a certain type of exercises or subjects but average otherwise and typically do not stand out in any list. However they have the potential to become true geniuses in the area of their interest if encouraged\n10. Could graph based models be exploited better?\n11. How can we better leverage the attention weights? Luong et al. show that there is a great benefit in passing along the attention weights to the next timestep so that subsequent steps have an idea of what worked in the past. Of course that was in the field of NLP. Would it help here? We have seen in one of the papers above how they relate exercises to one another. In that case, does it makes sense to share past attention weight history at least among similar exercises in future timesteps?\n\nThis article is a result of a quick weekend binge of reading various papers on the subject and my hurried interpretation of them. Please excuse typos/other lapses if any. I will review and correct as we go along.\n\nNote - I have not yet explored the actual Riiid dataset as of now because I wanted to have an unbiased look at existing literature and well, to be frank, my interests lie more in the application of technology rather than the technology  or the competition itself. Next weekend, I will try to create a more specific EDA and analysis on Riiid, but that may not be needed as I already see many great kernels...\n\nThis has gone beyond what I set out to write but the papers were pretty exciting and I couldn't resist sharing my interpretation and thoughts on them. We are also about a month away from the deadline and I hope the couple of code sources I have pointed out will augment the existing Kaggle kernels in a positive way and not cause any type of shakeups.\n\nAll the very best to all participants!\n\nRemaining series:\n\nFrom Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203184 - Hidden features and possible architectures\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185 - Some additional clarifications on SAKT/SAINT\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584 - A small discussion on position embeddings for those interested.\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719 - On lectures, the art of forgetting and why I retired hurt\n",
      "votes": 98
    },
    {
      "id": 1720516,
      "postDate": "2022-03-12T21:44:43.963Z",
      "content": "<p>What an awesome summary!</p>",
      "rawMarkdown": "What an awesome summary!",
      "votes": 1
    },
    {
      "id": 1106977,
      "postDate": "2020-12-09T09:09:05.833Z",
      "content": "<p>Thank you, this is great! I dipped my toes into this stuff myself, but then got distracted by other things.</p>\n<p>One interesting note (although maybe only theoretically interesting, in the context of this competition): Intelligent Tutoring Systems as I'm used to them often don't actually use this kind of Deep Knowledge Tracing stuff; in part because they try to do a lot more varied and sometimes subjective things than predicting whether the student will answer the next question correctly (e.g. synthesizing and giving hints, detecting misconceptions, finding bugs in student code…), and in part because they often need a cohesive and interpretable student model which can be reasoned about and applied to all of these things that the ITS needs to do.</p>\n<p>One cool direction that split off from BKT is Learning Factor Analysis (and later also Performance Factor Analysis), based on the idea that the amount of difficulty a student has with a concept (e.g. odds of making an error, or the time it takes to complete a task) decreases according to a power function of the amount of practice they have. This can be used not only tot predict the chance of success, but also to analyze and refine a proposed set of knowledge concepts (e.g. if there are \"humps\" in the learning curve, maybe it's not just one concept…)</p>\n<p><a href=\"https://books.google.com/books?id=wfxYPwQ3A20C&amp;lpg=PA103&amp;lr&amp;pg=PA106\" target=\"_blank\">At least one source claims</a> (last paragraph of page 106) that this power function model is more consistent with real-world data than the learning curve that comes out of the KT model used in BKT (which is fundamentally a geometric curve, and won't necessarily start looking like a power curve even if you add other things like forgetting on top of it).</p>",
      "rawMarkdown": "Thank you, this is great! I dipped my toes into this stuff myself, but then got distracted by other things.\n\nOne interesting note (although maybe only theoretically interesting, in the context of this competition): Intelligent Tutoring Systems as I'm used to them often don't actually use this kind of Deep Knowledge Tracing stuff; in part because they try to do a lot more varied and sometimes subjective things than predicting whether the student will answer the next question correctly (e.g. synthesizing and giving hints, detecting misconceptions, finding bugs in student code...), and in part because they often need a cohesive and interpretable student model which can be reasoned about and applied to all of these things that the ITS needs to do.\n\nOne cool direction that split off from BKT is Learning Factor Analysis (and later also Performance Factor Analysis), based on the idea that the amount of difficulty a student has with a concept (e.g. odds of making an error, or the time it takes to complete a task) decreases according to a power function of the amount of practice they have. This can be used not only tot predict the chance of success, but also to analyze and refine a proposed set of knowledge concepts (e.g. if there are \"humps\" in the learning curve, maybe it's not just one concept...)\n\n[At least one source claims](https://books.google.com/books?id=wfxYPwQ3A20C&lpg=PA103&lr&pg=PA106) (last paragraph of page 106) that this power function model is more consistent with real-world data than the learning curve that comes out of the KT model used in BKT (which is fundamentally a geometric curve, and won't necessarily start looking like a power curve even if you add other things like forgetting on top of it).",
      "votes": 4,
      "replies": [
        {
          "id": 1107226,
          "postDate": "2020-12-09T13:56:25.587Z",
          "content": "<p>Thank you for your comments <a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a> The link seems interesting. With the usage of NN coupled with attention, the recent models(past 24 months) have become much much more powerful than BKT and may be able to model the student's knowledge better. But there are still many areas that can be further improved. This is a field that is going to evolve a lot in the coming couple of years, I guess..</p>",
          "rawMarkdown": "Thank you for your comments @yanamal The link seems interesting. With the usage of NN coupled with attention, the recent models(past 24 months) have become much much more powerful than BKT and may be able to model the student's knowledge better. But there are still many areas that can be further improved. This is a field that is going to evolve a lot in the coming couple of years, I guess..",
          "votes": 1
        },
        {
          "id": 1107444,
          "postDate": "2020-12-09T17:38:31.773Z",
          "content": "<p>I agree  - the limitations of BKT probably don't apply to the NN approaches that evolved from it, especially the most recent ones. I just think it's interesting to think about what would happen if, say, these kinds of approaches evolved using LFA as a starting point instead of BKT.</p>",
          "rawMarkdown": "I agree  - the limitations of BKT probably don't apply to the NN approaches that evolved from it, especially the most recent ones. I just think it's interesting to think about what would happen if, say, these kinds of approaches evolved using LFA as a starting point instead of BKT.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1107948,
      "postDate": "2020-12-10T05:53:02.047Z",
      "content": "<p>Greate note. Good to see your notes here again after OpenVaccine :)</p>",
      "rawMarkdown": "Greate note. Good to see your notes here again after OpenVaccine :)",
      "votes": 1,
      "replies": [
        {
          "id": 1107971,
          "postDate": "2020-12-10T06:17:03.457Z",
          "content": "<p>yea..that was fun Vijay :)</p>",
          "rawMarkdown": "yea..that was fun Vijay :)"
        }
      ]
    },
    {
      "id": 1106561,
      "postDate": "2020-12-08T23:19:56.033Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/Allohvk\" target=\"_blank\">@Allohvk</a>, thanks for sharing your work and also thanks for the work put into this brilliant write up. Keep up with the Great work and weel done.</p>",
      "rawMarkdown": "Hi @Allohvk, thanks for sharing your work and also thanks for the work put into this brilliant write up. Keep up with the Great work and weel done.",
      "votes": 1
    },
    {
      "id": 1106366,
      "postDate": "2020-12-08T18:42:29.087Z",
      "content": "<p>Wonderful sharing!<br>\nBut I am confused that what is the terminology of KC means(being equal to  interactions or representing addtional information like the lecture information in this competition)? Thank you!</p>",
      "rawMarkdown": "Wonderful sharing!\nBut I am confused that what is the terminology of KC means(being equal to  interactions or representing addtional information like the lecture information in this competition)? Thank you!",
      "votes": 1,
      "replies": [
        {
          "id": 1106946,
          "postDate": "2020-12-09T08:43:22.227Z",
          "content": "<p>KC generally stands for \"Knowledge Component\" - like a particular concept or skill which  (in the BKT model) the student either knows or does not know. In the competition, these might be roughly equivalent to the \"tags\" that questions and lectures have. Though we don't know much about tags, so it's not clear whether and how well they correlate to KCs.</p>",
          "rawMarkdown": "KC generally stands for \"Knowledge Component\" - like a particular concept or skill which  (in the BKT model) the student either knows or does not know. In the competition, these might be roughly equivalent to the \"tags\" that questions and lectures have. Though we don't know much about tags, so it's not clear whether and how well they correlate to KCs.",
          "votes": 2,
          "replies": [
            {
              "id": 1107411,
              "postDate": "2020-12-09T17:12:16.110Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 1107412,
              "postDate": "2020-12-09T17:12:45.220Z",
              "content": "<p><a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a> <a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> Thanks for your explanations！<br>\nSo, it means that I just consider it as an concept of how much knowledge students have studied in the domain of Knowledge Tracing. And in this competiton it still need to do more feature engineering to explain it  completely. 👍<br>\nHoping more thoughts about KCs and information can be used in embedding.</p>",
              "rawMarkdown": "@yanamal @allohvk Thanks for your explanations！\nSo, it means that I just consider it as an concept of how much knowledge students have studied in the domain of Knowledge Tracing. And in this competiton it still need to do more feature engineering to explain it  completely. 👍\nHoping more thoughts about KCs and information can be used in embedding.",
              "votes": 1
            }
          ]
        },
        {
          "id": 1107370,
          "postDate": "2020-12-09T16:22:14.727Z",
          "content": "<p><a href=\"https://www.kaggle.com/barcarum\" target=\"_blank\">@barcarum</a> As Yana mentioned, KC loosely maps to skill-tags. In the SAKT paper they have considered exercises as individual KCs to keep things simple. So these semantics could be slightly different fro model to model. I haven't yet had a chance to go thru' the dataset or the kernels of this competition though in the past couple of days I have been going thru some of the other discussions. Next weekend, I plan to go thru the couple of kernels and the dataset and will share my thoughts if I do so..</p>",
          "rawMarkdown": "@barcarum As Yana mentioned, KC loosely maps to skill-tags. In the SAKT paper they have considered exercises as individual KCs to keep things simple. So these semantics could be slightly different fro model to model. I haven't yet had a chance to go thru' the dataset or the kernels of this competition though in the past couple of days I have been going thru some of the other discussions. Next weekend, I plan to go thru the couple of kernels and the dataset and will share my thoughts if I do so..",
          "votes": 1
        }
      ]
    },
    {
      "id": 3084786,
      "postDate": "2024-12-31T09:49:26.597Z",
      "content": "<p>I'm feeling so validated after reading the general observations at the end. My thoughts were exactly the same, a recommendation system featuring collaborative filtering can be game changer in personalised learning. <br>\nAnd not just forgetting, the guessing should also be considered. After going through the dataset, I'm positive that a correlation can be found in between the guessing parameter and the time spent. Exciting times!</p>",
      "rawMarkdown": "I'm feeling so validated after reading the general observations at the end. My thoughts were exactly the same, a recommendation system featuring collaborative filtering can be game changer in personalised learning. \nAnd not just forgetting, the guessing should also be considered. After going through the dataset, I'm positive that a correlation can be found in between the guessing parameter and the time spent. Exciting times!"
    }
  ],
  "comments": [
    {
      "id": 1720516,
      "author_name": "Jay Ahn",
      "author_url": "",
      "post_date": "2022-03-12T21:44:43.963000",
      "content": "<p>What an awesome summary!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1106977,
      "author_name": "Yana Malysheva",
      "author_url": "",
      "post_date": "2020-12-09T09:09:05.833000",
      "content": "<p>Thank you, this is great! I dipped my toes into this stuff myself, but then got distracted by other things.</p>\n<p>One interesting note (although maybe only theoretically interesting, in the context of this competition): Intelligent Tutoring Systems as I'm used to them often don't actually use this kind of Deep Knowledge Tracing stuff; in part because they try to do a lot more varied and sometimes subjective things than predicting whether the student will answer the next question correctly (e.g. synthesizing and giving hints, detecting misconceptions, finding bugs in student code…), and in part because they often need a cohesive and interpretable student model which can be reasoned about and applied to all of these things that the ITS needs to do.</p>\n<p>One cool direction that split off from BKT is Learning Factor Analysis (and later also Performance Factor Analysis), based on the idea that the amount of difficulty a student has with a concept (e.g. odds of making an error, or the time it takes to complete a task) decreases according to a power function of the amount of practice they have. This can be used not only tot predict the chance of success, but also to analyze and refine a proposed set of knowledge concepts (e.g. if there are \"humps\" in the learning curve, maybe it's not just one concept…)</p>\n<p><a href=\"https://books.google.com/books?id=wfxYPwQ3A20C&amp;lpg=PA103&amp;lr&amp;pg=PA106\" target=\"_blank\">At least one source claims</a> (last paragraph of page 106) that this power function model is more consistent with real-world data than the learning curve that comes out of the KT model used in BKT (which is fundamentally a geometric curve, and won't necessarily start looking like a power curve even if you add other things like forgetting on top of it).</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1107226,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-09T13:56:25.587000",
          "content": "<p>Thank you for your comments <a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a> The link seems interesting. With the usage of NN coupled with attention, the recent models(past 24 months) have become much much more powerful than BKT and may be able to model the student's knowledge better. But there are still many areas that can be further improved. This is a field that is going to evolve a lot in the coming couple of years, I guess..</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107444,
          "author_name": "Yana Malysheva",
          "author_url": "",
          "post_date": "2020-12-09T17:38:31.773000",
          "content": "<p>I agree  - the limitations of BKT probably don't apply to the NN approaches that evolved from it, especially the most recent ones. I just think it's interesting to think about what would happen if, say, these kinds of approaches evolved using LFA as a starting point instead of BKT.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1107948,
      "author_name": "Vee",
      "author_url": "",
      "post_date": "2020-12-10T05:53:02.047000",
      "content": "<p>Greate note. Good to see your notes here again after OpenVaccine :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1107971,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-10T06:17:03.457000",
          "content": "<p>yea..that was fun Vijay :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1106561,
      "author_name": "Patrick Uzuwe",
      "author_url": "",
      "post_date": "2020-12-08T23:19:56.033000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/Allohvk\" target=\"_blank\">@Allohvk</a>, thanks for sharing your work and also thanks for the work put into this brilliant write up. Keep up with the Great work and weel done.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1106366,
      "author_name": "barcarum",
      "author_url": "",
      "post_date": "2020-12-08T18:42:29.087000",
      "content": "<p>Wonderful sharing!<br>\nBut I am confused that what is the terminology of KC means(being equal to  interactions or representing addtional information like the lecture information in this competition)? Thank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1106946,
          "author_name": "Yana Malysheva",
          "author_url": "",
          "post_date": "2020-12-09T08:43:22.227000",
          "content": "<p>KC generally stands for \"Knowledge Component\" - like a particular concept or skill which  (in the BKT model) the student either knows or does not know. In the competition, these might be roughly equivalent to the \"tags\" that questions and lectures have. Though we don't know much about tags, so it's not clear whether and how well they correlate to KCs.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 1107411,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-12-09T17:12:16.110000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1107412,
              "author_name": "barcarum",
              "author_url": "",
              "post_date": "2020-12-09T17:12:45.220000",
              "content": "<p><a href=\"https://www.kaggle.com/yanamal\" target=\"_blank\">@yanamal</a> <a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> Thanks for your explanations！<br>\nSo, it means that I just consider it as an concept of how much knowledge students have studied in the domain of Knowledge Tracing. And in this competiton it still need to do more feature engineering to explain it  completely. 👍<br>\nHoping more thoughts about KCs and information can be used in embedding.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1107370,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-09T16:22:14.727000",
          "content": "<p><a href=\"https://www.kaggle.com/barcarum\" target=\"_blank\">@barcarum</a> As Yana mentioned, KC loosely maps to skill-tags. In the SAKT paper they have considered exercises as individual KCs to keep things simple. So these semantics could be slightly different fro model to model. I haven't yet had a chance to go thru' the dataset or the kernels of this competition though in the past couple of days I have been going thru some of the other discussions. Next weekend, I plan to go thru the couple of kernels and the dataset and will share my thoughts if I do so..</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3084786,
      "author_name": "shambles",
      "author_url": "",
      "post_date": "2024-12-31T09:49:26.597000",
      "content": "<p>I'm feeling so validated after reading the general observations at the end. My thoughts were exactly the same, a recommendation system featuring collaborative filtering can be game changer in personalised learning. <br>\nAnd not just forgetting, the guessing should also be considered. After going through the dataset, I'm positive that a correlation can be found in between the guessing parameter and the time spent. Exciting times!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1102712": "- A general intro to the domain\n- Quick summary of key models in plain English\n- My interpretation of the salient approaches in top 5 papers on this domain, few code links & a good video I found\n- Few concluding remarks\n\nTarget Audience - Newcomers like me who have no clue on the domain & want a 15-min 'starter-kit' before going on to tackle the competition\n\n<br><br>\n\nGENERAL INTRO TO THE DOMAIN\n-----------------------------------\nKnowledge Tracing (KT) is the art of modeling the knowledge of a student as she interacts with her course work. It is a core part of Intelligent Tutoring Systems (ITS) which aim to provide a personalized learning experience to students. The data for Knowledge Tracing (KT) comes from 'interactions' which happen when the students perform 'exercise's, and this allows the prediction of their performance on future exercises. Interactions contains the student 'response's. These are binary values indicating whether the student answered the exercise correctly or not and the student's knowledge is updated accordingly. In certain models, exercises are tagged to 'Knowledge-concepts' or 'Skill-tags' by domain experts and these additional tags help the models perform better. \n\nSo KT is all about:\nStudent interactions (exams) -> Knowledge model(skill proficiency) --> Prediction (on new exams)\n\nIn general, most models right from Bayesian ones in early nineties have performed reasonably well and helped pedagogy evolve. The explosion of data in recent times naturally benefited the deep net models and though traditional models made a valiant comeback in between, most of the papers in the past 2 years have all been about deep models. The Covid situation has pushed even more students towards online education which has resulted in further data explosion and 2021 could be the year of mining all this data. Like how 2018 belonged to NLP, 2021 could well belong to Knowledge Tracing (KT) and we can definitely expect more such competitions. Co-incidentally, almost all SOTA models in the past 2 years use at least some part of the 'Attention is all you need' architecture by Vaswani et al. which changed the field of NLP in 2018.\n\n\n<br><br>\nQUICK SUMMARY OF KEY MODELS IN PLAIN ENGLISH\n----------------------------------------------------------\nBayesian Knowledge Tracing (BKT) models knowledge state as a set of binaries - each representing the understanding or non-understanding of a single KC (Knowledge Concept). So it is essentially a set of tuples. Hidden Markov Models are used to update these, as the student answers exercises. The model assumes the learner never forgets and every new question has a fixed probability of helping the student understand the KC. BKT models were proposed as early as 1972 but refined in 95 by Corbett & Anderson. It held sway for a good 2 decades with minor changes along the way (incorporating the possibility of guess-work, silly-mistakes etc). It was challenged by DL methods in 2015 and made a brilliant comeback by incorporating features like forgetting etc and defeated the benchmarks set by DL models. Nonetheless with the explosion of data and the proliferation of 'attention' mechanism, this model seems to have receded to the background in the past 2 years. A pity.\n\nUnlike BKT which assumed that the knowledge state at any time-step depends only on the previous step state, RNNs were able to capture the complexity and diversity of data over time effectively and rode the next wave. Also, with deep nets - domain expertise no longer became a mandatory requirement for modeling (though it is still a huge added advantage). So any Tom, Dick or Allohvk could now use a large dataset, a GPU and a few lines of Keras code and make an attempt to model student knowledge. Deep knowledge tracing (DKT) systems use an LSTM which at each time-step, takes the 'interaction' tuple as input. The tuple is simply encoded using a one-hot vector. The output is a vector of length equal to the number of questions, each element representing the predicted probability that the student would correctly answer the same. When it came out, the results of DKT were significantly better than BKT, though it was later found to that there were some data preprocessing issues (an inadvertent mistake) which resulted in higher scores which tempered some of the jubilation.\n\nNonetheless, this success set the tone for coming waves. DKT was followed by Dynamic Key-Value Memory Networks (DKVMN) to improve DKT’s structure. These are also LSTM based but recognize the fact that the hidden state in the LSTM network has limited power and augment this with memory matrices to (better) store information. DKVMN maps higher level skill-tags to exercises and can observe which skill tags are needed for which exercise, find out which areas the student needs to be trained on to improve her/his performance and get better scores or tell which skill-tag was lacking in the student's output due to which the scores were lower. They need to be used in combination with domain specific annotations to fully exploit the power. The ability to predict which specific skill tag a student needs to be trained upon to improve the upcoming (exam/competition) scores seems to be a dream for students/parents/teachers and wannabe Kaggle GMs alike :) \n\nMost original implementations of KT used one-hot encoding mechanisms to encode interactions. With the explosion of data, some moved onto word embeddings while others insisted that embeddings are an overkill as in this domain, the distance between different samples have no major correlation unlike in NLP. They use OHE but reduce the dimensionality leveraging certain techniques. I believe this is a minority voice now as most 2018-2020 papers that I saw were all about embeddings. Secondly, in many papers exercises have been grouped into KCs and all exercises in one KC are treated as the same exercise. Basically all exercises covering the same concept are treated as a single exercise. So Q = C and Qt = Ct. This is done to keep things simple.\n\nThe last couple of years have followed a predictable path with Attention being introduced followed by self-attention and transformer-like architectures. DKVMN took a back seat to transformer-like architectures which is again a pity as I feel DKVMN's scope had not fully been exploited. \n\nExercise-Enhanced Recurrent Neural Network (EERNN) use a Bi-LSTM network to obtain the text embedding of each question, and concatenates the embedding with that of the corresponding student-interaction tuple. The concatenated embeddings now contain the text feature of the questions and are fed into a LSTM network. EERNN also uses attention mechanism to aggregate all the hidden states of the past LSTM units. \n\nThe self-attentive knowledge tracing (SAKT) method introduced self-attention to KT and is inspired by various pieces of the transformer architecture from Vaswani's paper. The query is an exercise embedding vector, and key and value are interaction embedding vectors. The code is at: https://github.com/shalini1194/SAKT/tree/master/2019-EDM. \n\nTo put it very briefly, SAKT identifies relevant Knowledge Concepts (KCs) from the past interactions and then predicts student’s performance based on her performance on those KCs by 'attend'ing to them. Note that they use the exercises themselves as KCs and the semantics can get a bit confusing due to this. There is an embedding layer for the exercises and an embedding layer for the interactions. When predicting the response to an exercise, the model uses the exercise embedding as a query and the interaction embeddings as keys and values and builds a set of attention weights which is then used for predictions. Note that at every time stamp t+1, the current exercise et+1's embedding is used to query the all past interactions (xt) which act as both the key and value. The other way to look at it would be to consider the model with inputs x1,x2,...,xt−1 and shift the exercise sequence one position ahead, e2,e3,..., et and the output being the correctness of the response r2,r3,...,rt. \n\nThis is a simple enough model and the code link shared above can be taken as a starting point to the competition.\n\nWhen it came out, SAKT's results were exceptionally good. However, another paper did point out a small un-intentional bug in SAKT (see https://arxiv.org/pdf/2008.01169.pdf) which may have inflated the AUC a bit. Nonetheless, this paper was the first to usher in the age of transformers in KT universe and deserves praise. They use 'self-attention' and do away with RNNs altogether. For encoding temporal info, they have 'position-encoding' as in the transformer paper to maintain sequence knowledge, residual connections to propagate the lower layer features to the higher layers, the feed forward layer to bring in an element of non-linearity and a normalizing layer to help stabilize and accelerate...all these were inspired by the original transformer paper.  The most important was the use of multi-head attention. Multiple heads help in capturing the attention weights in different subspaces. It was bit hard for me to grasp this intuitively - why MHA works here? The whole idea of MHA in the original 'attention' paper was to capture different types of relationship between words. Possibly here, it may be capturing relations across different intervals of time? Anyway, as per the SAKT paper, the authors clearly demonstrate that performance drops when single attention heads were used. \n\nThough this is not a discussion on 'attention', it may be pertinent to define it in 1 line because nearly all the SOTA models use Attention in one form or the other. When presented with a question for which we want a prediction, the 'attention' model identifies relevant past interactions – it 'attends' to past interactions – giving higher weightage to interactions that matter while drowning out interactions that don't matter and then predicts future performance from these interactions. Which interactions to drown out and which to amplify is determined by a simple feed-forward network. The concept of 'self-attention' takes it one step further where every step in the sequence attends to every other step in the same sequence and attempts to build relationships within the sequence (before doing the inter-sequence Attention as mentioned above). So self-attention could be used to better embed interactions and exercises individually before a final attention layer where they connect to each other. Vaswani et al. also decided to use separate embeddings in the form of key, query and values instead of using just one embedding as in the original Bahdanau Attention paper in 2014. \n\nFor a more detailed explanation of 'Attention' and an introduction to the Key, Query, Value concepts, please refer to: Craft your own Attention layer in 6 lines- The essence of Attention across all its intoxicating flavors ( https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e ). The transformer explained in just 1 line is - a bunch of components like self-attention coupled with a few neat hacks like multi-head attention (repeat the Key, Query, Value  generation multiple times with different initial weights so as to more richly capture all relationships), position encoding (we are doing away with RNN so retain some way of encoding positions), residuals (to propagate the lower layer features to the higher layers), the feed forward layer(to bring in non-linearity) and a normalizing layers(stabilize and speeden). These blocks are there at both the encoder and decoder end and of course there is a connection between them. Almost all SOTA papers use some or all components of this architecture. All models require a masking mechanism that prevents the current position from 'attend'ing to subsequent positions. \n\n\n\n\n<br><br>\nMY INTERPRETATION OF THE SALIENT FEATURES OF TOP 5 SOTA MODELS IN THIS DOMAIN\n------------------------------------------------------------------------------\nDisclaimer - I evaluated the top dozen odd papers with public access which were returned by a Google search. I am sure there are many other good papers as well\n\nMost of the below papers have come out in the last year or so around the same time. This has 2 interesting ramifications. They don't benchmark against one another. Also some of the claimed 'innovations' made have been replicated in another paper. Of course this is perfectly explainable since papers take a long time to get peer-reviewed and get published so there could be multiple groups working on same 'innovations' at a time.\n\nWe will discuss them very briefly in NO PARTICULAR ORDER. I will only point out the salient items. I will continue to use the terminology - interactions, exercises, responses and Knowledge Concepts (KC) across papers even though the authors may have used different terminology. In some places, I may have inadvertently used the term 'question' instead of 'exercise'.\n\n<br>\nDeep Knowledge Tracing with Transformers: \n--------------\nhttp://link-springer-com-443.webvpn.fjmu.edu.cn/chapter/10.1007%2F978-3-030-52240-7_46. \nThis is a 2020 paper. They introduce 2 specific innovations. Instead of directly encoding the questions, they define a separate set of weights which relate Knowledge Concepts(KC) to questions and then apply that to transform and come up with final representation of the exercises. Secondly they introduce a time decay function - they allow the attention weights to decay - this basically brings in the element of 'forgetting'. \n\nThe model works like this - Take the interaction and pass it to an interaction-embedding layer. The output is the interaction-embedding (big surprise :) !) which is passed to the transformer block. We will come to the transformer block layer but first let us see what this interaction embedding layer is doing. Here, every KC is represented by a vector. A weight matrix is learnt during raining and this that associates the exercise in the interaction to every KC available. A weighted sum of this association(weight*KC vector) is taken and will represent the interaction. 1 point came to my mind - The interaction consists of the student response in addition to the exercise. Would it not make more sense to just see how the 'exercise' relates to KC vectors instead of seeing how the 'interaction' relates to the KC vectors?\n\nLet us now check out the 'transformer block'. This is not really a complete transformer as such but includes a few key components of the original transformer architecture - a self-attention layer, a FeedForward to bring in non-linearity and normalization layers. The authors talk of residual connections but I couldn't see this in their image. Anyway this transformer block takes the 'interaction embedding' and creates a Key, query, value for each. Based on that it calculates the attention weights and then the 'interaction-context' for each time-step which is then used for prediction. \n\nHowever there seem to be few missing info here - It looks like a simplified arch is shown and not the final one. For one thing, I feel almost certain that a multi-head attention would have been used...possibly positional encodings as well. The actual conf paper hyperlink does not work. I also couldn't find the source code for this paper. They also partially explain the model. In the block diagram provided, I could see a similar interaction-embedding layer and a transformer block which takes only the exercises as input (aah.. to some extent, my doubt above is cleared...) So looking at this image, I am guessing that the same process of embedding is carried out for the 'exercise's also (they really should title it exercise-embedding layer and not interaction-embedding). This exercise-embedding is now fed to another transformer block ALONG WITH THE 'interaction-context' to get the final context - one for each time-step. \n\nFinally a dense layer is used for predictions. \n\nSo to summarize, salient features would be:\n- A learnable weight matrix to map exercises to KCs - now we start out with better embeddings as we have used 'Attention' while preparing the initial embeddings itself (note this does not seem to be self-attention)\n- A time-decay function for attention weights to model forgetting\n- An encoder-decoder type of architecture where interaction-embeddings of t-1 steps and exercise-embeddings of t steps interact via attention mechanism (using Q, K, V concepts from the transformer paper) to generate a final context which is then used for predictions via a Dense layer\n\n\n<br>\n\nLet us now look at SAINT - Separated Self-AttentIve Neural Knowledge Tracing: \n--------------\nhttps://arxiv.org/pdf/2002.07033.pdf. This is Gold standard and probably came out before the above paper as they clearly mention that they are the first to introduce encoder-decoder transformer-like structure. They also introduce innovations in discovering the right key, query, values. \n\nThe encoder takes the sequence of exercise embeddings as queries, keys and values, and produces output through a repeated self-attention mechanism. The decoder takes response embeddings (shifted right and prefixed with a start-token) as queries, keys and values then alternately applies self-attention and attention layers to the encoder output, thereby somewhat mimicking an actual transformer. This clean separation of input allows them to stack attention layers multiple times. In a nutshell, SAINT encoder takes the exercise embeddings and feeds the processed output to the decoder. The decoder meanwhile is ready with a first self-attention layer whose input is the response embeddings (along with some meta info) and which attends to itself (key, query, value all the same) and prepares a response embedding. Now the next attention block kicks in where the 'query' is the output of the first decoder layer. The keys and values are the encoder output. Finally, a prediction layer, consisting of a linear transformation layer followed by a sigmoid operation, is applied to the output of last layer so that the final decoder output is a series of probability values.\n\nThey have additional meta data associated with questions and responses - things like response time, type of exercise etc - all of which is part of the respective embedding. Apart from this, they use multi-head attention, FF and residuals just as in the original transformer architecture.\n\nFew salient points:\n- SOTA model - Has set the benchmark and seems to be quite thorough\n- Deeper than other models - Stacking self-attention blocks improves performance unlike in other models like SAKT where perf deteriorates\n- First to introduce an encoder-decode type of arch. Encoder feed is the exercises. Decoder feeds are the responses and the encoder output\n- Key point - They do not generate interaction-embeddings but instead generate separate exercise-embeddings and response-embeddings. These are the encoder/decoder respectively. They do generate interaction-embeddings in their paper just to show in comparison studies that it does not add much value\n- Use of meta tags along with question and responses\n- Neat Experiments with masking and different key, query, values combinations before arriving at their final model\n- Important: It must be noted how they incorporate temporal features only on the decoder end and show that this decision achieves the best AUC compared to incorporating them in the encoder input OR both the encoder/decoder\n- I didn't see an explicit positional encoding. May not make a big difference\n- I didn't see any 'forgetting' feature - This is interesting. Some sort of time-decay could have been applied to the weights. Ideally this should improve the scores and I am fairly sure this would have been experimented with before the paper was published. So there could be a good reason to give this a skip. Could the model be using some of the existing features and produce attention weights that have factored in some sort of time-decay already? Or is it just that 'forgetting' is at best a tricky area to tackle and in reality is far more complicated that what is typically modeled? \n\nThe best performing model of SAINT has 4 layers and a latent space dimension of 512. I couldn't find associated source code. \n\n\n\n<br>\nDeep Knowledge Tracing with Convolutions: \n--------------\nhttps://arxiv.org/pdf/2008.01169.pdf.\nThis is called the Convolutional Knowledge Tracing (CKT) model and was the first to introduce a convolutional network in the knowledge tracing. In addition to modeling the long-term effect of the entire question-answer sequence, CKT also strengthens the short-term effect of recent questions using 3D convolutions, thereby more effectively modeling the forgetting curve in the learning process. This is important because many studies have shown that in as little as 20 minutes after learning, only 58% of memory is retained, and this number drops to 25.4% in 6 days. They reshape the interaction-embeddings into a matrix and then use convolutions to extract important spatial patterns.\n\nNote that they use RNNs. The LSTM network outputs a hidden representation at every step t, which represents the hidden state of the processed interactions in the sequence so far. They now reshape each of the previous k embeddings (k is the window - so it takes embeddings from step t − k + 1 to t) into a matrix and then stack the k matrices to form a 3D tensor. They now use 3D convolutional network to learn from the tensor, which eventually outputs a hidden representation with length equal to that of the hidden representation generated by the LSTM network. They 'fuse' the two hidden representations and this fused representation is transformed to predict the student’s response.\n\nThis is a refreshingly different approach. This model seems slightly older than the others nonetheless I liked its idea of using convolutions to capture spatial patterns. The way they fuse the 2 representations is also interesting. How does this way of adding temporal information compare to self-attention? Can self-attention be augmented further by convolutions? So instead of fusing the LSTM state with the convolution representation does it make sense to fuse it with the embeddings generated from self-attention (and of course do away with LSTMs altogether)? I guess it may not give better results because self-attention combined with MHA (and the whole lot of paraphernalia from the transformer - residuals, FF, norm layers, positional encoding etc) do completely account for temporal information as well as bring out deep representations...but is there any other way convolutions can benefit KT? This is a 100K dollar question\n\n\n\n<br>\nRelation-aware self-attention model for Knowledge Tracing\n--------------\nOne of the author of SAKT has come out with a new paper a couple of months back and introduces - RKT - Relation-aware self-attention model for Knowledge Tracing. This strengthens relationship between exercises using their textual content AS WELL AS student interaction data AND the forget behavior information (by modeling an exponentially decaying kernel function). Initially I did not understand why they needed student interaction data to capture relationship between exercises. I guess this is needed because apart from the textual similarity, many exercises need common skills - if two exercises are related, then the performance on one affects the other. Putting this in the reverse way, performance in interactions can identify relationships between exercises. We can say that if two exercises are textually very close to each other, then it makes sense to check student performance in interactions for these exercises and use it to further strengthen the relationship between those 2 exercises. So two exercise texts which are very similar and student interactions yield the same response for both (say both are correct) could be given higher weightage over two exercises that have similar text but student interactions show different responses (one correct, one wrong). In that sense, student interactions can be taken into account when building relationships between exercises. They use Phi coefficients (a measure of association for two binary variables - in this case the student responses) to build the exercise relationship and then add the textual similarity (provided it is within a certain distance) to derive the relation coefficients else make it zero. It seemed to be a neat innovative approach to me. Initially when reading the paper you may feel they seem to be prioritizing interaction performance over textual similarity but that is not the case. Relationships are only built if the questions are textually near to each other else the weight is zero!\n\nThis exercise-relation information is represented by a set of relation coefficients. As before, they use the self-attention mechanism to learn the attention weights corresponding to the previous interaction for predicting whether a student will provide correct answer to the next exercise. These weights are now augmented by the relation coefficients to provide even better predictions. \n\nAnother interesting point to note would be - They use the textual content of exercises to create simple embeddings for each word (not BERT). They use word2vec and use TEX tokens to transform equations in Maths exams. The exercise embedding is then a weighted combination of the embedding of all words present in the text of the exercise leveraging Smooth Inverse Frequency (SIF). SIF downgrades unimportant words and keeps relevant info only. The distance between exercises is then computed. I think they use cosine similarity.\n\nLastly let us look at the ablations in their study. It is mighty interesting - particularly the one related to exercise relationship building. They compared model performances using a) only textual similarities b) only student interaction performance and c) using only KC tags (so all exercises in a same KC tag are related (weight=1) else they are not (weight=0) with d - the combination of a and b. The observations are as follows - 'c' performs the worst! This is expected as exercises can span KCs. 'b' seems to perform better than 'a' while 'd' is the best. Let us analyze 'b' again. The authors feel that even if textual content of two exercises are not similar, the association of knowledge involved in solving the two exercises could be high. Maybe I am totally wrong, but my understanding is this - With 'b' we are just clustering exercises into groups based on complexity. On one spectrum are exercises which are consistently solved correctly by most students (or is it just 1 student?), the other end of the spectrum consists of exercises that nobody could solve and the rest fall somewhere in between. While this may help in prediction, will this help in pedagogy? However a combination of 'a' and 'b' definitely makes sense and this is what the authors have chosen to go with and this is what gets the best results in the study. All in all, a fascinating paper.\n\n\n\n<br>\nContext-Aware Attentive Knowledge Tracing\n--------------\nhttps://arxiv.org/pdf/2007.12324.pdf.\nAttentive knowledge tracing (AKT) brings in the following innovations: \n- Instead of using raw question and response embeddings, they use context-aware representations of past questions and responses by taking a learner’s practice history into account using a modified version of attention. These modified representations reflect each learner’s actual comprehension of the question and the knowledge they actually acquire, given their personal response history. \n- They modified the scaled dot product attention and have the attention weights decay exponentially based on the context-aware relative distance measure\n- Use the Rasch model to bring out similarity & differences in exercises within the same KC. Thus questions labeled as covering the same concept should not be treated as the same as they could have important individual differences\n\nThere are four components: two self-attentive encoders, one for exercises and one for knowledge acquisition (I believe in plain words this is the interaction-embedding layer), a single attention-based knowledge retriever, and a feed-forward response prediction model.\n- The two self-attentive encoders learn context-aware representations of the exercises and interactions. The first is the exercise encoder, which produces modified, contextualized representations of each exercise, given the sequence of exercises the learner has previously practiced on. The context-aware embedding of each exercise depends on both itself and the past exercises. The second is the knowledge encoder, which produces modified, contextualized representations of the knowledge the learner acquired while responding to past questions. This is the interaction-embedding \n- The knowledge retriever, which retrieves knowledge acquired in the past that is relevant to the current question using an attention mechanism\n- The response prediction model predicts the learner’s response to the current question using the retrieved knowledge\n\nBoth encoders employ the self-attention mechanism. The knowledge retriever, on the other hand, uses the embedding of the current exercise as query, the keys are all the past exercises embeddings and values would be the interaction-embeddings. Notice how keys and values are different unlike in normal attention scenarios. For e.g. SAKT uses exercise embeddings as queries and interaction-embeddings for key as well as values. In AKT, exercise embeddings are used as queries and keys and values are the interaction embeddings. This is an interesting change and they claim that this method is more effective.\n\nThe other interesting bit is the modifications made to the scaled dot product attention that is used in all other models. They say that scaled dot-product attention mechanism is not going to model memories decay effectively nor strengthen recent interactions. They add a multiplicative exponential decay term to the attention scores and a learnable decay rate parameter. So now the attention weights for the current question on a past question depends not only on the similarity between the corresponding query and key, but also on the relative number of time steps between them.\n\nThe source code is available at: https://github.com/arghosh/AKT\n\n\n<br>\nAdditional notes: I later discovered a successor SAINT+ released a couple of months back which further pushes the benchmark by incorporating two additional temporal feature embeddings into the response embeddings: elapsed time, the time taken for a student to answer, and lag time, the time interval between adjacent learning activities. The lag time is used to build in 'forget'fulness (yay!). They show that the elapsed time is better represented as a continuous embedding instead of categorical which is somewhat counter-intuitive. Ideally 3-4 categories based on Z-score should have sufficed but of course this is just my opinion. Saint+ performs slightly better than Saint and is the new benchmark!\n\nAdditional links:\n- https://github.com/thosgt/edm_main_algorithms: Independent repository containing some models like DKT, SAKT etc\n- Similar to above - contains a bunch of models - https://github.com/seewoo5/KT\n- A good introductory video: https://www.youtube.com/watch?v=CzRmRZNpB1Y\n\n\n\n<br><br>\n\nGENERAL OBSERVATIONS:\n-----------------------\nMy domain knowledge has accrued over a grand total of (the past) 36 hours... but perhaps viewing this domain as a rank outsider has its own advantages. My random thoughts (note - this is about the domain not the competition):\n1. 1. While the current focus is on personalized recommendations, could there be plenty to exploit by generalization? Basically clustering of exercises or clustering of KCs, student profiles etc? The results could further aid personalization. But they have other implications. We could have signals that can be leveraged by tutors instead of students. If a set of questions are consistently answererd incorrectly, this could be a signal that the KCs related to those may need to be augmented with better course contents? \n2. While self-attention and transformer-like architectures are all the rage currently, can we leverage recommendation systems? Basically by using k-nearest on student profiles, we can predict how a student(/s) can perform by just looking at a similar profile from a past student who has competed the full course. Then use elements of recommendation systems to design the best course-path for this student?\n3. Can we bring in elements of reinforcement learning?\n4. Can we model the art of forgetting better? We forget a lot in the first 24-48 hours. After that the decay slows down. Beyond a certain point, knowledge does not decay! Also teh decay rate could vary across KCs. Maybe we could 'learn' the decay rate of cluster of KCs?\n5. Can we generate relevant KCs from exercises using DL methods? Some kind of topic modeling? Could this help? In general, question tagging can be further explored. This has nothing to do with individual student responses and is more generic\n6. Relevance of other cues could matter much more. For e.g. engagement of the student during the course is a much better predictor..tracking of eye movements etc?\n7. Performance is so much dependent on external factors. Brilliant students who have done well in the past  may start deteriorating if they have personal issues (death in the family, divorce, drugs, bad company, abuse etc)..The opposite too could happen. The shift does not happen overnight. Can we capture patterns to identify these shifts before they fully materialize? Interventions like counseling could help immensely if done at the right time\n8. How do we spot the super-performers? Most of the work is aimed at improving the scores of the majority. Super-performers may need a different kind of mentoring and customization\n9. How do we identify hidden talents - This is different from super-performers who are all-rounders. Hidden talents may be exceptionally good at only a certain type of exercises or subjects but average otherwise and typically do not stand out in any list. However they have the potential to become true geniuses in the area of their interest if encouraged\n10. Could graph based models be exploited better?\n11. How can we better leverage the attention weights? Luong et al. show that there is a great benefit in passing along the attention weights to the next timestep so that subsequent steps have an idea of what worked in the past. Of course that was in the field of NLP. Would it help here? We have seen in one of the papers above how they relate exercises to one another. In that case, does it makes sense to share past attention weight history at least among similar exercises in future timesteps?\n\nThis article is a result of a quick weekend binge of reading various papers on the subject and my hurried interpretation of them. Please excuse typos/other lapses if any. I will review and correct as we go along.\n\nNote - I have not yet explored the actual Riiid dataset as of now because I wanted to have an unbiased look at existing literature and well, to be frank, my interests lie more in the application of technology rather than the technology  or the competition itself. Next weekend, I will try to create a more specific EDA and analysis on Riiid, but that may not be needed as I already see many great kernels...\n\nThis has gone beyond what I set out to write but the papers were pretty exciting and I couldn't resist sharing my interpretation and thoughts on them. We are also about a month away from the deadline and I hope the couple of code sources I have pointed out will augment the existing Kaggle kernels in a positive way and not cause any type of shakeups.\n\nAll the very best to all participants!\n\nRemaining series:\n\nFrom Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203184 - Hidden features and possible architectures\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185 - Some additional clarifications on SAKT/SAINT\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584 - A small discussion on position embeddings for those interested.\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719 - On lectures, the art of forgetting and why I retired hurt\n",
    "1720516": "What an awesome summary!",
    "1106977": "Thank you, this is great! I dipped my toes into this stuff myself, but then got distracted by other things.\n\nOne interesting note (although maybe only theoretically interesting, in the context of this competition): Intelligent Tutoring Systems as I'm used to them often don't actually use this kind of Deep Knowledge Tracing stuff; in part because they try to do a lot more varied and sometimes subjective things than predicting whether the student will answer the next question correctly (e.g. synthesizing and giving hints, detecting misconceptions, finding bugs in student code...), and in part because they often need a cohesive and interpretable student model which can be reasoned about and applied to all of these things that the ITS needs to do.\n\nOne cool direction that split off from BKT is Learning Factor Analysis (and later also Performance Factor Analysis), based on the idea that the amount of difficulty a student has with a concept (e.g. odds of making an error, or the time it takes to complete a task) decreases according to a power function of the amount of practice they have. This can be used not only tot predict the chance of success, but also to analyze and refine a proposed set of knowledge concepts (e.g. if there are \"humps\" in the learning curve, maybe it's not just one concept...)\n\n[At least one source claims](https://books.google.com/books?id=wfxYPwQ3A20C&lpg=PA103&lr&pg=PA106) (last paragraph of page 106) that this power function model is more consistent with real-world data than the learning curve that comes out of the KT model used in BKT (which is fundamentally a geometric curve, and won't necessarily start looking like a power curve even if you add other things like forgetting on top of it).",
    "1107948": "Greate note. Good to see your notes here again after OpenVaccine :)",
    "1106561": "Hi @Allohvk, thanks for sharing your work and also thanks for the work put into this brilliant write up. Keep up with the Great work and weel done.",
    "1106366": "Wonderful sharing!\nBut I am confused that what is the terminology of KC means(being equal to  interactions or representing addtional information like the lecture information in this competition)? Thank you!",
    "3084786": "I'm feeling so validated after reading the general observations at the end. My thoughts were exactly the same, a recommendation system featuring collaborative filtering can be game changer in personalised learning. \nAnd not just forgetting, the guessing should also be considered. After going through the dataset, I'm positive that a correlation can be found in between the guessing parameter and the time spent. Exciting times!"
  }
}