{
  "id": 203184,
  "title": "Hidden features & thoughts on relevant architectures",
  "url": "/competitions/riiid-test-answer-prediction/discussion/203184",
  "author_name": "",
  "post_date": "2020-12-14T06:18:54.033780100Z",
  "votes": 104,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Note: Part 1 is here: From Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481</a></p>\n<p>Part 3: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185</a> - Some additional clarifications on SAKT/SAINT</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584</a> - A small discussion on position embeddings for those interested.</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719</a> - On lectures, the art of forgetting and why I retired hurt</p>\n<p>For many reasons, this is an interesting competition. The sheer volume of data, the process for submission, the scope for online learning and the imposed constraints - all mimic real-world situations. This is a gold-mine for aspiring data scientists (newbies) because there is so much to learn. Rarely does someone give this kind of rich data. Such learning experiences are unique in case you are not already on this ship - do hop on for a fascinating ride. </p>\n<p>Last weekend, I had shared my interpretation of the general domain and some of the interesting developments in this area recent times. I had also shared my interpretations of the top 5 SOTA papers in this area. Let us now quickly look at the actual data shared by RiiiD and focus on the interesting columns. We will jot down some intuitions along the way and then test those against actual data (I am afraid the 'testing' part will have to wait till next weekend)..<br>\n<br></p>\n<h2>1. Timestamp</h2>\n<p>time in milliseconds between this user interaction and the first event completion from that user</p>\n<ul>\n<li>The most useful feature in my view</li>\n<li>Dividing this by the number of interactions gives an idea of the frequency of student interactions.  I would rate 'continuous learning' as the single biggest prediction of success, if we can find an effective way to capture this</li>\n<li>This will also give the time interval between user interactions which can be a very useful feature. For e.g. many exercises are completed in one sitting. But if there is a gap of more than few hours, then it can be considered to be a next sitting. 12 hour gap would mean that the last exercise was done the previous day and the user is starting fresh for the day - This has lot of interesting connotations - too many to list here. Just to give one example - someone talked about the power of listening to lectures in predicting success. This effect will be all the more pronounced if we determine whether the user listened to the lecture in the same day. After a few hours memory declines by almost a half. Another interesting application of this feature - if the data needs to be crunched, it might not be a bad idea to group all exercises performed in one session as one single exercise and take the average score as the response. How about - Num of interactions the user has had in the past 1 month as another feature - sounds interesting right? Many more..<br>\n<br></li>\n</ul>\n<h2>2. Interactions</h2>\n<ul>\n<li>In RiiiD, interactions can be exercises as well as lectures. content_type_id determines this. This is interesting because I didn't see lectures being captured as interactions in their Saint paper or for that matter any other SOTA paper. How do we make use of this feature? Each lecture is mapped to a tag and it is obvious that if the exercise in question (which needs to be predicted) is associated with a tag and the student has listened to the corresponding lecture in recent past, it would definitely predict success. That could be a separate feature as discussed above. But more importantly we need to check how we can incorporate the lecture interactions within the time-series data as they will be very very useful in determining attention weights at the decoder end. If I am not wrong, most of the SOTA papers currently only have exercises on the encoder side. Plugging in lectures might need some thoughts/additional code. The architecture is not a 1-1 sequence to sequence any more. There are lectures interleaved in between. This is not entirely a new scenario. In NLP world, we have situations where a 20 word sentence on the encoder end is translated to 12 words at the decoder end. With that thought, let us quickly move on…<br>\n<br></li>\n</ul>\n<h2>3. Bundle and task_container_id</h2>\n<p>It took me a while to understand this. The exercises are structured in such a way that sometimes there could be an image or a video displayed and then a set of exercises that follow - all of which are related to the same. This combined set of related exercises is called a bundle… all share the same bundle id. Because exercises can be repeated multiple times over the course of a students interactions, they are identified each time by a unique task_container_id, which reflects the order in which the user starts the tasks. So it is an important positional information and definitely a strong feature. One of the notebooks already proved the intuition that 'when questions are repeated rate, the success rate is higher'. That example showed many cases where questions were repeated even if user got them right the last time around. This could be because the user may be learning a new concept and that new concept is based on 2-3 old concepts. Probably if the user gets the new concept wrong, the AI in the system (yes! I am sure AI would already be built in the system…so in a way, we are modelling an AI model :) Now u see why this competition is so fascinating) would probably display the exercises of the associated concepts (the old ones) to refresh the student memory. Why am I mentioning this? This 'might' be useful in clustering exercises. Clustering exercises is going to play an imp role and while one way to do it would be using tags, another would be using the temporal information (if user answers Q4 and Q7 correctly, he is very likely to get Q17 also correct) and the third way would be the above<br>\n<br></p>\n<h2>4. prior_question_elapsed_time</h2>\n<p>Average time taken to answer each question in the previous exercise bundle. If the exercise was not a part of a bundle but an individual entity, then it is just the time taken to answer the prev question. </p>\n<ul>\n<li>One obvious ramification is that this is null for a user's first question..So we can identify new users this way. UPDATE - Timestamp could be better leveraged for this</li>\n<li>Eric shows in his EDA that this feature does not seem to make much of a difference but he also does not drop this because this shows up as a useful feature in the model. Let us analyse this a bit to see why this could be so (at least theoretically). Firstly, this is the time taken for the prev exercise, not the current one. Most of the studies including SAINT+ seem to suggest that this feature provides a clue to whether the user is going to answer the next exercise correctly. If we take the average time for right responses and compare it with the average time for wrong answers, there does not seem to be any visible difference. But maybe avg time is not the right measure to look at and there is some other relationship which is far more complex. Maybe there is a whole category of students who blaze their way across exams not really caring much about results and excluding this category may show a clear demarcated boundary. For e.g. it would be interesting to see the means of correct/wrong responses if we eliminate all responses (right or wrong) that took less than (say) 5 seconds? Or we could see the plots for all +ve versus -ve interactions where response time is &gt; 2 times the mean for THAT question. It is logical to assume we could see a greater %age of wrong answers in this list. One potential way to make use of this feature is to make categories..circling around the mean time for the exercise. Since we are talking about exercises here, we can take the mean for the full data without worrying about rolling averages. (Of course over time the data might change but that will be too slow and should not affect).<br>\nAnyway this seems like a really useful feature intuitively… Saint+ seems to use a learnable weight to embed this info into a vector and use the continuous feature (not categorical) but I feel using a z-score here and categorising will bring in more benefits)<br>\n<br></li>\n</ul>\n<h2>5. prior_question_had_explanation: </h2>\n<p>This is defined as \"Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle\". <br>\nThe above seems to indicate that after answering a question, the student sees the correct response and an explanation for the same. The user can choose to read this explanation. Possibly this could be a hyperlink given below the correct answer. If the user clicks the hyperlink, then probably this boolean flag is recorded. This immediately strikes me as a very good feature. If the user is correct, then he already knows about the concept and is probably further strengthening his memory. Sincere students who were wrong, would definitely want to see where they are going wrong and correct themselves. IF an user gives a wrong answer AND this flag is FALSE is almost certainly going to get subsequent exercises (related to this topic) as incorrect. Firstly, I hope my understanding of this feature is correct. Please correct me if I am wrong. The next interesting question is how do we vectorize this interesting feature. We can retain it as it is and let the model automatically discover its relation with answer correctness in predicting the labels or we can experiment if adding this type of OHE helps:<br>\n1- Answer correct and flag is false: 1 0 0<br>\n2- Answer correct and flag is true: 0 1 0 <br>\n3- Answer wrong and flag is false: 0 0 1<br>\n4- Answer wrong and flag is true: 1 0 0<br>\nUPDATE: Yana has shown that field is mainly for diagnostic questions. If so then it is better to discard this field for predictions</p>\n<p>Now let us look at the interesting fields in the other 2 datasets:</p>\n<ul>\n<li>There are no question texts given. So relationship between questions can be only gauged by - (a) do they share the same bundle (b) do they share same or similar tags (c) LSTM relation and if possible the 4th clustering mechanism we discussed in one of the points above</li>\n<li>Tags seems to be loosely mapped to skill tags or knowledge concepts. They are annotated by domain experts in the EDNET database and hence can be used to build relationship between questions. I feel this is an important aspect to be analysed. Yana has done some excellent work in analysing tags and their relation to questions. This is something we need to definitely think about in great detail when building the model</li>\n</ul>\n<p>Then we have lectures - each lecture is mapped to 1 tag and most likely strengthens the students knowledge for that tag.</p>\n<p>Lastly we have Parts - This may be too generic but could still be useful. We could take the student average correctness for each Part rather than their performance as a whole. This is because students may be stronger in certain 'Part's of the exam compared to other parts. Kkiller took the harmonic mean of student correctness and exercise correctness and gets around 74%. I am sure this can be improved by taking student correctness for each of the 6-7 parts of the exam. </p>\n<p>Let us see what else is important:</p>\n<ul>\n<li>'Consistency' is the best indicator of success. Nothing else beats it (IMHO). Students who regularly login and spend time can be expected to have better scores. So frequency of interactions is a good measure. Num of interactions of the student per week seems like an excellent feature to me to test. This is also a nice way to cluster students. Clustering student is going to be very important</li>\n<li>Users with very large gaps (&gt; 1 month?) between interactions can be marked as new users? Can we drop these records considering the processing limitations? The challenge is the skills acquired quickly revert to that of a new user but the aptitude of the student will not change. So there is definitely useful information there. Maybe we could drop it from the NN ensemble and retain it for non-NN ones?</li>\n<li>Remove users with limited interactions (&lt;30..I believe this is the diagnostic test count?). Will this save a lot of space? Intuitively it should since there are many folks who join in and drop out at an early stage</li>\n<li>What is the typical flow of bundle IDs as the student flows across the course from beg to end? This is useful in itself and can also be used for clustering students</li>\n<li>Lecture frequency of student? Per week? Could be another feature apart from the interaction frequency</li>\n<li>Remove outlier students - bottom n top</li>\n<li>Needless to say the part-based student rolling avg and the exercise whole average are very important features</li>\n<li>We saw exercises can repeat itself based on student's progress. This could be an indicator to model 'forgetfulness'. The more instances of repeated questions we see, the more we can conclude that the student is of a forgetful nature. Possibly this can be an input to decay attention weights at a faster rate</li>\n<li>Handling cold starts - In the interest of space and time, I will not cover this. Please do refer to my recent discussions with kkiller on this</li>\n<li>The SAKT paper makes an interesting observation and I quote - \"\"We can see that the based on the attention weights, we are able to achieve the perfect clustering of the exercise tags based on the hidden concepts from which they are derived\". Reversing this sentence we can conclude that tags are going to play a key role in determining attention weights. Use them wisely as a key feature esp paying attention to how you design the encoder/decoder inputs</li>\n</ul>\n<p>We have covered all possible features that we could intuitively think of. For more we need to do an elaborate EDA. We also need to explore the relationship between exercise-lecture-tags in a deeper way.</p>\n<p>Lastly, let us come to the time-series aspect of the data and this could guide us regarding the possible architectures. It was a bit hard for me to intuitively grasp the architectures in the SOTA papers, so I hope this below discussion helps newbies like me who are the intended audience. Surprisingly I couldn't find too many explanations of using transformer like architectures in Time-series situations..beyond the actual papers themselves. My focus is on the intuitive aspects of the architecture.</p>\n<p>It is possible to fit a Time-series model into a encoder, decoder type of architecture &amp; by that extension into a transformer architecture. In case of a simple sequence classification (like the covid mRNA Kaggle challenge), we needed only the encoder part of the transformer. But in case we want to predict a series (say the next 1 month of the financial index), then this can be a seq-seq exercise. In case we want to predict the entire 1 month at one shot, we can still restrict ourselves to the encoder architecture but in case we want to predict something and then use the predicted value to predict the next day's index we will need a decoder as well. The decoder uses the 'just decoded value' from the previous step and the prev hidden state (in case of LSTM) to predict the next word. In case we add attention, it will also look at the encoder outputs to determine which of the outputs could have a strong 'say' in predicting the label and give weightage to it accordingly. Discussing attention or transformer architecture itself is not the scope here, but those interested for a background can refer to: <a href=\"https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e\" target=\"_blank\">https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e</a>.</p>\n<p>There are minor differences between NLP and Time-series Seq to Seq and it would be good to analyse those because most of us are familiar with NLP Seq-Seq architectures and we just need to 'adapt' that model to other time-series situations. In NLP there is a limit to which the sentence or the para or even the document can stretch backwards. In Time-series, we can have financial indices that are a century old. So determining a 'window' where we are going to focus our 'attention' is going to be the key. Local and soft attention techniques are better here. In summary, to produce a labeled dataset, we need to use a fixed-length sliding time window approach to construct X, Y pairs for model training and evaluation. Having such a window means we may need to pad up inputs with dummy data. Could the diagnostic set of 30 odd questions act like some sort of a padding? This is just a thought. Obviously we have to do look-ahead masking and one-position offset between the decoder input and target output in the decoder to ensure that prediction of a time series data point will only depend on previous data points. Most importantly, for a time-series model, the input is a vector of continuous numbers already. There is no need for an NLP like embedding layer. HOWEVER, we need to ensure that the data is richly represented. We could make use of a (learnable) weight matrix to transform the one-dimension vector into a certain dimension vector. Alternatively we could use self-attention to embed the input. Self-attention with all the fancy multiheads, residuals and all the standard tricks result in capturing the richness of the data well. Using self-attention will require positional encodings as the embedding has no knowledge of sequence order. There are some standard well-defined hacks to incorporate this and we will not go into the details. This is a quick generic summary of how a time-series data can leverage the transformer architecture. But I felt there are two unanswered questions regarding this. We will come to that later but first let us look at the specificities of this competition.</p>\n<p>Let us assume there is only one student. We have a sequence of all her interactions 'I' from I1 to I1000. She viewed 100 lectures and solved 900 exercises. She has reached the fag end of the course now and another 100 exercises are left. We have to predict how she performs in these 100 exercises. I assume there could be a couple of lectures embedded here also, let us just ignore those for now as we dont have to predict those. The way we could train now is to break the data into time chunks. We could have Q1 to Q1000 fed to the encoder. The decoder will consist of responses to Q2 to Q1000. Let us now predict Q2 using just the questions Q1, Q2 and Q1's response by the student. Now you may see the justification for masks on both the encoder and the decoder. When predicting Q2, we just know the question Q1, the question Q2 and we just know the student's response to Q1 (and of course all the associated meta data). Possibly this sentence makes better sense now - \"The key is to do look-ahead masking and one-position offset between the decoder input and target output in the decoder to ensure that prediction of a time series data point will only depend on previous data points\".</p>\n<p>Now repeat the above, this time the input on encoder side will be the questions Q1, Q2, Q3 and the decoder side will be the response to Q1, Q2 and our goal is to predict the response for Q3. Similarly assume we have reached Q100. We notice a pattern now. The predictions of certain question responses depend on the response to some of the previous questions. Maybe it is because some questions are repeated or maybe it is because questions with same tags are somewhat similar and if a student has answered one she will almost certainly answer the other. So your model determines (after going thru' many such student interactions) that Q100's response depends on the student's past response to Q78, Q99 and Q46. This is attention logic at work. The decoder is paying attention to all the states of the encoder output (Q1 to Q100) while predicting Q100th response (in addition to considering the decoder inputs. There will also be student specific temporal information which will be captured by the model and thus the model is trained. Every additional student contributes to strengthening the models weights and biases and now the model is ready to predict. Notice that for the 1000th question, it will be computationally expensive to go all the way back to Q1, so we take a sliding window and carve out time slices of student interaction data. If the window is 100, we will have Q1-Q100 questions and responses as 1 sliced row input and Q101 to be predicted. Then we will have Q2-Q101 and Q102 to be predicted…and all the way till Q901-Q1000 questions and responses as inputs and Q 1001 to be predicted.</p>\n<p>So far so good. But where do we feed what? If we take the case of a language translator for whom the original encoder-decoder arch was conceived, the source language inputs are fed to the encoder and the destination language to the decoder. The situation in time-series  is such that we can feed the inputs to the encoder and the outputs to the decoder. So Q1-Q100 goes to encoder and also Q101. The responses of Q1 to Q100 goes to decoder and now we have to predict the response for Q101. We can assume that each response is like the translated sentence of the concerned question.</p>\n<p>So all questions and question related information could be fed to the encoder. This includes the exercises, the tags, the task_container_id's and other derived features related to QUESTIONS etc etc. The response related meta data along with the response themselves go into the decoder. This includes data like prev question explanation, elapsed time, lag time, saw lecture and all derived features related to response….basically everything other than what is put in the encoder. I like this style of segregation between encoder and decoder and I believe this is what is done in the SAINT paper. HOWEVER this need not be the only solution. Some papers have fed the coupled interactions (exercise+response) on decoder end. Others have tried many other things. This could be an area that will need a bit of experimentation. I feel it might not altogether be a bad idea to add tags meta-data in the responses?</p>\n<p>The rest of the paraphernalia is the same as in any standard transformer architecture and there are source codes in every framework for the same which you can re-use.</p>\n<p>Other possible considerations </p>\n<ul>\n<li>For forming language representations focusing on the closest word makes a lot of sense. However, this is much more variable with time series data, in certain time series sequences causality can come from steps much further back. Additionally, in the real-world scenario, the exercises which occur close in the sequence tend to belong to the same concept. Thus, we expect that the attention weights biased towards the exercises that occur recently in the interaction sequence.Can we have multiple windows into which to look into? Or when we are choosing time slices can we pick up the last time slice based on the tags?</li>\n<li>The attention weights can be augmented by having a set of relation coefficients between each exercise. Shalini Pandey's latest SOTA paper does this, but there she uses text information from questions to build relationships. Here we could use tags.</li>\n</ul>\n<p>Lastly, coming to the two things which I couldn't understand intuitively. I can see the following issues when using transformers (or encoder-decoder) on time-series data and I hope someone can guide me on where my understanding is incorrect:</p>\n<p>Issue 1: In NLP every translation has one and only one correct answer. While the inputs may change to form different sequences, the output will always be the same for a given sequence of words. The model converges towards that output…it cannot deviate astray with more data. In a Time-series data, there could be different outputs for a given input sequence. In this case it is based on the student profile. The two problem statements are similar but not the same and using the same architecture for both may not be optimal</p>\n<p>Issue 2: Somewhat linked to Issue 1. Let us think of it in terms of financial indices. We want to determine the NYSE index future. Let us say NYSE index is represented by a proportion of every stock traded on NYSE. One option is to take the NYSE index itself over time and only use that to model. The other option is to look at all the constituents. However modelling based on the constituents could be tricky. Each constituent follows a different path. Every stock listed in the NYSE stock exchange could behave differently. Using their daily data to predict the NYSE index futures is at best a tricky job. Considering the problem at hand, we see each different student as a individual stock being traded on NYSE. There is no equivalent of NYSE index here. Instead we want to predict how these stocks continue to move and also we want to predict how new stocks added to the stock exchange would move. We do have lots and lots of data but doesn't this introduce tons of noise also? One obvious solution is to cluster the constituents. So in case of the NYSE scenario, we could have a bunch of stocks getting represented by a single index - the tech index. We could have brick n mortar companies as another cluster, retail, housing and form dozen odd clusters. Calculate the movements for each cluster separately…as they could have different weights and balances. Now we are left with time-series data of a dozen odd indices - all of which will help determine the NYSE index futures more predictably. If we apply this sort of a scenario to RiiiD, this means that clusters of students need to be created and then it makes sense to have separate sets of weights and biases for each category? Let existing student continue in the cluster they already are in and the new students are moved to one of the clusters. But I haven't seen this kind of approach in any time-series model leveraging encpder/decoders.</p>\n<p>We still haven't talked of validation strategies, thought about what could be the best combination of keys queries and values or the importance of online learning in this competition or usage of reductions and many others. But we have covered a fair bit of theory otherwise. While kernels are shared abundantly, I felt some theoretical discussion on the features and architectures could help newcomers like me and hence quickly jotted a few points that stuck me as I went thru' a few of the beautiful kernels.</p>\n<p>In Silogram's golden words, this competition is not just a test of AI skills, but is also a test of our engineering skills. How creatively we make use of the imposed constraints will decide the winner. With unlimited processing power and memory, the NN models would have won hands down but with the imposed constraints a blend of NN and non-NN models may hold the key.</p>\n<p>As with all features and architectures, there are risks/rewards. </p>\n<p>Choose wisely. All the very best to all!</p>",
  "messages": [
    {
      "id": "1111913",
      "postDate": "12/14/2020 06:18:54",
      "content": "<p>Note: Part 1 is here: From Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481</a></p>\n<p>Part 3: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185</a> - Some additional clarifications on SAKT/SAINT</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584</a> - A small discussion on position embeddings for those interested.</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719</a> - On lectures, the art of forgetting and why I retired hurt</p>\n<p>For many reasons, this is an interesting competition. The sheer volume of data, the process for submission, the scope for online learning and the imposed constraints - all mimic real-world situations. This is a gold-mine for aspiring data scientists (newbies) because there is so much to learn. Rarely does someone give this kind of rich data. Such learning experiences are unique in case you are not already on this ship - do hop on for a fascinating ride. </p>\n<p>Last weekend, I had shared my interpretation of the general domain and some of the interesting developments in this area recent times. I had also shared my interpretations of the top 5 SOTA papers in this area. Let us now quickly look at the actual data shared by RiiiD and focus on the interesting columns. We will jot down some intuitions along the way and then test those against actual data (I am afraid the 'testing' part will have to wait till next weekend)..<br>\n<br></p>\n<h2>1. Timestamp</h2>\n<p>time in milliseconds between this user interaction and the first event completion from that user</p>\n<ul>\n<li>The most useful feature in my view</li>\n<li>Dividing this by the number of interactions gives an idea of the frequency of student interactions.  I would rate 'continuous learning' as the single biggest prediction of success, if we can find an effective way to capture this</li>\n<li>This will also give the time interval between user interactions which can be a very useful feature. For e.g. many exercises are completed in one sitting. But if there is a gap of more than few hours, then it can be considered to be a next sitting. 12 hour gap would mean that the last exercise was done the previous day and the user is starting fresh for the day - This has lot of interesting connotations - too many to list here. Just to give one example - someone talked about the power of listening to lectures in predicting success. This effect will be all the more pronounced if we determine whether the user listened to the lecture in the same day. After a few hours memory declines by almost a half. Another interesting application of this feature - if the data needs to be crunched, it might not be a bad idea to group all exercises performed in one session as one single exercise and take the average score as the response. How about - Num of interactions the user has had in the past 1 month as another feature - sounds interesting right? Many more..<br>\n<br></li>\n</ul>\n<h2>2. Interactions</h2>\n<ul>\n<li>In RiiiD, interactions can be exercises as well as lectures. content_type_id determines this. This is interesting because I didn't see lectures being captured as interactions in their Saint paper or for that matter any other SOTA paper. How do we make use of this feature? Each lecture is mapped to a tag and it is obvious that if the exercise in question (which needs to be predicted) is associated with a tag and the student has listened to the corresponding lecture in recent past, it would definitely predict success. That could be a separate feature as discussed above. But more importantly we need to check how we can incorporate the lecture interactions within the time-series data as they will be very very useful in determining attention weights at the decoder end. If I am not wrong, most of the SOTA papers currently only have exercises on the encoder side. Plugging in lectures might need some thoughts/additional code. The architecture is not a 1-1 sequence to sequence any more. There are lectures interleaved in between. This is not entirely a new scenario. In NLP world, we have situations where a 20 word sentence on the encoder end is translated to 12 words at the decoder end. With that thought, let us quickly move on…<br>\n<br></li>\n</ul>\n<h2>3. Bundle and task_container_id</h2>\n<p>It took me a while to understand this. The exercises are structured in such a way that sometimes there could be an image or a video displayed and then a set of exercises that follow - all of which are related to the same. This combined set of related exercises is called a bundle… all share the same bundle id. Because exercises can be repeated multiple times over the course of a students interactions, they are identified each time by a unique task_container_id, which reflects the order in which the user starts the tasks. So it is an important positional information and definitely a strong feature. One of the notebooks already proved the intuition that 'when questions are repeated rate, the success rate is higher'. That example showed many cases where questions were repeated even if user got them right the last time around. This could be because the user may be learning a new concept and that new concept is based on 2-3 old concepts. Probably if the user gets the new concept wrong, the AI in the system (yes! I am sure AI would already be built in the system…so in a way, we are modelling an AI model :) Now u see why this competition is so fascinating) would probably display the exercises of the associated concepts (the old ones) to refresh the student memory. Why am I mentioning this? This 'might' be useful in clustering exercises. Clustering exercises is going to play an imp role and while one way to do it would be using tags, another would be using the temporal information (if user answers Q4 and Q7 correctly, he is very likely to get Q17 also correct) and the third way would be the above<br>\n<br></p>\n<h2>4. prior_question_elapsed_time</h2>\n<p>Average time taken to answer each question in the previous exercise bundle. If the exercise was not a part of a bundle but an individual entity, then it is just the time taken to answer the prev question. </p>\n<ul>\n<li>One obvious ramification is that this is null for a user's first question..So we can identify new users this way. UPDATE - Timestamp could be better leveraged for this</li>\n<li>Eric shows in his EDA that this feature does not seem to make much of a difference but he also does not drop this because this shows up as a useful feature in the model. Let us analyse this a bit to see why this could be so (at least theoretically). Firstly, this is the time taken for the prev exercise, not the current one. Most of the studies including SAINT+ seem to suggest that this feature provides a clue to whether the user is going to answer the next exercise correctly. If we take the average time for right responses and compare it with the average time for wrong answers, there does not seem to be any visible difference. But maybe avg time is not the right measure to look at and there is some other relationship which is far more complex. Maybe there is a whole category of students who blaze their way across exams not really caring much about results and excluding this category may show a clear demarcated boundary. For e.g. it would be interesting to see the means of correct/wrong responses if we eliminate all responses (right or wrong) that took less than (say) 5 seconds? Or we could see the plots for all +ve versus -ve interactions where response time is &gt; 2 times the mean for THAT question. It is logical to assume we could see a greater %age of wrong answers in this list. One potential way to make use of this feature is to make categories..circling around the mean time for the exercise. Since we are talking about exercises here, we can take the mean for the full data without worrying about rolling averages. (Of course over time the data might change but that will be too slow and should not affect).<br>\nAnyway this seems like a really useful feature intuitively… Saint+ seems to use a learnable weight to embed this info into a vector and use the continuous feature (not categorical) but I feel using a z-score here and categorising will bring in more benefits)<br>\n<br></li>\n</ul>\n<h2>5. prior_question_had_explanation: </h2>\n<p>This is defined as \"Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle\". <br>\nThe above seems to indicate that after answering a question, the student sees the correct response and an explanation for the same. The user can choose to read this explanation. Possibly this could be a hyperlink given below the correct answer. If the user clicks the hyperlink, then probably this boolean flag is recorded. This immediately strikes me as a very good feature. If the user is correct, then he already knows about the concept and is probably further strengthening his memory. Sincere students who were wrong, would definitely want to see where they are going wrong and correct themselves. IF an user gives a wrong answer AND this flag is FALSE is almost certainly going to get subsequent exercises (related to this topic) as incorrect. Firstly, I hope my understanding of this feature is correct. Please correct me if I am wrong. The next interesting question is how do we vectorize this interesting feature. We can retain it as it is and let the model automatically discover its relation with answer correctness in predicting the labels or we can experiment if adding this type of OHE helps:<br>\n1- Answer correct and flag is false: 1 0 0<br>\n2- Answer correct and flag is true: 0 1 0 <br>\n3- Answer wrong and flag is false: 0 0 1<br>\n4- Answer wrong and flag is true: 1 0 0<br>\nUPDATE: Yana has shown that field is mainly for diagnostic questions. If so then it is better to discard this field for predictions</p>\n<p>Now let us look at the interesting fields in the other 2 datasets:</p>\n<ul>\n<li>There are no question texts given. So relationship between questions can be only gauged by - (a) do they share the same bundle (b) do they share same or similar tags (c) LSTM relation and if possible the 4th clustering mechanism we discussed in one of the points above</li>\n<li>Tags seems to be loosely mapped to skill tags or knowledge concepts. They are annotated by domain experts in the EDNET database and hence can be used to build relationship between questions. I feel this is an important aspect to be analysed. Yana has done some excellent work in analysing tags and their relation to questions. This is something we need to definitely think about in great detail when building the model</li>\n</ul>\n<p>Then we have lectures - each lecture is mapped to 1 tag and most likely strengthens the students knowledge for that tag.</p>\n<p>Lastly we have Parts - This may be too generic but could still be useful. We could take the student average correctness for each Part rather than their performance as a whole. This is because students may be stronger in certain 'Part's of the exam compared to other parts. Kkiller took the harmonic mean of student correctness and exercise correctness and gets around 74%. I am sure this can be improved by taking student correctness for each of the 6-7 parts of the exam. </p>\n<p>Let us see what else is important:</p>\n<ul>\n<li>'Consistency' is the best indicator of success. Nothing else beats it (IMHO). Students who regularly login and spend time can be expected to have better scores. So frequency of interactions is a good measure. Num of interactions of the student per week seems like an excellent feature to me to test. This is also a nice way to cluster students. Clustering student is going to be very important</li>\n<li>Users with very large gaps (&gt; 1 month?) between interactions can be marked as new users? Can we drop these records considering the processing limitations? The challenge is the skills acquired quickly revert to that of a new user but the aptitude of the student will not change. So there is definitely useful information there. Maybe we could drop it from the NN ensemble and retain it for non-NN ones?</li>\n<li>Remove users with limited interactions (&lt;30..I believe this is the diagnostic test count?). Will this save a lot of space? Intuitively it should since there are many folks who join in and drop out at an early stage</li>\n<li>What is the typical flow of bundle IDs as the student flows across the course from beg to end? This is useful in itself and can also be used for clustering students</li>\n<li>Lecture frequency of student? Per week? Could be another feature apart from the interaction frequency</li>\n<li>Remove outlier students - bottom n top</li>\n<li>Needless to say the part-based student rolling avg and the exercise whole average are very important features</li>\n<li>We saw exercises can repeat itself based on student's progress. This could be an indicator to model 'forgetfulness'. The more instances of repeated questions we see, the more we can conclude that the student is of a forgetful nature. Possibly this can be an input to decay attention weights at a faster rate</li>\n<li>Handling cold starts - In the interest of space and time, I will not cover this. Please do refer to my recent discussions with kkiller on this</li>\n<li>The SAKT paper makes an interesting observation and I quote - \"\"We can see that the based on the attention weights, we are able to achieve the perfect clustering of the exercise tags based on the hidden concepts from which they are derived\". Reversing this sentence we can conclude that tags are going to play a key role in determining attention weights. Use them wisely as a key feature esp paying attention to how you design the encoder/decoder inputs</li>\n</ul>\n<p>We have covered all possible features that we could intuitively think of. For more we need to do an elaborate EDA. We also need to explore the relationship between exercise-lecture-tags in a deeper way.</p>\n<p>Lastly, let us come to the time-series aspect of the data and this could guide us regarding the possible architectures. It was a bit hard for me to intuitively grasp the architectures in the SOTA papers, so I hope this below discussion helps newbies like me who are the intended audience. Surprisingly I couldn't find too many explanations of using transformer like architectures in Time-series situations..beyond the actual papers themselves. My focus is on the intuitive aspects of the architecture.</p>\n<p>It is possible to fit a Time-series model into a encoder, decoder type of architecture &amp; by that extension into a transformer architecture. In case of a simple sequence classification (like the covid mRNA Kaggle challenge), we needed only the encoder part of the transformer. But in case we want to predict a series (say the next 1 month of the financial index), then this can be a seq-seq exercise. In case we want to predict the entire 1 month at one shot, we can still restrict ourselves to the encoder architecture but in case we want to predict something and then use the predicted value to predict the next day's index we will need a decoder as well. The decoder uses the 'just decoded value' from the previous step and the prev hidden state (in case of LSTM) to predict the next word. In case we add attention, it will also look at the encoder outputs to determine which of the outputs could have a strong 'say' in predicting the label and give weightage to it accordingly. Discussing attention or transformer architecture itself is not the scope here, but those interested for a background can refer to: <a href=\"https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e\" target=\"_blank\">https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e</a>.</p>\n<p>There are minor differences between NLP and Time-series Seq to Seq and it would be good to analyse those because most of us are familiar with NLP Seq-Seq architectures and we just need to 'adapt' that model to other time-series situations. In NLP there is a limit to which the sentence or the para or even the document can stretch backwards. In Time-series, we can have financial indices that are a century old. So determining a 'window' where we are going to focus our 'attention' is going to be the key. Local and soft attention techniques are better here. In summary, to produce a labeled dataset, we need to use a fixed-length sliding time window approach to construct X, Y pairs for model training and evaluation. Having such a window means we may need to pad up inputs with dummy data. Could the diagnostic set of 30 odd questions act like some sort of a padding? This is just a thought. Obviously we have to do look-ahead masking and one-position offset between the decoder input and target output in the decoder to ensure that prediction of a time series data point will only depend on previous data points. Most importantly, for a time-series model, the input is a vector of continuous numbers already. There is no need for an NLP like embedding layer. HOWEVER, we need to ensure that the data is richly represented. We could make use of a (learnable) weight matrix to transform the one-dimension vector into a certain dimension vector. Alternatively we could use self-attention to embed the input. Self-attention with all the fancy multiheads, residuals and all the standard tricks result in capturing the richness of the data well. Using self-attention will require positional encodings as the embedding has no knowledge of sequence order. There are some standard well-defined hacks to incorporate this and we will not go into the details. This is a quick generic summary of how a time-series data can leverage the transformer architecture. But I felt there are two unanswered questions regarding this. We will come to that later but first let us look at the specificities of this competition.</p>\n<p>Let us assume there is only one student. We have a sequence of all her interactions 'I' from I1 to I1000. She viewed 100 lectures and solved 900 exercises. She has reached the fag end of the course now and another 100 exercises are left. We have to predict how she performs in these 100 exercises. I assume there could be a couple of lectures embedded here also, let us just ignore those for now as we dont have to predict those. The way we could train now is to break the data into time chunks. We could have Q1 to Q1000 fed to the encoder. The decoder will consist of responses to Q2 to Q1000. Let us now predict Q2 using just the questions Q1, Q2 and Q1's response by the student. Now you may see the justification for masks on both the encoder and the decoder. When predicting Q2, we just know the question Q1, the question Q2 and we just know the student's response to Q1 (and of course all the associated meta data). Possibly this sentence makes better sense now - \"The key is to do look-ahead masking and one-position offset between the decoder input and target output in the decoder to ensure that prediction of a time series data point will only depend on previous data points\".</p>\n<p>Now repeat the above, this time the input on encoder side will be the questions Q1, Q2, Q3 and the decoder side will be the response to Q1, Q2 and our goal is to predict the response for Q3. Similarly assume we have reached Q100. We notice a pattern now. The predictions of certain question responses depend on the response to some of the previous questions. Maybe it is because some questions are repeated or maybe it is because questions with same tags are somewhat similar and if a student has answered one she will almost certainly answer the other. So your model determines (after going thru' many such student interactions) that Q100's response depends on the student's past response to Q78, Q99 and Q46. This is attention logic at work. The decoder is paying attention to all the states of the encoder output (Q1 to Q100) while predicting Q100th response (in addition to considering the decoder inputs. There will also be student specific temporal information which will be captured by the model and thus the model is trained. Every additional student contributes to strengthening the models weights and biases and now the model is ready to predict. Notice that for the 1000th question, it will be computationally expensive to go all the way back to Q1, so we take a sliding window and carve out time slices of student interaction data. If the window is 100, we will have Q1-Q100 questions and responses as 1 sliced row input and Q101 to be predicted. Then we will have Q2-Q101 and Q102 to be predicted…and all the way till Q901-Q1000 questions and responses as inputs and Q 1001 to be predicted.</p>\n<p>So far so good. But where do we feed what? If we take the case of a language translator for whom the original encoder-decoder arch was conceived, the source language inputs are fed to the encoder and the destination language to the decoder. The situation in time-series  is such that we can feed the inputs to the encoder and the outputs to the decoder. So Q1-Q100 goes to encoder and also Q101. The responses of Q1 to Q100 goes to decoder and now we have to predict the response for Q101. We can assume that each response is like the translated sentence of the concerned question.</p>\n<p>So all questions and question related information could be fed to the encoder. This includes the exercises, the tags, the task_container_id's and other derived features related to QUESTIONS etc etc. The response related meta data along with the response themselves go into the decoder. This includes data like prev question explanation, elapsed time, lag time, saw lecture and all derived features related to response….basically everything other than what is put in the encoder. I like this style of segregation between encoder and decoder and I believe this is what is done in the SAINT paper. HOWEVER this need not be the only solution. Some papers have fed the coupled interactions (exercise+response) on decoder end. Others have tried many other things. This could be an area that will need a bit of experimentation. I feel it might not altogether be a bad idea to add tags meta-data in the responses?</p>\n<p>The rest of the paraphernalia is the same as in any standard transformer architecture and there are source codes in every framework for the same which you can re-use.</p>\n<p>Other possible considerations </p>\n<ul>\n<li>For forming language representations focusing on the closest word makes a lot of sense. However, this is much more variable with time series data, in certain time series sequences causality can come from steps much further back. Additionally, in the real-world scenario, the exercises which occur close in the sequence tend to belong to the same concept. Thus, we expect that the attention weights biased towards the exercises that occur recently in the interaction sequence.Can we have multiple windows into which to look into? Or when we are choosing time slices can we pick up the last time slice based on the tags?</li>\n<li>The attention weights can be augmented by having a set of relation coefficients between each exercise. Shalini Pandey's latest SOTA paper does this, but there she uses text information from questions to build relationships. Here we could use tags.</li>\n</ul>\n<p>Lastly, coming to the two things which I couldn't understand intuitively. I can see the following issues when using transformers (or encoder-decoder) on time-series data and I hope someone can guide me on where my understanding is incorrect:</p>\n<p>Issue 1: In NLP every translation has one and only one correct answer. While the inputs may change to form different sequences, the output will always be the same for a given sequence of words. The model converges towards that output…it cannot deviate astray with more data. In a Time-series data, there could be different outputs for a given input sequence. In this case it is based on the student profile. The two problem statements are similar but not the same and using the same architecture for both may not be optimal</p>\n<p>Issue 2: Somewhat linked to Issue 1. Let us think of it in terms of financial indices. We want to determine the NYSE index future. Let us say NYSE index is represented by a proportion of every stock traded on NYSE. One option is to take the NYSE index itself over time and only use that to model. The other option is to look at all the constituents. However modelling based on the constituents could be tricky. Each constituent follows a different path. Every stock listed in the NYSE stock exchange could behave differently. Using their daily data to predict the NYSE index futures is at best a tricky job. Considering the problem at hand, we see each different student as a individual stock being traded on NYSE. There is no equivalent of NYSE index here. Instead we want to predict how these stocks continue to move and also we want to predict how new stocks added to the stock exchange would move. We do have lots and lots of data but doesn't this introduce tons of noise also? One obvious solution is to cluster the constituents. So in case of the NYSE scenario, we could have a bunch of stocks getting represented by a single index - the tech index. We could have brick n mortar companies as another cluster, retail, housing and form dozen odd clusters. Calculate the movements for each cluster separately…as they could have different weights and balances. Now we are left with time-series data of a dozen odd indices - all of which will help determine the NYSE index futures more predictably. If we apply this sort of a scenario to RiiiD, this means that clusters of students need to be created and then it makes sense to have separate sets of weights and biases for each category? Let existing student continue in the cluster they already are in and the new students are moved to one of the clusters. But I haven't seen this kind of approach in any time-series model leveraging encpder/decoders.</p>\n<p>We still haven't talked of validation strategies, thought about what could be the best combination of keys queries and values or the importance of online learning in this competition or usage of reductions and many others. But we have covered a fair bit of theory otherwise. While kernels are shared abundantly, I felt some theoretical discussion on the features and architectures could help newcomers like me and hence quickly jotted a few points that stuck me as I went thru' a few of the beautiful kernels.</p>\n<p>In Silogram's golden words, this competition is not just a test of AI skills, but is also a test of our engineering skills. How creatively we make use of the imposed constraints will decide the winner. With unlimited processing power and memory, the NN models would have won hands down but with the imposed constraints a blend of NN and non-NN models may hold the key.</p>\n<p>As with all features and architectures, there are risks/rewards. </p>\n<p>Choose wisely. All the very best to all!</p>",
      "rawMarkdown": "Note: Part 1 is here: From Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\n\nPart 3: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185 - Some additional clarifications on SAKT/SAINT\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584 - A small discussion on position embeddings for those interested.\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719 - On lectures, the art of forgetting and why I retired hurt\n\nFor many reasons, this is an interesting competition. The sheer volume of data, the process for submission, the scope for online learning and the imposed constraints - all mimic real-world situations. This is a gold-mine for aspiring data scientists (newbies) because there is so much to learn. Rarely does someone give this kind of rich data. Such learning experiences are unique in case you are not already on this ship - do hop on for a fascinating ride. \n\nLast weekend, I had shared my interpretation of the general domain and some of the interesting developments in this area recent times. I had also shared my interpretations of the top 5 SOTA papers in this area. Let us now quickly look at the actual data shared by RiiiD and focus on the interesting columns. We will jot down some intuitions along the way and then test those against actual data (I am afraid the 'testing' part will have to wait till next weekend)..\n<br>\n1. Timestamp\n--------------------\ntime in milliseconds between this user interaction and the first event completion from that user\n- The most useful feature in my view\n- Dividing this by the number of interactions gives an idea of the frequency of student interactions.  I would rate 'continuous learning' as the single biggest prediction of success, if we can find an effective way to capture this\n- This will also give the time interval between user interactions which can be a very useful feature. For e.g. many exercises are completed in one sitting. But if there is a gap of more than few hours, then it can be considered to be a next sitting. 12 hour gap would mean that the last exercise was done the previous day and the user is starting fresh for the day - This has lot of interesting connotations - too many to list here. Just to give one example - someone talked about the power of listening to lectures in predicting success. This effect will be all the more pronounced if we determine whether the user listened to the lecture in the same day. After a few hours memory declines by almost a half. Another interesting application of this feature - if the data needs to be crunched, it might not be a bad idea to group all exercises performed in one session as one single exercise and take the average score as the response. How about - Num of interactions the user has had in the past 1 month as another feature - sounds interesting right? Many more..\n<br>\n2. Interactions\n------------------------\n - In RiiiD, interactions can be exercises as well as lectures. content_type_id determines this. This is interesting because I didn't see lectures being captured as interactions in their Saint paper or for that matter any other SOTA paper. How do we make use of this feature? Each lecture is mapped to a tag and it is obvious that if the exercise in question (which needs to be predicted) is associated with a tag and the student has listened to the corresponding lecture in recent past, it would definitely predict success. That could be a separate feature as discussed above. But more importantly we need to check how we can incorporate the lecture interactions within the time-series data as they will be very very useful in determining attention weights at the decoder end. If I am not wrong, most of the SOTA papers currently only have exercises on the encoder side. Plugging in lectures might need some thoughts/additional code. The architecture is not a 1-1 sequence to sequence any more. There are lectures interleaved in between. This is not entirely a new scenario. In NLP world, we have situations where a 20 word sentence on the encoder end is translated to 12 words at the decoder end. With that thought, let us quickly move on...\n<br>\n3. Bundle and task_container_id\n------------------------\n It took me a while to understand this. The exercises are structured in such a way that sometimes there could be an image or a video displayed and then a set of exercises that follow - all of which are related to the same. This combined set of related exercises is called a bundle... all share the same bundle id. Because exercises can be repeated multiple times over the course of a students interactions, they are identified each time by a unique task_container_id, which reflects the order in which the user starts the tasks. So it is an important positional information and definitely a strong feature. One of the notebooks already proved the intuition that 'when questions are repeated rate, the success rate is higher'. That example showed many cases where questions were repeated even if user got them right the last time around. This could be because the user may be learning a new concept and that new concept is based on 2-3 old concepts. Probably if the user gets the new concept wrong, the AI in the system (yes! I am sure AI would already be built in the system...so in a way, we are modelling an AI model :) Now u see why this competition is so fascinating) would probably display the exercises of the associated concepts (the old ones) to refresh the student memory. Why am I mentioning this? This 'might' be useful in clustering exercises. Clustering exercises is going to play an imp role and while one way to do it would be using tags, another would be using the temporal information (if user answers Q4 and Q7 correctly, he is very likely to get Q17 also correct) and the third way would be the above\n<br>\n4. prior_question_elapsed_time\n------------------------\n Average time taken to answer each question in the previous exercise bundle. If the exercise was not a part of a bundle but an individual entity, then it is just the time taken to answer the prev question. \n- One obvious ramification is that this is null for a user's first question..So we can identify new users this way. UPDATE - Timestamp could be better leveraged for this\n- Eric shows in his EDA that this feature does not seem to make much of a difference but he also does not drop this because this shows up as a useful feature in the model. Let us analyse this a bit to see why this could be so (at least theoretically). Firstly, this is the time taken for the prev exercise, not the current one. Most of the studies including SAINT+ seem to suggest that this feature provides a clue to whether the user is going to answer the next exercise correctly. If we take the average time for right responses and compare it with the average time for wrong answers, there does not seem to be any visible difference. But maybe avg time is not the right measure to look at and there is some other relationship which is far more complex. Maybe there is a whole category of students who blaze their way across exams not really caring much about results and excluding this category may show a clear demarcated boundary. For e.g. it would be interesting to see the means of correct/wrong responses if we eliminate all responses (right or wrong) that took less than (say) 5 seconds? Or we could see the plots for all +ve versus -ve interactions where response time is > 2 times the mean for THAT question. It is logical to assume we could see a greater %age of wrong answers in this list. One potential way to make use of this feature is to make categories..circling around the mean time for the exercise. Since we are talking about exercises here, we can take the mean for the full data without worrying about rolling averages. (Of course over time the data might change but that will be too slow and should not affect).\nAnyway this seems like a really useful feature intuitively... Saint+ seems to use a learnable weight to embed this info into a vector and use the continuous feature (not categorical) but I feel using a z-score here and categorising will bring in more benefits)\n<br>\n5. prior_question_had_explanation: \n------------------------\nThis is defined as \"Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle\". \nThe above seems to indicate that after answering a question, the student sees the correct response and an explanation for the same. The user can choose to read this explanation. Possibly this could be a hyperlink given below the correct answer. If the user clicks the hyperlink, then probably this boolean flag is recorded. This immediately strikes me as a very good feature. If the user is correct, then he already knows about the concept and is probably further strengthening his memory. Sincere students who were wrong, would definitely want to see where they are going wrong and correct themselves. IF an user gives a wrong answer AND this flag is FALSE is almost certainly going to get subsequent exercises (related to this topic) as incorrect. Firstly, I hope my understanding of this feature is correct. Please correct me if I am wrong. The next interesting question is how do we vectorize this interesting feature. We can retain it as it is and let the model automatically discover its relation with answer correctness in predicting the labels or we can experiment if adding this type of OHE helps:\n1- Answer correct and flag is false: 1 0 0\n2- Answer correct and flag is true: 0 1 0 \n3- Answer wrong and flag is false: 0 0 1\n4- Answer wrong and flag is true: 1 0 0\nUPDATE: Yana has shown that field is mainly for diagnostic questions. If so then it is better to discard this field for predictions\n\nNow let us look at the interesting fields in the other 2 datasets:\n- There are no question texts given. So relationship between questions can be only gauged by - (a) do they share the same bundle (b) do they share same or similar tags (c) LSTM relation and if possible the 4th clustering mechanism we discussed in one of the points above\n- Tags seems to be loosely mapped to skill tags or knowledge concepts. They are annotated by domain experts in the EDNET database and hence can be used to build relationship between questions. I feel this is an important aspect to be analysed. Yana has done some excellent work in analysing tags and their relation to questions. This is something we need to definitely think about in great detail when building the model\n\nThen we have lectures - each lecture is mapped to 1 tag and most likely strengthens the students knowledge for that tag.\n\nLastly we have Parts - This may be too generic but could still be useful. We could take the student average correctness for each Part rather than their performance as a whole. This is because students may be stronger in certain 'Part's of the exam compared to other parts. Kkiller took the harmonic mean of student correctness and exercise correctness and gets around 74%. I am sure this can be improved by taking student correctness for each of the 6-7 parts of the exam. \n\nLet us see what else is important:\n- 'Consistency' is the best indicator of success. Nothing else beats it (IMHO). Students who regularly login and spend time can be expected to have better scores. So frequency of interactions is a good measure. Num of interactions of the student per week seems like an excellent feature to me to test. This is also a nice way to cluster students. Clustering student is going to be very important\n- Users with very large gaps (> 1 month?) between interactions can be marked as new users? Can we drop these records considering the processing limitations? The challenge is the skills acquired quickly revert to that of a new user but the aptitude of the student will not change. So there is definitely useful information there. Maybe we could drop it from the NN ensemble and retain it for non-NN ones?\n- Remove users with limited interactions (<30..I believe this is the diagnostic test count?). Will this save a lot of space? Intuitively it should since there are many folks who join in and drop out at an early stage\n- What is the typical flow of bundle IDs as the student flows across the course from beg to end? This is useful in itself and can also be used for clustering students\n- Lecture frequency of student? Per week? Could be another feature apart from the interaction frequency\n- Remove outlier students - bottom n top\n- Needless to say the part-based student rolling avg and the exercise whole average are very important features\n- We saw exercises can repeat itself based on student's progress. This could be an indicator to model 'forgetfulness'. The more instances of repeated questions we see, the more we can conclude that the student is of a forgetful nature. Possibly this can be an input to decay attention weights at a faster rate\n- Handling cold starts - In the interest of space and time, I will not cover this. Please do refer to my recent discussions with kkiller on this\n- The SAKT paper makes an interesting observation and I quote - \"\"We can see that the based on the attention weights, we are able to achieve the perfect clustering of the exercise tags based on the hidden concepts from which they are derived\". Reversing this sentence we can conclude that tags are going to play a key role in determining attention weights. Use them wisely as a key feature esp paying attention to how you design the encoder/decoder inputs\n\nWe have covered all possible features that we could intuitively think of. For more we need to do an elaborate EDA. We also need to explore the relationship between exercise-lecture-tags in a deeper way.\n\nLastly, let us come to the time-series aspect of the data and this could guide us regarding the possible architectures. It was a bit hard for me to intuitively grasp the architectures in the SOTA papers, so I hope this below discussion helps newbies like me who are the intended audience. Surprisingly I couldn't find too many explanations of using transformer like architectures in Time-series situations..beyond the actual papers themselves. My focus is on the intuitive aspects of the architecture.\n\nIt is possible to fit a Time-series model into a encoder, decoder type of architecture & by that extension into a transformer architecture. In case of a simple sequence classification (like the covid mRNA Kaggle challenge), we needed only the encoder part of the transformer. But in case we want to predict a series (say the next 1 month of the financial index), then this can be a seq-seq exercise. In case we want to predict the entire 1 month at one shot, we can still restrict ourselves to the encoder architecture but in case we want to predict something and then use the predicted value to predict the next day's index we will need a decoder as well. The decoder uses the 'just decoded value' from the previous step and the prev hidden state (in case of LSTM) to predict the next word. In case we add attention, it will also look at the encoder outputs to determine which of the outputs could have a strong 'say' in predicting the label and give weightage to it accordingly. Discussing attention or transformer architecture itself is not the scope here, but those interested for a background can refer to: https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e.\n\nThere are minor differences between NLP and Time-series Seq to Seq and it would be good to analyse those because most of us are familiar with NLP Seq-Seq architectures and we just need to 'adapt' that model to other time-series situations. In NLP there is a limit to which the sentence or the para or even the document can stretch backwards. In Time-series, we can have financial indices that are a century old. So determining a 'window' where we are going to focus our 'attention' is going to be the key. Local and soft attention techniques are better here. In summary, to produce a labeled dataset, we need to use a fixed-length sliding time window approach to construct X, Y pairs for model training and evaluation. Having such a window means we may need to pad up inputs with dummy data. Could the diagnostic set of 30 odd questions act like some sort of a padding? This is just a thought. Obviously we have to do look-ahead masking and one-position offset between the decoder input and target output in the decoder to ensure that prediction of a time series data point will only depend on previous data points. Most importantly, for a time-series model, the input is a vector of continuous numbers already. There is no need for an NLP like embedding layer. HOWEVER, we need to ensure that the data is richly represented. We could make use of a (learnable) weight matrix to transform the one-dimension vector into a certain dimension vector. Alternatively we could use self-attention to embed the input. Self-attention with all the fancy multiheads, residuals and all the standard tricks result in capturing the richness of the data well. Using self-attention will require positional encodings as the embedding has no knowledge of sequence order. There are some standard well-defined hacks to incorporate this and we will not go into the details. This is a quick generic summary of how a time-series data can leverage the transformer architecture. But I felt there are two unanswered questions regarding this. We will come to that later but first let us look at the specificities of this competition.\n\nLet us assume there is only one student. We have a sequence of all her interactions 'I' from I1 to I1000. She viewed 100 lectures and solved 900 exercises. She has reached the fag end of the course now and another 100 exercises are left. We have to predict how she performs in these 100 exercises. I assume there could be a couple of lectures embedded here also, let us just ignore those for now as we dont have to predict those. The way we could train now is to break the data into time chunks. We could have Q1 to Q1000 fed to the encoder. The decoder will consist of responses to Q2 to Q1000. Let us now predict Q2 using just the questions Q1, Q2 and Q1's response by the student. Now you may see the justification for masks on both the encoder and the decoder. When predicting Q2, we just know the question Q1, the question Q2 and we just know the student's response to Q1 (and of course all the associated meta data). Possibly this sentence makes better sense now - \"The key is to do look-ahead masking and one-position offset between the decoder input and target output in the decoder to ensure that prediction of a time series data point will only depend on previous data points\".\n\nNow repeat the above, this time the input on encoder side will be the questions Q1, Q2, Q3 and the decoder side will be the response to Q1, Q2 and our goal is to predict the response for Q3. Similarly assume we have reached Q100. We notice a pattern now. The predictions of certain question responses depend on the response to some of the previous questions. Maybe it is because some questions are repeated or maybe it is because questions with same tags are somewhat similar and if a student has answered one she will almost certainly answer the other. So your model determines (after going thru' many such student interactions) that Q100's response depends on the student's past response to Q78, Q99 and Q46. This is attention logic at work. The decoder is paying attention to all the states of the encoder output (Q1 to Q100) while predicting Q100th response (in addition to considering the decoder inputs. There will also be student specific temporal information which will be captured by the model and thus the model is trained. Every additional student contributes to strengthening the models weights and biases and now the model is ready to predict. Notice that for the 1000th question, it will be computationally expensive to go all the way back to Q1, so we take a sliding window and carve out time slices of student interaction data. If the window is 100, we will have Q1-Q100 questions and responses as 1 sliced row input and Q101 to be predicted. Then we will have Q2-Q101 and Q102 to be predicted...and all the way till Q901-Q1000 questions and responses as inputs and Q 1001 to be predicted.\n\nSo far so good. But where do we feed what? If we take the case of a language translator for whom the original encoder-decoder arch was conceived, the source language inputs are fed to the encoder and the destination language to the decoder. The situation in time-series  is such that we can feed the inputs to the encoder and the outputs to the decoder. So Q1-Q100 goes to encoder and also Q101. The responses of Q1 to Q100 goes to decoder and now we have to predict the response for Q101. We can assume that each response is like the translated sentence of the concerned question.\n\nSo all questions and question related information could be fed to the encoder. This includes the exercises, the tags, the task_container_id's and other derived features related to QUESTIONS etc etc. The response related meta data along with the response themselves go into the decoder. This includes data like prev question explanation, elapsed time, lag time, saw lecture and all derived features related to response....basically everything other than what is put in the encoder. I like this style of segregation between encoder and decoder and I believe this is what is done in the SAINT paper. HOWEVER this need not be the only solution. Some papers have fed the coupled interactions (exercise+response) on decoder end. Others have tried many other things. This could be an area that will need a bit of experimentation. I feel it might not altogether be a bad idea to add tags meta-data in the responses?\n\nThe rest of the paraphernalia is the same as in any standard transformer architecture and there are source codes in every framework for the same which you can re-use.\n\nOther possible considerations \n- For forming language representations focusing on the closest word makes a lot of sense. However, this is much more variable with time series data, in certain time series sequences causality can come from steps much further back. Additionally, in the real-world scenario, the exercises which occur close in the sequence tend to belong to the same concept. Thus, we expect that the attention weights biased towards the exercises that occur recently in the interaction sequence.Can we have multiple windows into which to look into? Or when we are choosing time slices can we pick up the last time slice based on the tags?\n- The attention weights can be augmented by having a set of relation coefficients between each exercise. Shalini Pandey's latest SOTA paper does this, but there she uses text information from questions to build relationships. Here we could use tags.\n\nLastly, coming to the two things which I couldn't understand intuitively. I can see the following issues when using transformers (or encoder-decoder) on time-series data and I hope someone can guide me on where my understanding is incorrect:\n\nIssue 1: In NLP every translation has one and only one correct answer. While the inputs may change to form different sequences, the output will always be the same for a given sequence of words. The model converges towards that output...it cannot deviate astray with more data. In a Time-series data, there could be different outputs for a given input sequence. In this case it is based on the student profile. The two problem statements are similar but not the same and using the same architecture for both may not be optimal\n\nIssue 2: Somewhat linked to Issue 1. Let us think of it in terms of financial indices. We want to determine the NYSE index future. Let us say NYSE index is represented by a proportion of every stock traded on NYSE. One option is to take the NYSE index itself over time and only use that to model. The other option is to look at all the constituents. However modelling based on the constituents could be tricky. Each constituent follows a different path. Every stock listed in the NYSE stock exchange could behave differently. Using their daily data to predict the NYSE index futures is at best a tricky job. Considering the problem at hand, we see each different student as a individual stock being traded on NYSE. There is no equivalent of NYSE index here. Instead we want to predict how these stocks continue to move and also we want to predict how new stocks added to the stock exchange would move. We do have lots and lots of data but doesn't this introduce tons of noise also? One obvious solution is to cluster the constituents. So in case of the NYSE scenario, we could have a bunch of stocks getting represented by a single index - the tech index. We could have brick n mortar companies as another cluster, retail, housing and form dozen odd clusters. Calculate the movements for each cluster separately...as they could have different weights and balances. Now we are left with time-series data of a dozen odd indices - all of which will help determine the NYSE index futures more predictably. If we apply this sort of a scenario to RiiiD, this means that clusters of students need to be created and then it makes sense to have separate sets of weights and biases for each category? Let existing student continue in the cluster they already are in and the new students are moved to one of the clusters. But I haven't seen this kind of approach in any time-series model leveraging encpder/decoders.\n\nWe still haven't talked of validation strategies, thought about what could be the best combination of keys queries and values or the importance of online learning in this competition or usage of reductions and many others. But we have covered a fair bit of theory otherwise. While kernels are shared abundantly, I felt some theoretical discussion on the features and architectures could help newcomers like me and hence quickly jotted a few points that stuck me as I went thru' a few of the beautiful kernels.\n\nIn Silogram's golden words, this competition is not just a test of AI skills, but is also a test of our engineering skills. How creatively we make use of the imposed constraints will decide the winner. With unlimited processing power and memory, the NN models would have won hands down but with the imposed constraints a blend of NN and non-NN models may hold the key.\n\nAs with all features and architectures, there are risks/rewards. \n\nChoose wisely. All the very best to all!",
      "votes": null
    },
    {
      "id": "1112034",
      "postDate": "12/14/2020 08:18:47",
      "content": "<p>I have benefited from every discussion you have had and look forward to next week's discussion.I will check my feature engineering</p>",
      "rawMarkdown": "I have benefited from every discussion you have had and look forward to next week's discussion.I will check my feature engineering",
      "votes": null
    },
    {
      "id": "1112108",
      "postDate": "12/14/2020 09:30:54",
      "content": "<p>Hey !</p>\n<p>Thanks for the messages.<br>\nSome thoughts from me:</p>\n<ul>\n<li><p>I prefere to use <strong>Timestamp</strong> rather than <strong>prior_question_elapsed_time</strong> to define if a user is new or not (it is precised that timestamp is set to 0 for every first interaction of a user.</p></li>\n<li><p>Something very interesting but that is hard to capture is the <strong>lag_time between two exercices</strong>. It can gives indication of a student taking notes of the last question for example.</p></li>\n<li><p>Still in the timestamp aspect, I think it can be interesting to check last time a user had an interaction with a section of the toeic. If a user start to improve on a topic, but then don't go to it for days, he might have some troubles redoing the exercices of that section.</p></li>\n</ul>",
      "rawMarkdown": "Hey !\n\nThanks for the messages.\nSome thoughts from me:\n\n- I prefere to use **Timestamp** rather than **prior_question_elapsed_time** to define if a user is new or not (it is precised that timestamp is set to 0 for every first interaction of a user.\n\n- Something very interesting but that is hard to capture is the **lag_time between two exercices**. It can gives indication of a student taking notes of the last question for example.\n\n- Still in the timestamp aspect, I think it can be interesting to check last time a user had an interaction with a section of the toeic. If a user start to improve on a topic, but then don't go to it for days, he might have some troubles redoing the exercices of that section.",
      "votes": null
    },
    {
      "id": "1112189",
      "postDate": "12/14/2020 11:16:41",
      "content": "<p>Thanks Qiaqia. All the very best to you for the competition!!!<br>\nAs only couple of weekends left now, I may not be able to post anything tangible…but all the best once again!</p>",
      "rawMarkdown": "Thanks Qiaqia. All the very best to you for the competition!!!\nAs only couple of weekends left now, I may not be able to post anything tangible...but all the best once again!",
      "votes": null
    },
    {
      "id": "1112403",
      "postDate": "12/14/2020 14:58:40",
      "content": "<p>Point 1 - Yes<br>\nPoint 2 - The difference between the two consecutive timestamps for a student should give an indication of lag time? So like in the Saint+ paper, we do have both the elapsed time and the lag time here.. Of course the elapsed time is for the prev question…Still it does help apparently…</p>\n<p>just to be clear, their definition is as follows:<br>\nelapsed time - the time taken for a student to answer (the previous question)<br>\nlag time - the time interval between adjacent learning activities.</p>\n<p>From this, I guess we can calculate the difference between times to see if the student was taking some notes after exercises if that is what is needed. I dont know how strong that feature would be though. However this would indeed be a very strong feature for lectures - amount of time student watched the lecture versus average time lecture is typically watched but as far as I can see they dont give elapsed time for lectures. Still we can try to guess try to play around with timestamp differences and make some guesses based on lag times (of course we have to assume that student is doing all this in one continuous session and hasn't opened a new window to play their favorite video game :) could be a noisy feature but worth a try</p>",
      "rawMarkdown": "Point 1 - Yes\nPoint 2 - The difference between the two consecutive timestamps for a student should give an indication of lag time? So like in the Saint+ paper, we do have both the elapsed time and the lag time here.. Of course the elapsed time is for the prev question...Still it does help apparently...\n\njust to be clear, their definition is as follows:\nelapsed time - the time taken for a student to answer (the previous question)\nlag time - the time interval between adjacent learning activities.\n\nFrom this, I guess we can calculate the difference between times to see if the student was taking some notes after exercises if that is what is needed. I dont know how strong that feature would be though. However this would indeed be a very strong feature for lectures - amount of time student watched the lecture versus average time lecture is typically watched but as far as I can see they dont give elapsed time for lectures. Still we can try to guess try to play around with timestamp differences and make some guesses based on lag times (of course we have to assume that student is doing all this in one continuous session and hasn't opened a new window to play their favorite video game :) could be a noisy feature but worth a try",
      "votes": null
    },
    {
      "id": "1139368",
      "postDate": "01/05/2021 10:52:14",
      "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> thank you for compiling such a helpful post!! actually all of your discussions are worth a lot -:) I'll have to read many times to digest.. regarding the following point:</p>\n<blockquote>\n  <p>HOWEVER, we need to ensure that the data is richly represented. We could make use of a (learnable) weight matrix to transform the one-dimension vector into a certain dimension vector. Alternatively we could use self-attention to embed the input. Self-attention with all the fancy multiheads, residuals and all the standard tricks result in capturing the richness of the data well. Using self-attention will require positional encodings as the embedding has no knowledge of sequence order. There are some standard well-defined hacks to incorporate this and we will not go into the details. </p>\n</blockquote>\n<p>could you please point me to any example code/paper that uses self-att to capture the input? I don't remember to have come across smth like that - I'll give a try anyway but any source with info will help. Thank you and good luck to the comp </p>",
      "rawMarkdown": "allohvk thank you for compiling such a helpful post!! actually all of your discussions are worth a lot -:) I'll have to read many times to digest.. regarding the following point:\n> HOWEVER, we need to ensure that the data is richly represented. We could make use of a (learnable) weight matrix to transform the one-dimension vector into a certain dimension vector. Alternatively we could use self-attention to embed the input. Self-attention with all the fancy multiheads, residuals and all the standard tricks result in capturing the richness of the data well. Using self-attention will require positional encodings as the embedding has no knowledge of sequence order. There are some standard well-defined hacks to incorporate this and we will not go into the details. \n\n could you please point me to any example code/paper that uses self-att to capture the input? I don't remember to have come across smth like that - I'll give a try anyway but any source with info will help. Thank you and good luck to the comp",
      "votes": null
    },
    {
      "id": "1139426",
      "postDate": "01/05/2021 11:39:06",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> :)<br>\nThis was written at the very beginning of the competition. So there could be slight changes here  and there. Definitely self-attention should help. There are couple of Saint implementations and I believe they use self-attention because that is part of Saint architecture by default. Have you checked those. Unfortnuately my Kaggle ID had an issue till last week and only last weekend I took the SAKT kernel from Wang and trying to make changes to make it work. I am still stuck at basic changes..</p>",
      "rawMarkdown": "Thanks @imeintanis :)\nThis was written at the very beginning of the competition. So there could be slight changes here  and there. Definitely self-attention should help. There are couple of Saint implementations and I believe they use self-attention because that is part of Saint architecture by default. Have you checked those. Unfortnuately my Kaggle ID had an issue till last week and only last weekend I took the SAKT kernel from Wang and trying to make changes to make it work. I am still stuck at basic changes..",
      "votes": null
    },
    {
      "id": "1139442",
      "postDate": "01/05/2021 11:54:33",
      "content": "<p>Thank you I got it (actually I had misunderstood your point). Happy that you resolved your issues &amp; you 're back again<br>\nps: I started from the same SAKT kernel but cannot surpass the 0780 stage.. we'll see </p>",
      "rawMarkdown": "Thank you I got it (actually I had misunderstood your point). Happy that you resolved your issues & you 're back again\nps: I started from the same SAKT kernel but cannot surpass the 0780 stage.. we'll see",
      "votes": null
    },
    {
      "id": "1139568",
      "postDate": "01/05/2021 13:52:21",
      "content": "<p>prior_question_elapsed_time is <code>&lt;NA&gt;</code> for lectures, regrettably. The full dataset has it; but not the version they prepared for our Kaggle.</p>",
      "rawMarkdown": "prior_question_elapsed_time is `<NA>` for lectures, regrettably. The full dataset has it; but not the version they prepared for our Kaggle.",
      "votes": null
    },
    {
      "id": "1139600",
      "postDate": "01/05/2021 14:12:06",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> did u find a way to create a moving mask for the question bundles so that they dont attend to answers in same bundle? I can share my thought if u r looking for a solution</p>",
      "rawMarkdown": "authman did u find a way to create a moving mask for the question bundles so that they dont attend to answers in same bundle? I can share my thought if u r looking for a solution",
      "votes": null
    },
    {
      "id": "1139647",
      "postDate": "01/05/2021 14:42:49",
      "content": "<p>Not satisfactorily =\\</p>",
      "rawMarkdown": "Not satisfactorily =\\",
      "votes": null
    },
    {
      "id": "1139745",
      "postDate": "01/05/2021 15:39:29",
      "content": "<p>see if this works..</p>\n<p>The trick here is to create separate masks for interaction sequence and question sequence in each row (if u r using SAKT) and create multiple rows<br>\nQ1 Q2 Q3 Q4 Q5 Q6, Q7<br>\nI0 I1 I2 I3 I4 I5 I6</p>\n<p>Let us say Q3, Q4, Q5 are a bundle, then:<br>\nI0 I1 I2 0 0 0 0 <br>\nQ1 Q2 Q3 Q4 Q5 0 0</p>\n<p>and<br>\nI0 I1 I2 I3 I4 I5 I6<br>\n0 0 0 0 0 Q6 Q7</p>\n<p>Repeat for all bundles in row</p>",
      "rawMarkdown": "see if this works..\n\nThe trick here is to create separate masks for interaction sequence and question sequence in each row (if u r using SAKT) and create multiple rows\nQ1 Q2 Q3 Q4 Q5 Q6, Q7\nI0 I1 I2 I3 I4 I5 I6\n\nLet us say Q3, Q4, Q5 are a bundle, then:\nI0 I1 I2 0 0 0 0 \nQ1 Q2 Q3 Q4 Q5 0 0\n\nand\nI0 I1 I2 I3 I4 I5 I6\n0 0 0 0 0 Q6 Q7\n\nRepeat for all bundles in row",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1112034,
      "author_name": "yangxiaoshuai",
      "author_url": "",
      "post_date": "12/14/2020 08:18:47",
      "content": "<p>I have benefited from every discussion you have had and look forward to next week's discussion.I will check my feature engineering</p>",
      "votes": null,
      "replies": [
        {
          "id": 1112189,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "12/14/2020 11:16:41",
          "content": "<p>Thanks Qiaqia. All the very best to you for the competition!!!<br>\nAs only couple of weekends left now, I may not be able to post anything tangible…but all the best once again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1112108,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "12/14/2020 09:30:54",
      "content": "<p>Hey !</p>\n<p>Thanks for the messages.<br>\nSome thoughts from me:</p>\n<ul>\n<li><p>I prefere to use <strong>Timestamp</strong> rather than <strong>prior_question_elapsed_time</strong> to define if a user is new or not (it is precised that timestamp is set to 0 for every first interaction of a user.</p></li>\n<li><p>Something very interesting but that is hard to capture is the <strong>lag_time between two exercices</strong>. It can gives indication of a student taking notes of the last question for example.</p></li>\n<li><p>Still in the timestamp aspect, I think it can be interesting to check last time a user had an interaction with a section of the toeic. If a user start to improve on a topic, but then don't go to it for days, he might have some troubles redoing the exercices of that section.</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1112403,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "12/14/2020 14:58:40",
          "content": "<p>Point 1 - Yes<br>\nPoint 2 - The difference between the two consecutive timestamps for a student should give an indication of lag time? So like in the Saint+ paper, we do have both the elapsed time and the lag time here.. Of course the elapsed time is for the prev question…Still it does help apparently…</p>\n<p>just to be clear, their definition is as follows:<br>\nelapsed time - the time taken for a student to answer (the previous question)<br>\nlag time - the time interval between adjacent learning activities.</p>\n<p>From this, I guess we can calculate the difference between times to see if the student was taking some notes after exercises if that is what is needed. I dont know how strong that feature would be though. However this would indeed be a very strong feature for lectures - amount of time student watched the lecture versus average time lecture is typically watched but as far as I can see they dont give elapsed time for lectures. Still we can try to guess try to play around with timestamp differences and make some guesses based on lag times (of course we have to assume that student is doing all this in one continuous session and hasn't opened a new window to play their favorite video game :) could be a noisy feature but worth a try</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1139568,
          "author_name": "authman",
          "author_url": "",
          "post_date": "01/05/2021 13:52:21",
          "content": "<p>prior_question_elapsed_time is <code>&lt;NA&gt;</code> for lectures, regrettably. The full dataset has it; but not the version they prepared for our Kaggle.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1139600,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "01/05/2021 14:12:06",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> did u find a way to create a moving mask for the question bundles so that they dont attend to answers in same bundle? I can share my thought if u r looking for a solution</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1139647,
          "author_name": "authman",
          "author_url": "",
          "post_date": "01/05/2021 14:42:49",
          "content": "<p>Not satisfactorily =\\</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1139745,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "01/05/2021 15:39:29",
          "content": "<p>see if this works..</p>\n<p>The trick here is to create separate masks for interaction sequence and question sequence in each row (if u r using SAKT) and create multiple rows<br>\nQ1 Q2 Q3 Q4 Q5 Q6, Q7<br>\nI0 I1 I2 I3 I4 I5 I6</p>\n<p>Let us say Q3, Q4, Q5 are a bundle, then:<br>\nI0 I1 I2 0 0 0 0 <br>\nQ1 Q2 Q3 Q4 Q5 0 0</p>\n<p>and<br>\nI0 I1 I2 I3 I4 I5 I6<br>\n0 0 0 0 0 Q6 Q7</p>\n<p>Repeat for all bundles in row</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1139368,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "01/05/2021 10:52:14",
      "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> thank you for compiling such a helpful post!! actually all of your discussions are worth a lot -:) I'll have to read many times to digest.. regarding the following point:</p>\n<blockquote>\n  <p>HOWEVER, we need to ensure that the data is richly represented. We could make use of a (learnable) weight matrix to transform the one-dimension vector into a certain dimension vector. Alternatively we could use self-attention to embed the input. Self-attention with all the fancy multiheads, residuals and all the standard tricks result in capturing the richness of the data well. Using self-attention will require positional encodings as the embedding has no knowledge of sequence order. There are some standard well-defined hacks to incorporate this and we will not go into the details. </p>\n</blockquote>\n<p>could you please point me to any example code/paper that uses self-att to capture the input? I don't remember to have come across smth like that - I'll give a try anyway but any source with info will help. Thank you and good luck to the comp </p>",
      "votes": null,
      "replies": [
        {
          "id": 1139426,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "01/05/2021 11:39:06",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> :)<br>\nThis was written at the very beginning of the competition. So there could be slight changes here  and there. Definitely self-attention should help. There are couple of Saint implementations and I believe they use self-attention because that is part of Saint architecture by default. Have you checked those. Unfortnuately my Kaggle ID had an issue till last week and only last weekend I took the SAKT kernel from Wang and trying to make changes to make it work. I am still stuck at basic changes..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1139442,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "01/05/2021 11:54:33",
          "content": "<p>Thank you I got it (actually I had misunderstood your point). Happy that you resolved your issues &amp; you 're back again<br>\nps: I started from the same SAKT kernel but cannot surpass the 0780 stage.. we'll see </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1111913": "Note: Part 1 is here: From Bayesian to Transformers - Tracing the 'Knowledge Tracing' models over time: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201481\n\nPart 3: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206185 - Some additional clarifications on SAKT/SAINT\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584 - A small discussion on position embeddings for those interested.\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719 - On lectures, the art of forgetting and why I retired hurt\n\nFor many reasons, this is an interesting competition. The sheer volume of data, the process for submission, the scope for online learning and the imposed constraints - all mimic real-world situations. This is a gold-mine for aspiring data scientists (newbies) because there is so much to learn. Rarely does someone give this kind of rich data. Such learning experiences are unique in case you are not already on this ship - do hop on for a fascinating ride. \n\nLast weekend, I had shared my interpretation of the general domain and some of the interesting developments in this area recent times. I had also shared my interpretations of the top 5 SOTA papers in this area. Let us now quickly look at the actual data shared by RiiiD and focus on the interesting columns. We will jot down some intuitions along the way and then test those against actual data (I am afraid the 'testing' part will have to wait till next weekend)..\n<br>\n1. Timestamp\n--------------------\ntime in milliseconds between this user interaction and the first event completion from that user\n- The most useful feature in my view\n- Dividing this by the number of interactions gives an idea of the frequency of student interactions.  I would rate 'continuous learning' as the single biggest prediction of success, if we can find an effective way to capture this\n- This will also give the time interval between user interactions which can be a very useful feature. For e.g. many exercises are completed in one sitting. But if there is a gap of more than few hours, then it can be considered to be a next sitting. 12 hour gap would mean that the last exercise was done the previous day and the user is starting fresh for the day - This has lot of interesting connotations - too many to list here. Just to give one example - someone talked about the power of listening to lectures in predicting success. This effect will be all the more pronounced if we determine whether the user listened to the lecture in the same day. After a few hours memory declines by almost a half. Another interesting application of this feature - if the data needs to be crunched, it might not be a bad idea to group all exercises performed in one session as one single exercise and take the average score as the response. How about - Num of interactions the user has had in the past 1 month as another feature - sounds interesting right? Many more..\n<br>\n2. Interactions\n------------------------\n - In RiiiD, interactions can be exercises as well as lectures. content_type_id determines this. This is interesting because I didn't see lectures being captured as interactions in their Saint paper or for that matter any other SOTA paper. How do we make use of this feature? Each lecture is mapped to a tag and it is obvious that if the exercise in question (which needs to be predicted) is associated with a tag and the student has listened to the corresponding lecture in recent past, it would definitely predict success. That could be a separate feature as discussed above. But more importantly we need to check how we can incorporate the lecture interactions within the time-series data as they will be very very useful in determining attention weights at the decoder end. If I am not wrong, most of the SOTA papers currently only have exercises on the encoder side. Plugging in lectures might need some thoughts/additional code. The architecture is not a 1-1 sequence to sequence any more. There are lectures interleaved in between. This is not entirely a new scenario. In NLP world, we have situations where a 20 word sentence on the encoder end is translated to 12 words at the decoder end. With that thought, let us quickly move on...\n<br>\n3. Bundle and task_container_id\n------------------------\n It took me a while to understand this. The exercises are structured in such a way that sometimes there could be an image or a video displayed and then a set of exercises that follow - all of which are related to the same. This combined set of related exercises is called a bundle... all share the same bundle id. Because exercises can be repeated multiple times over the course of a students interactions, they are identified each time by a unique task_container_id, which reflects the order in which the user starts the tasks. So it is an important positional information and definitely a strong feature. One of the notebooks already proved the intuition that 'when questions are repeated rate, the success rate is higher'. That example showed many cases where questions were repeated even if user got them right the last time around. This could be because the user may be learning a new concept and that new concept is based on 2-3 old concepts. Probably if the user gets the new concept wrong, the AI in the system (yes! I am sure AI would already be built in the system...so in a way, we are modelling an AI model :) Now u see why this competition is so fascinating) would probably display the exercises of the associated concepts (the old ones) to refresh the student memory. Why am I mentioning this? This 'might' be useful in clustering exercises. Clustering exercises is going to play an imp role and while one way to do it would be using tags, another would be using the temporal information (if user answers Q4 and Q7 correctly, he is very likely to get Q17 also correct) and the third way would be the above\n<br>\n4. prior_question_elapsed_time\n------------------------\n Average time taken to answer each question in the previous exercise bundle. If the exercise was not a part of a bundle but an individual entity, then it is just the time taken to answer the prev question. \n- One obvious ramification is that this is null for a user's first question..So we can identify new users this way. UPDATE - Timestamp could be better leveraged for this\n- Eric shows in his EDA that this feature does not seem to make much of a difference but he also does not drop this because this shows up as a useful feature in the model. Let us analyse this a bit to see why this could be so (at least theoretically). Firstly, this is the time taken for the prev exercise, not the current one. Most of the studies including SAINT+ seem to suggest that this feature provides a clue to whether the user is going to answer the next exercise correctly. If we take the average time for right responses and compare it with the average time for wrong answers, there does not seem to be any visible difference. But maybe avg time is not the right measure to look at and there is some other relationship which is far more complex. Maybe there is a whole category of students who blaze their way across exams not really caring much about results and excluding this category may show a clear demarcated boundary. For e.g. it would be interesting to see the means of correct/wrong responses if we eliminate all responses (right or wrong) that took less than (say) 5 seconds? Or we could see the plots for all +ve versus -ve interactions where response time is > 2 times the mean for THAT question. It is logical to assume we could see a greater %age of wrong answers in this list. One potential way to make use of this feature is to make categories..circling around the mean time for the exercise. Since we are talking about exercises here, we can take the mean for the full data without worrying about rolling averages. (Of course over time the data might change but that will be too slow and should not affect).\nAnyway this seems like a really useful feature intuitively... Saint+ seems to use a learnable weight to embed this info into a vector and use the continuous feature (not categorical) but I feel using a z-score here and categorising will bring in more benefits)\n<br>\n5. prior_question_had_explanation: \n------------------------\nThis is defined as \"Whether or not the user saw an explanation and the correct response(s) after answering the previous question bundle\". \nThe above seems to indicate that after answering a question, the student sees the correct response and an explanation for the same. The user can choose to read this explanation. Possibly this could be a hyperlink given below the correct answer. If the user clicks the hyperlink, then probably this boolean flag is recorded. This immediately strikes me as a very good feature. If the user is correct, then he already knows about the concept and is probably further strengthening his memory. Sincere students who were wrong, would definitely want to see where they are going wrong and correct themselves. IF an user gives a wrong answer AND this flag is FALSE is almost certainly going to get subsequent exercises (related to this topic) as incorrect. Firstly, I hope my understanding of this feature is correct. Please correct me if I am wrong. The next interesting question is how do we vectorize this interesting feature. We can retain it as it is and let the model automatically discover its relation with answer correctness in predicting the labels or we can experiment if adding this type of OHE helps:\n1- Answer correct and flag is false: 1 0 0\n2- Answer correct and flag is true: 0 1 0 \n3- Answer wrong and flag is false: 0 0 1\n4- Answer wrong and flag is true: 1 0 0\nUPDATE: Yana has shown that field is mainly for diagnostic questions. If so then it is better to discard this field for predictions\n\nNow let us look at the interesting fields in the other 2 datasets:\n- There are no question texts given. So relationship between questions can be only gauged by - (a) do they share the same bundle (b) do they share same or similar tags (c) LSTM relation and if possible the 4th clustering mechanism we discussed in one of the points above\n- Tags seems to be loosely mapped to skill tags or knowledge concepts. They are annotated by domain experts in the EDNET database and hence can be used to build relationship between questions. I feel this is an important aspect to be analysed. Yana has done some excellent work in analysing tags and their relation to questions. This is something we need to definitely think about in great detail when building the model\n\nThen we have lectures - each lecture is mapped to 1 tag and most likely strengthens the students knowledge for that tag.\n\nLastly we have Parts - This may be too generic but could still be useful. We could take the student average correctness for each Part rather than their performance as a whole. This is because students may be stronger in certain 'Part's of the exam compared to other parts. Kkiller took the harmonic mean of student correctness and exercise correctness and gets around 74%. I am sure this can be improved by taking student correctness for each of the 6-7 parts of the exam. \n\nLet us see what else is important:\n- 'Consistency' is the best indicator of success. Nothing else beats it (IMHO). Students who regularly login and spend time can be expected to have better scores. So frequency of interactions is a good measure. Num of interactions of the student per week seems like an excellent feature to me to test. This is also a nice way to cluster students. Clustering student is going to be very important\n- Users with very large gaps (> 1 month?) between interactions can be marked as new users? Can we drop these records considering the processing limitations? The challenge is the skills acquired quickly revert to that of a new user but the aptitude of the student will not change. So there is definitely useful information there. Maybe we could drop it from the NN ensemble and retain it for non-NN ones?\n- Remove users with limited interactions (<30..I believe this is the diagnostic test count?). Will this save a lot of space? Intuitively it should since there are many folks who join in and drop out at an early stage\n- What is the typical flow of bundle IDs as the student flows across the course from beg to end? This is useful in itself and can also be used for clustering students\n- Lecture frequency of student? Per week? Could be another feature apart from the interaction frequency\n- Remove outlier students - bottom n top\n- Needless to say the part-based student rolling avg and the exercise whole average are very important features\n- We saw exercises can repeat itself based on student's progress. This could be an indicator to model 'forgetfulness'. The more instances of repeated questions we see, the more we can conclude that the student is of a forgetful nature. Possibly this can be an input to decay attention weights at a faster rate\n- Handling cold starts - In the interest of space and time, I will not cover this. Please do refer to my recent discussions with kkiller on this\n- The SAKT paper makes an interesting observation and I quote - \"\"We can see that the based on the attention weights, we are able to achieve the perfect clustering of the exercise tags based on the hidden concepts from which they are derived\". Reversing this sentence we can conclude that tags are going to play a key role in determining attention weights. Use them wisely as a key feature esp paying attention to how you design the encoder/decoder inputs\n\nWe have covered all possible features that we could intuitively think of. For more we need to do an elaborate EDA. We also need to explore the relationship between exercise-lecture-tags in a deeper way.\n\nLastly, let us come to the time-series aspect of the data and this could guide us regarding the possible architectures. It was a bit hard for me to intuitively grasp the architectures in the SOTA papers, so I hope this below discussion helps newbies like me who are the intended audience. Surprisingly I couldn't find too many explanations of using transformer like architectures in Time-series situations..beyond the actual papers themselves. My focus is on the intuitive aspects of the architecture.\n\nIt is possible to fit a Time-series model into a encoder, decoder type of architecture & by that extension into a transformer architecture. In case of a simple sequence classification (like the covid mRNA Kaggle challenge), we needed only the encoder part of the transformer. But in case we want to predict a series (say the next 1 month of the financial index), then this can be a seq-seq exercise. In case we want to predict the entire 1 month at one shot, we can still restrict ourselves to the encoder architecture but in case we want to predict something and then use the predicted value to predict the next day's index we will need a decoder as well. The decoder uses the 'just decoded value' from the previous step and the prev hidden state (in case of LSTM) to predict the next word. In case we add attention, it will also look at the encoder outputs to determine which of the outputs could have a strong 'say' in predicting the label and give weightage to it accordingly. Discussing attention or transformer architecture itself is not the scope here, but those interested for a background can refer to: https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e.\n\nThere are minor differences between NLP and Time-series Seq to Seq and it would be good to analyse those because most of us are familiar with NLP Seq-Seq architectures and we just need to 'adapt' that model to other time-series situations. In NLP there is a limit to which the sentence or the para or even the document can stretch backwards. In Time-series, we can have financial indices that are a century old. So determining a 'window' where we are going to focus our 'attention' is going to be the key. Local and soft attention techniques are better here. In summary, to produce a labeled dataset, we need to use a fixed-length sliding time window approach to construct X, Y pairs for model training and evaluation. Having such a window means we may need to pad up inputs with dummy data. Could the diagnostic set of 30 odd questions act like some sort of a padding? This is just a thought. Obviously we have to do look-ahead masking and one-position offset between the decoder input and target output in the decoder to ensure that prediction of a time series data point will only depend on previous data points. Most importantly, for a time-series model, the input is a vector of continuous numbers already. There is no need for an NLP like embedding layer. HOWEVER, we need to ensure that the data is richly represented. We could make use of a (learnable) weight matrix to transform the one-dimension vector into a certain dimension vector. Alternatively we could use self-attention to embed the input. Self-attention with all the fancy multiheads, residuals and all the standard tricks result in capturing the richness of the data well. Using self-attention will require positional encodings as the embedding has no knowledge of sequence order. There are some standard well-defined hacks to incorporate this and we will not go into the details. This is a quick generic summary of how a time-series data can leverage the transformer architecture. But I felt there are two unanswered questions regarding this. We will come to that later but first let us look at the specificities of this competition.\n\nLet us assume there is only one student. We have a sequence of all her interactions 'I' from I1 to I1000. She viewed 100 lectures and solved 900 exercises. She has reached the fag end of the course now and another 100 exercises are left. We have to predict how she performs in these 100 exercises. I assume there could be a couple of lectures embedded here also, let us just ignore those for now as we dont have to predict those. The way we could train now is to break the data into time chunks. We could have Q1 to Q1000 fed to the encoder. The decoder will consist of responses to Q2 to Q1000. Let us now predict Q2 using just the questions Q1, Q2 and Q1's response by the student. Now you may see the justification for masks on both the encoder and the decoder. When predicting Q2, we just know the question Q1, the question Q2 and we just know the student's response to Q1 (and of course all the associated meta data). Possibly this sentence makes better sense now - \"The key is to do look-ahead masking and one-position offset between the decoder input and target output in the decoder to ensure that prediction of a time series data point will only depend on previous data points\".\n\nNow repeat the above, this time the input on encoder side will be the questions Q1, Q2, Q3 and the decoder side will be the response to Q1, Q2 and our goal is to predict the response for Q3. Similarly assume we have reached Q100. We notice a pattern now. The predictions of certain question responses depend on the response to some of the previous questions. Maybe it is because some questions are repeated or maybe it is because questions with same tags are somewhat similar and if a student has answered one she will almost certainly answer the other. So your model determines (after going thru' many such student interactions) that Q100's response depends on the student's past response to Q78, Q99 and Q46. This is attention logic at work. The decoder is paying attention to all the states of the encoder output (Q1 to Q100) while predicting Q100th response (in addition to considering the decoder inputs. There will also be student specific temporal information which will be captured by the model and thus the model is trained. Every additional student contributes to strengthening the models weights and biases and now the model is ready to predict. Notice that for the 1000th question, it will be computationally expensive to go all the way back to Q1, so we take a sliding window and carve out time slices of student interaction data. If the window is 100, we will have Q1-Q100 questions and responses as 1 sliced row input and Q101 to be predicted. Then we will have Q2-Q101 and Q102 to be predicted...and all the way till Q901-Q1000 questions and responses as inputs and Q 1001 to be predicted.\n\nSo far so good. But where do we feed what? If we take the case of a language translator for whom the original encoder-decoder arch was conceived, the source language inputs are fed to the encoder and the destination language to the decoder. The situation in time-series  is such that we can feed the inputs to the encoder and the outputs to the decoder. So Q1-Q100 goes to encoder and also Q101. The responses of Q1 to Q100 goes to decoder and now we have to predict the response for Q101. We can assume that each response is like the translated sentence of the concerned question.\n\nSo all questions and question related information could be fed to the encoder. This includes the exercises, the tags, the task_container_id's and other derived features related to QUESTIONS etc etc. The response related meta data along with the response themselves go into the decoder. This includes data like prev question explanation, elapsed time, lag time, saw lecture and all derived features related to response....basically everything other than what is put in the encoder. I like this style of segregation between encoder and decoder and I believe this is what is done in the SAINT paper. HOWEVER this need not be the only solution. Some papers have fed the coupled interactions (exercise+response) on decoder end. Others have tried many other things. This could be an area that will need a bit of experimentation. I feel it might not altogether be a bad idea to add tags meta-data in the responses?\n\nThe rest of the paraphernalia is the same as in any standard transformer architecture and there are source codes in every framework for the same which you can re-use.\n\nOther possible considerations \n- For forming language representations focusing on the closest word makes a lot of sense. However, this is much more variable with time series data, in certain time series sequences causality can come from steps much further back. Additionally, in the real-world scenario, the exercises which occur close in the sequence tend to belong to the same concept. Thus, we expect that the attention weights biased towards the exercises that occur recently in the interaction sequence.Can we have multiple windows into which to look into? Or when we are choosing time slices can we pick up the last time slice based on the tags?\n- The attention weights can be augmented by having a set of relation coefficients between each exercise. Shalini Pandey's latest SOTA paper does this, but there she uses text information from questions to build relationships. Here we could use tags.\n\nLastly, coming to the two things which I couldn't understand intuitively. I can see the following issues when using transformers (or encoder-decoder) on time-series data and I hope someone can guide me on where my understanding is incorrect:\n\nIssue 1: In NLP every translation has one and only one correct answer. While the inputs may change to form different sequences, the output will always be the same for a given sequence of words. The model converges towards that output...it cannot deviate astray with more data. In a Time-series data, there could be different outputs for a given input sequence. In this case it is based on the student profile. The two problem statements are similar but not the same and using the same architecture for both may not be optimal\n\nIssue 2: Somewhat linked to Issue 1. Let us think of it in terms of financial indices. We want to determine the NYSE index future. Let us say NYSE index is represented by a proportion of every stock traded on NYSE. One option is to take the NYSE index itself over time and only use that to model. The other option is to look at all the constituents. However modelling based on the constituents could be tricky. Each constituent follows a different path. Every stock listed in the NYSE stock exchange could behave differently. Using their daily data to predict the NYSE index futures is at best a tricky job. Considering the problem at hand, we see each different student as a individual stock being traded on NYSE. There is no equivalent of NYSE index here. Instead we want to predict how these stocks continue to move and also we want to predict how new stocks added to the stock exchange would move. We do have lots and lots of data but doesn't this introduce tons of noise also? One obvious solution is to cluster the constituents. So in case of the NYSE scenario, we could have a bunch of stocks getting represented by a single index - the tech index. We could have brick n mortar companies as another cluster, retail, housing and form dozen odd clusters. Calculate the movements for each cluster separately...as they could have different weights and balances. Now we are left with time-series data of a dozen odd indices - all of which will help determine the NYSE index futures more predictably. If we apply this sort of a scenario to RiiiD, this means that clusters of students need to be created and then it makes sense to have separate sets of weights and biases for each category? Let existing student continue in the cluster they already are in and the new students are moved to one of the clusters. But I haven't seen this kind of approach in any time-series model leveraging encpder/decoders.\n\nWe still haven't talked of validation strategies, thought about what could be the best combination of keys queries and values or the importance of online learning in this competition or usage of reductions and many others. But we have covered a fair bit of theory otherwise. While kernels are shared abundantly, I felt some theoretical discussion on the features and architectures could help newcomers like me and hence quickly jotted a few points that stuck me as I went thru' a few of the beautiful kernels.\n\nIn Silogram's golden words, this competition is not just a test of AI skills, but is also a test of our engineering skills. How creatively we make use of the imposed constraints will decide the winner. With unlimited processing power and memory, the NN models would have won hands down but with the imposed constraints a blend of NN and non-NN models may hold the key.\n\nAs with all features and architectures, there are risks/rewards. \n\nChoose wisely. All the very best to all!",
    "1112034": "I have benefited from every discussion you have had and look forward to next week's discussion.I will check my feature engineering",
    "1112108": "Hey !\n\nThanks for the messages.\nSome thoughts from me:\n\n- I prefere to use **Timestamp** rather than **prior_question_elapsed_time** to define if a user is new or not (it is precised that timestamp is set to 0 for every first interaction of a user.\n\n- Something very interesting but that is hard to capture is the **lag_time between two exercices**. It can gives indication of a student taking notes of the last question for example.\n\n- Still in the timestamp aspect, I think it can be interesting to check last time a user had an interaction with a section of the toeic. If a user start to improve on a topic, but then don't go to it for days, he might have some troubles redoing the exercices of that section.",
    "1112189": "Thanks Qiaqia. All the very best to you for the competition!!!\nAs only couple of weekends left now, I may not be able to post anything tangible...but all the best once again!",
    "1112403": "Point 1 - Yes\nPoint 2 - The difference between the two consecutive timestamps for a student should give an indication of lag time? So like in the Saint+ paper, we do have both the elapsed time and the lag time here.. Of course the elapsed time is for the prev question...Still it does help apparently...\n\njust to be clear, their definition is as follows:\nelapsed time - the time taken for a student to answer (the previous question)\nlag time - the time interval between adjacent learning activities.\n\nFrom this, I guess we can calculate the difference between times to see if the student was taking some notes after exercises if that is what is needed. I dont know how strong that feature would be though. However this would indeed be a very strong feature for lectures - amount of time student watched the lecture versus average time lecture is typically watched but as far as I can see they dont give elapsed time for lectures. Still we can try to guess try to play around with timestamp differences and make some guesses based on lag times (of course we have to assume that student is doing all this in one continuous session and hasn't opened a new window to play their favorite video game :) could be a noisy feature but worth a try",
    "1139368": "allohvk thank you for compiling such a helpful post!! actually all of your discussions are worth a lot -:) I'll have to read many times to digest.. regarding the following point:\n> HOWEVER, we need to ensure that the data is richly represented. We could make use of a (learnable) weight matrix to transform the one-dimension vector into a certain dimension vector. Alternatively we could use self-attention to embed the input. Self-attention with all the fancy multiheads, residuals and all the standard tricks result in capturing the richness of the data well. Using self-attention will require positional encodings as the embedding has no knowledge of sequence order. There are some standard well-defined hacks to incorporate this and we will not go into the details. \n\n could you please point me to any example code/paper that uses self-att to capture the input? I don't remember to have come across smth like that - I'll give a try anyway but any source with info will help. Thank you and good luck to the comp",
    "1139426": "Thanks @imeintanis :)\nThis was written at the very beginning of the competition. So there could be slight changes here  and there. Definitely self-attention should help. There are couple of Saint implementations and I believe they use self-attention because that is part of Saint architecture by default. Have you checked those. Unfortnuately my Kaggle ID had an issue till last week and only last weekend I took the SAKT kernel from Wang and trying to make changes to make it work. I am still stuck at basic changes..",
    "1139442": "Thank you I got it (actually I had misunderstood your point). Happy that you resolved your issues & you 're back again\nps: I started from the same SAKT kernel but cannot surpass the 0780 stage.. we'll see",
    "1139568": "prior_question_elapsed_time is `<NA>` for lectures, regrettably. The full dataset has it; but not the version they prepared for our Kaggle.",
    "1139600": "authman did u find a way to create a moving mask for the question bundles so that they dont attend to answers in same bundle? I can share my thought if u r looking for a solution",
    "1139647": "Not satisfactorily =\\",
    "1139745": "see if this works..\n\nThe trick here is to create separate masks for interaction sequence and question sequence in each row (if u r using SAKT) and create multiple rows\nQ1 Q2 Q3 Q4 Q5 Q6, Q7\nI0 I1 I2 I3 I4 I5 I6\n\nLet us say Q3, Q4, Q5 are a bundle, then:\nI0 I1 I2 0 0 0 0 \nQ1 Q2 Q3 Q4 Q5 0 0\n\nand\nI0 I1 I2 I3 I4 I5 I6\n0 0 0 0 0 Q6 Q7\n\nRepeat for all bundles in row"
  },
  "source": "meta"
}