{
  "id": 209798,
  "title": "16th Place - Single Model",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209798",
  "author_name": "MPWARE",
  "post_date": "2021-01-08T16:05:25.672000",
  "votes": 69,
  "comment_count": 54,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>Here is an insight of our 2 solutions that both scored public 0.812/private 0.815 and that reached 16th gold place.</p>\n<p>This competition was both ML and engineering optimization to make everything work in 9h with 13GB RAM/16GB GPU. We spent almost 30% of time on optimization to keep the last 512 interactions per users +  per content attempts in memory + required for our features.</p>\n<p>We would like to thank Kaggle and RIIID organizers for this great competition! Congratulations to the top teams and all competitors for their motivation all along the challenge.<br>\nI would like to thank my teammates <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a>, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> and <a href=\"https://www.kaggle.com/matthiasanderer\" target=\"_blank\">@matthiasanderer</a>. You've been amazing, I've learnt a lot from you. I really enjoyed this competition.</p>\n<h2>Solution 1: Single transformer model</h2>\n<p>The SAINT+ model is described here <a href=\"https://arxiv.org/pdf/2010.12042.pdf\" target=\"_blank\">https://arxiv.org/pdf/2010.12042.pdf</a><br>\nThe code for our SAINT+ adaptation is available here <a href=\"https://github.com/rafiko1/Riiid-sharing\" target=\"_blank\">https://github.com/rafiko1/Riiid-sharing</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fa47293277b9e989ab5c81c269e6187a8%2Fdoc_saint.png?generation=1610121546239700&amp;alt=media\" alt=\"\"></p>\n<p>Our single model SAINT++ achieved, CV: 0.812 Public LB: 0.812, Private LB: 0.815</p>\n<p>We trained with 95% of users first then fine tuned with all data using smart window technique (see diagram below for SAKT). The model is simple in terms of features. It only contains the four features of SAINT+ (pictured above), with one additional feature - the number of attempts of a user for specific content (hence SAINT++):</p>\n<ul>\n<li>Content id</li>\n<li>Lag time</li>\n<li>Prior question elapsed time</li>\n<li>Previous responses</li>\n<li>Number of attempts </li>\n</ul>\n<p>The greatest improvement in features compared to SAINT+ came from grouping lag time into seconds, unlike minutes as done in the paper. <br>\nThen, we went bigger and bigger on the architecture and burned some GPU power 🔥. We increased on parameters of the model, most importantly the sequence length and number of layers. Final parameters of the model are as follows: </p>\n<ol>\n<li>Input Sequence length: 512</li>\n<li>Encoding layers: 4</li>\n<li>Decoding Layers: 4</li>\n<li>Embedding size: 288</li>\n<li>Dense Layer: 768</li>\n<li>heads: 8</li>\n<li>Dropout: 0.20</li>\n</ol>\n<p>We used the Noam learning rate scheduler: with initial warmup and exponential decrease down to 2e-5. </p>\n<p>Final improvement came from our <strong><em>recursive trick</em></strong> during inference. Here, we rounded predictions that came from the same bundle to <strong><em>0 or 1 </em></strong>- as their true response is unknown in time yet. The rounded predictions are then fed back to the model to predict the next response within the same bundle. This trick boosts CV LB +0.0025, but requires a batch size of 1, so we couldn’t ensemble multiple transformer models.</p>\n<h3>Solution 2: Ensemble of transformer, modified SAKT and LGB</h3>\n<ul>\n<li><p>LightGBM model scored CV=0.793, public LB=0.792 with 44 features.  <br>\nOur main features: <br>\nQuestion correctness per content and per user, tags 1 and 2, part, elapsed time, had explanation, number of attempts, multiple lags, running average (answer), multiple rolling means/median (answers, lags, elapsed time) + weighted mean, mean after/before 30 interactions, multiple momentums (lag, answers), per part correctness, per session (8 hours split) running average. Only 3 categories: part, tags1, tags2. Train/valid split from <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">Tito</a>. </p></li>\n<li><p>Pytorch SAKT modified scored CV=0.786, LB=0.789 with additional features. Training procedure with a smart window. </p></li>\n<li><p>When a user’s sequence length is larger than model input, i.e. N&gt;W, then using random crops gives +0.002 CV versus tiled crops. And using smart random crops gives +0.003 CV versus tiled crops. Basic random crops have a low probability of selecting the early or late questions from a user’s sequence whereas smart random crops have an equally likely probability of selecting all questions from a user.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fa1239fda69769669432879b2b1a4ae39%2Fdoc_window.png?generation=1610121604322940&amp;alt=media\" alt=\"\"></p></li>\n\n\n\n\n<li><p>TensorFlow Transformer model alone scored CV=0.811, LB=0.811</p>\n<p>Same as solution#1 but with sequence length = 256</p></li>\n</ul>\n<h3>What did not work:</h3>\n<ul>\n<li>TabNet</li>\n<li>Features with lectures for LGB. It worked on CV but not on LB (might be an issue in inference).</li>\n<li>Post processing using absolute position of question aka. question sequence number. Plotting mean(answered_correctly) vs question number looked like the image below. We can see that the 30 first questions have a different distribution compared with the rest. Also looks like there are subsequently batches of 30 questions (becomes visible if we zoom in the plot below). PP using that information worked in CV improving by around 0.0009, but didn’t work on LB.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe4b8abca67c11258a438304b8fc67db5%2Fdoc_pp.png?generation=1610121577542081&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h3>What worked partially:</h3>\n<p>But was not applicable for us within the 9h runtime limit:</p>\n<ul>\n<li>More than 3 models ensemble</li>\n<li>Level 2 model (XGB) could boost by +0.001</li>\n</ul>\n<h3>Lessons learnt:</h3>\n<ul>\n<li>Start inference Kernel as soon as possible when you need to deal with an API.</li>\n<li>Try to simulate API locally to understand how data will be handled. <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">Tito</a>’s simulator was perfect for that purpose.</li>\n<li>Push your inference (with the simulator) to the limits to debug it, it will avoid the frustrating “submission scoring error”. </li>\n<li>Team-up at some point, your teammates always have good ideas.</li>\n</ul>\n<p>One additional word to Kaggle <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I loved your API and the way it hides private data, it’s more realistic as in real world usage/production and it avoided chaotic blending. Congratulations for that, however, even if I guess you want to prevent probing, you should find a solution to provide better error feedback. If it is not possible (the more error codes the more probing) then you need to provide a simulator and guidelines allowing competitors to troubleshoot locally.</p>",
  "messages": [
    {
      "id": 1144696,
      "postDate": "2021-01-08T16:05:25.673Z",
      "content": "<p>Hi all,</p>\n<p>Here is an insight of our 2 solutions that both scored public 0.812/private 0.815 and that reached 16th gold place.</p>\n<p>This competition was both ML and engineering optimization to make everything work in 9h with 13GB RAM/16GB GPU. We spent almost 30% of time on optimization to keep the last 512 interactions per users +  per content attempts in memory + required for our features.</p>\n<p>We would like to thank Kaggle and RIIID organizers for this great competition! Congratulations to the top teams and all competitors for their motivation all along the challenge.<br>\nI would like to thank my teammates <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a>, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> and <a href=\"https://www.kaggle.com/matthiasanderer\" target=\"_blank\">@matthiasanderer</a>. You've been amazing, I've learnt a lot from you. I really enjoyed this competition.</p>\n<h2>Solution 1: Single transformer model</h2>\n<p>The SAINT+ model is described here <a href=\"https://arxiv.org/pdf/2010.12042.pdf\" target=\"_blank\">https://arxiv.org/pdf/2010.12042.pdf</a><br>\nThe code for our SAINT+ adaptation is available here <a href=\"https://github.com/rafiko1/Riiid-sharing\" target=\"_blank\">https://github.com/rafiko1/Riiid-sharing</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fa47293277b9e989ab5c81c269e6187a8%2Fdoc_saint.png?generation=1610121546239700&amp;alt=media\" alt=\"\"></p>\n<p>Our single model SAINT++ achieved, CV: 0.812 Public LB: 0.812, Private LB: 0.815</p>\n<p>We trained with 95% of users first then fine tuned with all data using smart window technique (see diagram below for SAKT). The model is simple in terms of features. It only contains the four features of SAINT+ (pictured above), with one additional feature - the number of attempts of a user for specific content (hence SAINT++):</p>\n<ul>\n<li>Content id</li>\n<li>Lag time</li>\n<li>Prior question elapsed time</li>\n<li>Previous responses</li>\n<li>Number of attempts </li>\n</ul>\n<p>The greatest improvement in features compared to SAINT+ came from grouping lag time into seconds, unlike minutes as done in the paper. <br>\nThen, we went bigger and bigger on the architecture and burned some GPU power 🔥. We increased on parameters of the model, most importantly the sequence length and number of layers. Final parameters of the model are as follows: </p>\n<ol>\n<li>Input Sequence length: 512</li>\n<li>Encoding layers: 4</li>\n<li>Decoding Layers: 4</li>\n<li>Embedding size: 288</li>\n<li>Dense Layer: 768</li>\n<li>heads: 8</li>\n<li>Dropout: 0.20</li>\n</ol>\n<p>We used the Noam learning rate scheduler: with initial warmup and exponential decrease down to 2e-5. </p>\n<p>Final improvement came from our <strong><em>recursive trick</em></strong> during inference. Here, we rounded predictions that came from the same bundle to <strong><em>0 or 1 </em></strong>- as their true response is unknown in time yet. The rounded predictions are then fed back to the model to predict the next response within the same bundle. This trick boosts CV LB +0.0025, but requires a batch size of 1, so we couldn’t ensemble multiple transformer models.</p>\n<h3>Solution 2: Ensemble of transformer, modified SAKT and LGB</h3>\n<ul>\n<li><p>LightGBM model scored CV=0.793, public LB=0.792 with 44 features.  <br>\nOur main features: <br>\nQuestion correctness per content and per user, tags 1 and 2, part, elapsed time, had explanation, number of attempts, multiple lags, running average (answer), multiple rolling means/median (answers, lags, elapsed time) + weighted mean, mean after/before 30 interactions, multiple momentums (lag, answers), per part correctness, per session (8 hours split) running average. Only 3 categories: part, tags1, tags2. Train/valid split from <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">Tito</a>. </p></li>\n<li><p>Pytorch SAKT modified scored CV=0.786, LB=0.789 with additional features. Training procedure with a smart window. </p></li>\n<li><p>When a user’s sequence length is larger than model input, i.e. N&gt;W, then using random crops gives +0.002 CV versus tiled crops. And using smart random crops gives +0.003 CV versus tiled crops. Basic random crops have a low probability of selecting the early or late questions from a user’s sequence whereas smart random crops have an equally likely probability of selecting all questions from a user.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fa1239fda69769669432879b2b1a4ae39%2Fdoc_window.png?generation=1610121604322940&amp;alt=media\" alt=\"\"></p></li>\n\n\n\n\n<li><p>TensorFlow Transformer model alone scored CV=0.811, LB=0.811</p>\n<p>Same as solution#1 but with sequence length = 256</p></li>\n</ul>\n<h3>What did not work:</h3>\n<ul>\n<li>TabNet</li>\n<li>Features with lectures for LGB. It worked on CV but not on LB (might be an issue in inference).</li>\n<li>Post processing using absolute position of question aka. question sequence number. Plotting mean(answered_correctly) vs question number looked like the image below. We can see that the 30 first questions have a different distribution compared with the rest. Also looks like there are subsequently batches of 30 questions (becomes visible if we zoom in the plot below). PP using that information worked in CV improving by around 0.0009, but didn’t work on LB.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe4b8abca67c11258a438304b8fc67db5%2Fdoc_pp.png?generation=1610121577542081&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h3>What worked partially:</h3>\n<p>But was not applicable for us within the 9h runtime limit:</p>\n<ul>\n<li>More than 3 models ensemble</li>\n<li>Level 2 model (XGB) could boost by +0.001</li>\n</ul>\n<h3>Lessons learnt:</h3>\n<ul>\n<li>Start inference Kernel as soon as possible when you need to deal with an API.</li>\n<li>Try to simulate API locally to understand how data will be handled. <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">Tito</a>’s simulator was perfect for that purpose.</li>\n<li>Push your inference (with the simulator) to the limits to debug it, it will avoid the frustrating “submission scoring error”. </li>\n<li>Team-up at some point, your teammates always have good ideas.</li>\n</ul>\n<p>One additional word to Kaggle <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I loved your API and the way it hides private data, it’s more realistic as in real world usage/production and it avoided chaotic blending. Congratulations for that, however, even if I guess you want to prevent probing, you should find a solution to provide better error feedback. If it is not possible (the more error codes the more probing) then you need to provide a simulator and guidelines allowing competitors to troubleshoot locally.</p>",
      "rawMarkdown": "Hi all,\n\n\nHere is an insight of our 2 solutions that both scored public 0.812/private 0.815 and that reached 16th gold place.\n\n\nThis competition was both ML and engineering optimization to make everything work in 9h with 13GB RAM/16GB GPU. We spent almost 30% of time on optimization to keep the last 512 interactions per users +  per content attempts in memory + required for our features.\n\nWe would like to thank Kaggle and RIIID organizers for this great competition! Congratulations to the top teams and all competitors for their motivation all along the challenge.\nI would like to thank my teammates @rafiko1, @cdeotte @titericz and @matthiasanderer. You've been amazing, I've learnt a lot from you. I really enjoyed this competition.\n\n\n## Solution 1: Single transformer model\n\nThe SAINT+ model is described here https://arxiv.org/pdf/2010.12042.pdf\nThe code for our SAINT+ adaptation is available here https://github.com/rafiko1/Riiid-sharing.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fa47293277b9e989ab5c81c269e6187a8%2Fdoc_saint.png?generation=1610121546239700&alt=media)\n\n\n\n\nOur single model SAINT++ achieved, CV: 0.812 Public LB: 0.812, Private LB: 0.815\n\nWe trained with 95% of users first then fine tuned with all data using smart window technique (see diagram below for SAKT). The model is simple in terms of features. It only contains the four features of SAINT+ (pictured above), with one additional feature - the number of attempts of a user for specific content (hence SAINT++):\n\n\n\n*   Content id\n*   Lag time\n*   Prior question elapsed time\n*   Previous responses\n*   Number of attempts \n\nThe greatest improvement in features compared to SAINT+ came from grouping lag time into seconds, unlike minutes as done in the paper. \nThen, we went bigger and bigger on the architecture and burned some GPU power 🔥. We increased on parameters of the model, most importantly the sequence length and number of layers. Final parameters of the model are as follows: \n\n\n\n\n1. Input Sequence length: 512\n2. Encoding layers: 4\n3. Decoding Layers: 4\n4. Embedding size: 288\n5. Dense Layer: 768\n6. heads: 8\n7. Dropout: 0.20\n\nWe used the Noam learning rate scheduler: with initial warmup and exponential decrease down to 2e-5. \n \nFinal improvement came from our **_recursive trick_** during inference. Here, we rounded predictions that came from the same bundle to **_0 or 1 _**- as their true response is unknown in time yet. The rounded predictions are then fed back to the model to predict the next response within the same bundle. This trick boosts CV LB +0.0025, but requires a batch size of 1, so we couldn’t ensemble multiple transformer models.\n\n\n### Solution 2: Ensemble of transformer, modified SAKT and LGB\n\n\n\n*   LightGBM model scored CV=0.793, public LB=0.792 with 44 features.  \nOur main features: \nQuestion correctness per content and per user, tags 1 and 2, part, elapsed time, had explanation, number of attempts, multiple lags, running average (answer), multiple rolling means/median (answers, lags, elapsed time) + weighted mean, mean after/before 30 interactions, multiple momentums (lag, answers), per part correctness, per session (8 hours split) running average. Only 3 categories: part, tags1, tags2. Train/valid split from [Tito](https://www.kaggle.com/its7171/cv-strategy). \n\n\n*   Pytorch SAKT modified scored CV=0.786, LB=0.789 with additional features. Training procedure with a smart window. \n*   When a user’s sequence length is larger than model input, i.e. N>W, then using random crops gives +0.002 CV versus tiled crops. And using smart random crops gives +0.003 CV versus tiled crops. Basic random crops have a low probability of selecting the early or late questions from a user’s sequence whereas smart random crops have an equally likely probability of selecting all questions from a user.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fa1239fda69769669432879b2b1a4ae39%2Fdoc_window.png?generation=1610121604322940&alt=media)\n    \n\n\n \n\n\n*   TensorFlow Transformer model alone scored CV=0.811, LB=0.811\n\n    Same as solution#1 but with sequence length = 256\n\n\n\n### What did not work:\n\n\n\n*   TabNet\n*   Features with lectures for LGB. It worked on CV but not on LB (might be an issue in inference).\n*   Post processing using absolute position of question aka. question sequence number. Plotting mean(answered_correctly) vs question number looked like the image below. We can see that the 30 first questions have a different distribution compared with the rest. Also looks like there are subsequently batches of 30 questions (becomes visible if we zoom in the plot below). PP using that information worked in CV improving by around 0.0009, but didn’t work on LB.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe4b8abca67c11258a438304b8fc67db5%2Fdoc_pp.png?generation=1610121577542081&alt=media)\n    \n\n\n\n\n\n### What worked partially:\n\n \nBut was not applicable for us within the 9h runtime limit:\n\n\n\n*   More than 3 models ensemble\n*   Level 2 model (XGB) could boost by +0.001\n\n\n### Lessons learnt:\n\n\n\n*   Start inference Kernel as soon as possible when you need to deal with an API.\n*   Try to simulate API locally to understand how data will be handled. [Tito](https://www.kaggle.com/its7171/time-series-api-iter-test-emulator)’s simulator was perfect for that purpose.\n*   Push your inference (with the simulator) to the limits to debug it, it will avoid the frustrating “submission scoring error”. \n*   Team-up at some point, your teammates always have good ideas.\n\nOne additional word to Kaggle @sohier I loved your API and the way it hides private data, it’s more realistic as in real world usage/production and it avoided chaotic blending. Congratulations for that, however, even if I guess you want to prevent probing, you should find a solution to provide better error feedback. If it is not possible (the more error codes the more probing) then you need to provide a simulator and guidelines allowing competitors to troubleshoot locally.\n",
      "votes": 68
    },
    {
      "id": 1146920,
      "postDate": "2021-01-10T07:05:16.133Z",
      "content": "<p><a href=\"https://www.kaggle.com/MPWare\" target=\"_blank\">@MPWare</a> Thanks for sharing the \"Simple\" Pytorch code, but why does it score so much less than Tensorflow?</p>\n<ol>\n<li>The Pytorch model uses 2 layers instead of Tensorflow uses 4 layers</li>\n<li>The Pytorch model uses 100 sequence length instead of Tensorflow uses 512 sequence length</li>\n<li>The Pytorch model uses 256 embedding size instead of Tensorflow uses 288</li>\n<li>The Pytorch model uses dropout = 0.1 but Tensorflow uses dropout = 0.2</li>\n<li>The Pytorch model only uses exercise_id  in encoder, but Tensorflow uses and response, but Tensorflow uses exercise_id, Lag time, prior question elapsed time, previous responses, number of attempts. Both models just use response_correct embedding in the decoder</li>\n</ol>\n<p>Are there any other differences? Would the Pytorch model be as good if you changed it to use all these features? Ty, I will really cherish that \"simple\" pytorch demo notebook you provided.</p>",
      "rawMarkdown": "@MPWare Thanks for sharing the \"Simple\" Pytorch code, but why does it score so much less than Tensorflow?\n\n1. The Pytorch model uses 2 layers instead of Tensorflow uses 4 layers\n2. The Pytorch model uses 100 sequence length instead of Tensorflow uses 512 sequence length\n3. The Pytorch model uses 256 embedding size instead of Tensorflow uses 288\n4. The Pytorch model uses dropout = 0.1 but Tensorflow uses dropout = 0.2\n5. The Pytorch model only uses exercise_id  in encoder, but Tensorflow uses and response, but Tensorflow uses exercise_id, Lag time, prior question elapsed time, previous responses, number of attempts. Both models just use response_correct embedding in the decoder\n\nAre there any other differences? Would the Pytorch model be as good if you changed it to use all these features? Ty, I will really cherish that \"simple\" pytorch demo notebook you provided.",
      "votes": 3,
      "replies": [
        {
          "id": 1147039,
          "postDate": "2021-01-10T08:50:14.530Z",
          "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> As you've noticed, the main difference is that I'm not using exactly the same features. I'm not using <code>attempts</code> at all. The configuration to get LB=0.795 with Pytorch model is below, it requires <code>seq_len=256</code>. We've noticed that <code>seq_len</code> is a quite important to get a better model. The training procedure is also important, I'm not using the smart window described by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. I've spent some time to fix the <code>NaN</code> loss issue with Pytorch <code>MultiHeadAttention</code> + padding masks. Root cause was right padding instead of left padding. I've stopped to try improving it once we got better results with <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> TensorFlow model which uses <code>seq_len=512</code>. Other difference is <code>sigmoid</code> vs <code>softmax</code> but I don't think it's really important.</p>\n<pre><code>class raw_conf:\n\n    pad_mode = \"token\"\n    pad_right = False\n    flatten = True\n    sampler = None # \"prob\" # Option to give long user's sequence higher probability\n\n    seq_len = 256 # 100\n    embedding_dim = 256\n    exercices_id_size = 13523\n    exercices_part_size = 7\n    response_size = 2 \n    elapsed_time_cat = True\n    elapsed_time_size = 73 # Categories after binning\n    lag_time_cat = True\n    lag_time_size = 366 # Categories after binning\n    explanation_size = 2\n\n    # Model\n    nhead = 8 \n    num_encoder_layers = 4\n    num_decoder_layers = 4\n    dim_feedforward = 2048\n    dropout = 0.1\n    activation = None\n    num_classes = 1\n    loss = MaskedBCEWithLogitsLoss(num_classes)\n    post_activation = \"sigmoid\"\n\n    optimizer = \"Noam\" # \"Adam\"\n    scheduler = \"Cosine\" if optimizer == \"Adam\" else None\n    lr = 0.0001\n    min_lr = 0.00005\n    beta1 = 0.9\n\n    BATCH_SIZE = 128\n</code></pre>",
          "rawMarkdown": "@returnofsputnik As you've noticed, the main difference is that I'm not using exactly the same features. I'm not using `attempts` at all. The configuration to get LB=0.795 with Pytorch model is below, it requires `seq_len=256`. We've noticed that `seq_len` is a quite important to get a better model. The training procedure is also important, I'm not using the smart window described by @cdeotte. I've spent some time to fix the `NaN` loss issue with Pytorch `MultiHeadAttention` + padding masks. Root cause was right padding instead of left padding. I've stopped to try improving it once we got better results with @rafiko1 TensorFlow model which uses `seq_len=512`. Other difference is `sigmoid` vs `softmax` but I don't think it's really important.\n\n\n```\nclass raw_conf:\n\n    pad_mode = \"token\"\n    pad_right = False\n    flatten = True\n    sampler = None # \"prob\" # Option to give long user's sequence higher probability\n\n    seq_len = 256 # 100\n    embedding_dim = 256\n    exercices_id_size = 13523\n    exercices_part_size = 7\n    response_size = 2 \n    elapsed_time_cat = True\n    elapsed_time_size = 73 # Categories after binning\n    lag_time_cat = True\n    lag_time_size = 366 # Categories after binning\n    explanation_size = 2\n    \n    # Model\n    nhead = 8 \n    num_encoder_layers = 4\n    num_decoder_layers = 4\n    dim_feedforward = 2048\n    dropout = 0.1\n    activation = None\n    num_classes = 1\n    loss = MaskedBCEWithLogitsLoss(num_classes)\n    post_activation = \"sigmoid\"\n\n    optimizer = \"Noam\" # \"Adam\"\n    scheduler = \"Cosine\" if optimizer == \"Adam\" else None\n    lr = 0.0001\n    min_lr = 0.00005\n    beta1 = 0.9\n\n    BATCH_SIZE = 128\n```",
          "votes": 3
        },
        {
          "id": 1147157,
          "postDate": "2021-01-10T10:08:37.647Z",
          "content": "<p>We're going to upload the full pytorch model on GitHub with the weights.</p>",
          "rawMarkdown": "We're going to upload the full pytorch model on GitHub with the weights.",
          "votes": 3
        },
        {
          "id": 1147512,
          "postDate": "2021-01-10T14:59:40.747Z",
          "content": "<p>Thank you. Since I know PT better than TF, I want to use your PT Transformer model for many future competitions since to me it is the simplest since it invokes nn.Transformer.</p>",
          "rawMarkdown": "Thank you. Since I know PT better than TF, I want to use your PT Transformer model for many future competitions since to me it is the simplest since it invokes nn.Transformer.",
          "votes": 1
        },
        {
          "id": 1149848,
          "postDate": "2021-01-12T07:50:08.993Z",
          "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> We've uploaded Pytorch full code here:<br>\n<a href=\"https://github.com/rafiko1/Riiid-sharing/tree/main/pytorch\" target=\"_blank\">https://github.com/rafiko1/Riiid-sharing/tree/main/pytorch</a><br>\nIt's MODEL-PT-v14</p>",
          "rawMarkdown": "@returnofsputnik We've uploaded Pytorch full code here:\nhttps://github.com/rafiko1/Riiid-sharing/tree/main/pytorch\nIt's MODEL-PT-v14",
          "votes": 2
        },
        {
          "id": 1149949,
          "postDate": "2021-01-12T09:07:52.317Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> thanks. I will go through it and will try to reproduce your solution.</p>",
          "rawMarkdown": "@mpware thanks. I will go through it and will try to reproduce your solution."
        },
        {
          "id": 1149996,
          "postDate": "2021-01-12T09:33:39.397Z",
          "content": "<p>The best model is with TensorFlow (<code>seq_len=512</code>), Public LB=0.812, Pytorch (<code>seq_len=256</code>) one is only Public LB=0.795.</p>",
          "rawMarkdown": "The best model is with TensorFlow (`seq_len=512`), Public LB=0.812, Pytorch (`seq_len=256`) one is only Public LB=0.795.",
          "votes": 1
        },
        {
          "id": 1150265,
          "postDate": "2021-01-12T13:37:10.160Z",
          "content": "<p>Thanks for the complete src!</p>",
          "rawMarkdown": "Thanks for the complete src!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1146671,
      "postDate": "2021-01-09T23:46:57.540Z",
      "content": "<p>Congrats &amp; thank you for sharing your approach and codes.</p>",
      "rawMarkdown": "Congrats & thank you for sharing your approach and codes.",
      "votes": 3
    },
    {
      "id": 1146361,
      "postDate": "2021-01-09T18:15:33.017Z",
      "content": "<p>Congratulations and thanks for sharing your solution, I and many others will be able to learn a lot from it. I also used Tensorflow and implemented Saint+ however my code to sample sequences/users was not robust so I intend to see how the model would have performed using the sampling code you have shared.</p>\n<p>I have read through the code in demo_riiid_train.ipynb and think I understand how you have set up training and validation sets but was hoping you could confirm my understanding?</p>\n<ol>\n<li>You allocate distinct users to both the training and validation set</li>\n<li>You then assign a probability to each user based on sequence length. For say user 115 in the training set their probability is their sequence length divided by the sum of all sequence lengths in training set.</li>\n<li>You then generate a one off validation set where you randomly sample users with replacement and take <strong>random training crops</strong> for each row. Based on the 'select_window_size' function</li>\n<li>For the training dataset you do the same but repeat step 3 each epoch</li>\n</ol>\n<p>You sample using N_SELECT_PER_EPOCH = 100000, was this just a hyperparameter you tuned?</p>\n<p>Thank you </p>",
      "rawMarkdown": "Congratulations and thanks for sharing your solution, I and many others will be able to learn a lot from it. I also used Tensorflow and implemented Saint+ however my code to sample sequences/users was not robust so I intend to see how the model would have performed using the sampling code you have shared.\n\nI have read through the code in demo_riiid_train.ipynb and think I understand how you have set up training and validation sets but was hoping you could confirm my understanding?\n\n1. You allocate distinct users to both the training and validation set\n2. You then assign a probability to each user based on sequence length. For say user 115 in the training set their probability is their sequence length divided by the sum of all sequence lengths in training set.\n3. You then generate a one off validation set where you randomly sample users with replacement and take **random training crops** for each row. Based on the 'select_window_size' function\n4. For the training dataset you do the same but repeat step 3 each epoch\n\nYou sample using N_SELECT_PER_EPOCH = 100000, was this just a hyperparameter you tuned?\n\n\nThank you ",
      "votes": 3,
      "replies": [
        {
          "id": 1146978,
          "postDate": "2021-01-10T08:15:12.450Z",
          "content": "<ol>\n<li>Correct. Distinct users gives a simple and reliable validation.</li>\n<li>Correct. It will be used inside <code>select_window_size</code> as the equivalent of <code>WeightedRandomSampler</code> in Pytorch</li>\n<li>Correct. This is important for training. For validation it's the convenient choice, but not the best choice. Better is to take all sequences for validation instead.</li>\n<li>Correct.</li>\n</ol>\n<p><code>N_SELECT_PER_EPOCH</code> is defining how many samples to pass within each epoch. It's not really a hyperparameter to tune. Just how long you'd like each epoch to be by specifying number of samples.</p>",
          "rawMarkdown": "1. Correct. Distinct users gives a simple and reliable validation.\n2. Correct. It will be used inside `select_window_size` as the equivalent of `WeightedRandomSampler` in Pytorch\n3. Correct. This is important for training. For validation it's the convenient choice, but not the best choice. Better is to take all sequences for validation instead.\n4. Correct.\n\n`N_SELECT_PER_EPOCH` is defining how many samples to pass within each epoch. It's not really a hyperparameter to tune. Just how long you'd like each epoch to be by specifying number of samples.",
          "votes": 1
        },
        {
          "id": 1147277,
          "postDate": "2021-01-10T11:58:37.433Z",
          "content": "<p>Thank you, will do a late submission and see how my score would have changed.</p>",
          "rawMarkdown": "Thank you, will do a late submission and see how my score would have changed.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1146392,
      "postDate": "2021-01-09T18:35:27.793Z",
      "content": "<p>Congrats, single 0.812/0.815 with SAINT is really great!</p>",
      "rawMarkdown": "Congrats, single 0.812/0.815 with SAINT is really great!",
      "votes": 4,
      "replies": [
        {
          "id": 1146980,
          "postDate": "2021-01-10T08:17:43.870Z",
          "content": "<p>Congrats on your amazing 2nd place <a href=\"https://www.kaggle.com/mamas\" target=\"_blank\">@mamas</a>, well-deserved!</p>",
          "rawMarkdown": "Congrats on your amazing 2nd place @mamas, well-deserved!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1144929,
      "postDate": "2021-01-08T18:49:27.457Z",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> congrats on gold medal and thanks for sharing your approach. Would you like to share your code ?</p>",
      "rawMarkdown": "@mpware congrats on gold medal and thanks for sharing your approach. Would you like to share your code ?",
      "votes": 1,
      "replies": [
        {
          "id": 1144950,
          "postDate": "2021-01-08T19:12:24.127Z",
          "content": "<p>Thanks! Our best single model + features is with TensorFlow. <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> if you get a chance to share it …</p>\n<p>In the other <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632\" target=\"_blank\">thread</a>, it was a Pytorch model (with simple features) that only reached public LB 0.795. Here it is if it can help. </p>\n<pre><code># Model\nclass RIIIDModel(nn.Module):\n    def __init__(self, cfg, verbose=False):\n        super().__init__()\n        self.response_size = cfg.response_size\n        self.lag_time_size = cfg.lag_time_size\n        self.elapsed_time_size = cfg.elapsed_time_size\n        self.explanation_size = cfg.explanation_size\n        self.attempt_size = cfg.attempt_size\n        self.seq_len = cfg.seq_len\n        self.embedding_dim = cfg.embedding_dim\n        self.elapsed_time_cat = cfg.elapsed_time_cat\n        self.lag_time_cat = cfg.lag_time_cat\n        self.num_classes = cfg.num_classes\n        self.verbose = verbose\n        self.pad_mode = cfg.pad_mode\n\n        self.pos_encoder1 = None\n        self.pos_encoder2 = None\n\n        additional_token_dim = 1 if self.pad_mode == \"token\" else 0\n\n        # Exercices embeddings\n        self.exercices_id_embedding = nn.Embedding(cfg.exercices_id_size + additional_token_dim, self.embedding_dim)\n        self.exercices_part_embedding = nn.Embedding(cfg.exercices_part_size + additional_token_dim, self.embedding_dim) if cfg.exercices_part_size is not None else None\n\n        # Response embeddings\n        self.response_embedding = nn.Embedding(cfg.response_size + 1 + additional_token_dim, self.embedding_dim) # +1 to include start token\n\n        if self.elapsed_time_cat is True:\n            self.elapsed_time_embedding = nn.Embedding(cfg.elapsed_time_size + 1 + additional_token_dim, self.embedding_dim) if cfg.elapsed_time_size is not None else None # +1 to include start token\n        else:\n            self.elapsed_time_embedding = nn.Linear(1, self.embedding_dim, bias=False) if cfg.elapsed_time_size is not None else None # Continuous embedding\n\n        if self.lag_time_cat is True:\n            self.lag_time_embedding = nn.Embedding(cfg.lag_time_size + 1 + additional_token_dim, self.embedding_dim) if cfg.lag_time_size is not None else None # +1 to include start token\n        else:\n            self.lag_time_embedding = nn.Linear(1, self.embedding_dim, bias=False) if cfg.lag_time_size is not None else None # Continuous embedding\n\n        self.explanation_embedding = nn.Embedding(cfg.explanation_size + 1 + additional_token_dim, self.embedding_dim) if cfg.explanation_size is not None else None # +1 to include start token\n\n        input_features_dim = self.embedding_dim\n\n        # Position encoder (relative or absolute position of the tokens in the sequence)\n        if cfg.position_encoding_enabled is True:\n            self.pos_encoder1 = PositionalEncoding(input_features_dim, cfg.dropout)\n            self.pos_encoder2 = self.pos_encoder1\n\n        # Transformer with default encoder/decoder        \n        self.transformer = nn.Transformer(d_model=input_features_dim, \n                                          nhead=cfg.nhead, \n                                          num_encoder_layers=cfg.num_encoder_layers,\n                                          num_decoder_layers=cfg.num_decoder_layers, \n                                          dim_feedforward=cfg.dim_feedforward, \n                                          dropout=cfg.dropout, \n                                          activation='relu', \n                                          custom_encoder = None, \n                                          custom_decoder = None)\n\n        # Decoder\n        self.fc = nn.Linear(input_features_dim, self.num_classes)\n\n\n    # If a BoolTensor is provided, the positions with the value of True will be ignored while the position with the value of False will be unchanged.\n    # tensor([[False,  True,  True,  True],\n    #         [False, False,  True,  True],\n    #         [False, False, False,  True],\n    #         [False, False, False, False]])    \n    def generate_mask(self, size, diagonal=1):        \n        return torch.triu(torch.ones(size, size)==1, diagonal=diagonal)\n\n    def forward(self, data, src_mask=None, tgt_mask=None, mem_mask=None, src_key_padding_mask=None, tgt_key_padding_mask=None, memory_key_padding_mask=None):\n\n        # Each input is (BS, seq_len)\n        # Content\n        data_content_id = data[CONTENT_ID].long()\n        # Answers\n        data_response = data[TARGET].long()        \n\n        # Optional features\n        data_part = data[PART].long() if self.exercices_part_embedding is not None else None\n        if self.elapsed_time_cat is True:\n            data_elapsed_time = data[PRIOR_QUESTION_ELAPSED_TIME].long() if self.elapsed_time_embedding is not None else None\n        else:\n            data_elapsed_time = data[PRIOR_QUESTION_ELAPSED_TIME].float().unsqueeze(2) if self.elapsed_time_embedding is not None else None\n        if self.lag_time_cat is True:\n            data_lag_time = data[LAG].long() if self.lag_time_embedding is not None else None\n        else:\n            data_lag_time = data[LAG].float().unsqueeze(2) if self.lag_time_embedding is not None else None\n\n        data_explanation = data[PRIOR_QUESTION_HAD_EXPLANATION].long() if self.explanation_embedding is not None else None\n\n        # Start token(s)\n\n        # Add start token to correctness\n        data_response = torch.roll(data_response, shifts=(0, 1), dims=(0, 1)) # Shift right the sequence\n        data_response[:,0] = self.response_size # Start token (2)\n\n        # Add start token to lag time\n        if data_lag_time is not None:\n            data_lag_time = torch.roll(data_lag_time, shifts=(0, 1), dims=(0, 1)) # Shift right the sequence\n            data_lag_time[:,0] = self.lag_time_size # Start token\n\n        # Add start token to elapsed time\n        if data_elapsed_time is not None:\n            data_elapsed_time = torch.roll(data_elapsed_time, shifts=(0, 1), dims=(0, 1)) # Shift right the sequence\n            if self.elapsed_time_cat is True:\n                data_elapsed_time[:,0] = self.elapsed_time_size # Start token\n            else:\n                data_elapsed_time[:,0] = 0.0\n\n        # Add start token to explanation\n        if data_explanation is not None:\n            data_explanation = torch.roll(data_explanation, shifts=(0, 1), dims=(0, 1)) # Shift right the sequence\n            data_explanation[:,0] = self.explanation_size # Start token\n\n        # Questions, Part, Elapsed time, Lag embeddings\n        x_content_id = self.exercices_id_embedding(data_content_id) # (BS, seq_len, embedding_dim)\n\n        x_exercices_part = self.exercices_part_embedding(data_part) if self.exercices_part_embedding is not None else None # (BS, seq_len, embedding_dim)\n        x_elapsed_time = self.elapsed_time_embedding(data_elapsed_time) if self.elapsed_time_embedding is not None else None # (BS, seq_len, embedding_dim)\n        x_lag_time = self.lag_time_embedding(data_lag_time) if self.lag_time_embedding is not None else None # (BS, seq_len, embedding_dim)\n        x_explanation = self.explanation_embedding(data_explanation) if self.explanation_embedding is not None else None # (BS, seq_len, embedding_dim)\n\n        # Response embeddings\n        x_correctness = self.response_embedding(data_response) # (BS, seq_len, embedding_dim)\n\n        x_position = None\n\n        # Ei (sum of embeddings)\n        x_exercices = x_content_id\n\n        if x_exercices_part is not None:\n            x_exercices = x_exercices + x_exercices_part  # (BS, seq_len, embedding_dim)\n\n        x_position_exercices = self.pos_encoder1(x_exercices) if self.pos_encoder1 is not None else x_exercices # (BS, seq_len, embedding_dim)\n\n        # Ri (sum of embeddings) [S, R1, Rk-1], S is start token\n\n        x_responses = x_correctness \n\n        if x_lag_time is not None:\n            x_responses = x_responses + x_lag_time\n\n        if x_elapsed_time is not None:\n            x_responses = x_responses + x_elapsed_time # (BS, seq_len, embedding_dim)\n\n        if x_explanation is not None:\n            x_responses = x_responses + x_explanation # (BS, seq_len, embedding_dim)\n\n        if x_attempt is not None:\n            x_responses = x_responses + x_attempt # (BS, seq_len, embedding_dim)            \n\n        if x_exercices_task is not None:\n            x_responses = x_responses + x_exercices_task # (BS, seq_len, embedding_dim)\n\n\n        x_position_responses = self.pos_encoder2(x_responses) if self.pos_encoder2 is not None else x_responses # (BS, seq_len, embedding_dim)\n\n        # Transformer src: (S,N,E), tgt:(T,N,E), src_mask:(S,S), tgt_mask:(T,T)\n        # where S is the source sequence length, T is the target sequence length, N is the batch size, E is the feature number\n        # output: (T,N,E)\n        # src_key_padding_mask: (N,S), tgt_key_padding_mask: (N,T), memory_key_padding_mask: (N,S)    \n        x_position_exercices = x_position_exercices.transpose(1,0) # (seq_len, BS, embedding_dim)\n        x_position_responses = x_position_responses.transpose(1,0) # (seq_len, BS, embedding_dim)\n\n        x_transformer = self.transformer(src=x_position_exercices, tgt=x_position_responses, src_mask=src_mask, tgt_mask=tgt_mask, memory_mask=mem_mask, \n                                         src_key_padding_mask=src_key_padding_mask, tgt_key_padding_mask=tgt_key_padding_mask, memory_key_padding_mask=memory_key_padding_mask) # (seq_len, BS, embedding_dim)\n        x_transformer = x_transformer.transpose(1,0) # (BS, seq_len, embedding_dim)\n\n        output = self.fc(x_transformer)\n        output = output.squeeze(dim=2)\n\n        return output\n</code></pre>",
          "rawMarkdown": "Thanks! Our best single model + features is with TensorFlow. @rafiko1 if you get a chance to share it ...\n\nIn the other [thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632), it was a Pytorch model (with simple features) that only reached public LB 0.795. Here it is if it can help. \n\n```\n# Model\nclass RIIIDModel(nn.Module):\n    def __init__(self, cfg, verbose=False):\n        super().__init__()\n        self.response_size = cfg.response_size\n        self.lag_time_size = cfg.lag_time_size\n        self.elapsed_time_size = cfg.elapsed_time_size\n        self.explanation_size = cfg.explanation_size\n        self.attempt_size = cfg.attempt_size\n        self.seq_len = cfg.seq_len\n        self.embedding_dim = cfg.embedding_dim\n        self.elapsed_time_cat = cfg.elapsed_time_cat\n        self.lag_time_cat = cfg.lag_time_cat\n        self.num_classes = cfg.num_classes\n        self.verbose = verbose\n        self.pad_mode = cfg.pad_mode\n\n        self.pos_encoder1 = None\n        self.pos_encoder2 = None\n\n        additional_token_dim = 1 if self.pad_mode == \"token\" else 0\n\n        # Exercices embeddings\n        self.exercices_id_embedding = nn.Embedding(cfg.exercices_id_size + additional_token_dim, self.embedding_dim)\n        self.exercices_part_embedding = nn.Embedding(cfg.exercices_part_size + additional_token_dim, self.embedding_dim) if cfg.exercices_part_size is not None else None\n\n        # Response embeddings\n        self.response_embedding = nn.Embedding(cfg.response_size + 1 + additional_token_dim, self.embedding_dim) # +1 to include start token\n        \n        if self.elapsed_time_cat is True:\n            self.elapsed_time_embedding = nn.Embedding(cfg.elapsed_time_size + 1 + additional_token_dim, self.embedding_dim) if cfg.elapsed_time_size is not None else None # +1 to include start token\n        else:\n            self.elapsed_time_embedding = nn.Linear(1, self.embedding_dim, bias=False) if cfg.elapsed_time_size is not None else None # Continuous embedding\n\n        if self.lag_time_cat is True:\n            self.lag_time_embedding = nn.Embedding(cfg.lag_time_size + 1 + additional_token_dim, self.embedding_dim) if cfg.lag_time_size is not None else None # +1 to include start token\n        else:\n            self.lag_time_embedding = nn.Linear(1, self.embedding_dim, bias=False) if cfg.lag_time_size is not None else None # Continuous embedding\n            \n        self.explanation_embedding = nn.Embedding(cfg.explanation_size + 1 + additional_token_dim, self.embedding_dim) if cfg.explanation_size is not None else None # +1 to include start token\n        \n        input_features_dim = self.embedding_dim\n\n        # Position encoder (relative or absolute position of the tokens in the sequence)\n        if cfg.position_encoding_enabled is True:\n            self.pos_encoder1 = PositionalEncoding(input_features_dim, cfg.dropout)\n            self.pos_encoder2 = self.pos_encoder1\n\n        # Transformer with default encoder/decoder        \n        self.transformer = nn.Transformer(d_model=input_features_dim, \n                                          nhead=cfg.nhead, \n                                          num_encoder_layers=cfg.num_encoder_layers,\n                                          num_decoder_layers=cfg.num_decoder_layers, \n                                          dim_feedforward=cfg.dim_feedforward, \n                                          dropout=cfg.dropout, \n                                          activation='relu', \n                                          custom_encoder = None, \n                                          custom_decoder = None)\n                \n        # Decoder\n        self.fc = nn.Linear(input_features_dim, self.num_classes)\n\n\n    # If a BoolTensor is provided, the positions with the value of True will be ignored while the position with the value of False will be unchanged.\n    # tensor([[False,  True,  True,  True],\n    #         [False, False,  True,  True],\n    #         [False, False, False,  True],\n    #         [False, False, False, False]])    \n    def generate_mask(self, size, diagonal=1):        \n        return torch.triu(torch.ones(size, size)==1, diagonal=diagonal)\n\n    def forward(self, data, src_mask=None, tgt_mask=None, mem_mask=None, src_key_padding_mask=None, tgt_key_padding_mask=None, memory_key_padding_mask=None):\n        \n        # Each input is (BS, seq_len)\n        # Content\n        data_content_id = data[CONTENT_ID].long()\n        # Answers\n        data_response = data[TARGET].long()        \n\n        # Optional features\n        data_part = data[PART].long() if self.exercices_part_embedding is not None else None\n        if self.elapsed_time_cat is True:\n            data_elapsed_time = data[PRIOR_QUESTION_ELAPSED_TIME].long() if self.elapsed_time_embedding is not None else None\n        else:\n            data_elapsed_time = data[PRIOR_QUESTION_ELAPSED_TIME].float().unsqueeze(2) if self.elapsed_time_embedding is not None else None\n        if self.lag_time_cat is True:\n            data_lag_time = data[LAG].long() if self.lag_time_embedding is not None else None\n        else:\n            data_lag_time = data[LAG].float().unsqueeze(2) if self.lag_time_embedding is not None else None\n        \n        data_explanation = data[PRIOR_QUESTION_HAD_EXPLANATION].long() if self.explanation_embedding is not None else None\n\n        # Start token(s)\n\n        # Add start token to correctness\n        data_response = torch.roll(data_response, shifts=(0, 1), dims=(0, 1)) # Shift right the sequence\n        data_response[:,0] = self.response_size # Start token (2)\n\n        # Add start token to lag time\n        if data_lag_time is not None:\n            data_lag_time = torch.roll(data_lag_time, shifts=(0, 1), dims=(0, 1)) # Shift right the sequence\n            data_lag_time[:,0] = self.lag_time_size # Start token\n\n        # Add start token to elapsed time\n        if data_elapsed_time is not None:\n            data_elapsed_time = torch.roll(data_elapsed_time, shifts=(0, 1), dims=(0, 1)) # Shift right the sequence\n            if self.elapsed_time_cat is True:\n                data_elapsed_time[:,0] = self.elapsed_time_size # Start token\n            else:\n                data_elapsed_time[:,0] = 0.0\n\n        # Add start token to explanation\n        if data_explanation is not None:\n            data_explanation = torch.roll(data_explanation, shifts=(0, 1), dims=(0, 1)) # Shift right the sequence\n            data_explanation[:,0] = self.explanation_size # Start token\n        \n        # Questions, Part, Elapsed time, Lag embeddings\n        x_content_id = self.exercices_id_embedding(data_content_id) # (BS, seq_len, embedding_dim)\n\n        x_exercices_part = self.exercices_part_embedding(data_part) if self.exercices_part_embedding is not None else None # (BS, seq_len, embedding_dim)\n        x_elapsed_time = self.elapsed_time_embedding(data_elapsed_time) if self.elapsed_time_embedding is not None else None # (BS, seq_len, embedding_dim)\n        x_lag_time = self.lag_time_embedding(data_lag_time) if self.lag_time_embedding is not None else None # (BS, seq_len, embedding_dim)\n        x_explanation = self.explanation_embedding(data_explanation) if self.explanation_embedding is not None else None # (BS, seq_len, embedding_dim)\n\n        # Response embeddings\n        x_correctness = self.response_embedding(data_response) # (BS, seq_len, embedding_dim)\n\n        x_position = None\n\n        # Ei (sum of embeddings)\n        x_exercices = x_content_id\n        \n        if x_exercices_part is not None:\n            x_exercices = x_exercices + x_exercices_part  # (BS, seq_len, embedding_dim)\n        \n        x_position_exercices = self.pos_encoder1(x_exercices) if self.pos_encoder1 is not None else x_exercices # (BS, seq_len, embedding_dim)\n\n        # Ri (sum of embeddings) [S, R1, Rk-1], S is start token\n\n        x_responses = x_correctness \n        \n        if x_lag_time is not None:\n            x_responses = x_responses + x_lag_time\n\n        if x_elapsed_time is not None:\n            x_responses = x_responses + x_elapsed_time # (BS, seq_len, embedding_dim)\n\n        if x_explanation is not None:\n            x_responses = x_responses + x_explanation # (BS, seq_len, embedding_dim)\n\n        if x_attempt is not None:\n            x_responses = x_responses + x_attempt # (BS, seq_len, embedding_dim)            \n\n        if x_exercices_task is not None:\n            x_responses = x_responses + x_exercices_task # (BS, seq_len, embedding_dim)\n        \n\n        x_position_responses = self.pos_encoder2(x_responses) if self.pos_encoder2 is not None else x_responses # (BS, seq_len, embedding_dim)\n\n        # Transformer src: (S,N,E), tgt:(T,N,E), src_mask:(S,S), tgt_mask:(T,T)\n        # where S is the source sequence length, T is the target sequence length, N is the batch size, E is the feature number\n        # output: (T,N,E)\n        # src_key_padding_mask: (N,S), tgt_key_padding_mask: (N,T), memory_key_padding_mask: (N,S)    \n        x_position_exercices = x_position_exercices.transpose(1,0) # (seq_len, BS, embedding_dim)\n        x_position_responses = x_position_responses.transpose(1,0) # (seq_len, BS, embedding_dim)\n        \n        x_transformer = self.transformer(src=x_position_exercices, tgt=x_position_responses, src_mask=src_mask, tgt_mask=tgt_mask, memory_mask=mem_mask, \n                                         src_key_padding_mask=src_key_padding_mask, tgt_key_padding_mask=tgt_key_padding_mask, memory_key_padding_mask=memory_key_padding_mask) # (seq_len, BS, embedding_dim)\n        x_transformer = x_transformer.transpose(1,0) # (BS, seq_len, embedding_dim)\n     \n        output = self.fc(x_transformer)\n        output = output.squeeze(dim=2)\n\n        return output\n```",
          "votes": 4
        },
        {
          "id": 1144981,
          "postDate": "2021-01-08T19:45:16.803Z",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrehman245\" target=\"_blank\">@abdurrehman245</a>  You can find the code for our Tensorflow model <a href=\"https://github.com/rafiko1/Riiid-sharing\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "@abdurrehman245  You can find the code for our Tensorflow model [here](https://github.com/rafiko1/Riiid-sharing)",
          "votes": 2
        },
        {
          "id": 1145005,
          "postDate": "2021-01-08T19:57:33.353Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Thanks for sharing. </p>\n<p>Did your best single model in tensorflow have same architecture as of pytorch model above apart from the features which achieved <code>Private LB 0.815</code>?</p>\n<p>Also, what type of embedding <code>(continuous or categorical)</code>did you used for <code>lag_time</code> and <code>Prior question elapsed time</code> ?</p>\n<p>Did you guys used any other feature in your best single model (tensorflow model) other than which you mentioned above as ?</p>\n<pre><code>Content id\nLag time\nPrior question elapsed time\nPrevious responses\nNumber of attempts \n</code></pre>\n<p>It would be great if you can share some training details like lr_scheduler, epochs for model convergence and anything special if you guys used.</p>",
          "rawMarkdown": "@mpware Thanks for sharing. \n\nDid your best single model in tensorflow have same architecture as of pytorch model above apart from the features which achieved `Private LB 0.815`?\n\nAlso, what type of embedding `(continuous or categorical) `did you used for `lag_time` and `Prior question elapsed time` ?\n\nDid you guys used any other feature in your best single model (tensorflow model) other than which you mentioned above as ?\n\n```\nContent id\nLag time\nPrior question elapsed time\nPrevious responses\nNumber of attempts \n```\n\nIt would be great if you can share some training details like lr_scheduler, epochs for model convergence and anything special if you guys used.",
          "votes": 1
        },
        {
          "id": 1145009,
          "postDate": "2021-01-08T20:00:11.827Z",
          "content": "<p><a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> thanks for sharing the solution.</p>",
          "rawMarkdown": "@rafiko1 thanks for sharing the solution.",
          "votes": 2
        },
        {
          "id": 1145017,
          "postDate": "2021-01-08T20:11:19.707Z",
          "content": "<p>For the Tensorflow model that achieved the <code>private LB 0.815</code>:<br>\nWe tried other features as well, but those were by far the most important for our model and others didn't really improve our CV much.<br>\nEmbedding were all categorical, as continuous didn't work out well for us.<br>\nWe transformed <code>lag_time</code> to 173 categories, which gave a nice boost to the model.  We used the help of pandas <code>qcut</code>, like so:</p>\n<pre><code>N_ltg = 400 # number of groups 400 -&gt; 173\ntrain[\"ltg\"] = train.groupby(\"user_id\")[\"timestamp\"].shift()\n\n# Lag in seconds\ntrain[\"ltg\"] = ((train[\"timestamp\"] - train[\"ltg\"])/(1000.0))\ntrain[\"ltg\"] = round(train[\"ltg\"])\n\ntrain[\"ltg\"], bins = pd.qcut(train[\"ltg\"], N_ltg, duplicates=\"drop\", retbins=True) # duplicated -&gt; reduce about 1/2 of N_ltg\ntrain[\"ltg\"] = train[\"ltg\"].cat.codes # codes\n\n# Replace values\nN_ltg = train[\"ltg\"].nunique() # 173\ntrain[\"ltg\"] = train[\"ltg\"].replace(-1, N_ltg) # Replace -1 = NaN id \n</code></pre>\n<p>For the lr scheduler, we used Noam like in SAINT papers, and trained until convergence of the model, can be around 40-50 epochs of full training data per epoch.</p>",
          "rawMarkdown": "For the Tensorflow model that achieved the `private LB 0.815`:\nWe tried other features as well, but those were by far the most important for our model and others didn't really improve our CV much.\nEmbedding were all categorical, as continuous didn't work out well for us.\nWe transformed `lag_time` to 173 categories, which gave a nice boost to the model.  We used the help of pandas `qcut`, like so:\n```\nN_ltg = 400 # number of groups 400 -> 173\ntrain[\"ltg\"] = train.groupby(\"user_id\")[\"timestamp\"].shift()\n    \n# Lag in seconds\ntrain[\"ltg\"] = ((train[\"timestamp\"] - train[\"ltg\"])/(1000.0))\ntrain[\"ltg\"] = round(train[\"ltg\"])\n\ntrain[\"ltg\"], bins = pd.qcut(train[\"ltg\"], N_ltg, duplicates=\"drop\", retbins=True) # duplicated -> reduce about 1/2 of N_ltg\ntrain[\"ltg\"] = train[\"ltg\"].cat.codes # codes\n    \n# Replace values\nN_ltg = train[\"ltg\"].nunique() # 173\ntrain[\"ltg\"] = train[\"ltg\"].replace(-1, N_ltg) # Replace -1 = NaN id \n   \n```\n\nFor the lr scheduler, we used Noam like in SAINT papers, and trained until convergence of the model, can be around 40-50 epochs of full training data per epoch.",
          "votes": 3
        },
        {
          "id": 1145072,
          "postDate": "2021-01-08T21:22:49.230Z",
          "content": "<p><a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> thanks.</p>\n<p>The model in the train notebook have achieved AUC 0.8001 even after 68 epochs. What's the reason of that even the model is using all features ? </p>",
          "rawMarkdown": "@rafiko1 thanks.\n\nThe model in the train notebook have achieved AUC 0.8001 even after 68 epochs. What's the reason of that even the model is using all features ? ",
          "votes": 1
        },
        {
          "id": 1145099,
          "postDate": "2021-01-08T22:10:22.567Z",
          "content": "<p>Mainly because this is the \"demo\" version. You'll need to pass in the parameters of the transformers written in the description in order to get the higher score.</p>",
          "rawMarkdown": "Mainly because this is the \"demo\" version. You'll need to pass in the parameters of the transformers written in the description in order to get the higher score.",
          "votes": 2
        },
        {
          "id": 1145115,
          "postDate": "2021-01-08T22:49:33.730Z",
          "content": "<p>yeah you are right. I did not notice the hyperparams.</p>",
          "rawMarkdown": "yeah you are right. I did not notice the hyperparams.",
          "votes": 1
        },
        {
          "id": 1145264,
          "postDate": "2021-01-09T03:04:37.637Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> thanks for sharing, your team did a great job.</p>\n<p>Did you try to calculate lagged time by task_container_id? I did something like below to calculate lagged time. I wonder what is the good way for calculating lagged time? If you have tried it, could you tell me what is better and how much does it boost your model?</p>\n<pre><code>class UserLaggedTimeEncoder:\n    def __init__(self):\n        self.last_timestamp = defaultdict(float)\n        self.last_task_ts = defaultdict(float)\n        self.last_task_id = defaultdict(int)\n\n    def update(self, user_id, timestamp, task_id):\n        if task_id &gt; self.last_task_id[user_id]:\n            self.last_task_id[user_id] = task_id\n            self.last_task_ts[user_id] = self.last_timestamp[user_id]\n            self.last_timestamp[user_id] = timestamp\n\n    def encode(self, df):\n        lagged_time = np.zeros((df.shape[0],), dtype=np.float32)\n        for i, row in enumerate(tqdm(zip(df['user_id'].to_numpy(),\n                                         df['timestamp'].to_numpy(),\n                                         df['task_container_id'].to_numpy()),\n                                     total=df.shape[0])):\n            self.update(row[0], row[1], row[2])\n            lagged_time[i] = row[1] - self.last_task_ts[row[0]]\n        return lagged_time\n</code></pre>",
          "rawMarkdown": "@mpware @rafiko1 thanks for sharing, your team did a great job.\n\nDid you try to calculate lagged time by task_container_id? I did something like below to calculate lagged time. I wonder what is the good way for calculating lagged time? If you have tried it, could you tell me what is better and how much does it boost your model?\n\n```python\nclass UserLaggedTimeEncoder:\n    def __init__(self):\n        self.last_timestamp = defaultdict(float)\n        self.last_task_ts = defaultdict(float)\n        self.last_task_id = defaultdict(int)\n    \n    def update(self, user_id, timestamp, task_id):\n        if task_id > self.last_task_id[user_id]:\n            self.last_task_id[user_id] = task_id\n            self.last_task_ts[user_id] = self.last_timestamp[user_id]\n            self.last_timestamp[user_id] = timestamp\n    \n    def encode(self, df):\n        lagged_time = np.zeros((df.shape[0],), dtype=np.float32)\n        for i, row in enumerate(tqdm(zip(df['user_id'].to_numpy(),\n                                         df['timestamp'].to_numpy(),\n                                         df['task_container_id'].to_numpy()),\n                                     total=df.shape[0])):\n            self.update(row[0], row[1], row[2])\n            lagged_time[i] = row[1] - self.last_task_ts[row[0]]\n        return lagged_time\n```",
          "votes": 1
        },
        {
          "id": 1145646,
          "postDate": "2021-01-09T09:22:12.210Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/nvhbk16k53\" target=\"_blank\">@nvhbk16k53</a>. I noticed it's better to have a specific category for questions within a bundle. The lag time for questions within a bundle becomes 0 when you subtract.<br>\nYou can see how <code>lag time</code> was calculated in my code above.</p>",
          "rawMarkdown": "Thank you @nvhbk16k53. I noticed it's better to have a specific category for questions within a bundle. The lag time for questions within a bundle becomes 0 when you subtract.\nYou can see how `lag time` was calculated in my code above.",
          "votes": 2
        },
        {
          "id": 1145746,
          "postDate": "2021-01-09T10:11:56.537Z",
          "content": "<p>Thank you for your reply <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> . I saw the way you calculate <code>lag time</code>, it's be all zeros for all questions that have the same <code>task_container_id</code> except the first one. Mine is questions have the same <code>task_container_id</code> will have the same <code>lag time</code> value. What is the advantage of your compare to mine?</p>\n<p>I also experiment with LGBM that when calculate <code>lag time</code> like I did will give higher score.</p>",
          "rawMarkdown": "Thank you for your reply @rafiko1 . I saw the way you calculate `lag time`, it's be all zeros for all questions that have the same `task_container_id` except the first one. Mine is questions have the same `task_container_id` will have the same `lag time` value. What is the advantage of your compare to mine?\n\nI also experiment with LGBM that when calculate `lag time` like I did will give higher score.",
          "votes": 1
        },
        {
          "id": 1145767,
          "postDate": "2021-01-09T10:22:46.280Z",
          "content": "<p>I see. I haven't tried the same for <code>lag time</code> like you did. So I can't tell what has more advantage, your method might work better also for transformers.</p>",
          "rawMarkdown": "I see. I haven't tried the same for `lag time` like you did. So I can't tell what has more advantage, your method might work better also for transformers.",
          "votes": 2
        },
        {
          "id": 1145869,
          "postDate": "2021-01-09T11:46:36.867Z",
          "content": "<p><a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> thanks for sharing all the code and thoughts.</p>\n<p>I am asking a bit irrelevant question but it would be great if you can share your thoughts. Since you have done awesome with transformers so I just want to know that can we apply transformers on any time-series data or it depends on some characteristics of data to apply transformers on them.</p>\n<p>For example, if we want to apply the transformers to predict the fraud transaction of a user given its history of transactions(similarly we have history of user interactions in riid dataset), do you think transformers will be a good approach to apply on that dataset. I will appreciate your suggestions.</p>",
          "rawMarkdown": "@rafiko1 thanks for sharing all the code and thoughts.\n\nI am asking a bit irrelevant question but it would be great if you can share your thoughts. Since you have done awesome with transformers so I just want to know that can we apply transformers on any time-series data or it depends on some characteristics of data to apply transformers on them.\n\nFor example, if we want to apply the transformers to predict the fraud transaction of a user given its history of transactions(similarly we have history of user interactions in riid dataset), do you think transformers will be a good approach to apply on that dataset. I will appreciate your suggestions.",
          "votes": 2
        },
        {
          "id": 1146960,
          "postDate": "2021-01-10T08:00:26.043Z",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrehman245\" target=\"_blank\">@abdurrehman245</a> <br>\nI have a feeling transformers can be applied to more use-cases of time series, including fraud detection.<br>\nCan't tell with uncertainty until one experiments with it.</p>",
          "rawMarkdown": "@abdurrehman245 \nI have a feeling transformers can be applied to more use-cases of time series, including fraud detection.\nCan't tell with uncertainty until one experiments with it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1147263,
      "postDate": "2021-01-10T11:50:15.543Z",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>About</p>\n<pre><code>And @mpware has a brilliant way of storing each user's attempt history using only 3 bits per time step to save memory. (3 bits can represent 0 thru 7 in binary).\n</code></pre>\n<p>Would it be possible to share how you perform this? At the very end of this competition, I needed to deal with the time/memory issue once I introduced the performance features. And it took time, and I didn't have much time to fix the other inference bugs due to the new features.</p>\n<p>It would be great to learn from your great management of the memory, hope I could hear from you, at least a bit.</p>",
      "rawMarkdown": "@mpware @cdeotte \n\nAbout\n\n```\nAnd @mpware has a brilliant way of storing each user's attempt history using only 3 bits per time step to save memory. (3 bits can represent 0 thru 7 in binary).\n```\n\nWould it be possible to share how you perform this? At the very end of this competition, I needed to deal with the time/memory issue once I introduced the performance features. And it took time, and I didn't have much time to fix the other inference bugs due to the new features.\n\nIt would be great to learn from your great management of the memory, hope I could hear from you, at least a bit.",
      "votes": 2,
      "replies": [
        {
          "id": 1148710,
          "postDate": "2021-01-11T10:33:09.883Z",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> I'm maintaining the following stucture in memory. I've been inspired from different threads in forum and benchmarked multiple solutions before keeping this one. Per user and per content attempts dictionary costs too much memory if stored as integers (even <code>np.int8</code>) because it could be up to 393k x 13k.  To lower memory usage I'm using 3xbitarray (index is content_id) to address 2x2x2=8 attempts. As we're using attempts are categories we have 0-7 values and all attempts beyond 7 fall into last category which is indeed 7 or higher. </p>\n<p>We also need to store per user history for:</p>\n<ul>\n<li>Questions</li>\n<li>Answers</li>\n<li>Lags</li>\n<li>Elapsed time</li>\n<li>Attempts</li>\n<li>Part (optional)</li>\n<li>Had explanation (optional)</li>\n</ul>\n<p>Attempts history per user is computed on-fly based on the  <code>cseen</code> values containing only the total (that's the trick).</p>\n<p>Our different models (LGBM, Transfomer) can share this dictionary.</p>\n<p><code>per_user_dict = defaultdict(CustomDictNP)</code></p>\n<pre><code>def bit_array():\n    b = bitarray(13530, endian='little') # Higher than totals question if any new question in test set\n    b.setall(False)\n    return b\n\nclass CustomDictNP:\n    def __init__(self):\n        # Count/sum for further average\n        self.qc = 0 # answers count\n        self.qs = 0 # answers sum (correct)\n        self.s = 0 # session id\n        self.sc = 1 # session count\n        self.ss = 0 # total sessions\n        self.qes = 0.0 # total elapsed time\n\n        # History\n        self.ha = np.array([], dtype=np.bool) # answers\n        self.hp = np.array([], dtype=np.int8) # parts\n        self.hq = np.array([], dtype=np.int16) # questions\n        self.he = np.array([], dtype=np.float32) # elapsed time\n        self.hx = np.array([], dtype=np.bool) # prior question had explanation\n        self.hl = np.array([], dtype=np.float32) # lag\n\n        self.hat = np.array([], dtype=np.int8) # attempts\n\n        # Last timestamps\n        self.q = 0 # Questions\n        self.timestamp_u = []\n        self.timestamp_u_correct = []\n\n        # content_id seen        \n        self.cseen0 = bit_array()\n        self.cseen1 = bit_array()\n        self.cseen2 = bit_array()\n\n\ndef add_attempt(custom, cid):\n    new_attempt = get_attempts(custom, cid) + 1\n    if new_attempt &lt; 8:\n        if new_attempt == 1: \n            custom.cseen0[cid] = True\n            custom.cseen1[cid] = False\n            custom.cseen2[cid] = False\n        elif new_attempt == 2: \n            custom.cseen0[cid] = False\n            custom.cseen1[cid] = True\n            custom.cseen2[cid] = False\n        elif new_attempt == 3: \n            custom.cseen0[cid] = True\n            custom.cseen1[cid] = True\n            custom.cseen2[cid] = False\n        elif new_attempt == 4: \n            custom.cseen0[cid] = False\n            custom.cseen1[cid] = False\n            custom.cseen2[cid] = True\n        elif new_attempt == 5: \n            custom.cseen0[cid] = True\n            custom.cseen1[cid] = False\n            custom.cseen2[cid] = True\n        elif new_attempt == 6: \n            custom.cseen0[cid] = False\n            custom.cseen1[cid] = True\n            custom.cseen2[cid] = True\n        else: \n            custom.cseen0[cid] = True\n            custom.cseen1[cid] = True\n            custom.cseen2[cid] = True\n\ndef get_attempts(custom, cid):\n    return custom.cseen0[cid] + 2*custom.cseen1[cid] + 4*custom.cseen2[cid]\n</code></pre>",
          "rawMarkdown": "@yihdarshieh I'm maintaining the following stucture in memory. I've been inspired from different threads in forum and benchmarked multiple solutions before keeping this one. Per user and per content attempts dictionary costs too much memory if stored as integers (even `np.int8`) because it could be up to 393k x 13k.  To lower memory usage I'm using 3xbitarray (index is content_id) to address 2x2x2=8 attempts. As we're using attempts are categories we have 0-7 values and all attempts beyond 7 fall into last category which is indeed 7 or higher. \n\nWe also need to store per user history for:\n- Questions\n- Answers\n- Lags\n- Elapsed time\n- Attempts\n- Part (optional)\n- Had explanation (optional)\n\nAttempts history per user is computed on-fly based on the  `cseen` values containing only the total (that's the trick).\n\nOur different models (LGBM, Transfomer) can share this dictionary.\n\n`per_user_dict = defaultdict(CustomDictNP)`\n\n\n```\ndef bit_array():\n    b = bitarray(13530, endian='little') # Higher than totals question if any new question in test set\n    b.setall(False)\n    return b\n\t\nclass CustomDictNP:\n    def __init__(self):\n        # Count/sum for further average\n        self.qc = 0 # answers count\n        self.qs = 0 # answers sum (correct)\n        self.s = 0 # session id\n        self.sc = 1 # session count\n        self.ss = 0 # total sessions\n        self.qes = 0.0 # total elapsed time\n\n        # History\n        self.ha = np.array([], dtype=np.bool) # answers\n        self.hp = np.array([], dtype=np.int8) # parts\n        self.hq = np.array([], dtype=np.int16) # questions\n        self.he = np.array([], dtype=np.float32) # elapsed time\n        self.hx = np.array([], dtype=np.bool) # prior question had explanation\n        self.hl = np.array([], dtype=np.float32) # lag\n\t\t\n        self.hat = np.array([], dtype=np.int8) # attempts\n\n        # Last timestamps\n        self.q = 0 # Questions\n        self.timestamp_u = []\n        self.timestamp_u_correct = []\n        \n        # content_id seen        \n        self.cseen0 = bit_array()\n        self.cseen1 = bit_array()\n        self.cseen2 = bit_array()\n\n\ndef add_attempt(custom, cid):\n    new_attempt = get_attempts(custom, cid) + 1\n    if new_attempt < 8:\n        if new_attempt == 1: \n            custom.cseen0[cid] = True\n            custom.cseen1[cid] = False\n            custom.cseen2[cid] = False\n        elif new_attempt == 2: \n            custom.cseen0[cid] = False\n            custom.cseen1[cid] = True\n            custom.cseen2[cid] = False\n        elif new_attempt == 3: \n            custom.cseen0[cid] = True\n            custom.cseen1[cid] = True\n            custom.cseen2[cid] = False\n        elif new_attempt == 4: \n            custom.cseen0[cid] = False\n            custom.cseen1[cid] = False\n            custom.cseen2[cid] = True\n        elif new_attempt == 5: \n            custom.cseen0[cid] = True\n            custom.cseen1[cid] = False\n            custom.cseen2[cid] = True\n        elif new_attempt == 6: \n            custom.cseen0[cid] = False\n            custom.cseen1[cid] = True\n            custom.cseen2[cid] = True\n        else: \n            custom.cseen0[cid] = True\n            custom.cseen1[cid] = True\n            custom.cseen2[cid] = True\n\ndef get_attempts(custom, cid):\n    return custom.cseen0[cid] + 2*custom.cseen1[cid] + 4*custom.cseen2[cid]\n```\n",
          "votes": 4
        },
        {
          "id": 1148739,
          "postDate": "2021-01-11T11:04:35.493Z",
          "content": "<p><a href=\"https://www.kaggle.com/MPWARE\" target=\"_blank\">@MPWARE</a>, thank you very much for being kindness to share, really appreciated. I haven't read it in detail to fully understand, but seems a very elegant yet efficient approach to the problem.</p>\n<p>Personally, I was able to fix this memory issue ( Per user and per content performance history ) in a short time (yet not a very clean way) by the following idea:</p>\n<p>Quite a lot of (user, question) pairs are in fact with (n_attempt, n_correctness) being (1, 1) and (1, 0). The pairs with <code>attempt &gt;= 2</code> is quite rare. Therefore, I have 2 nested dictionaries:</p>\n<pre><code>1. The 1st dict is:  user_id -&gt; question_id -&gt; the performances for those n_attempt &gt;= 2 (quite small dict)\n\n2. The 2nd dict: (for those pairs with `n_attemp = 1`)\n\n    {\n        user_id: \n           { \n               'correct': [q_id_x, q_id_y, ...]\n               'incorrect': [q_id_m, q_in_, ...]\n           }\n    }\n</code></pre>\n<p>If  (user, question) can't be found in the above 2 dictionaries, it means that question is not seen by the user.</p>\n<p>The idea is kind similar to sparse matrix, where we only store the content with values.</p>\n<p>Despite this effort, it still take 4G or 5G memory when loading into memory. I will try to compare your approach later.</p>\n<p>Comments to my approaches  are appreciated (if any). Thanks again for sharing.</p>",
          "rawMarkdown": "@MPWARE, thank you very much for being kindness to share, really appreciated. I haven't read it in detail to fully understand, but seems a very elegant yet efficient approach to the problem.\n\nPersonally, I was able to fix this memory issue ( Per user and per content performance history ) in a short time (yet not a very clean way) by the following idea:\n\nQuite a lot of (user, question) pairs are in fact with (n_attempt, n_correctness) being (1, 1) and (1, 0). The pairs with `attempt >= 2` is quite rare. Therefore, I have 2 nested dictionaries:\n\n    1. The 1st dict is:  user_id -> question_id -> the performances for those n_attempt >= 2 (quite small dict)\n    \n    2. The 2nd dict: (for those pairs with `n_attemp = 1`)\n\n        {\n            user_id: \n               { \n                   'correct': [q_id_x, q_id_y, ...]\n                   'incorrect': [q_id_m, q_in_, ...]\n               }\n        }\n\nIf  (user, question) can't be found in the above 2 dictionaries, it means that question is not seen by the user.\n\nThe idea is kind similar to sparse matrix, where we only store the content with values.\n\nDespite this effort, it still take 4G or 5G memory when loading into memory. I will try to compare your approach later.\n\nComments to my approaches  are appreciated (if any). Thanks again for sharing.",
          "votes": 2
        },
        {
          "id": 1149507,
          "postDate": "2021-01-11T22:23:04.897Z",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> I understand your way of memory management now. Great.</p>\n<p>I am wondering why your team don't make a feature of the nb of correctness per user per content, or better, the ratio, if you already have such good way to manage memory. It think you might even get better by doing so, although in my case, I can't see the positive effect on LB due to the input inconsistency bug in my pipeline. I am just wondering the reason your team don't use it</p>",
          "rawMarkdown": "@mpware I understand your way of memory management now. Great.\n\nI am wondering why your team don't make a feature of the nb of correctness per user per content, or better, the ratio, if you already have such good way to manage memory. It think you might even get better by doing so, although in my case, I can't see the positive effect on LB due to the input inconsistency bug in my pipeline. I am just wondering the reason your team don't use it",
          "votes": 1
        },
        {
          "id": 1149648,
          "postDate": "2021-01-12T03:27:22.690Z",
          "content": "<p>Really cool use of bit arrays! I didn't extend it to do the counting part as well! Sweet!</p>\n<pre><code>def get_attempts(custom, cid):\n    return custom.cseen0[cid] + 2*custom.cseen1[cid] + 4*custom.cseen2[cid]\n</code></pre>\n<p>This is basically bit shifting, correct? (Converting the bits back to decimal radix)</p>\n<pre><code>        # content_id seen        \n        self.cseen0 = bit_array()\n        self.cseen1 = bit_array()\n        self.cseen2 = bit_array()\n</code></pre>\n<p>This basically represents the 3 bits needed to count till 7 i believe.[ with <code>cseen2</code> as the MSB]</p>\n<p>Another use of the same bitarray's is <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266#1069347\" target=\"_blank\">here</a> <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>! Basically a bool flag to detect whether use has seen that content in past or not.</p>",
          "rawMarkdown": "Really cool use of bit arrays! I didn't extend it to do the counting part as well! Sweet!\n\n```\ndef get_attempts(custom, cid):\n    return custom.cseen0[cid] + 2*custom.cseen1[cid] + 4*custom.cseen2[cid]\n```\n\nThis is basically bit shifting, correct? (Converting the bits back to decimal radix)\n\n\n```\n        # content_id seen        \n        self.cseen0 = bit_array()\n        self.cseen1 = bit_array()\n        self.cseen2 = bit_array()\n```\nThis basically represents the 3 bits needed to count till 7 i believe.[ with `cseen2` as the MSB]\n\nAnother use of the same bitarray's is [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266#1069347) @yihdarshieh! Basically a bool flag to detect whether use has seen that content in past or not.",
          "votes": 2
        },
        {
          "id": 1149990,
          "postDate": "2021-01-12T09:31:09.773Z",
          "content": "<p>We've tried it but results were close (little bit lower) than without.</p>",
          "rawMarkdown": "We've tried it but results were close (little bit lower) than without.",
          "votes": 1
        },
        {
          "id": 1150173,
          "postDate": "2021-01-12T12:30:57.373Z",
          "content": "<p>Alright, that is quite strange. Do anyone in your team have some insight about why this is not working?</p>",
          "rawMarkdown": "Alright, that is quite strange. Do anyone in your team have some insight about why this is not working?"
        },
        {
          "id": 1150174,
          "postDate": "2021-01-12T12:32:56.903Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> I saw that thread, but didn't take the time to read it carefully and find the shown code, I was so depressed about my inference bug at that time. Anyway, now I know there is bitarray. </p>",
          "rawMarkdown": "@adityaecdrid I saw that thread, but didn't take the time to read it carefully and find the shown code, I was so depressed about my inference bug at that time. Anyway, now I know there is bitarray. "
        }
      ]
    },
    {
      "id": 1145911,
      "postDate": "2021-01-09T12:23:13.433Z",
      "content": "<p>Congratz to you all !</p>",
      "rawMarkdown": "Congratz to you all !",
      "votes": 2
    },
    {
      "id": 1145432,
      "postDate": "2021-01-09T06:24:04.433Z",
      "content": "<p>thanks for sharing! I am a beginer to transformer model. I think I need to learn the model before seeing your share.</p>",
      "rawMarkdown": "thanks for sharing! I am a beginer to transformer model. I think I need to learn the model before seeing your share.",
      "votes": 2
    },
    {
      "id": 1145294,
      "postDate": "2021-01-09T03:52:07.940Z",
      "content": "<p>Thanks for your sharing and congratulations to your team.</p>\n<p>Would you like to explain how to add the \"Number of attempts\" feature into encoder?<br>\nIs it convert to embedding?<br>\nThank you!!</p>",
      "rawMarkdown": "Thanks for your sharing and congratulations to your team.\n\nWould you like to explain how to add the \"Number of attempts\" feature into encoder?\nIs it convert to embedding?\nThank you!!",
      "votes": 2,
      "replies": [
        {
          "id": 1145326,
          "postDate": "2021-01-09T04:25:36.937Z",
          "content": "<p>Yes as an embedding and add it to other encoder input embeddings. </p>\n<p>The encoder input begins as <code>x = tf.keras.layers.Embedding(number of unique content id, 288)(CONTENT ID SEQUENCE)</code>. Then you add positional encoding, <code>x += positional_encoding(window size, 288)</code>. And last you add \"Number of attempts\" encoding. <code>x += tf.keras.layers.Embedding(max number of attempts + 1, 288)(NUMBER OF ATTEMPTS SEQUENCE)</code>.</p>\n<p>The maximum number of attempts is clipped with <code>NUMBER OF ATTEMPTS SEQUENCE = np.clip(NUMBER OF ATTEMPTS SEQUENCE, 0, 7)</code>.</p>\n<p>And <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> has a brilliant way of storing each user's attempt history using only 3 bits per time step to save memory. (3 bits can represent 0 thru 7 in binary).</p>",
          "rawMarkdown": "Yes as an embedding and add it to other encoder input embeddings. \n\nThe encoder input begins as `x = tf.keras.layers.Embedding(number of unique content id, 288)(CONTENT ID SEQUENCE)`. Then you add positional encoding, `x += positional_encoding(window size, 288)`. And last you add \"Number of attempts\" encoding. `x += tf.keras.layers.Embedding(max number of attempts + 1, 288)(NUMBER OF ATTEMPTS SEQUENCE)`.\n\nThe maximum number of attempts is clipped with `NUMBER OF ATTEMPTS SEQUENCE = np.clip(NUMBER OF ATTEMPTS SEQUENCE, 0, 7)`.\n\nAnd @mpware has a brilliant way of storing each user's attempt history using only 3 bits per time step to save memory. (3 bits can represent 0 thru 7 in binary).",
          "votes": 3
        },
        {
          "id": 1145376,
          "postDate": "2021-01-09T05:17:44.343Z",
          "content": "<p>Very impressive. Thanks for your reply.</p>",
          "rawMarkdown": "Very impressive. Thanks for your reply.",
          "votes": 2
        },
        {
          "id": 1145897,
          "postDate": "2021-01-09T12:17:48.453Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I can see that <code>99%</code> of of attempts on content_id is &lt;=7. Is that the reason of clipping the max value to 7?  </p>\n<p><code>x += tf.keras.layers.Embedding(max number of attempts + 1, 288)(NUMBER OF ATTEMPTS SEQUENCE)</code></p>\n<p>so <code>max number of attempts</code> will be 7 then and why are you adding 1 more to number of attempts ?</p>",
          "rawMarkdown": "@cdeotte I can see that `99%` of of attempts on content_id is <=7. Is that the reason of clipping the max value to 7?  \n\n`x += tf.keras.layers.Embedding(max number of attempts + 1, 288)(NUMBER OF ATTEMPTS SEQUENCE)`\n\nso `max number of attempts` will be 7 then and why are you adding 1 more to number of attempts ?",
          "votes": 1
        },
        {
          "id": 1146604,
          "postDate": "2021-01-09T21:55:56.690Z",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrehman245\" target=\"_blank\">@abdurrehman245</a> The real reason was to limit memory usage for inference. We limit attempts from 0 to 7, everything above 7 falls in the last category which is indeed 7 and beyond.</p>",
          "rawMarkdown": "@abdurrehman245 The real reason was to limit memory usage for inference. We limit attempts from 0 to 7, everything above 7 falls in the last category which is indeed 7 and beyond.",
          "votes": 1
        },
        {
          "id": 1148196,
          "postDate": "2021-01-11T01:43:04.583Z",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrehman245\" target=\"_blank\">@abdurrehman245</a></p>\n<blockquote>\n  <p>why are you adding 1 more to number of attempts?</p>\n</blockquote>\n<p>Since the max is 7, there are 8 unique values, i.e. 0,1,2,3,4,5,6,7. So the embedding needs to handle a categorical variable of cardinality 8.</p>",
          "rawMarkdown": "@abdurrehman245\n>why are you adding 1 more to number of attempts?\n\nSince the max is 7, there are 8 unique values, i.e. 0,1,2,3,4,5,6,7. So the embedding needs to handle a categorical variable of cardinality 8.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1145155,
      "postDate": "2021-01-09T00:10:55.967Z",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Congrats on the gold medal and nice single model solution !</p>",
      "rawMarkdown": "@mpware Congrats on the gold medal and nice single model solution !",
      "votes": 2,
      "replies": [
        {
          "id": 1145680,
          "postDate": "2021-01-09T09:39:16.443Z",
          "content": "<p><a href=\"https://www.kaggle.com/tikutiku\" target=\"_blank\">@tikutiku</a> Thank you! </p>",
          "rawMarkdown": " @tikutiku Thank you! "
        }
      ]
    },
    {
      "id": 1145144,
      "postDate": "2021-01-08T23:44:57.310Z",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Congratulations to your team, very impressed that single RAINT+ with these 5 features can reach such a high score. My <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209596\" target=\"_blank\">best RAINT+ model</a> is only 0.802 on private</p>\n<p>Would like to know more details about which of these features are on the encoder side and which are on the decoder side?</p>\n<pre><code>Content id\nLag time\nPrior question elapsed time\nPrevious responses\nNumber of attempts\n</code></pre>\n<p><br>\nThere's no doubt on the content id and response. For other features, could you please explain a little bit about which parts did you put the feature in?</p>",
      "rawMarkdown": "@mpware Congratulations to your team, very impressed that single RAINT+ with these 5 features can reach such a high score. My [best RAINT+ model](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209596) is only 0.802 on private\n\nWould like to know more details about which of these features are on the encoder side and which are on the decoder side?\n```\nContent id\nLag time\nPrior question elapsed time\nPrevious responses\nNumber of attempts\n``` \nThere's no doubt on the content id and response. For other features, could you please explain a little bit about which parts did you put the feature in?",
      "votes": 2,
      "replies": [
        {
          "id": 1145160,
          "postDate": "2021-01-09T00:14:19.633Z",
          "content": "<p>Thanks. Encoder had Content id and Number of attempts. Decoder had Lag time, Prior question elapsed time, and Previous responses.</p>",
          "rawMarkdown": "Thanks. Encoder had Content id and Number of attempts. Decoder had Lag time, Prior question elapsed time, and Previous responses.",
          "votes": 2
        },
        {
          "id": 1145175,
          "postDate": "2021-01-09T00:31:15.080Z",
          "content": "<p>Got it, thanks! </p>",
          "rawMarkdown": "Got it, thanks! ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1144738,
      "postDate": "2021-01-08T16:32:19.443Z",
      "content": "<p>Congratulations on 16th place and gold medal <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and team. Good work and strongly transformer model</p>",
      "rawMarkdown": "Congratulations on 16th place and gold medal @mpware @cdeotte and team. Good work and strongly transformer model",
      "votes": 2
    },
    {
      "id": 1147557,
      "postDate": "2021-01-10T15:21:41.433Z",
      "content": "<p>How about the question's tags, it's a VarLen feature. Do you have any idea of adding VarLen feature in the RAINT+? What I did is sorting the tags in each question then combine them and map to an integer. But it's not an optimal way, since the information of each individual tag is lost.</p>",
      "rawMarkdown": "How about the question's tags, it's a VarLen feature. Do you have any idea of adding VarLen feature in the RAINT+? What I did is sorting the tags in each question then combine them and map to an integer. But it's not an optimal way, since the information of each individual tag is lost.",
      "replies": [
        {
          "id": 1148508,
          "postDate": "2021-01-11T07:28:17.583Z",
          "content": "<p>I use a tag embedding, i.e. map each integer in [-1, 188) to vectors of a fixed dim (in my case 256). Then these embedding are averaged (the embedding of -1 is for padding for  the tags, as you said, it is varlen, and it is ignored during the averaging). I don't know how the team of <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> did it though.</p>",
          "rawMarkdown": "I use a tag embedding, i.e. map each integer in [-1, 188) to vectors of a fixed dim (in my case 256). Then these embedding are averaged (the embedding of -1 is for padding for  the tags, as you said, it is varlen, and it is ignored during the averaging). I don't know how the team of @mpware did it though."
        }
      ]
    },
    {
      "id": 1145433,
      "postDate": "2021-01-09T06:24:04.433Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1146920,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2021-01-10T07:05:16.133000",
      "content": "<p><a href=\"https://www.kaggle.com/MPWare\" target=\"_blank\">@MPWare</a> Thanks for sharing the \"Simple\" Pytorch code, but why does it score so much less than Tensorflow?</p>\n<ol>\n<li>The Pytorch model uses 2 layers instead of Tensorflow uses 4 layers</li>\n<li>The Pytorch model uses 100 sequence length instead of Tensorflow uses 512 sequence length</li>\n<li>The Pytorch model uses 256 embedding size instead of Tensorflow uses 288</li>\n<li>The Pytorch model uses dropout = 0.1 but Tensorflow uses dropout = 0.2</li>\n<li>The Pytorch model only uses exercise_id  in encoder, but Tensorflow uses and response, but Tensorflow uses exercise_id, Lag time, prior question elapsed time, previous responses, number of attempts. Both models just use response_correct embedding in the decoder</li>\n</ol>\n<p>Are there any other differences? Would the Pytorch model be as good if you changed it to use all these features? Ty, I will really cherish that \"simple\" pytorch demo notebook you provided.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1147039,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-10T08:50:14.530000",
          "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> As you've noticed, the main difference is that I'm not using exactly the same features. I'm not using <code>attempts</code> at all. The configuration to get LB=0.795 with Pytorch model is below, it requires <code>seq_len=256</code>. We've noticed that <code>seq_len</code> is a quite important to get a better model. The training procedure is also important, I'm not using the smart window described by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. I've spent some time to fix the <code>NaN</code> loss issue with Pytorch <code>MultiHeadAttention</code> + padding masks. Root cause was right padding instead of left padding. I've stopped to try improving it once we got better results with <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> TensorFlow model which uses <code>seq_len=512</code>. Other difference is <code>sigmoid</code> vs <code>softmax</code> but I don't think it's really important.</p>\n<pre><code>class raw_conf:\n\n    pad_mode = \"token\"\n    pad_right = False\n    flatten = True\n    sampler = None # \"prob\" # Option to give long user's sequence higher probability\n\n    seq_len = 256 # 100\n    embedding_dim = 256\n    exercices_id_size = 13523\n    exercices_part_size = 7\n    response_size = 2 \n    elapsed_time_cat = True\n    elapsed_time_size = 73 # Categories after binning\n    lag_time_cat = True\n    lag_time_size = 366 # Categories after binning\n    explanation_size = 2\n\n    # Model\n    nhead = 8 \n    num_encoder_layers = 4\n    num_decoder_layers = 4\n    dim_feedforward = 2048\n    dropout = 0.1\n    activation = None\n    num_classes = 1\n    loss = MaskedBCEWithLogitsLoss(num_classes)\n    post_activation = \"sigmoid\"\n\n    optimizer = \"Noam\" # \"Adam\"\n    scheduler = \"Cosine\" if optimizer == \"Adam\" else None\n    lr = 0.0001\n    min_lr = 0.00005\n    beta1 = 0.9\n\n    BATCH_SIZE = 128\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1147157,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-10T10:08:37.647000",
          "content": "<p>We're going to upload the full pytorch model on GitHub with the weights.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1147512,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2021-01-10T14:59:40.747000",
          "content": "<p>Thank you. Since I know PT better than TF, I want to use your PT Transformer model for many future competitions since to me it is the simplest since it invokes nn.Transformer.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1149848,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-12T07:50:08.993000",
          "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> We've uploaded Pytorch full code here:<br>\n<a href=\"https://github.com/rafiko1/Riiid-sharing/tree/main/pytorch\" target=\"_blank\">https://github.com/rafiko1/Riiid-sharing/tree/main/pytorch</a><br>\nIt's MODEL-PT-v14</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1149949,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2021-01-12T09:07:52.317000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> thanks. I will go through it and will try to reproduce your solution.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1149996,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-12T09:33:39.397000",
          "content": "<p>The best model is with TensorFlow (<code>seq_len=512</code>), Public LB=0.812, Pytorch (<code>seq_len=256</code>) one is only Public LB=0.795.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1150265,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2021-01-12T13:37:10.160000",
          "content": "<p>Thanks for the complete src!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1146671,
      "author_name": "u++",
      "author_url": "",
      "post_date": "2021-01-09T23:46:57.540000",
      "content": "<p>Congrats &amp; thank you for sharing your approach and codes.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1146361,
      "author_name": "Darren Lahr",
      "author_url": "",
      "post_date": "2021-01-09T18:15:33.017000",
      "content": "<p>Congratulations and thanks for sharing your solution, I and many others will be able to learn a lot from it. I also used Tensorflow and implemented Saint+ however my code to sample sequences/users was not robust so I intend to see how the model would have performed using the sampling code you have shared.</p>\n<p>I have read through the code in demo_riiid_train.ipynb and think I understand how you have set up training and validation sets but was hoping you could confirm my understanding?</p>\n<ol>\n<li>You allocate distinct users to both the training and validation set</li>\n<li>You then assign a probability to each user based on sequence length. For say user 115 in the training set their probability is their sequence length divided by the sum of all sequence lengths in training set.</li>\n<li>You then generate a one off validation set where you randomly sample users with replacement and take <strong>random training crops</strong> for each row. Based on the 'select_window_size' function</li>\n<li>For the training dataset you do the same but repeat step 3 each epoch</li>\n</ol>\n<p>You sample using N_SELECT_PER_EPOCH = 100000, was this just a hyperparameter you tuned?</p>\n<p>Thank you </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1146978,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2021-01-10T08:15:12.450000",
          "content": "<ol>\n<li>Correct. Distinct users gives a simple and reliable validation.</li>\n<li>Correct. It will be used inside <code>select_window_size</code> as the equivalent of <code>WeightedRandomSampler</code> in Pytorch</li>\n<li>Correct. This is important for training. For validation it's the convenient choice, but not the best choice. Better is to take all sequences for validation instead.</li>\n<li>Correct.</li>\n</ol>\n<p><code>N_SELECT_PER_EPOCH</code> is defining how many samples to pass within each epoch. It's not really a hyperparameter to tune. Just how long you'd like each epoch to be by specifying number of samples.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1147277,
          "author_name": "Darren Lahr",
          "author_url": "",
          "post_date": "2021-01-10T11:58:37.433000",
          "content": "<p>Thank you, will do a late submission and see how my score would have changed.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1146392,
      "author_name": "mamas",
      "author_url": "",
      "post_date": "2021-01-09T18:35:27.793000",
      "content": "<p>Congrats, single 0.812/0.815 with SAINT is really great!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1146980,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2021-01-10T08:17:43.870000",
          "content": "<p>Congrats on your amazing 2nd place <a href=\"https://www.kaggle.com/mamas\" target=\"_blank\">@mamas</a>, well-deserved!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1144929,
      "author_name": "Abdur Rehman",
      "author_url": "",
      "post_date": "2021-01-08T18:49:27.457000",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> congrats on gold medal and thanks for sharing your approach. Would you like to share your code ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1144950,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-08T19:12:24.127000",
          "content": "<p>Thanks! Our best single model + features is with TensorFlow. <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> if you get a chance to share it …</p>\n<p>In the other <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632\" target=\"_blank\">thread</a>, it was a Pytorch model (with simple features) that only reached public LB 0.795. Here it is if it can help. </p>\n<pre><code># Model\nclass RIIIDModel(nn.Module):\n    def __init__(self, cfg, verbose=False):\n        super().__init__()\n        self.response_size = cfg.response_size\n        self.lag_time_size = cfg.lag_time_size\n        self.elapsed_time_size = cfg.elapsed_time_size\n        self.explanation_size = cfg.explanation_size\n        self.attempt_size = cfg.attempt_size\n        self.seq_len = cfg.seq_len\n        self.embedding_dim = cfg.embedding_dim\n        self.elapsed_time_cat = cfg.elapsed_time_cat\n        self.lag_time_cat = cfg.lag_time_cat\n        self.num_classes = cfg.num_classes\n        self.verbose = verbose\n        self.pad_mode = cfg.pad_mode\n\n        self.pos_encoder1 = None\n        self.pos_encoder2 = None\n\n        additional_token_dim = 1 if self.pad_mode == \"token\" else 0\n\n        # Exercices embeddings\n        self.exercices_id_embedding = nn.Embedding(cfg.exercices_id_size + additional_token_dim, self.embedding_dim)\n        self.exercices_part_embedding = nn.Embedding(cfg.exercices_part_size + additional_token_dim, self.embedding_dim) if cfg.exercices_part_size is not None else None\n\n        # Response embeddings\n        self.response_embedding = nn.Embedding(cfg.response_size + 1 + additional_token_dim, self.embedding_dim) # +1 to include start token\n\n        if self.elapsed_time_cat is True:\n            self.elapsed_time_embedding = nn.Embedding(cfg.elapsed_time_size + 1 + additional_token_dim, self.embedding_dim) if cfg.elapsed_time_size is not None else None # +1 to include start token\n        else:\n            self.elapsed_time_embedding = nn.Linear(1, self.embedding_dim, bias=False) if cfg.elapsed_time_size is not None else None # Continuous embedding\n\n        if self.lag_time_cat is True:\n            self.lag_time_embedding = nn.Embedding(cfg.lag_time_size + 1 + additional_token_dim, self.embedding_dim) if cfg.lag_time_size is not None else None # +1 to include start token\n        else:\n            self.lag_time_embedding = nn.Linear(1, self.embedding_dim, bias=False) if cfg.lag_time_size is not None else None # Continuous embedding\n\n        self.explanation_embedding = nn.Embedding(cfg.explanation_size + 1 + additional_token_dim, self.embedding_dim) if cfg.explanation_size is not None else None # +1 to include start token\n\n        input_features_dim = self.embedding_dim\n\n        # Position encoder (relative or absolute position of the tokens in the sequence)\n        if cfg.position_encoding_enabled is True:\n            self.pos_encoder1 = PositionalEncoding(input_features_dim, cfg.dropout)\n            self.pos_encoder2 = self.pos_encoder1\n\n        # Transformer with default encoder/decoder        \n        self.transformer = nn.Transformer(d_model=input_features_dim, \n                                          nhead=cfg.nhead, \n                                          num_encoder_layers=cfg.num_encoder_layers,\n                                          num_decoder_layers=cfg.num_decoder_layers, \n                                          dim_feedforward=cfg.dim_feedforward, \n                                          dropout=cfg.dropout, \n                                          activation='relu', \n                                          custom_encoder = None, \n                                          custom_decoder = None)\n\n        # Decoder\n        self.fc = nn.Linear(input_features_dim, self.num_classes)\n\n\n    # If a BoolTensor is provided, the positions with the value of True will be ignored while the position with the value of False will be unchanged.\n    # tensor([[False,  True,  True,  True],\n    #         [False, False,  True,  True],\n    #         [False, False, False,  True],\n    #         [False, False, False, False]])    \n    def generate_mask(self, size, diagonal=1):        \n        return torch.triu(torch.ones(size, size)==1, diagonal=diagonal)\n\n    def forward(self, data, src_mask=None, tgt_mask=None, mem_mask=None, src_key_padding_mask=None, tgt_key_padding_mask=None, memory_key_padding_mask=None):\n\n        # Each input is (BS, seq_len)\n        # Content\n        data_content_id = data[CONTENT_ID].long()\n        # Answers\n        data_response = data[TARGET].long()        \n\n        # Optional features\n        data_part = data[PART].long() if self.exercices_part_embedding is not None else None\n        if self.elapsed_time_cat is True:\n            data_elapsed_time = data[PRIOR_QUESTION_ELAPSED_TIME].long() if self.elapsed_time_embedding is not None else None\n        else:\n            data_elapsed_time = data[PRIOR_QUESTION_ELAPSED_TIME].float().unsqueeze(2) if self.elapsed_time_embedding is not None else None\n        if self.lag_time_cat is True:\n            data_lag_time = data[LAG].long() if self.lag_time_embedding is not None else None\n        else:\n            data_lag_time = data[LAG].float().unsqueeze(2) if self.lag_time_embedding is not None else None\n\n        data_explanation = data[PRIOR_QUESTION_HAD_EXPLANATION].long() if self.explanation_embedding is not None else None\n\n        # Start token(s)\n\n        # Add start token to correctness\n        data_response = torch.roll(data_response, shifts=(0, 1), dims=(0, 1)) # Shift right the sequence\n        data_response[:,0] = self.response_size # Start token (2)\n\n        # Add start token to lag time\n        if data_lag_time is not None:\n            data_lag_time = torch.roll(data_lag_time, shifts=(0, 1), dims=(0, 1)) # Shift right the sequence\n            data_lag_time[:,0] = self.lag_time_size # Start token\n\n        # Add start token to elapsed time\n        if data_elapsed_time is not None:\n            data_elapsed_time = torch.roll(data_elapsed_time, shifts=(0, 1), dims=(0, 1)) # Shift right the sequence\n            if self.elapsed_time_cat is True:\n                data_elapsed_time[:,0] = self.elapsed_time_size # Start token\n            else:\n                data_elapsed_time[:,0] = 0.0\n\n        # Add start token to explanation\n        if data_explanation is not None:\n            data_explanation = torch.roll(data_explanation, shifts=(0, 1), dims=(0, 1)) # Shift right the sequence\n            data_explanation[:,0] = self.explanation_size # Start token\n\n        # Questions, Part, Elapsed time, Lag embeddings\n        x_content_id = self.exercices_id_embedding(data_content_id) # (BS, seq_len, embedding_dim)\n\n        x_exercices_part = self.exercices_part_embedding(data_part) if self.exercices_part_embedding is not None else None # (BS, seq_len, embedding_dim)\n        x_elapsed_time = self.elapsed_time_embedding(data_elapsed_time) if self.elapsed_time_embedding is not None else None # (BS, seq_len, embedding_dim)\n        x_lag_time = self.lag_time_embedding(data_lag_time) if self.lag_time_embedding is not None else None # (BS, seq_len, embedding_dim)\n        x_explanation = self.explanation_embedding(data_explanation) if self.explanation_embedding is not None else None # (BS, seq_len, embedding_dim)\n\n        # Response embeddings\n        x_correctness = self.response_embedding(data_response) # (BS, seq_len, embedding_dim)\n\n        x_position = None\n\n        # Ei (sum of embeddings)\n        x_exercices = x_content_id\n\n        if x_exercices_part is not None:\n            x_exercices = x_exercices + x_exercices_part  # (BS, seq_len, embedding_dim)\n\n        x_position_exercices = self.pos_encoder1(x_exercices) if self.pos_encoder1 is not None else x_exercices # (BS, seq_len, embedding_dim)\n\n        # Ri (sum of embeddings) [S, R1, Rk-1], S is start token\n\n        x_responses = x_correctness \n\n        if x_lag_time is not None:\n            x_responses = x_responses + x_lag_time\n\n        if x_elapsed_time is not None:\n            x_responses = x_responses + x_elapsed_time # (BS, seq_len, embedding_dim)\n\n        if x_explanation is not None:\n            x_responses = x_responses + x_explanation # (BS, seq_len, embedding_dim)\n\n        if x_attempt is not None:\n            x_responses = x_responses + x_attempt # (BS, seq_len, embedding_dim)            \n\n        if x_exercices_task is not None:\n            x_responses = x_responses + x_exercices_task # (BS, seq_len, embedding_dim)\n\n\n        x_position_responses = self.pos_encoder2(x_responses) if self.pos_encoder2 is not None else x_responses # (BS, seq_len, embedding_dim)\n\n        # Transformer src: (S,N,E), tgt:(T,N,E), src_mask:(S,S), tgt_mask:(T,T)\n        # where S is the source sequence length, T is the target sequence length, N is the batch size, E is the feature number\n        # output: (T,N,E)\n        # src_key_padding_mask: (N,S), tgt_key_padding_mask: (N,T), memory_key_padding_mask: (N,S)    \n        x_position_exercices = x_position_exercices.transpose(1,0) # (seq_len, BS, embedding_dim)\n        x_position_responses = x_position_responses.transpose(1,0) # (seq_len, BS, embedding_dim)\n\n        x_transformer = self.transformer(src=x_position_exercices, tgt=x_position_responses, src_mask=src_mask, tgt_mask=tgt_mask, memory_mask=mem_mask, \n                                         src_key_padding_mask=src_key_padding_mask, tgt_key_padding_mask=tgt_key_padding_mask, memory_key_padding_mask=memory_key_padding_mask) # (seq_len, BS, embedding_dim)\n        x_transformer = x_transformer.transpose(1,0) # (BS, seq_len, embedding_dim)\n\n        output = self.fc(x_transformer)\n        output = output.squeeze(dim=2)\n\n        return output\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1144981,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2021-01-08T19:45:16.803000",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrehman245\" target=\"_blank\">@abdurrehman245</a>  You can find the code for our Tensorflow model <a href=\"https://github.com/rafiko1/Riiid-sharing\" target=\"_blank\">here</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1145005,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2021-01-08T19:57:33.353000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Thanks for sharing. </p>\n<p>Did your best single model in tensorflow have same architecture as of pytorch model above apart from the features which achieved <code>Private LB 0.815</code>?</p>\n<p>Also, what type of embedding <code>(continuous or categorical)</code>did you used for <code>lag_time</code> and <code>Prior question elapsed time</code> ?</p>\n<p>Did you guys used any other feature in your best single model (tensorflow model) other than which you mentioned above as ?</p>\n<pre><code>Content id\nLag time\nPrior question elapsed time\nPrevious responses\nNumber of attempts \n</code></pre>\n<p>It would be great if you can share some training details like lr_scheduler, epochs for model convergence and anything special if you guys used.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1145009,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2021-01-08T20:00:11.827000",
          "content": "<p><a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> thanks for sharing the solution.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1145017,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2021-01-08T20:11:19.707000",
          "content": "<p>For the Tensorflow model that achieved the <code>private LB 0.815</code>:<br>\nWe tried other features as well, but those were by far the most important for our model and others didn't really improve our CV much.<br>\nEmbedding were all categorical, as continuous didn't work out well for us.<br>\nWe transformed <code>lag_time</code> to 173 categories, which gave a nice boost to the model.  We used the help of pandas <code>qcut</code>, like so:</p>\n<pre><code>N_ltg = 400 # number of groups 400 -&gt; 173\ntrain[\"ltg\"] = train.groupby(\"user_id\")[\"timestamp\"].shift()\n\n# Lag in seconds\ntrain[\"ltg\"] = ((train[\"timestamp\"] - train[\"ltg\"])/(1000.0))\ntrain[\"ltg\"] = round(train[\"ltg\"])\n\ntrain[\"ltg\"], bins = pd.qcut(train[\"ltg\"], N_ltg, duplicates=\"drop\", retbins=True) # duplicated -&gt; reduce about 1/2 of N_ltg\ntrain[\"ltg\"] = train[\"ltg\"].cat.codes # codes\n\n# Replace values\nN_ltg = train[\"ltg\"].nunique() # 173\ntrain[\"ltg\"] = train[\"ltg\"].replace(-1, N_ltg) # Replace -1 = NaN id \n</code></pre>\n<p>For the lr scheduler, we used Noam like in SAINT papers, and trained until convergence of the model, can be around 40-50 epochs of full training data per epoch.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1145072,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2021-01-08T21:22:49.230000",
          "content": "<p><a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> thanks.</p>\n<p>The model in the train notebook have achieved AUC 0.8001 even after 68 epochs. What's the reason of that even the model is using all features ? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1145099,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2021-01-08T22:10:22.567000",
          "content": "<p>Mainly because this is the \"demo\" version. You'll need to pass in the parameters of the transformers written in the description in order to get the higher score.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1145115,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2021-01-08T22:49:33.730000",
          "content": "<p>yeah you are right. I did not notice the hyperparams.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1145264,
          "author_name": "Hiep Nguyen",
          "author_url": "",
          "post_date": "2021-01-09T03:04:37.637000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> thanks for sharing, your team did a great job.</p>\n<p>Did you try to calculate lagged time by task_container_id? I did something like below to calculate lagged time. I wonder what is the good way for calculating lagged time? If you have tried it, could you tell me what is better and how much does it boost your model?</p>\n<pre><code>class UserLaggedTimeEncoder:\n    def __init__(self):\n        self.last_timestamp = defaultdict(float)\n        self.last_task_ts = defaultdict(float)\n        self.last_task_id = defaultdict(int)\n\n    def update(self, user_id, timestamp, task_id):\n        if task_id &gt; self.last_task_id[user_id]:\n            self.last_task_id[user_id] = task_id\n            self.last_task_ts[user_id] = self.last_timestamp[user_id]\n            self.last_timestamp[user_id] = timestamp\n\n    def encode(self, df):\n        lagged_time = np.zeros((df.shape[0],), dtype=np.float32)\n        for i, row in enumerate(tqdm(zip(df['user_id'].to_numpy(),\n                                         df['timestamp'].to_numpy(),\n                                         df['task_container_id'].to_numpy()),\n                                     total=df.shape[0])):\n            self.update(row[0], row[1], row[2])\n            lagged_time[i] = row[1] - self.last_task_ts[row[0]]\n        return lagged_time\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1145646,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2021-01-09T09:22:12.210000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/nvhbk16k53\" target=\"_blank\">@nvhbk16k53</a>. I noticed it's better to have a specific category for questions within a bundle. The lag time for questions within a bundle becomes 0 when you subtract.<br>\nYou can see how <code>lag time</code> was calculated in my code above.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1145746,
          "author_name": "Hiep Nguyen",
          "author_url": "",
          "post_date": "2021-01-09T10:11:56.537000",
          "content": "<p>Thank you for your reply <a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> . I saw the way you calculate <code>lag time</code>, it's be all zeros for all questions that have the same <code>task_container_id</code> except the first one. Mine is questions have the same <code>task_container_id</code> will have the same <code>lag time</code> value. What is the advantage of your compare to mine?</p>\n<p>I also experiment with LGBM that when calculate <code>lag time</code> like I did will give higher score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1145767,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2021-01-09T10:22:46.280000",
          "content": "<p>I see. I haven't tried the same for <code>lag time</code> like you did. So I can't tell what has more advantage, your method might work better also for transformers.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1145869,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2021-01-09T11:46:36.867000",
          "content": "<p><a href=\"https://www.kaggle.com/rafiko1\" target=\"_blank\">@rafiko1</a> thanks for sharing all the code and thoughts.</p>\n<p>I am asking a bit irrelevant question but it would be great if you can share your thoughts. Since you have done awesome with transformers so I just want to know that can we apply transformers on any time-series data or it depends on some characteristics of data to apply transformers on them.</p>\n<p>For example, if we want to apply the transformers to predict the fraud transaction of a user given its history of transactions(similarly we have history of user interactions in riid dataset), do you think transformers will be a good approach to apply on that dataset. I will appreciate your suggestions.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1146960,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2021-01-10T08:00:26.043000",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrehman245\" target=\"_blank\">@abdurrehman245</a> <br>\nI have a feeling transformers can be applied to more use-cases of time series, including fraud detection.<br>\nCan't tell with uncertainty until one experiments with it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1147263,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2021-01-10T11:50:15.543000",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p>About</p>\n<pre><code>And @mpware has a brilliant way of storing each user's attempt history using only 3 bits per time step to save memory. (3 bits can represent 0 thru 7 in binary).\n</code></pre>\n<p>Would it be possible to share how you perform this? At the very end of this competition, I needed to deal with the time/memory issue once I introduced the performance features. And it took time, and I didn't have much time to fix the other inference bugs due to the new features.</p>\n<p>It would be great to learn from your great management of the memory, hope I could hear from you, at least a bit.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1148710,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-11T10:33:09.883000",
          "content": "<p><a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> I'm maintaining the following stucture in memory. I've been inspired from different threads in forum and benchmarked multiple solutions before keeping this one. Per user and per content attempts dictionary costs too much memory if stored as integers (even <code>np.int8</code>) because it could be up to 393k x 13k.  To lower memory usage I'm using 3xbitarray (index is content_id) to address 2x2x2=8 attempts. As we're using attempts are categories we have 0-7 values and all attempts beyond 7 fall into last category which is indeed 7 or higher. </p>\n<p>We also need to store per user history for:</p>\n<ul>\n<li>Questions</li>\n<li>Answers</li>\n<li>Lags</li>\n<li>Elapsed time</li>\n<li>Attempts</li>\n<li>Part (optional)</li>\n<li>Had explanation (optional)</li>\n</ul>\n<p>Attempts history per user is computed on-fly based on the  <code>cseen</code> values containing only the total (that's the trick).</p>\n<p>Our different models (LGBM, Transfomer) can share this dictionary.</p>\n<p><code>per_user_dict = defaultdict(CustomDictNP)</code></p>\n<pre><code>def bit_array():\n    b = bitarray(13530, endian='little') # Higher than totals question if any new question in test set\n    b.setall(False)\n    return b\n\nclass CustomDictNP:\n    def __init__(self):\n        # Count/sum for further average\n        self.qc = 0 # answers count\n        self.qs = 0 # answers sum (correct)\n        self.s = 0 # session id\n        self.sc = 1 # session count\n        self.ss = 0 # total sessions\n        self.qes = 0.0 # total elapsed time\n\n        # History\n        self.ha = np.array([], dtype=np.bool) # answers\n        self.hp = np.array([], dtype=np.int8) # parts\n        self.hq = np.array([], dtype=np.int16) # questions\n        self.he = np.array([], dtype=np.float32) # elapsed time\n        self.hx = np.array([], dtype=np.bool) # prior question had explanation\n        self.hl = np.array([], dtype=np.float32) # lag\n\n        self.hat = np.array([], dtype=np.int8) # attempts\n\n        # Last timestamps\n        self.q = 0 # Questions\n        self.timestamp_u = []\n        self.timestamp_u_correct = []\n\n        # content_id seen        \n        self.cseen0 = bit_array()\n        self.cseen1 = bit_array()\n        self.cseen2 = bit_array()\n\n\ndef add_attempt(custom, cid):\n    new_attempt = get_attempts(custom, cid) + 1\n    if new_attempt &lt; 8:\n        if new_attempt == 1: \n            custom.cseen0[cid] = True\n            custom.cseen1[cid] = False\n            custom.cseen2[cid] = False\n        elif new_attempt == 2: \n            custom.cseen0[cid] = False\n            custom.cseen1[cid] = True\n            custom.cseen2[cid] = False\n        elif new_attempt == 3: \n            custom.cseen0[cid] = True\n            custom.cseen1[cid] = True\n            custom.cseen2[cid] = False\n        elif new_attempt == 4: \n            custom.cseen0[cid] = False\n            custom.cseen1[cid] = False\n            custom.cseen2[cid] = True\n        elif new_attempt == 5: \n            custom.cseen0[cid] = True\n            custom.cseen1[cid] = False\n            custom.cseen2[cid] = True\n        elif new_attempt == 6: \n            custom.cseen0[cid] = False\n            custom.cseen1[cid] = True\n            custom.cseen2[cid] = True\n        else: \n            custom.cseen0[cid] = True\n            custom.cseen1[cid] = True\n            custom.cseen2[cid] = True\n\ndef get_attempts(custom, cid):\n    return custom.cseen0[cid] + 2*custom.cseen1[cid] + 4*custom.cseen2[cid]\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1148739,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-11T11:04:35.493000",
          "content": "<p><a href=\"https://www.kaggle.com/MPWARE\" target=\"_blank\">@MPWARE</a>, thank you very much for being kindness to share, really appreciated. I haven't read it in detail to fully understand, but seems a very elegant yet efficient approach to the problem.</p>\n<p>Personally, I was able to fix this memory issue ( Per user and per content performance history ) in a short time (yet not a very clean way) by the following idea:</p>\n<p>Quite a lot of (user, question) pairs are in fact with (n_attempt, n_correctness) being (1, 1) and (1, 0). The pairs with <code>attempt &gt;= 2</code> is quite rare. Therefore, I have 2 nested dictionaries:</p>\n<pre><code>1. The 1st dict is:  user_id -&gt; question_id -&gt; the performances for those n_attempt &gt;= 2 (quite small dict)\n\n2. The 2nd dict: (for those pairs with `n_attemp = 1`)\n\n    {\n        user_id: \n           { \n               'correct': [q_id_x, q_id_y, ...]\n               'incorrect': [q_id_m, q_in_, ...]\n           }\n    }\n</code></pre>\n<p>If  (user, question) can't be found in the above 2 dictionaries, it means that question is not seen by the user.</p>\n<p>The idea is kind similar to sparse matrix, where we only store the content with values.</p>\n<p>Despite this effort, it still take 4G or 5G memory when loading into memory. I will try to compare your approach later.</p>\n<p>Comments to my approaches  are appreciated (if any). Thanks again for sharing.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1149507,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-11T22:23:04.897000",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> I understand your way of memory management now. Great.</p>\n<p>I am wondering why your team don't make a feature of the nb of correctness per user per content, or better, the ratio, if you already have such good way to manage memory. It think you might even get better by doing so, although in my case, I can't see the positive effect on LB due to the input inconsistency bug in my pipeline. I am just wondering the reason your team don't use it</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1149648,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2021-01-12T03:27:22.690000",
          "content": "<p>Really cool use of bit arrays! I didn't extend it to do the counting part as well! Sweet!</p>\n<pre><code>def get_attempts(custom, cid):\n    return custom.cseen0[cid] + 2*custom.cseen1[cid] + 4*custom.cseen2[cid]\n</code></pre>\n<p>This is basically bit shifting, correct? (Converting the bits back to decimal radix)</p>\n<pre><code>        # content_id seen        \n        self.cseen0 = bit_array()\n        self.cseen1 = bit_array()\n        self.cseen2 = bit_array()\n</code></pre>\n<p>This basically represents the 3 bits needed to count till 7 i believe.[ with <code>cseen2</code> as the MSB]</p>\n<p>Another use of the same bitarray's is <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266#1069347\" target=\"_blank\">here</a> <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a>! Basically a bool flag to detect whether use has seen that content in past or not.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1149990,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-12T09:31:09.773000",
          "content": "<p>We've tried it but results were close (little bit lower) than without.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1150173,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-12T12:30:57.373000",
          "content": "<p>Alright, that is quite strange. Do anyone in your team have some insight about why this is not working?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1150174,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-12T12:32:56.903000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> I saw that thread, but didn't take the time to read it carefully and find the shown code, I was so depressed about my inference bug at that time. Anyway, now I know there is bitarray. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1145911,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2021-01-09T12:23:13.433000",
      "content": "<p>Congratz to you all !</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1145432,
      "author_name": "2981",
      "author_url": "",
      "post_date": "2021-01-09T06:24:04.433000",
      "content": "<p>thanks for sharing! I am a beginer to transformer model. I think I need to learn the model before seeing your share.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1145294,
      "author_name": "Dean",
      "author_url": "",
      "post_date": "2021-01-09T03:52:07.940000",
      "content": "<p>Thanks for your sharing and congratulations to your team.</p>\n<p>Would you like to explain how to add the \"Number of attempts\" feature into encoder?<br>\nIs it convert to embedding?<br>\nThank you!!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1145326,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-01-09T04:25:36.937000",
          "content": "<p>Yes as an embedding and add it to other encoder input embeddings. </p>\n<p>The encoder input begins as <code>x = tf.keras.layers.Embedding(number of unique content id, 288)(CONTENT ID SEQUENCE)</code>. Then you add positional encoding, <code>x += positional_encoding(window size, 288)</code>. And last you add \"Number of attempts\" encoding. <code>x += tf.keras.layers.Embedding(max number of attempts + 1, 288)(NUMBER OF ATTEMPTS SEQUENCE)</code>.</p>\n<p>The maximum number of attempts is clipped with <code>NUMBER OF ATTEMPTS SEQUENCE = np.clip(NUMBER OF ATTEMPTS SEQUENCE, 0, 7)</code>.</p>\n<p>And <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> has a brilliant way of storing each user's attempt history using only 3 bits per time step to save memory. (3 bits can represent 0 thru 7 in binary).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1145376,
          "author_name": "Dean",
          "author_url": "",
          "post_date": "2021-01-09T05:17:44.343000",
          "content": "<p>Very impressive. Thanks for your reply.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1145897,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2021-01-09T12:17:48.453000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I can see that <code>99%</code> of of attempts on content_id is &lt;=7. Is that the reason of clipping the max value to 7?  </p>\n<p><code>x += tf.keras.layers.Embedding(max number of attempts + 1, 288)(NUMBER OF ATTEMPTS SEQUENCE)</code></p>\n<p>so <code>max number of attempts</code> will be 7 then and why are you adding 1 more to number of attempts ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1146604,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-09T21:55:56.690000",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrehman245\" target=\"_blank\">@abdurrehman245</a> The real reason was to limit memory usage for inference. We limit attempts from 0 to 7, everything above 7 falls in the last category which is indeed 7 and beyond.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1148196,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-01-11T01:43:04.583000",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrehman245\" target=\"_blank\">@abdurrehman245</a></p>\n<blockquote>\n  <p>why are you adding 1 more to number of attempts?</p>\n</blockquote>\n<p>Since the max is 7, there are 8 unique values, i.e. 0,1,2,3,4,5,6,7. So the embedding needs to handle a categorical variable of cardinality 8.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1145155,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2021-01-09T00:10:55.967000",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Congrats on the gold medal and nice single model solution !</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1145680,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-01-09T09:39:16.443000",
          "content": "<p><a href=\"https://www.kaggle.com/tikutiku\" target=\"_blank\">@tikutiku</a> Thank you! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1145144,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2021-01-08T23:44:57.310000",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> Congratulations to your team, very impressed that single RAINT+ with these 5 features can reach such a high score. My <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209596\" target=\"_blank\">best RAINT+ model</a> is only 0.802 on private</p>\n<p>Would like to know more details about which of these features are on the encoder side and which are on the decoder side?</p>\n<pre><code>Content id\nLag time\nPrior question elapsed time\nPrevious responses\nNumber of attempts\n</code></pre>\n<p><br>\nThere's no doubt on the content id and response. For other features, could you please explain a little bit about which parts did you put the feature in?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1145160,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-01-09T00:14:19.633000",
          "content": "<p>Thanks. Encoder had Content id and Number of attempts. Decoder had Lag time, Prior question elapsed time, and Previous responses.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1145175,
          "author_name": "william.wu",
          "author_url": "",
          "post_date": "2021-01-09T00:31:15.080000",
          "content": "<p>Got it, thanks! </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1144738,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-01-08T16:32:19.443000",
      "content": "<p>Congratulations on 16th place and gold medal <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and team. Good work and strongly transformer model</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1147557,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2021-01-10T15:21:41.433000",
      "content": "<p>How about the question's tags, it's a VarLen feature. Do you have any idea of adding VarLen feature in the RAINT+? What I did is sorting the tags in each question then combine them and map to an integer. But it's not an optimal way, since the information of each individual tag is lost.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1148508,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2021-01-11T07:28:17.583000",
          "content": "<p>I use a tag embedding, i.e. map each integer in [-1, 188) to vectors of a fixed dim (in my case 256). Then these embedding are averaged (the embedding of -1 is for padding for  the tags, as you said, it is varlen, and it is ignored during the averaging). I don't know how the team of <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> did it though.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1145433,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-09T06:24:04.433000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1144696": "Hi all,\n\n\nHere is an insight of our 2 solutions that both scored public 0.812/private 0.815 and that reached 16th gold place.\n\n\nThis competition was both ML and engineering optimization to make everything work in 9h with 13GB RAM/16GB GPU. We spent almost 30% of time on optimization to keep the last 512 interactions per users +  per content attempts in memory + required for our features.\n\nWe would like to thank Kaggle and RIIID organizers for this great competition! Congratulations to the top teams and all competitors for their motivation all along the challenge.\nI would like to thank my teammates @rafiko1, @cdeotte @titericz and @matthiasanderer. You've been amazing, I've learnt a lot from you. I really enjoyed this competition.\n\n\n## Solution 1: Single transformer model\n\nThe SAINT+ model is described here https://arxiv.org/pdf/2010.12042.pdf\nThe code for our SAINT+ adaptation is available here https://github.com/rafiko1/Riiid-sharing.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fa47293277b9e989ab5c81c269e6187a8%2Fdoc_saint.png?generation=1610121546239700&alt=media)\n\n\n\n\nOur single model SAINT++ achieved, CV: 0.812 Public LB: 0.812, Private LB: 0.815\n\nWe trained with 95% of users first then fine tuned with all data using smart window technique (see diagram below for SAKT). The model is simple in terms of features. It only contains the four features of SAINT+ (pictured above), with one additional feature - the number of attempts of a user for specific content (hence SAINT++):\n\n\n\n*   Content id\n*   Lag time\n*   Prior question elapsed time\n*   Previous responses\n*   Number of attempts \n\nThe greatest improvement in features compared to SAINT+ came from grouping lag time into seconds, unlike minutes as done in the paper. \nThen, we went bigger and bigger on the architecture and burned some GPU power 🔥. We increased on parameters of the model, most importantly the sequence length and number of layers. Final parameters of the model are as follows: \n\n\n\n\n1. Input Sequence length: 512\n2. Encoding layers: 4\n3. Decoding Layers: 4\n4. Embedding size: 288\n5. Dense Layer: 768\n6. heads: 8\n7. Dropout: 0.20\n\nWe used the Noam learning rate scheduler: with initial warmup and exponential decrease down to 2e-5. \n \nFinal improvement came from our **_recursive trick_** during inference. Here, we rounded predictions that came from the same bundle to **_0 or 1 _**- as their true response is unknown in time yet. The rounded predictions are then fed back to the model to predict the next response within the same bundle. This trick boosts CV LB +0.0025, but requires a batch size of 1, so we couldn’t ensemble multiple transformer models.\n\n\n### Solution 2: Ensemble of transformer, modified SAKT and LGB\n\n\n\n*   LightGBM model scored CV=0.793, public LB=0.792 with 44 features.  \nOur main features: \nQuestion correctness per content and per user, tags 1 and 2, part, elapsed time, had explanation, number of attempts, multiple lags, running average (answer), multiple rolling means/median (answers, lags, elapsed time) + weighted mean, mean after/before 30 interactions, multiple momentums (lag, answers), per part correctness, per session (8 hours split) running average. Only 3 categories: part, tags1, tags2. Train/valid split from [Tito](https://www.kaggle.com/its7171/cv-strategy). \n\n\n*   Pytorch SAKT modified scored CV=0.786, LB=0.789 with additional features. Training procedure with a smart window. \n*   When a user’s sequence length is larger than model input, i.e. N>W, then using random crops gives +0.002 CV versus tiled crops. And using smart random crops gives +0.003 CV versus tiled crops. Basic random crops have a low probability of selecting the early or late questions from a user’s sequence whereas smart random crops have an equally likely probability of selecting all questions from a user.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fa1239fda69769669432879b2b1a4ae39%2Fdoc_window.png?generation=1610121604322940&alt=media)\n    \n\n\n \n\n\n*   TensorFlow Transformer model alone scored CV=0.811, LB=0.811\n\n    Same as solution#1 but with sequence length = 256\n\n\n\n### What did not work:\n\n\n\n*   TabNet\n*   Features with lectures for LGB. It worked on CV but not on LB (might be an issue in inference).\n*   Post processing using absolute position of question aka. question sequence number. Plotting mean(answered_correctly) vs question number looked like the image below. We can see that the 30 first questions have a different distribution compared with the rest. Also looks like there are subsequently batches of 30 questions (becomes visible if we zoom in the plot below). PP using that information worked in CV improving by around 0.0009, but didn’t work on LB.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe4b8abca67c11258a438304b8fc67db5%2Fdoc_pp.png?generation=1610121577542081&alt=media)\n    \n\n\n\n\n\n### What worked partially:\n\n \nBut was not applicable for us within the 9h runtime limit:\n\n\n\n*   More than 3 models ensemble\n*   Level 2 model (XGB) could boost by +0.001\n\n\n### Lessons learnt:\n\n\n\n*   Start inference Kernel as soon as possible when you need to deal with an API.\n*   Try to simulate API locally to understand how data will be handled. [Tito](https://www.kaggle.com/its7171/time-series-api-iter-test-emulator)’s simulator was perfect for that purpose.\n*   Push your inference (with the simulator) to the limits to debug it, it will avoid the frustrating “submission scoring error”. \n*   Team-up at some point, your teammates always have good ideas.\n\nOne additional word to Kaggle @sohier I loved your API and the way it hides private data, it’s more realistic as in real world usage/production and it avoided chaotic blending. Congratulations for that, however, even if I guess you want to prevent probing, you should find a solution to provide better error feedback. If it is not possible (the more error codes the more probing) then you need to provide a simulator and guidelines allowing competitors to troubleshoot locally.\n",
    "1146920": "@MPWare Thanks for sharing the \"Simple\" Pytorch code, but why does it score so much less than Tensorflow?\n\n1. The Pytorch model uses 2 layers instead of Tensorflow uses 4 layers\n2. The Pytorch model uses 100 sequence length instead of Tensorflow uses 512 sequence length\n3. The Pytorch model uses 256 embedding size instead of Tensorflow uses 288\n4. The Pytorch model uses dropout = 0.1 but Tensorflow uses dropout = 0.2\n5. The Pytorch model only uses exercise_id  in encoder, but Tensorflow uses and response, but Tensorflow uses exercise_id, Lag time, prior question elapsed time, previous responses, number of attempts. Both models just use response_correct embedding in the decoder\n\nAre there any other differences? Would the Pytorch model be as good if you changed it to use all these features? Ty, I will really cherish that \"simple\" pytorch demo notebook you provided.",
    "1146671": "Congrats & thank you for sharing your approach and codes.",
    "1146361": "Congratulations and thanks for sharing your solution, I and many others will be able to learn a lot from it. I also used Tensorflow and implemented Saint+ however my code to sample sequences/users was not robust so I intend to see how the model would have performed using the sampling code you have shared.\n\nI have read through the code in demo_riiid_train.ipynb and think I understand how you have set up training and validation sets but was hoping you could confirm my understanding?\n\n1. You allocate distinct users to both the training and validation set\n2. You then assign a probability to each user based on sequence length. For say user 115 in the training set their probability is their sequence length divided by the sum of all sequence lengths in training set.\n3. You then generate a one off validation set where you randomly sample users with replacement and take **random training crops** for each row. Based on the 'select_window_size' function\n4. For the training dataset you do the same but repeat step 3 each epoch\n\nYou sample using N_SELECT_PER_EPOCH = 100000, was this just a hyperparameter you tuned?\n\n\nThank you ",
    "1146392": "Congrats, single 0.812/0.815 with SAINT is really great!",
    "1144929": "@mpware congrats on gold medal and thanks for sharing your approach. Would you like to share your code ?",
    "1147263": "@mpware @cdeotte \n\nAbout\n\n```\nAnd @mpware has a brilliant way of storing each user's attempt history using only 3 bits per time step to save memory. (3 bits can represent 0 thru 7 in binary).\n```\n\nWould it be possible to share how you perform this? At the very end of this competition, I needed to deal with the time/memory issue once I introduced the performance features. And it took time, and I didn't have much time to fix the other inference bugs due to the new features.\n\nIt would be great to learn from your great management of the memory, hope I could hear from you, at least a bit.",
    "1145911": "Congratz to you all !",
    "1145432": "thanks for sharing! I am a beginer to transformer model. I think I need to learn the model before seeing your share.",
    "1145294": "Thanks for your sharing and congratulations to your team.\n\nWould you like to explain how to add the \"Number of attempts\" feature into encoder?\nIs it convert to embedding?\nThank you!!",
    "1145155": "@mpware Congrats on the gold medal and nice single model solution !",
    "1145144": "@mpware Congratulations to your team, very impressed that single RAINT+ with these 5 features can reach such a high score. My [best RAINT+ model](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209596) is only 0.802 on private\n\nWould like to know more details about which of these features are on the encoder side and which are on the decoder side?\n```\nContent id\nLag time\nPrior question elapsed time\nPrevious responses\nNumber of attempts\n``` \nThere's no doubt on the content id and response. For other features, could you please explain a little bit about which parts did you put the feature in?",
    "1144738": "Congratulations on 16th place and gold medal @mpware @cdeotte and team. Good work and strongly transformer model",
    "1147557": "How about the question's tags, it's a VarLen feature. Do you have any idea of adding VarLen feature in the RAINT+? What I did is sorting the tags in each question then combine them and map to an integer. But it's not an optimal way, since the information of each individual tag is lost.",
    "1145433": ""
  }
}