{
  "id": 209303,
  "title": "SAINT Debug Tips | Last Day & All the best for everyone",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209303",
  "author_name": "Manikanth Reddy",
  "post_date": "2021-01-07T01:25:56.313000",
  "votes": 31,
  "comment_count": 36,
  "views": 0,
  "content": "<p>We have finally reached the last day of the competition. It was a long journey for me but was worth it. Looking back at the last 1 month, I gained so much from the community. A huge thanks to all of you who participated in this competition especially to the ones who were actively sharing their ideas in discussion forums and notebooks. I can't thank you people enough. I hope that Kaggle becomes better in the coming days for us. </p>\n<hr>\n<p>I have gained  400 ranks since yesterday. I focused more on building my SAINT model. I wasted around 1 week or about 15 submissions debugging why my inference is giving an AUC of 0.5-0.6 on Public LB. It was very frustrating but I finally managed to find my error. My single best model of SAINT has a public LB of 0.792. Here are my tips if you have worked on SAINT and couldn't fix your inference kernel.</p>\n<h2>Check your Pipeline</h2>\n<ul>\n<li>Make sure that you recheck your training data creation and inference kernel creation. I made a mistake in creating lagtime feature(multiple discussions on this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107391\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207434\" target=\"_blank\">here</a>). I was calculating it properly in my inference but wrong during my data creation part. Once I fixed this, my AUC went back to normal. My current code is as follows. Even if I use the lagtime1 feature below, it is giving a good score for me.</li>\n</ul>\n<pre><code>df['last_timestamp'] = df[['user_id', 'timestamp']].groupby(['user_id'])['timestamp'].shift(1, fill_value=0)\ndf['last_timestamp'] = df[['user_id', 'task_container_id', 'last_timestamp']].groupby(['user_id', 'task_container_id'])['last_timestamp'].transform('first')\n# df['lagtime1'] = df['timestamp'] - df['last_timestamp']\n\ndf['last_task_container_size'] = df[['user_id', 'task_container_id']].groupby(['user_id', 'task_container_id'])['task_container_id'].transform('size')\ndf['last_task_container_size'] = df[['user_id', 'last_task_container_size']].groupby(['user_id'])['last_task_container_size'].shift(1, fill_value=0)\ndf['last_task_container_size'] = df[['user_id', 'task_container_id', 'last_task_container_size']].groupby(['user_id', 'task_container_id'])['last_task_container_size'].transform('first')\n\ndf['lagtime'] = df['timestamp'] - df['last_timestamp'] - (df['prior_question_elapsed_time'] * df['last_task_container_size'])\n</code></pre>\n<ul>\n<li>If you are not able to create group data due to memory issues, try to <a href=\"https://www.kaggle.com/yanamal/pandas-performance-hack-data-split-by-user\" target=\"_blank\">split data by user</a>. It really helped me out a lot.</li>\n<li>Make an inference kernel ASAP before you even try to optimize your model. It would be a waste of time to spend on something that doesn't work at the end. </li>\n</ul>\n<h2>SAINT Architecture</h2>\n<ul>\n<li>Make sure to use the masking properly. I am using masking suggested by <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a>. The code is as follows:</li>\n</ul>\n<pre><code>def create_task_mask(tasks):\n    seq_length = len(tasks)\n    future_mask = np.triu(np.ones((seq_length, seq_length)), k=1).astype('bool')\n    container_mask = np.ones((seq_length, seq_length))\n    container_mask = (container_mask * tasks.reshape(1,-1)) == (container_mask * tasks.reshape(-1,1))\n    future_mask = future_mask + container_mask\n    np.fill_diagonal(future_mask, 0)\n    return future_mask\n</code></pre>\n<p>This one works well for me.</p>\n<ul>\n<li>If you are not able to create Encoder and Decoder by yourself, you should use nn.Transformer class. </li>\n</ul>\n<pre><code>class SAINTModel(nn.Module):\n    def __init__(self, *, max_seq=128, embed_dim=128, d_model=128, dropout=0.0, feedforward_dim=128, enc_layers=1, dec_layers=1, heads=8):\n        super(SAINTModel, self).__init__()\n\n        self.enc_layers = enc_layers\n        self.dec_layers = dec_layers\n        self.heads = heads\n\n        self.enc_embed = EncoderEmbed(*)\n        self.dec_embed = DecoderEmbed(*)\n        self.transformer = torch.nn.Transformer(d_model, heads, enc_layers, dec_layers, feedforward_dim, dropout)\n        self.pred = nn.Linear(embed_dim, 1)\n\n    def forward(self, *encoder_inputs, *decoder_inputs, mask):\n        enc_embed = self.enc_embed(*encoder_inputs).permute(1, 0, 2)\n        dec_embed = self.dec_embed(*decoder_inputs).permute(1, 0, 2)\n\n        x = self.transformer(\n            enc_embed, dec_embed,\n            src_mask=mask.repeat(self.heads, 1, 1), \n            tgt_mask=mask.repeat(self.heads, 1, 1), \n            memory_mask=mask.repeat(self.heads, 1, 1)\n        )\n\n        x = x.permute(1, 0, 2)\n        x = self.pred(x)\n        return x.squeeze(-1)\n</code></pre>\n<ul>\n<li>* and *encoder_inputs and *decoder_inputs are placeholders. Replace them with your features. Encoder Inputs should be related to the question the user is solving now (like content_id, part, …), Decoder Inputs could be related to the previous attempts (Response, Lagtime, …). You only have to implement EncoderEmbed and DecoderEmbed classes. Even though this code is very simple, this works like a charm for me.</li>\n</ul>\n<h2>Inference</h2>\n<ul>\n<li>If you are ensembling with LightGBM with a lot of features, memory will be a big issue especially if you want to use ram. In that case, make sure to only load the last SEQ_LEN of interactions for each user instead of loading the whole dataset. I used to use joblib for everything but lately, I have been facing a lot of RAM issues due to it. Since then I have converted to using CPickle. It is much efficient for me. </li>\n<li>Make sure to optimize code as much as possible. My code is not that optimal. Takes me about 2+ hours for inference. I will try to optimize it before ensembling it.</li>\n</ul>\n<h2>Quick Tip</h2>\n<ul>\n<li>If you haven't used SAINT or SAKT, consider using <a href=\"https://www.kaggle.com/gilfernandes/riiid-self-attention-transformer\" target=\"_blank\">this one</a> or <a href=\"https://www.kaggle.com/manikanthr5/riiid-sakt-model-inference-public\" target=\"_blank\">my model</a> for your ensemble. Few top public ones use them and there is no reason you shouldn't try to use them if you are aiming for a good rank. The pretrained dataset is <a href=\"https://www.kaggle.com/manikanthr5/riiid-sakt-model-dataset-public\" target=\"_blank\">available here</a>. The data has been <a href=\"https://www.kaggle.com/manikanthr5/riiid-sakt-model-dataset-public/activity\" target=\"_blank\">downloaded 300+ times</a>. </li>\n</ul>\n<p>I will add today if I remember anything else. Peace. </p>",
  "messages": [
    {
      "id": 1141898,
      "postDate": "2021-01-07T01:25:56.313Z",
      "content": "<p>We have finally reached the last day of the competition. It was a long journey for me but was worth it. Looking back at the last 1 month, I gained so much from the community. A huge thanks to all of you who participated in this competition especially to the ones who were actively sharing their ideas in discussion forums and notebooks. I can't thank you people enough. I hope that Kaggle becomes better in the coming days for us. </p>\n<hr>\n<p>I have gained  400 ranks since yesterday. I focused more on building my SAINT model. I wasted around 1 week or about 15 submissions debugging why my inference is giving an AUC of 0.5-0.6 on Public LB. It was very frustrating but I finally managed to find my error. My single best model of SAINT has a public LB of 0.792. Here are my tips if you have worked on SAINT and couldn't fix your inference kernel.</p>\n<h2>Check your Pipeline</h2>\n<ul>\n<li>Make sure that you recheck your training data creation and inference kernel creation. I made a mistake in creating lagtime feature(multiple discussions on this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107391\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207434\" target=\"_blank\">here</a>). I was calculating it properly in my inference but wrong during my data creation part. Once I fixed this, my AUC went back to normal. My current code is as follows. Even if I use the lagtime1 feature below, it is giving a good score for me.</li>\n</ul>\n<pre><code>df['last_timestamp'] = df[['user_id', 'timestamp']].groupby(['user_id'])['timestamp'].shift(1, fill_value=0)\ndf['last_timestamp'] = df[['user_id', 'task_container_id', 'last_timestamp']].groupby(['user_id', 'task_container_id'])['last_timestamp'].transform('first')\n# df['lagtime1'] = df['timestamp'] - df['last_timestamp']\n\ndf['last_task_container_size'] = df[['user_id', 'task_container_id']].groupby(['user_id', 'task_container_id'])['task_container_id'].transform('size')\ndf['last_task_container_size'] = df[['user_id', 'last_task_container_size']].groupby(['user_id'])['last_task_container_size'].shift(1, fill_value=0)\ndf['last_task_container_size'] = df[['user_id', 'task_container_id', 'last_task_container_size']].groupby(['user_id', 'task_container_id'])['last_task_container_size'].transform('first')\n\ndf['lagtime'] = df['timestamp'] - df['last_timestamp'] - (df['prior_question_elapsed_time'] * df['last_task_container_size'])\n</code></pre>\n<ul>\n<li>If you are not able to create group data due to memory issues, try to <a href=\"https://www.kaggle.com/yanamal/pandas-performance-hack-data-split-by-user\" target=\"_blank\">split data by user</a>. It really helped me out a lot.</li>\n<li>Make an inference kernel ASAP before you even try to optimize your model. It would be a waste of time to spend on something that doesn't work at the end. </li>\n</ul>\n<h2>SAINT Architecture</h2>\n<ul>\n<li>Make sure to use the masking properly. I am using masking suggested by <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a>. The code is as follows:</li>\n</ul>\n<pre><code>def create_task_mask(tasks):\n    seq_length = len(tasks)\n    future_mask = np.triu(np.ones((seq_length, seq_length)), k=1).astype('bool')\n    container_mask = np.ones((seq_length, seq_length))\n    container_mask = (container_mask * tasks.reshape(1,-1)) == (container_mask * tasks.reshape(-1,1))\n    future_mask = future_mask + container_mask\n    np.fill_diagonal(future_mask, 0)\n    return future_mask\n</code></pre>\n<p>This one works well for me.</p>\n<ul>\n<li>If you are not able to create Encoder and Decoder by yourself, you should use nn.Transformer class. </li>\n</ul>\n<pre><code>class SAINTModel(nn.Module):\n    def __init__(self, *, max_seq=128, embed_dim=128, d_model=128, dropout=0.0, feedforward_dim=128, enc_layers=1, dec_layers=1, heads=8):\n        super(SAINTModel, self).__init__()\n\n        self.enc_layers = enc_layers\n        self.dec_layers = dec_layers\n        self.heads = heads\n\n        self.enc_embed = EncoderEmbed(*)\n        self.dec_embed = DecoderEmbed(*)\n        self.transformer = torch.nn.Transformer(d_model, heads, enc_layers, dec_layers, feedforward_dim, dropout)\n        self.pred = nn.Linear(embed_dim, 1)\n\n    def forward(self, *encoder_inputs, *decoder_inputs, mask):\n        enc_embed = self.enc_embed(*encoder_inputs).permute(1, 0, 2)\n        dec_embed = self.dec_embed(*decoder_inputs).permute(1, 0, 2)\n\n        x = self.transformer(\n            enc_embed, dec_embed,\n            src_mask=mask.repeat(self.heads, 1, 1), \n            tgt_mask=mask.repeat(self.heads, 1, 1), \n            memory_mask=mask.repeat(self.heads, 1, 1)\n        )\n\n        x = x.permute(1, 0, 2)\n        x = self.pred(x)\n        return x.squeeze(-1)\n</code></pre>\n<ul>\n<li>* and *encoder_inputs and *decoder_inputs are placeholders. Replace them with your features. Encoder Inputs should be related to the question the user is solving now (like content_id, part, …), Decoder Inputs could be related to the previous attempts (Response, Lagtime, …). You only have to implement EncoderEmbed and DecoderEmbed classes. Even though this code is very simple, this works like a charm for me.</li>\n</ul>\n<h2>Inference</h2>\n<ul>\n<li>If you are ensembling with LightGBM with a lot of features, memory will be a big issue especially if you want to use ram. In that case, make sure to only load the last SEQ_LEN of interactions for each user instead of loading the whole dataset. I used to use joblib for everything but lately, I have been facing a lot of RAM issues due to it. Since then I have converted to using CPickle. It is much efficient for me. </li>\n<li>Make sure to optimize code as much as possible. My code is not that optimal. Takes me about 2+ hours for inference. I will try to optimize it before ensembling it.</li>\n</ul>\n<h2>Quick Tip</h2>\n<ul>\n<li>If you haven't used SAINT or SAKT, consider using <a href=\"https://www.kaggle.com/gilfernandes/riiid-self-attention-transformer\" target=\"_blank\">this one</a> or <a href=\"https://www.kaggle.com/manikanthr5/riiid-sakt-model-inference-public\" target=\"_blank\">my model</a> for your ensemble. Few top public ones use them and there is no reason you shouldn't try to use them if you are aiming for a good rank. The pretrained dataset is <a href=\"https://www.kaggle.com/manikanthr5/riiid-sakt-model-dataset-public\" target=\"_blank\">available here</a>. The data has been <a href=\"https://www.kaggle.com/manikanthr5/riiid-sakt-model-dataset-public/activity\" target=\"_blank\">downloaded 300+ times</a>. </li>\n</ul>\n<p>I will add today if I remember anything else. Peace. </p>",
      "rawMarkdown": "We have finally reached the last day of the competition. It was a long journey for me but was worth it. Looking back at the last 1 month, I gained so much from the community. A huge thanks to all of you who participated in this competition especially to the ones who were actively sharing their ideas in discussion forums and notebooks. I can't thank you people enough. I hope that Kaggle becomes better in the coming days for us. \n\n---\n\nI have gained ~~300~~ 400 ranks since yesterday. I focused more on building my SAINT model. I wasted around 1 week or about 15 submissions debugging why my inference is giving an AUC of 0.5-0.6 on Public LB. It was very frustrating but I finally managed to find my error. My single best model of SAINT has a public LB of 0.792. Here are my tips if you have worked on SAINT and couldn't fix your inference kernel.\n\n## Check your Pipeline\n- Make sure that you recheck your training data creation and inference kernel creation. I made a mistake in creating lagtime feature(multiple discussions on this [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107391) and [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207434)). I was calculating it properly in my inference but wrong during my data creation part. Once I fixed this, my AUC went back to normal. My current code is as follows. Even if I use the lagtime1 feature below, it is giving a good score for me.\n```python\ndf['last_timestamp'] = df[['user_id', 'timestamp']].groupby(['user_id'])['timestamp'].shift(1, fill_value=0)\ndf['last_timestamp'] = df[['user_id', 'task_container_id', 'last_timestamp']].groupby(['user_id', 'task_container_id'])['last_timestamp'].transform('first')\n# df['lagtime1'] = df['timestamp'] - df['last_timestamp']\n\ndf['last_task_container_size'] = df[['user_id', 'task_container_id']].groupby(['user_id', 'task_container_id'])['task_container_id'].transform('size')\ndf['last_task_container_size'] = df[['user_id', 'last_task_container_size']].groupby(['user_id'])['last_task_container_size'].shift(1, fill_value=0)\ndf['last_task_container_size'] = df[['user_id', 'task_container_id', 'last_task_container_size']].groupby(['user_id', 'task_container_id'])['last_task_container_size'].transform('first')\n\ndf['lagtime'] = df['timestamp'] - df['last_timestamp'] - (df['prior_question_elapsed_time'] * df['last_task_container_size'])\n```\n- If you are not able to create group data due to memory issues, try to [split data by user](https://www.kaggle.com/yanamal/pandas-performance-hack-data-split-by-user). It really helped me out a lot.\n- Make an inference kernel ASAP before you even try to optimize your model. It would be a waste of time to spend on something that doesn't work at the end. \n\n## SAINT Architecture\n- Make sure to use the masking properly. I am using masking suggested by @shujun717. The code is as follows:\n```python\ndef create_task_mask(tasks):\n    seq_length = len(tasks)\n    future_mask = np.triu(np.ones((seq_length, seq_length)), k=1).astype('bool')\n    container_mask = np.ones((seq_length, seq_length))\n    container_mask = (container_mask * tasks.reshape(1,-1)) == (container_mask * tasks.reshape(-1,1))\n    future_mask = future_mask + container_mask\n    np.fill_diagonal(future_mask, 0)\n    return future_mask\n```\nThis one works well for me.\n- If you are not able to create Encoder and Decoder by yourself, you should use nn.Transformer class. \n```python\nclass SAINTModel(nn.Module):\n    def __init__(self, *, max_seq=128, embed_dim=128, d_model=128, dropout=0.0, feedforward_dim=128, enc_layers=1, dec_layers=1, heads=8):\n        super(SAINTModel, self).__init__()\n        \n        self.enc_layers = enc_layers\n        self.dec_layers = dec_layers\n        self.heads = heads\n        \n        self.enc_embed = EncoderEmbed(*)\n        self.dec_embed = DecoderEmbed(*)\n        self.transformer = torch.nn.Transformer(d_model, heads, enc_layers, dec_layers, feedforward_dim, dropout)\n        self.pred = nn.Linear(embed_dim, 1)\n        \n    def forward(self, *encoder_inputs, *decoder_inputs, mask):\n        enc_embed = self.enc_embed(*encoder_inputs).permute(1, 0, 2)\n        dec_embed = self.dec_embed(*decoder_inputs).permute(1, 0, 2)\n\n        x = self.transformer(\n            enc_embed, dec_embed,\n            src_mask=mask.repeat(self.heads, 1, 1), \n            tgt_mask=mask.repeat(self.heads, 1, 1), \n            memory_mask=mask.repeat(self.heads, 1, 1)\n        )\n        \n        x = x.permute(1, 0, 2)\n        x = self.pred(x)\n        return x.squeeze(-1)\n```\n- * and *encoder_inputs and *decoder_inputs are placeholders. Replace them with your features. Encoder Inputs should be related to the question the user is solving now (like content_id, part, ...), Decoder Inputs could be related to the previous attempts (Response, Lagtime, ...). You only have to implement EncoderEmbed and DecoderEmbed classes. Even though this code is very simple, this works like a charm for me.\n\n## Inference\n- If you are ensembling with LightGBM with a lot of features, memory will be a big issue especially if you want to use ram. In that case, make sure to only load the last SEQ_LEN of interactions for each user instead of loading the whole dataset. I used to use joblib for everything but lately, I have been facing a lot of RAM issues due to it. Since then I have converted to using CPickle. It is much efficient for me. \n- Make sure to optimize code as much as possible. My code is not that optimal. Takes me about 2+ hours for inference. I will try to optimize it before ensembling it.\n\n## Quick Tip\n- If you haven't used SAINT or SAKT, consider using [this one](https://www.kaggle.com/gilfernandes/riiid-self-attention-transformer) or [my model](https://www.kaggle.com/manikanthr5/riiid-sakt-model-inference-public) for your ensemble. Few top public ones use them and there is no reason you shouldn't try to use them if you are aiming for a good rank. The pretrained dataset is [available here](https://www.kaggle.com/manikanthr5/riiid-sakt-model-dataset-public). The data has been [downloaded 300+ times](https://www.kaggle.com/manikanthr5/riiid-sakt-model-dataset-public/activity). \n\nI will add today if I remember anything else. Peace. ",
      "votes": 29
    },
    {
      "id": 1142030,
      "postDate": "2021-01-07T05:08:56.060Z",
      "content": "<p>The task mask code is written by me 😃</p>",
      "rawMarkdown": "The task mask code is written by me 😃",
      "votes": 4,
      "replies": [
        {
          "id": 1142036,
          "postDate": "2021-01-07T05:12:22.513Z",
          "content": "<p>Thank you for the script. It helped me save a lot of time and get accurate results. </p>",
          "rawMarkdown": "Thank you for the script. It helped me save a lot of time and get accurate results. ",
          "votes": 1
        },
        {
          "id": 1142040,
          "postDate": "2021-01-07T05:16:47.797Z",
          "content": "<p>Thanks, glad it worked for you. Unfortunately it only made things worse for me for some reason, which might because I only use an encoder and that type of masking is too much</p>",
          "rawMarkdown": "Thanks, glad it worked for you. Unfortunately it only made things worse for me for some reason, which might because I only use an encoder and that type of masking is too much",
          "votes": 2
        },
        {
          "id": 1142056,
          "postDate": "2021-01-07T05:35:16.483Z",
          "content": "<p>I by default used it once my inference started working properly. I didn't get time to compare the results. </p>",
          "rawMarkdown": "I by default used it once my inference started working properly. I didn't get time to compare the results. ",
          "votes": 1
        },
        {
          "id": 1142326,
          "postDate": "2021-01-07T10:12:08.543Z",
          "content": "<p>Now that I think about it, I should have used normal future mask for encoder and task based future mask for decoder. All the information required for encoder is available in a batch. </p>",
          "rawMarkdown": "Now that I think about it, I should have used normal future mask for encoder and task based future mask for decoder. All the information required for encoder is available in a batch. ",
          "votes": 1
        },
        {
          "id": 1142341,
          "postDate": "2021-01-07T10:25:25.967Z",
          "content": "<p>Yeah that's also why I said I had worse results maybe because I only use an encoder, since all information has to be embedded together, task mask did not work well for me. </p>",
          "rawMarkdown": "Yeah that's also why I said I had worse results maybe because I only use an encoder, since all information has to be embedded together, task mask did not work well for me. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1142604,
      "postDate": "2021-01-07T13:48:05.173Z",
      "content": "<p>Thank you so much. Your works are very helpful.<br>\nI modified your model to blend with my GBT model for the final submission.</p>",
      "rawMarkdown": "Thank you so much. Your works are very helpful.\nI modified your model to blend with my GBT model for the final submission.",
      "votes": 1
    },
    {
      "id": 1142531,
      "postDate": "2021-01-07T12:56:28.897Z",
      "content": "<p>Good job and thanks for the tips and notebooks ! </p>\n<p>I didn't have the opportunity to dig transformers yet as I joined the competition quit late (end of MoA), but I'll sure explore more your work and your sharing when this end !</p>",
      "rawMarkdown": "Good job and thanks for the tips and notebooks ! \n\nI didn't have the opportunity to dig transformers yet as I joined the competition quit late (end of MoA), but I'll sure explore more your work and your sharing when this end !",
      "votes": 1,
      "replies": [
        {
          "id": 1142620,
          "postDate": "2021-01-07T13:58:09.250Z",
          "content": "<p>I should say I have learnt a lot from the discussions you have initiated. Your TabNet discussion helped me in MoA competition. Till now I have only been sharing modifications of public work. For my next public notebook, I want to make sure it is my genuine work. </p>",
          "rawMarkdown": "I should say I have learnt a lot from the discussions you have initiated. Your TabNet discussion helped me in MoA competition. Till now I have only been sharing modifications of public work. For my next public notebook, I want to make sure it is my genuine work. "
        }
      ]
    },
    {
      "id": 1142455,
      "postDate": "2021-01-07T11:52:46.167Z",
      "content": "<p>Huge THX to your SAKT notebook,it is my first kaggle competition and your work help me so much! </p>",
      "rawMarkdown": "Huge THX to your SAKT notebook,it is my first kaggle competition and your work help me so much! ",
      "votes": 1
    },
    {
      "id": 1142299,
      "postDate": "2021-01-07T09:51:48.860Z",
      "content": "<p>Thank you for sharing so much in this competition. I like your posts and notebooks 🙂</p>",
      "rawMarkdown": "Thank you for sharing so much in this competition. I like your posts and notebooks 🙂",
      "votes": 1,
      "replies": [
        {
          "id": 1142330,
          "postDate": "2021-01-07T10:13:39.003Z",
          "content": "<p>Thanks for your kind words, it means a lot. I have a lot in this competition from other people. Just trying to share a part of what I have learnt. </p>",
          "rawMarkdown": "Thanks for your kind words, it means a lot. I have a lot in this competition from other people. Just trying to share a part of what I have learnt. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1141904,
      "postDate": "2021-01-07T01:42:04.157Z",
      "content": "<p>You're so enthusiastic, even though I'm running out of time to understand it</p>",
      "rawMarkdown": "You're so enthusiastic, even though I'm running out of time to understand it",
      "votes": 1
    },
    {
      "id": 1142003,
      "postDate": "2021-01-07T04:24:33.427Z",
      "content": "<p>Great, thank you for share your knowledge!<br>\nI didn't get what the SAINT is until now, I focused on LGBM through this competition. <br>\nI plan to learn it after the competition finished. I appreciate that you shared your findings. </p>\n<p>By the way, I see a placeholder in your post in the beginning of the \"SAINT Architecture\" part.<br>\nI guess you want to say thank you a person? </p>",
      "rawMarkdown": "Great, thank you for share your knowledge!\nI didn't get what the SAINT is until now, I focused on LGBM through this competition. \nI plan to learn it after the competition finished. I appreciate that you shared your findings. \n\nBy the way, I see a placeholder in your post in the beginning of the \"SAINT Architecture\" part.\nI guess you want to say thank you a person? ",
      "votes": 2,
      "replies": [
        {
          "id": 1142020,
          "postDate": "2021-01-07T04:43:38.423Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kokitanisaka\" target=\"_blank\">@kokitanisaka</a>. With Just LGBM you have a really good rank. </p>\n<p>I was also confused about the SAINT architecture. After working on the public SAKT models, I got a good understanding of what was going on with the Transformer models. I went through the SAINT and SAINT+ paper multiple times to understand it. I think you will have a lot of learning tomorrow from top public notebooks. Any ways all the best for you. </p>\n<p>Yes I checked the discussions multiple times to find the person who actually posted the script. Will recheck and add him. It saved me a lot of time. </p>",
          "rawMarkdown": "Hi @kokitanisaka. With Just LGBM you have a really good rank. \n\nI was also confused about the SAINT architecture. After working on the public SAKT models, I got a good understanding of what was going on with the Transformer models. I went through the SAINT and SAINT+ paper multiple times to understand it. I think you will have a lot of learning tomorrow from top public notebooks. Any ways all the best for you. \n\nYes I checked the discussions multiple times to find the person who actually posted the script. Will recheck and add him. It saved me a lot of time. ",
          "votes": 2
        },
        {
          "id": 1142048,
          "postDate": "2021-01-07T05:25:58.777Z",
          "content": "<p>Our solution is ensemble of SAINT and LGBM. My team mate built a SAINT model. <br>\nSo start from SAKT might give better understanding in the model. <br>\nWill try, thanks! </p>",
          "rawMarkdown": "Our solution is ensemble of SAINT and LGBM. My team mate built a SAINT model. \nSo start from SAKT might give better understanding in the model. \nWill try, thanks! "
        },
        {
          "id": 1142057,
          "postDate": "2021-01-07T05:36:46.417Z",
          "content": "<p>Yes. SAKT only has encoder part. No decoder is present. SAKT is also not a Self Attention Model since the keys/queries and values are different. These are the stuff that SAINT added. The other major difference is how they are embedding responses. </p>",
          "rawMarkdown": "Yes. SAKT only has encoder part. No decoder is present. SAKT is also not a Self Attention Model since the keys/queries and values are different. These are the stuff that SAINT added. The other major difference is how they are embedding responses. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1141941,
      "postDate": "2021-01-07T02:51:44.243Z",
      "content": "<p>Congratulations in getting it working!</p>",
      "rawMarkdown": "Congratulations in getting it working!",
      "votes": 2,
      "replies": [
        {
          "id": 1141943,
          "postDate": "2021-01-07T02:56:38.087Z",
          "content": "<p>Thanks. Actually I was about to drop off thinking I might not be able to fix it. But then I remembered that you also posted it somewhere that you are also facing the same issue. I think that was the time major breakthrough for me. I learnt a lot from the discussions you have initiated in this competition. Thanks a lot for that. </p>",
          "rawMarkdown": "Thanks. Actually I was about to drop off thinking I might not be able to fix it. But then I remembered that you also posted it somewhere that you are also facing the same issue. I think that was the time major breakthrough for me. I learnt a lot from the discussions you have initiated in this competition. Thanks a lot for that. "
        },
        {
          "id": 1141946,
          "postDate": "2021-01-07T03:00:43.827Z",
          "content": "<p>The sad part is i am able to help others except myself :(; Nonetheless I am very happy! Good luck :)</p>",
          "rawMarkdown": "The sad part is i am able to help others except myself :(; Nonetheless I am very happy! Good luck :)",
          "votes": 1
        },
        {
          "id": 1141951,
          "postDate": "2021-01-07T03:07:14.453Z",
          "content": "<p>Don't be sad. You are doing great yourself. This competition may not have been up to your expectations but I am sure you will do great in your next competitions. You have been consistently participating in kaggle from last 1 year, that itself is a big feat. </p>",
          "rawMarkdown": "Don't be sad. You are doing great yourself. This competition may not have been up to your expectations but I am sure you will do great in your next competitions. You have been consistently participating in kaggle from last 1 year, that itself is a big feat. "
        },
        {
          "id": 1142504,
          "postDate": "2021-01-07T12:34:42.087Z",
          "content": "<blockquote>\n  <p>Don't be sad. You are doing great yourself. This competition may not have been up to your expectations but I am sure you will do great in your next competitions</p>\n</blockquote>\n<p>The comp has broken me now to a much greater extent… 🥺 Good luck in your future endeavours!</p>",
          "rawMarkdown": ">Don't be sad. You are doing great yourself. This competition may not have been up to your expectations but I am sure you will do great in your next competitions\n\nThe comp has broken me now to a much greater extent... 🥺 Good luck in your future endeavours!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1146771,
      "postDate": "2021-01-10T03:19:00.687Z",
      "content": "<p>Thanks for sharing!<br>\nI'm now learning what was wrong with my SAINT+ model setup.<br>\nThanks for referring to the mask stuff, I could have missed it.</p>",
      "rawMarkdown": "Thanks for sharing!\nI'm now learning what was wrong with my SAINT+ model setup.\nThanks for referring to the mask stuff, I could have missed it."
    },
    {
      "id": 1143259,
      "postDate": "2021-01-07T20:14:27.887Z",
      "content": "<pre><code>df['last_timestamp'] = df[['user_id', 'timestamp']].groupby(['user_id'])['timestamp'].shift(1, fill_value=0)\ndf['last_timestamp'] = df[['user_id', 'task_container_id', 'last_timestamp']].groupby(['user_id', 'task_container_id'])['last_timestamp'].transform('first')\n# df['lagtime1'] = df['timestamp'] - df['last_timestamp']\n\ndf['last_task_container_size'] = df[['user_id', 'task_container_id']].groupby(['user_id', 'task_container_id'])['task_container_id'].transform('size')\ndf['last_task_container_size'] = df[['user_id', 'last_task_container_size']].groupby(['user_id'])['last_task_container_size'].shift(1, fill_value=0)\ndf['last_task_container_size'] = df[['user_id', 'task_container_id', 'last_task_container_size']].groupby(['user_id', 'task_container_id'])['last_task_container_size'].transform('first')\n\ndf['lagtime'] = df['timestamp'] - df['last_timestamp'] - (df['prior_question_elapsed_time'] * df['last_task_container_size'])\n</code></pre>\n<p>Thank you for sharing this snippet and is really useful to me. <br>\nWould this code have desirable computation time in the state updates? I used this wonderful snippet to my final submission and hoping there will be no timeout,,😹</p>",
      "rawMarkdown": "```\ndf['last_timestamp'] = df[['user_id', 'timestamp']].groupby(['user_id'])['timestamp'].shift(1, fill_value=0)\ndf['last_timestamp'] = df[['user_id', 'task_container_id', 'last_timestamp']].groupby(['user_id', 'task_container_id'])['last_timestamp'].transform('first')\n# df['lagtime1'] = df['timestamp'] - df['last_timestamp']\n\ndf['last_task_container_size'] = df[['user_id', 'task_container_id']].groupby(['user_id', 'task_container_id'])['task_container_id'].transform('size')\ndf['last_task_container_size'] = df[['user_id', 'last_task_container_size']].groupby(['user_id'])['last_task_container_size'].shift(1, fill_value=0)\ndf['last_task_container_size'] = df[['user_id', 'task_container_id', 'last_task_container_size']].groupby(['user_id', 'task_container_id'])['last_task_container_size'].transform('first')\n\ndf['lagtime'] = df['timestamp'] - df['last_timestamp'] - (df['prior_question_elapsed_time'] * df['last_task_container_size'])\n```\nThank you for sharing this snippet and is really useful to me. \nWould this code have desirable computation time in the state updates? I used this wonderful snippet to my final submission and hoping there will be no timeout,,😹",
      "replies": [
        {
          "id": 1143824,
          "postDate": "2021-01-08T04:48:50.527Z",
          "content": "<p>Did it work for you? </p>",
          "rawMarkdown": "Did it work for you? "
        },
        {
          "id": 1143827,
          "postDate": "2021-01-08T04:50:58.297Z",
          "content": "<p>Yes, it worked and no timeout occurred. Thanks for the great code!</p>",
          "rawMarkdown": "Yes, it worked and no timeout occurred. Thanks for the great code!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1142118,
      "postDate": "2021-01-07T06:55:57.587Z",
      "content": "<p>tasks in create_task_mask fn mean seq of task id ?</p>",
      "rawMarkdown": "tasks in create_task_mask fn mean seq of task id ?",
      "replies": [
        {
          "id": 1142128,
          "postDate": "2021-01-07T07:10:06.767Z",
          "content": "<p>Yes. They are the values of Task Container Id</p>",
          "rawMarkdown": "Yes. They are the values of Task Container Id"
        }
      ]
    },
    {
      "id": 1142045,
      "postDate": "2021-01-07T05:24:31.480Z",
      "content": "<p>Thank you for your sharing. <br>\nChange from lagtime1 to lagtime, how much did you improve?</p>",
      "rawMarkdown": "Thank you for your sharing. \nChange from lagtime1 to lagtime, how much did you improve?",
      "replies": [
        {
          "id": 1142055,
          "postDate": "2021-01-07T05:34:01.283Z",
          "content": "<p>The CV was <strong>almost same</strong>. I am using lagtime1 only for my models since it is faster to compute. I didn't get to submit lagtime one to leaderboard since I have few submissions left. </p>",
          "rawMarkdown": "The CV was **almost same**. I am using lagtime1 only for my models since it is faster to compute. I didn't get to submit lagtime one to leaderboard since I have few submissions left. ",
          "votes": 1
        },
        {
          "id": 1142095,
          "postDate": "2021-01-07T06:27:50.307Z",
          "content": "<p>Thank you for replying. I also use lagtime1 as you, it can help me to get 0.800 for saint+ model.</p>",
          "rawMarkdown": "Thank you for replying. I also use lagtime1 as you, it can help me to get 0.800 for saint+ model.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1141984,
      "postDate": "2021-01-07T03:55:23.250Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 1141905,
      "postDate": "2021-01-07T01:42:27.993Z",
      "content": "<p>congratulations, I still debug lag time, my saint is 0.780, add lag time the score didn't increase first 5 epochs.my GPU time isn't enough yet.</p>",
      "rawMarkdown": "congratulations, I still debug lag time, my saint is 0.780, add lag time the score didn't increase first 5 epochs.my GPU time isn't enough yet.",
      "replies": [
        {
          "id": 1141944,
          "postDate": "2021-01-07T02:57:42.493Z",
          "content": "<p>I am using QuantileTransformer on my lagtime feature. Initially when I wasn't using that model was not getting converged. I don't if this is what is causing you the issue. Also check for the optimizer. </p>",
          "rawMarkdown": "I am using QuantileTransformer on my lagtime feature. Initially when I wasn't using that model was not getting converged. I don't if this is what is causing you the issue. Also check for the optimizer. "
        },
        {
          "id": 1141949,
          "postDate": "2021-01-07T03:06:26.443Z",
          "content": "<p>I use logtransform and scale [0,1],I will try run in TPU,Before adding lagtime my model could reach 0.78, after adding the first 5 epoch scores did not change, I have no GPU time to run through all epochs ,this is my last attempt,my lgbm model still can be improved</p>",
          "rawMarkdown": "I use logtransform and scale [0,1],I will try run in TPU,Before adding lagtime my model could reach 0.78, after adding the first 5 epoch scores did not change, I have no GPU time to run through all epochs ,this is my last attempt,my lgbm model still can be improved",
          "votes": 1
        }
      ]
    },
    {
      "id": 1142208,
      "postDate": "2021-01-07T08:28:55.820Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1142030,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2021-01-07T05:08:56.060000",
      "content": "<p>The task mask code is written by me 😃</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1142036,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-07T05:12:22.513000",
          "content": "<p>Thank you for the script. It helped me save a lot of time and get accurate results. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142040,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2021-01-07T05:16:47.797000",
          "content": "<p>Thanks, glad it worked for you. Unfortunately it only made things worse for me for some reason, which might because I only use an encoder and that type of masking is too much</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1142056,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-07T05:35:16.483000",
          "content": "<p>I by default used it once my inference started working properly. I didn't get time to compare the results. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142326,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-07T10:12:08.543000",
          "content": "<p>Now that I think about it, I should have used normal future mask for encoder and task based future mask for decoder. All the information required for encoder is available in a batch. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142341,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2021-01-07T10:25:25.967000",
          "content": "<p>Yeah that's also why I said I had worse results maybe because I only use an encoder, since all information has to be embedded together, task mask did not work well for me. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1142604,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2021-01-07T13:48:05.173000",
      "content": "<p>Thank you so much. Your works are very helpful.<br>\nI modified your model to blend with my GBT model for the final submission.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1142531,
      "author_name": "Jacky",
      "author_url": "",
      "post_date": "2021-01-07T12:56:28.897000",
      "content": "<p>Good job and thanks for the tips and notebooks ! </p>\n<p>I didn't have the opportunity to dig transformers yet as I joined the competition quit late (end of MoA), but I'll sure explore more your work and your sharing when this end !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1142620,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-07T13:58:09.250000",
          "content": "<p>I should say I have learnt a lot from the discussions you have initiated. Your TabNet discussion helped me in MoA competition. Till now I have only been sharing modifications of public work. For my next public notebook, I want to make sure it is my genuine work. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1142455,
      "author_name": "Tianyu Liu",
      "author_url": "",
      "post_date": "2021-01-07T11:52:46.167000",
      "content": "<p>Huge THX to your SAKT notebook,it is my first kaggle competition and your work help me so much! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1142299,
      "author_name": "Rodolphe Lampe",
      "author_url": "",
      "post_date": "2021-01-07T09:51:48.860000",
      "content": "<p>Thank you for sharing so much in this competition. I like your posts and notebooks 🙂</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1142330,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-07T10:13:39.003000",
          "content": "<p>Thanks for your kind words, it means a lot. I have a lot in this competition from other people. Just trying to share a part of what I have learnt. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1141904,
      "author_name": "Wain Wong",
      "author_url": "",
      "post_date": "2021-01-07T01:42:04.157000",
      "content": "<p>You're so enthusiastic, even though I'm running out of time to understand it</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1142003,
      "author_name": "Kouki",
      "author_url": "",
      "post_date": "2021-01-07T04:24:33.427000",
      "content": "<p>Great, thank you for share your knowledge!<br>\nI didn't get what the SAINT is until now, I focused on LGBM through this competition. <br>\nI plan to learn it after the competition finished. I appreciate that you shared your findings. </p>\n<p>By the way, I see a placeholder in your post in the beginning of the \"SAINT Architecture\" part.<br>\nI guess you want to say thank you a person? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1142020,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-07T04:43:38.423000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kokitanisaka\" target=\"_blank\">@kokitanisaka</a>. With Just LGBM you have a really good rank. </p>\n<p>I was also confused about the SAINT architecture. After working on the public SAKT models, I got a good understanding of what was going on with the Transformer models. I went through the SAINT and SAINT+ paper multiple times to understand it. I think you will have a lot of learning tomorrow from top public notebooks. Any ways all the best for you. </p>\n<p>Yes I checked the discussions multiple times to find the person who actually posted the script. Will recheck and add him. It saved me a lot of time. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1142048,
          "author_name": "Kouki",
          "author_url": "",
          "post_date": "2021-01-07T05:25:58.777000",
          "content": "<p>Our solution is ensemble of SAINT and LGBM. My team mate built a SAINT model. <br>\nSo start from SAKT might give better understanding in the model. <br>\nWill try, thanks! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142057,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-07T05:36:46.417000",
          "content": "<p>Yes. SAKT only has encoder part. No decoder is present. SAKT is also not a Self Attention Model since the keys/queries and values are different. These are the stuff that SAINT added. The other major difference is how they are embedding responses. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1141941,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2021-01-07T02:51:44.243000",
      "content": "<p>Congratulations in getting it working!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1141943,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-07T02:56:38.087000",
          "content": "<p>Thanks. Actually I was about to drop off thinking I might not be able to fix it. But then I remembered that you also posted it somewhere that you are also facing the same issue. I think that was the time major breakthrough for me. I learnt a lot from the discussions you have initiated in this competition. Thanks a lot for that. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1141946,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2021-01-07T03:00:43.827000",
          "content": "<p>The sad part is i am able to help others except myself :(; Nonetheless I am very happy! Good luck :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1141951,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-07T03:07:14.453000",
          "content": "<p>Don't be sad. You are doing great yourself. This competition may not have been up to your expectations but I am sure you will do great in your next competitions. You have been consistently participating in kaggle from last 1 year, that itself is a big feat. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142504,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2021-01-07T12:34:42.087000",
          "content": "<blockquote>\n  <p>Don't be sad. You are doing great yourself. This competition may not have been up to your expectations but I am sure you will do great in your next competitions</p>\n</blockquote>\n<p>The comp has broken me now to a much greater extent… 🥺 Good luck in your future endeavours!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1146771,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "2021-01-10T03:19:00.687000",
      "content": "<p>Thanks for sharing!<br>\nI'm now learning what was wrong with my SAINT+ model setup.<br>\nThanks for referring to the mask stuff, I could have missed it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1143259,
      "author_name": "Jihun Lorenzo Park",
      "author_url": "",
      "post_date": "2021-01-07T20:14:27.887000",
      "content": "<pre><code>df['last_timestamp'] = df[['user_id', 'timestamp']].groupby(['user_id'])['timestamp'].shift(1, fill_value=0)\ndf['last_timestamp'] = df[['user_id', 'task_container_id', 'last_timestamp']].groupby(['user_id', 'task_container_id'])['last_timestamp'].transform('first')\n# df['lagtime1'] = df['timestamp'] - df['last_timestamp']\n\ndf['last_task_container_size'] = df[['user_id', 'task_container_id']].groupby(['user_id', 'task_container_id'])['task_container_id'].transform('size')\ndf['last_task_container_size'] = df[['user_id', 'last_task_container_size']].groupby(['user_id'])['last_task_container_size'].shift(1, fill_value=0)\ndf['last_task_container_size'] = df[['user_id', 'task_container_id', 'last_task_container_size']].groupby(['user_id', 'task_container_id'])['last_task_container_size'].transform('first')\n\ndf['lagtime'] = df['timestamp'] - df['last_timestamp'] - (df['prior_question_elapsed_time'] * df['last_task_container_size'])\n</code></pre>\n<p>Thank you for sharing this snippet and is really useful to me. <br>\nWould this code have desirable computation time in the state updates? I used this wonderful snippet to my final submission and hoping there will be no timeout,,😹</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1143824,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-08T04:48:50.527000",
          "content": "<p>Did it work for you? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1143827,
          "author_name": "Jihun Lorenzo Park",
          "author_url": "",
          "post_date": "2021-01-08T04:50:58.297000",
          "content": "<p>Yes, it worked and no timeout occurred. Thanks for the great code!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1142118,
      "author_name": "qiaqia",
      "author_url": "",
      "post_date": "2021-01-07T06:55:57.587000",
      "content": "<p>tasks in create_task_mask fn mean seq of task id ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1142128,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-07T07:10:06.767000",
          "content": "<p>Yes. They are the values of Task Container Id</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1142045,
      "author_name": "chizhu",
      "author_url": "",
      "post_date": "2021-01-07T05:24:31.480000",
      "content": "<p>Thank you for your sharing. <br>\nChange from lagtime1 to lagtime, how much did you improve?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1142055,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-07T05:34:01.283000",
          "content": "<p>The CV was <strong>almost same</strong>. I am using lagtime1 only for my models since it is faster to compute. I didn't get to submit lagtime one to leaderboard since I have few submissions left. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1142095,
          "author_name": "chizhu",
          "author_url": "",
          "post_date": "2021-01-07T06:27:50.307000",
          "content": "<p>Thank you for replying. I also use lagtime1 as you, it can help me to get 0.800 for saint+ model.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1141984,
      "author_name": "zephyr666666",
      "author_url": "",
      "post_date": "2021-01-07T03:55:23.250000",
      "content": "<p>Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1141905,
      "author_name": "qiaqia",
      "author_url": "",
      "post_date": "2021-01-07T01:42:27.993000",
      "content": "<p>congratulations, I still debug lag time, my saint is 0.780, add lag time the score didn't increase first 5 epochs.my GPU time isn't enough yet.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1141944,
          "author_name": "Manikanth Reddy",
          "author_url": "",
          "post_date": "2021-01-07T02:57:42.493000",
          "content": "<p>I am using QuantileTransformer on my lagtime feature. Initially when I wasn't using that model was not getting converged. I don't if this is what is causing you the issue. Also check for the optimizer. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1141949,
          "author_name": "qiaqia",
          "author_url": "",
          "post_date": "2021-01-07T03:06:26.443000",
          "content": "<p>I use logtransform and scale [0,1],I will try run in TPU,Before adding lagtime my model could reach 0.78, after adding the first 5 epoch scores did not change, I have no GPU time to run through all epochs ,this is my last attempt,my lgbm model still can be improved</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1142208,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-07T08:28:55.820000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1141898": "We have finally reached the last day of the competition. It was a long journey for me but was worth it. Looking back at the last 1 month, I gained so much from the community. A huge thanks to all of you who participated in this competition especially to the ones who were actively sharing their ideas in discussion forums and notebooks. I can't thank you people enough. I hope that Kaggle becomes better in the coming days for us. \n\n---\n\nI have gained ~~300~~ 400 ranks since yesterday. I focused more on building my SAINT model. I wasted around 1 week or about 15 submissions debugging why my inference is giving an AUC of 0.5-0.6 on Public LB. It was very frustrating but I finally managed to find my error. My single best model of SAINT has a public LB of 0.792. Here are my tips if you have worked on SAINT and couldn't fix your inference kernel.\n\n## Check your Pipeline\n- Make sure that you recheck your training data creation and inference kernel creation. I made a mistake in creating lagtime feature(multiple discussions on this [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1107391) and [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207434)). I was calculating it properly in my inference but wrong during my data creation part. Once I fixed this, my AUC went back to normal. My current code is as follows. Even if I use the lagtime1 feature below, it is giving a good score for me.\n```python\ndf['last_timestamp'] = df[['user_id', 'timestamp']].groupby(['user_id'])['timestamp'].shift(1, fill_value=0)\ndf['last_timestamp'] = df[['user_id', 'task_container_id', 'last_timestamp']].groupby(['user_id', 'task_container_id'])['last_timestamp'].transform('first')\n# df['lagtime1'] = df['timestamp'] - df['last_timestamp']\n\ndf['last_task_container_size'] = df[['user_id', 'task_container_id']].groupby(['user_id', 'task_container_id'])['task_container_id'].transform('size')\ndf['last_task_container_size'] = df[['user_id', 'last_task_container_size']].groupby(['user_id'])['last_task_container_size'].shift(1, fill_value=0)\ndf['last_task_container_size'] = df[['user_id', 'task_container_id', 'last_task_container_size']].groupby(['user_id', 'task_container_id'])['last_task_container_size'].transform('first')\n\ndf['lagtime'] = df['timestamp'] - df['last_timestamp'] - (df['prior_question_elapsed_time'] * df['last_task_container_size'])\n```\n- If you are not able to create group data due to memory issues, try to [split data by user](https://www.kaggle.com/yanamal/pandas-performance-hack-data-split-by-user). It really helped me out a lot.\n- Make an inference kernel ASAP before you even try to optimize your model. It would be a waste of time to spend on something that doesn't work at the end. \n\n## SAINT Architecture\n- Make sure to use the masking properly. I am using masking suggested by @shujun717. The code is as follows:\n```python\ndef create_task_mask(tasks):\n    seq_length = len(tasks)\n    future_mask = np.triu(np.ones((seq_length, seq_length)), k=1).astype('bool')\n    container_mask = np.ones((seq_length, seq_length))\n    container_mask = (container_mask * tasks.reshape(1,-1)) == (container_mask * tasks.reshape(-1,1))\n    future_mask = future_mask + container_mask\n    np.fill_diagonal(future_mask, 0)\n    return future_mask\n```\nThis one works well for me.\n- If you are not able to create Encoder and Decoder by yourself, you should use nn.Transformer class. \n```python\nclass SAINTModel(nn.Module):\n    def __init__(self, *, max_seq=128, embed_dim=128, d_model=128, dropout=0.0, feedforward_dim=128, enc_layers=1, dec_layers=1, heads=8):\n        super(SAINTModel, self).__init__()\n        \n        self.enc_layers = enc_layers\n        self.dec_layers = dec_layers\n        self.heads = heads\n        \n        self.enc_embed = EncoderEmbed(*)\n        self.dec_embed = DecoderEmbed(*)\n        self.transformer = torch.nn.Transformer(d_model, heads, enc_layers, dec_layers, feedforward_dim, dropout)\n        self.pred = nn.Linear(embed_dim, 1)\n        \n    def forward(self, *encoder_inputs, *decoder_inputs, mask):\n        enc_embed = self.enc_embed(*encoder_inputs).permute(1, 0, 2)\n        dec_embed = self.dec_embed(*decoder_inputs).permute(1, 0, 2)\n\n        x = self.transformer(\n            enc_embed, dec_embed,\n            src_mask=mask.repeat(self.heads, 1, 1), \n            tgt_mask=mask.repeat(self.heads, 1, 1), \n            memory_mask=mask.repeat(self.heads, 1, 1)\n        )\n        \n        x = x.permute(1, 0, 2)\n        x = self.pred(x)\n        return x.squeeze(-1)\n```\n- * and *encoder_inputs and *decoder_inputs are placeholders. Replace them with your features. Encoder Inputs should be related to the question the user is solving now (like content_id, part, ...), Decoder Inputs could be related to the previous attempts (Response, Lagtime, ...). You only have to implement EncoderEmbed and DecoderEmbed classes. Even though this code is very simple, this works like a charm for me.\n\n## Inference\n- If you are ensembling with LightGBM with a lot of features, memory will be a big issue especially if you want to use ram. In that case, make sure to only load the last SEQ_LEN of interactions for each user instead of loading the whole dataset. I used to use joblib for everything but lately, I have been facing a lot of RAM issues due to it. Since then I have converted to using CPickle. It is much efficient for me. \n- Make sure to optimize code as much as possible. My code is not that optimal. Takes me about 2+ hours for inference. I will try to optimize it before ensembling it.\n\n## Quick Tip\n- If you haven't used SAINT or SAKT, consider using [this one](https://www.kaggle.com/gilfernandes/riiid-self-attention-transformer) or [my model](https://www.kaggle.com/manikanthr5/riiid-sakt-model-inference-public) for your ensemble. Few top public ones use them and there is no reason you shouldn't try to use them if you are aiming for a good rank. The pretrained dataset is [available here](https://www.kaggle.com/manikanthr5/riiid-sakt-model-dataset-public). The data has been [downloaded 300+ times](https://www.kaggle.com/manikanthr5/riiid-sakt-model-dataset-public/activity). \n\nI will add today if I remember anything else. Peace. ",
    "1142030": "The task mask code is written by me 😃",
    "1142604": "Thank you so much. Your works are very helpful.\nI modified your model to blend with my GBT model for the final submission.",
    "1142531": "Good job and thanks for the tips and notebooks ! \n\nI didn't have the opportunity to dig transformers yet as I joined the competition quit late (end of MoA), but I'll sure explore more your work and your sharing when this end !",
    "1142455": "Huge THX to your SAKT notebook,it is my first kaggle competition and your work help me so much! ",
    "1142299": "Thank you for sharing so much in this competition. I like your posts and notebooks 🙂",
    "1141904": "You're so enthusiastic, even though I'm running out of time to understand it",
    "1142003": "Great, thank you for share your knowledge!\nI didn't get what the SAINT is until now, I focused on LGBM through this competition. \nI plan to learn it after the competition finished. I appreciate that you shared your findings. \n\nBy the way, I see a placeholder in your post in the beginning of the \"SAINT Architecture\" part.\nI guess you want to say thank you a person? ",
    "1141941": "Congratulations in getting it working!",
    "1146771": "Thanks for sharing!\nI'm now learning what was wrong with my SAINT+ model setup.\nThanks for referring to the mask stuff, I could have missed it.",
    "1143259": "```\ndf['last_timestamp'] = df[['user_id', 'timestamp']].groupby(['user_id'])['timestamp'].shift(1, fill_value=0)\ndf['last_timestamp'] = df[['user_id', 'task_container_id', 'last_timestamp']].groupby(['user_id', 'task_container_id'])['last_timestamp'].transform('first')\n# df['lagtime1'] = df['timestamp'] - df['last_timestamp']\n\ndf['last_task_container_size'] = df[['user_id', 'task_container_id']].groupby(['user_id', 'task_container_id'])['task_container_id'].transform('size')\ndf['last_task_container_size'] = df[['user_id', 'last_task_container_size']].groupby(['user_id'])['last_task_container_size'].shift(1, fill_value=0)\ndf['last_task_container_size'] = df[['user_id', 'task_container_id', 'last_task_container_size']].groupby(['user_id', 'task_container_id'])['last_task_container_size'].transform('first')\n\ndf['lagtime'] = df['timestamp'] - df['last_timestamp'] - (df['prior_question_elapsed_time'] * df['last_task_container_size'])\n```\nThank you for sharing this snippet and is really useful to me. \nWould this code have desirable computation time in the state updates? I used this wonderful snippet to my final submission and hoping there will be no timeout,,😹",
    "1142118": "tasks in create_task_mask fn mean seq of task id ?",
    "1142045": "Thank you for your sharing. \nChange from lagtime1 to lagtime, how much did you improve?",
    "1141984": "Congratulations!",
    "1141905": "congratulations, I still debug lag time, my saint is 0.780, add lag time the score didn't increase first 5 epochs.my GPU time isn't enough yet.",
    "1142208": ""
  }
}