{
  "id": 209711,
  "title": "22nd Place Solution: Just Encoder Modules",
  "url": "/competitions/riiid-test-answer-prediction/writeups/hindsight-2020-22nd-place-solution-just-encoder-mo",
  "author_name": "",
  "post_date": "2021-01-08T10:51:56.690Z",
  "votes": 24,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Thanks for all the discussion and support from fellow Kagglers. I've learned a lot in the competition and tried a lot of things. Here a simple list of the stuff that I believe worked well. I'll try to cover the whole model but may make updates to clarify and add explanations later. </p>\n<p><strong><a href=\"https://www.kaggle.com/abdurrafae/22nd-solution-just-encoder-blocks\" target=\"_blank\">Link to notebook </a></strong></p>\n<p><strong>Architecture</strong><br>\nMy architecture was just 3 Encoder modules stacked together using 512 as d_model and 4 Encoder layers in each block.</p>\n<p>My Questions, Interactions and Response sequences were all padded on the left by a unique vector made using historical stats from the user. This ensured that all 3 sequences were aligned on the sequence number</p>\n<ol>\n<li><p>Questions Only Block (Q-Block):<br>\nThis one was just self attention over the questions.</p></li>\n<li><p>Interactions Only Block (I-Block):<br>\nThis one was just self attention over the interactions.</p></li>\n<li><p>Questions, Responses, Interaction Block (QRI Block):<br>\nI used output from Q-Block as query, I-Block as keys and Responses/(QRI Block) as values. The residual connection after attention module was made using value vector (instead of the default query vector). Self Attention Mask was used in first 3 layers and for the last layer I used a custom mask that only attended to prior interactions and didn't have the residual connection.</p></li>\n</ol>\n<p>Concatenated the output of Q-Block and QRI-Block, feed it into 3 Linear layers.</p>\n<p><strong>Data Usage</strong><br>\nI split the longer sequences into windows of 256 length and 128 overlap. (0-256, 128-384, 256-512).<br>\nI saved them in tf.records and during inference I took a uniform random 128 window from each 256 length sequences. This ensures any event would be placed at 0-127 location in my training sequence with equal probability (excluding the last 128 events in the last window of a user). For users having smaller sequence length I padded with 0 to make the length equal 128 and used those 128 each time.</p>\n<p><strong>Compute Resources</strong><br>\nI initially only used Kaggle GPUs to train the model. Started setting up a TPU pipeline in the last 3 weeks of the competition and in the end was able to use the TPUs as well. I think I've exhausted my complete TPU quota for last 3 weeks.</p>\n<p>I'd like to thank <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> (<strong>TPU Guru</strong>) for this <a href=\"https://www.kaggle.com/yihdarshieh/tpu-track-knowledge-states-of-1m-students\" target=\"_blank\">TPU notebook</a>. It helped a lot in setting up the TPU pipeline.</p>\n<p><strong>Sequential Encodings</strong><br>\nI used 2 encodings that captured the sequence of events. I subtracted the first value in each sequence from the rest to ensure that all encodings started with a 0.</p>\n<p><strong>Temporal Encoding</strong><br>\nFor this I converted the timestamp into minutes and used the same scheme as that of positional encoding with power (60 x 24 x 365 = 1 year). I believe this enabled the model to know how far apart in time were 2 questions/interactions. I believe this to be a better implementation of the lag time variable used in SAINT+ as it captures the difference in time between all of the events simultaneously rather than just between 2 adjacent events.</p>\n<p><strong>Positional Encoding</strong><br>\nFor this I used the task_container_id as the position and a power of 10,000.</p>\n<p><strong>Proxy for Knowledge</strong><br>\nInstead of just using 0/1 from responses I added an new feature which is a heuristic for knowledge in case the response was incorrect. If a user selects option 2 when 1 is the correct one. I would calculate (total number of times 2 was selected for that questions)/(total number of times the question has been answered incorrectly). This ratio was maintained and updated during inference as well. I checked the correlation of mean of proxy_knowledge for incorrect answers and overall user accuracy, it was around 0.4.</p>\n<p><strong>Question difficulty</strong><br>\nSimple feature that is calculated as  (total number of times the question has been answered correctly)/(total number of times the question has been answered). This ratio was maintained and updated during inference as well.</p>\n<p><strong>Custom Masks</strong><br>\nI used self attention mask that didn't attend on events of the same bundle.<br>\nFor the last layer of QRI block I removed the mask entries along the diagonal to ensure it only attends to prior values.</p>\n<p><strong>Starting Vector</strong><br>\nFor the starting vector I used counts of the time each tag was seen/answered correctly in questions and lectures. Just used a dense layer to encode this into the first vector of the sequence. </p>\n<p><strong>Other Details</strong></p>\n<ul>\n<li>Used Noam LR provided on the Transformers page on TF documentation. </li>\n<li>Batch size 1024 (Trained on TPUs - 1 Epoch took 4-6 mins)</li>\n<li>For Validation I just separated around 3.4% of users initially and used their sequences.</li>\n<li>Model converged around 20-30 Epochs. (4 Hours training time at max)</li>\n<li>Sequence length of 128 was used.</li>\n</ul>\n<p><strong>Question Embeddings</strong><br>\nI added up embeddings for content id, part id, (a weighted average of) tags ids and type_of (from lectures), then concatenated it with both Sequential embeddings and then into a dense layer with d_model dimensions.</p>\n<p><strong>Response Embeddings</strong><br>\nI concatenated answered correctly, proxy knowledge, question difficulty, time elapsed and question had explanation and feed into a dense layer with d_model dimensions.</p>\n<p><strong>Interaction Embeddings</strong><br>\nI just took the first d_model//2 units from both Questions/Response embeddings and concatenated them for this.</p>",
  "messages": [
    {
      "id": "1144237",
      "postDate": "01/08/2021 10:41:26",
      "content": "<p>Thanks for all the discussion and support from fellow Kagglers. I've learned a lot in the competition and tried a lot of things. Here a simple list of the stuff that I believe worked well. I'll try to cover the whole model but may make updates to clarify and add explanations later. </p>\n<p><strong><a href=\"https://www.kaggle.com/abdurrafae/22nd-solution-just-encoder-blocks\" target=\"_blank\">Link to notebook </a></strong></p>\n<p><strong>Architecture</strong><br>\nMy architecture was just 3 Encoder modules stacked together using 512 as d_model and 4 Encoder layers in each block.</p>\n<p>My Questions, Interactions and Response sequences were all padded on the left by a unique vector made using historical stats from the user. This ensured that all 3 sequences were aligned on the sequence number</p>\n<ol>\n<li><p>Questions Only Block (Q-Block):<br>\nThis one was just self attention over the questions.</p></li>\n<li><p>Interactions Only Block (I-Block):<br>\nThis one was just self attention over the interactions.</p></li>\n<li><p>Questions, Responses, Interaction Block (QRI Block):<br>\nI used output from Q-Block as query, I-Block as keys and Responses/(QRI Block) as values. The residual connection after attention module was made using value vector (instead of the default query vector). Self Attention Mask was used in first 3 layers and for the last layer I used a custom mask that only attended to prior interactions and didn't have the residual connection.</p></li>\n</ol>\n<p>Concatenated the output of Q-Block and QRI-Block, feed it into 3 Linear layers.</p>\n<p><strong>Data Usage</strong><br>\nI split the longer sequences into windows of 256 length and 128 overlap. (0-256, 128-384, 256-512).<br>\nI saved them in tf.records and during inference I took a uniform random 128 window from each 256 length sequences. This ensures any event would be placed at 0-127 location in my training sequence with equal probability (excluding the last 128 events in the last window of a user). For users having smaller sequence length I padded with 0 to make the length equal 128 and used those 128 each time.</p>\n<p><strong>Compute Resources</strong><br>\nI initially only used Kaggle GPUs to train the model. Started setting up a TPU pipeline in the last 3 weeks of the competition and in the end was able to use the TPUs as well. I think I've exhausted my complete TPU quota for last 3 weeks.</p>\n<p>I'd like to thank <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> (<strong>TPU Guru</strong>) for this <a href=\"https://www.kaggle.com/yihdarshieh/tpu-track-knowledge-states-of-1m-students\" target=\"_blank\">TPU notebook</a>. It helped a lot in setting up the TPU pipeline.</p>\n<p><strong>Sequential Encodings</strong><br>\nI used 2 encodings that captured the sequence of events. I subtracted the first value in each sequence from the rest to ensure that all encodings started with a 0.</p>\n<p><strong>Temporal Encoding</strong><br>\nFor this I converted the timestamp into minutes and used the same scheme as that of positional encoding with power (60 x 24 x 365 = 1 year). I believe this enabled the model to know how far apart in time were 2 questions/interactions. I believe this to be a better implementation of the lag time variable used in SAINT+ as it captures the difference in time between all of the events simultaneously rather than just between 2 adjacent events.</p>\n<p><strong>Positional Encoding</strong><br>\nFor this I used the task_container_id as the position and a power of 10,000.</p>\n<p><strong>Proxy for Knowledge</strong><br>\nInstead of just using 0/1 from responses I added an new feature which is a heuristic for knowledge in case the response was incorrect. If a user selects option 2 when 1 is the correct one. I would calculate (total number of times 2 was selected for that questions)/(total number of times the question has been answered incorrectly). This ratio was maintained and updated during inference as well. I checked the correlation of mean of proxy_knowledge for incorrect answers and overall user accuracy, it was around 0.4.</p>\n<p><strong>Question difficulty</strong><br>\nSimple feature that is calculated as  (total number of times the question has been answered correctly)/(total number of times the question has been answered). This ratio was maintained and updated during inference as well.</p>\n<p><strong>Custom Masks</strong><br>\nI used self attention mask that didn't attend on events of the same bundle.<br>\nFor the last layer of QRI block I removed the mask entries along the diagonal to ensure it only attends to prior values.</p>\n<p><strong>Starting Vector</strong><br>\nFor the starting vector I used counts of the time each tag was seen/answered correctly in questions and lectures. Just used a dense layer to encode this into the first vector of the sequence. </p>\n<p><strong>Other Details</strong></p>\n<ul>\n<li>Used Noam LR provided on the Transformers page on TF documentation. </li>\n<li>Batch size 1024 (Trained on TPUs - 1 Epoch took 4-6 mins)</li>\n<li>For Validation I just separated around 3.4% of users initially and used their sequences.</li>\n<li>Model converged around 20-30 Epochs. (4 Hours training time at max)</li>\n<li>Sequence length of 128 was used.</li>\n</ul>\n<p><strong>Question Embeddings</strong><br>\nI added up embeddings for content id, part id, (a weighted average of) tags ids and type_of (from lectures), then concatenated it with both Sequential embeddings and then into a dense layer with d_model dimensions.</p>\n<p><strong>Response Embeddings</strong><br>\nI concatenated answered correctly, proxy knowledge, question difficulty, time elapsed and question had explanation and feed into a dense layer with d_model dimensions.</p>\n<p><strong>Interaction Embeddings</strong><br>\nI just took the first d_model//2 units from both Questions/Response embeddings and concatenated them for this.</p>",
      "rawMarkdown": "Thanks for all the discussion and support from fellow Kagglers. I've learned a lot in the competition and tried a lot of things. Here a simple list of the stuff that I believe worked well. I'll try to cover the whole model but may make updates to clarify and add explanations later. \n\n**[Link to notebook ](https://www.kaggle.com/abdurrafae/22nd-solution-just-encoder-blocks)**\n\n**Architecture**\nMy architecture was just 3 Encoder modules stacked together using 512 as d_model and 4 Encoder layers in each block.\n\nMy Questions, Interactions and Response sequences were all padded on the left by a unique vector made using historical stats from the user. This ensured that all 3 sequences were aligned on the sequence number\n\n\n1. Questions Only Block (Q-Block):\nThis one was just self attention over the questions.\n\n2. Interactions Only Block (I-Block):\nThis one was just self attention over the interactions.\n\n3. Questions, Responses, Interaction Block (QRI Block):\nI used output from Q-Block as query, I-Block as keys and Responses/(QRI Block) as values. The residual connection after attention module was made using value vector (instead of the default query vector). Self Attention Mask was used in first 3 layers and for the last layer I used a custom mask that only attended to prior interactions and didn't have the residual connection.\n\nConcatenated the output of Q-Block and QRI-Block, feed it into 3 Linear layers.\n\n**Data Usage**\nI split the longer sequences into windows of 256 length and 128 overlap. (0-256, 128-384, 256-512).\nI saved them in tf.records and during inference I took a uniform random 128 window from each 256 length sequences. This ensures any event would be placed at 0-127 location in my training sequence with equal probability (excluding the last 128 events in the last window of a user). For users having smaller sequence length I padded with 0 to make the length equal 128 and used those 128 each time.\n\n**Compute Resources**\nI initially only used Kaggle GPUs to train the model. Started setting up a TPU pipeline in the last 3 weeks of the competition and in the end was able to use the TPUs as well. I think I've exhausted my complete TPU quota for last 3 weeks.\n\nI'd like to thank @yihdarshieh (**TPU Guru**) for this [TPU notebook](https://www.kaggle.com/yihdarshieh/tpu-track-knowledge-states-of-1m-students). It helped a lot in setting up the TPU pipeline.\n\n**Sequential Encodings**\nI used 2 encodings that captured the sequence of events. I subtracted the first value in each sequence from the rest to ensure that all encodings started with a 0.\n\n**Temporal Encoding**\nFor this I converted the timestamp into minutes and used the same scheme as that of positional encoding with power (60 x 24 x 365 = 1 year). I believe this enabled the model to know how far apart in time were 2 questions/interactions. I believe this to be a better implementation of the lag time variable used in SAINT+ as it captures the difference in time between all of the events simultaneously rather than just between 2 adjacent events.\n\n**Positional Encoding**\nFor this I used the task_container_id as the position and a power of 10,000.\n\n\n**Proxy for Knowledge**\nInstead of just using 0/1 from responses I added an new feature which is a heuristic for knowledge in case the response was incorrect. If a user selects option 2 when 1 is the correct one. I would calculate (total number of times 2 was selected for that questions)/(total number of times the question has been answered incorrectly). This ratio was maintained and updated during inference as well. I checked the correlation of mean of proxy_knowledge for incorrect answers and overall user accuracy, it was around 0.4.\n\n\n**Question difficulty**\nSimple feature that is calculated as  (total number of times the question has been answered correctly)/(total number of times the question has been answered). This ratio was maintained and updated during inference as well.\n\n\n**Custom Masks**\nI used self attention mask that didn't attend on events of the same bundle.\nFor the last layer of QRI block I removed the mask entries along the diagonal to ensure it only attends to prior values.\n\n**Starting Vector**\nFor the starting vector I used counts of the time each tag was seen/answered correctly in questions and lectures. Just used a dense layer to encode this into the first vector of the sequence. \n\n**Other Details**\n- Used Noam LR provided on the Transformers page on TF documentation. \n- Batch size 1024 (Trained on TPUs - 1 Epoch took 4-6 mins)\n- For Validation I just separated around 3.4% of users initially and used their sequences.\n- Model converged around 20-30 Epochs. (4 Hours training time at max)\n- Sequence length of 128 was used.\n\n**Question Embeddings**\nI added up embeddings for content id, part id, (a weighted average of) tags ids and type_of (from lectures), then concatenated it with both Sequential embeddings and then into a dense layer with d_model dimensions.\n\n\n**Response Embeddings**\nI concatenated answered correctly, proxy knowledge, question difficulty, time elapsed and question had explanation and feed into a dense layer with d_model dimensions.\n\n**Interaction Embeddings**\nI just took the first d_model//2 units from both Questions/Response embeddings and concatenated them for this.",
      "votes": null
    },
    {
      "id": "1144244",
      "postDate": "01/08/2021 10:45:54",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> , congratulation! Would it be possible to share your code later? And again, let's work together for a next interesting competition after some break, and get a Gold!</p>",
      "rawMarkdown": "abdurrafae , congratulation! Would it be possible to share your code later? And again, let's work together for a next interesting competition after some break, and get a Gold!",
      "votes": null
    },
    {
      "id": "1144248",
      "postDate": "01/08/2021 10:49:52",
      "content": "<p>Thanks and likewise.<br>\nI'll add the link to the notebook in this post in 10-15 mins. Just cleaning it up a bit. Haven't added comments in it though</p>",
      "rawMarkdown": "Thanks and likewise.\nI'll add the link to the notebook in this post in 10-15 mins. Just cleaning it up a bit. Haven't added comments in it though",
      "votes": null
    },
    {
      "id": "1144249",
      "postDate": "01/08/2021 10:52:24",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae/22nd-solution-just-encoder-blocks\" target=\"_blank\">https://www.kaggle.com/abdurrafae/22nd-solution-just-encoder-blocks</a></p>",
      "rawMarkdown": "https://www.kaggle.com/abdurrafae/22nd-solution-just-encoder-blocks",
      "votes": null
    },
    {
      "id": "1144250",
      "postDate": "01/08/2021 10:52:41",
      "content": "<p>Great - no problem, I just wonder why I couldn't get better results with smaller model. (But maybe it is because when I still worked with smaller model, I haven't used the good features I added later).</p>",
      "rawMarkdown": "Great - no problem, I just wonder why I couldn't get better results with smaller model. (But maybe it is because when I still worked with smaller model, I haven't used the good features I added later).",
      "votes": null
    },
    {
      "id": "1144254",
      "postDate": "01/08/2021 10:53:34",
      "content": "<p>Congrats, nice work, I'm looking forward to seeing the code.</p>",
      "rawMarkdown": "Congrats, nice work, I'm looking forward to seeing the code.",
      "votes": null
    },
    {
      "id": "1144257",
      "postDate": "01/08/2021 10:55:04",
      "content": "<p>Thanks mate. I've added the link in post now.</p>",
      "rawMarkdown": "Thanks mate. I've added the link in post now.",
      "votes": null
    },
    {
      "id": "1145253",
      "postDate": "01/09/2021 02:15:43",
      "content": "<p>Good job, it's impressive what you were able to achieve with such limited compute. Also glad to see someone else also using an encoder only model haha</p>",
      "rawMarkdown": "Good job, it's impressive what you were able to achieve with such limited compute. Also glad to see someone else also using an encoder only model haha",
      "votes": null
    },
    {
      "id": "1145750",
      "postDate": "01/09/2021 10:13:24",
      "content": "<p>It made more intuitive sense to only use encoders for this since we are only enhancing the vector space of each sequence (questions/interactions/responses). Plus my last encoder block simply put is saying \"having this question (query), find similarity with all the relevant past interactions (keys) and give a weighted average of the previous responses (values)\". </p>\n<p>The SAINT architecture didn't make intuitive sense to me at all.  </p>",
      "rawMarkdown": "It made more intuitive sense to only use encoders for this since we are only enhancing the vector space of each sequence (questions/interactions/responses). Plus my last encoder block simply put is saying \"having this question (query), find similarity with all the relevant past interactions (keys) and give a weighted average of the previous responses (values)\". \n\nThe SAINT architecture didn't make intuitive sense to me at all.",
      "votes": null
    },
    {
      "id": "1146845",
      "postDate": "01/10/2021 05:14:37",
      "content": "<p>Yeah I dont understand the SAINT architecture much either. In fact, it really confused me in the beginning</p>",
      "rawMarkdown": "Yeah I dont understand the SAINT architecture much either. In fact, it really confused me in the beginning",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1144244,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "01/08/2021 10:45:54",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> , congratulation! Would it be possible to share your code later? And again, let's work together for a next interesting competition after some break, and get a Gold!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144248,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "01/08/2021 10:49:52",
          "content": "<p>Thanks and likewise.<br>\nI'll add the link to the notebook in this post in 10-15 mins. Just cleaning it up a bit. Haven't added comments in it though</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144249,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "01/08/2021 10:52:24",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae/22nd-solution-just-encoder-blocks\" target=\"_blank\">https://www.kaggle.com/abdurrafae/22nd-solution-just-encoder-blocks</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144250,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/08/2021 10:52:41",
          "content": "<p>Great - no problem, I just wonder why I couldn't get better results with smaller model. (But maybe it is because when I still worked with smaller model, I haven't used the good features I added later).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144254,
      "author_name": "fomdata",
      "author_url": "",
      "post_date": "01/08/2021 10:53:34",
      "content": "<p>Congrats, nice work, I'm looking forward to seeing the code.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144257,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "01/08/2021 10:55:04",
          "content": "<p>Thanks mate. I've added the link in post now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1145253,
      "author_name": "shujun717",
      "author_url": "",
      "post_date": "01/09/2021 02:15:43",
      "content": "<p>Good job, it's impressive what you were able to achieve with such limited compute. Also glad to see someone else also using an encoder only model haha</p>",
      "votes": null,
      "replies": [
        {
          "id": 1145750,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "01/09/2021 10:13:24",
          "content": "<p>It made more intuitive sense to only use encoders for this since we are only enhancing the vector space of each sequence (questions/interactions/responses). Plus my last encoder block simply put is saying \"having this question (query), find similarity with all the relevant past interactions (keys) and give a weighted average of the previous responses (values)\". </p>\n<p>The SAINT architecture didn't make intuitive sense to me at all.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146845,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "01/10/2021 05:14:37",
          "content": "<p>Yeah I dont understand the SAINT architecture much either. In fact, it really confused me in the beginning</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1144237": "Thanks for all the discussion and support from fellow Kagglers. I've learned a lot in the competition and tried a lot of things. Here a simple list of the stuff that I believe worked well. I'll try to cover the whole model but may make updates to clarify and add explanations later. \n\n**[Link to notebook ](https://www.kaggle.com/abdurrafae/22nd-solution-just-encoder-blocks)**\n\n**Architecture**\nMy architecture was just 3 Encoder modules stacked together using 512 as d_model and 4 Encoder layers in each block.\n\nMy Questions, Interactions and Response sequences were all padded on the left by a unique vector made using historical stats from the user. This ensured that all 3 sequences were aligned on the sequence number\n\n\n1. Questions Only Block (Q-Block):\nThis one was just self attention over the questions.\n\n2. Interactions Only Block (I-Block):\nThis one was just self attention over the interactions.\n\n3. Questions, Responses, Interaction Block (QRI Block):\nI used output from Q-Block as query, I-Block as keys and Responses/(QRI Block) as values. The residual connection after attention module was made using value vector (instead of the default query vector). Self Attention Mask was used in first 3 layers and for the last layer I used a custom mask that only attended to prior interactions and didn't have the residual connection.\n\nConcatenated the output of Q-Block and QRI-Block, feed it into 3 Linear layers.\n\n**Data Usage**\nI split the longer sequences into windows of 256 length and 128 overlap. (0-256, 128-384, 256-512).\nI saved them in tf.records and during inference I took a uniform random 128 window from each 256 length sequences. This ensures any event would be placed at 0-127 location in my training sequence with equal probability (excluding the last 128 events in the last window of a user). For users having smaller sequence length I padded with 0 to make the length equal 128 and used those 128 each time.\n\n**Compute Resources**\nI initially only used Kaggle GPUs to train the model. Started setting up a TPU pipeline in the last 3 weeks of the competition and in the end was able to use the TPUs as well. I think I've exhausted my complete TPU quota for last 3 weeks.\n\nI'd like to thank @yihdarshieh (**TPU Guru**) for this [TPU notebook](https://www.kaggle.com/yihdarshieh/tpu-track-knowledge-states-of-1m-students). It helped a lot in setting up the TPU pipeline.\n\n**Sequential Encodings**\nI used 2 encodings that captured the sequence of events. I subtracted the first value in each sequence from the rest to ensure that all encodings started with a 0.\n\n**Temporal Encoding**\nFor this I converted the timestamp into minutes and used the same scheme as that of positional encoding with power (60 x 24 x 365 = 1 year). I believe this enabled the model to know how far apart in time were 2 questions/interactions. I believe this to be a better implementation of the lag time variable used in SAINT+ as it captures the difference in time between all of the events simultaneously rather than just between 2 adjacent events.\n\n**Positional Encoding**\nFor this I used the task_container_id as the position and a power of 10,000.\n\n\n**Proxy for Knowledge**\nInstead of just using 0/1 from responses I added an new feature which is a heuristic for knowledge in case the response was incorrect. If a user selects option 2 when 1 is the correct one. I would calculate (total number of times 2 was selected for that questions)/(total number of times the question has been answered incorrectly). This ratio was maintained and updated during inference as well. I checked the correlation of mean of proxy_knowledge for incorrect answers and overall user accuracy, it was around 0.4.\n\n\n**Question difficulty**\nSimple feature that is calculated as  (total number of times the question has been answered correctly)/(total number of times the question has been answered). This ratio was maintained and updated during inference as well.\n\n\n**Custom Masks**\nI used self attention mask that didn't attend on events of the same bundle.\nFor the last layer of QRI block I removed the mask entries along the diagonal to ensure it only attends to prior values.\n\n**Starting Vector**\nFor the starting vector I used counts of the time each tag was seen/answered correctly in questions and lectures. Just used a dense layer to encode this into the first vector of the sequence. \n\n**Other Details**\n- Used Noam LR provided on the Transformers page on TF documentation. \n- Batch size 1024 (Trained on TPUs - 1 Epoch took 4-6 mins)\n- For Validation I just separated around 3.4% of users initially and used their sequences.\n- Model converged around 20-30 Epochs. (4 Hours training time at max)\n- Sequence length of 128 was used.\n\n**Question Embeddings**\nI added up embeddings for content id, part id, (a weighted average of) tags ids and type_of (from lectures), then concatenated it with both Sequential embeddings and then into a dense layer with d_model dimensions.\n\n\n**Response Embeddings**\nI concatenated answered correctly, proxy knowledge, question difficulty, time elapsed and question had explanation and feed into a dense layer with d_model dimensions.\n\n**Interaction Embeddings**\nI just took the first d_model//2 units from both Questions/Response embeddings and concatenated them for this.",
    "1144244": "abdurrafae , congratulation! Would it be possible to share your code later? And again, let's work together for a next interesting competition after some break, and get a Gold!",
    "1144248": "Thanks and likewise.\nI'll add the link to the notebook in this post in 10-15 mins. Just cleaning it up a bit. Haven't added comments in it though",
    "1144249": "https://www.kaggle.com/abdurrafae/22nd-solution-just-encoder-blocks",
    "1144250": "Great - no problem, I just wonder why I couldn't get better results with smaller model. (But maybe it is because when I still worked with smaller model, I haven't used the good features I added later).",
    "1144254": "Congrats, nice work, I'm looking forward to seeing the code.",
    "1144257": "Thanks mate. I've added the link in post now.",
    "1145253": "Good job, it's impressive what you were able to achieve with such limited compute. Also glad to see someone else also using an encoder only model haha",
    "1145750": "It made more intuitive sense to only use encoders for this since we are only enhancing the vector space of each sequence (questions/interactions/responses). Plus my last encoder block simply put is saying \"having this question (query), find similarity with all the relevant past interactions (keys) and give a weighted average of the previous responses (values)\". \n\nThe SAINT architecture didn't make intuitive sense to me at all.",
    "1146845": "Yeah I dont understand the SAINT architecture much either. In fact, it really confused me in the beginning"
  },
  "source": "meta"
}