{
  "id": 209694,
  "title": "11th Place Solution",
  "url": "/competitions/riiid-test-answer-prediction/writeups/edulab-11th-place-solution",
  "author_name": "",
  "post_date": "2021-01-08T15:13:11.923Z",
  "votes": 54,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Thanks to my teammate <a href=\"https://www.kaggle.com/jusco11\" target=\"_blank\">Akihiko</a> for competing with me on this, there was great learning from the community for both of us. And thanks to our hosts for a wonderful challenge. Congrats all who competed. </p>\n<p>Solution has heavily inspired by Bestfitting's <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262\" target=\"_blank\">TalkingData solution</a> which we noticed a lot of the people in this competition took part in :) </p>\n<p>We used around 35 features including the raw, in a 2 LSTM layer model, each layer single direction trained on sequence length 256, infer on 512. <br>\nBatchsize for training 2048. Hidden layer size of 512. <br>\nExample below.  </p>\n<p>First layer,</p>\n<ul>\n<li>Used features below in <code>embcatq</code>, no continuous features. </li>\n<li>No label in first layer, like in the SAINT paper.</li>\n<li>Added the difference of some of the embedding to the final embedding. This gives the model info on how similar each historical question was to the question in the sample.   </li>\n</ul>\n<p>Second layer, </p>\n<ul>\n<li>Outputs of first layer and included continuous features. </li>\n<li>Added embedding for interaction of question and chosen answer. </li>\n<li>Continuous features generated using <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> great book - this is why I would like to be able to give multiple upvotes.</li>\n<li>One important feature was the answer ratio, what percentage of students picked the same answer as the chosen answer.   </li>\n</ul>\n<p>How to handle the histories in memory was a problem, but there was plenty of space for it on the GPU, so loaded that first to numpy then to a torch tensor on GPU (history features took around 6GB); then loaded the other objects to RAM. <br>\nAttention did not work for us - looked promising on validation though - should have persisted with it.  We just took the final hidden cell from the LSTM as output.  <br>\nThe below got ~ 0.811 public, by make some changed to the architecture and bagging four models, lifted to 0.813 public, 0.816 private. </p>\n<pre><code>class LearnNet(nn.Module):\n    def __init__(self, modcols, contcols, padvals, extracols, \n                 dropout = 0.2, hidden = args.hidden):\n        super(LearnNet, self).__init__()\n\n        self.dropout = nn.Dropout(dropout)\n\n        self.modcols = modcols + extracols\n        self.contcols = contcols\n\n        self.emb_content_id = nn.Embedding(13526, 32)\n        self.emb_content_id_prior = nn.Embedding(13526*3, 32)\n        self.emb_bundle_id = nn.Embedding(13526, 32)\n        self.emb_part = nn.Embedding(9, 4)\n        self.emb_tag= nn.Embedding(190, 8)\n        self.emb_lpart = nn.Embedding(9, 4)\n        self.emb_prior = nn.Embedding(3, 2)\n        self.emb_ltag= nn.Embedding(190, 16)\n        self.emb_lag_time = nn.Embedding(301, 16)\n        self.emb_elapsed_time = nn.Embedding(301, 16)\n        self.emb_cont_user_answer = nn.Embedding(13526 * 4, 5)\n\n        self.tag_idx = torch.tensor(['tag' in i for i in self.modcols])\n        self.cont_wts = nn.Parameter( torch.ones(len(self.contcols)) )\n        self.cont_wts.requires_grad = True\n        self.cont_idx = [self.modcols.index(c) for c in self.contcols]\n\n        self.embedding_dropout = SpatialDropout(dropout)\n\n        self.diffsize = self.emb_content_id.embedding_dim + self.emb_part.embedding_dim + \\\n                        self.emb_bundle_id.embedding_dim + self.emb_tag.embedding_dim * 7 \n        IN_UNITSQ = self.diffsize * 2 + \\\n                    self.emb_lpart.embedding_dim + self.emb_ltag.embedding_dim + \\\n                        self.emb_prior.embedding_dim + self.emb_content_id_prior.embedding_dim + \\\n                        len(self.cont_idxcts)\n        IN_UNITSQA = ( self.emb_lag_time.embedding_dim + self.emb_elapsed_time.embedding_dim + \\\n                self.emb_cont_user_answer.embedding_dim) + len(self.contcols)\n        LSTM_UNITS = hidden \n        self.diffsize = self.emb_content_id.embedding_dim + self.emb_part.embedding_dim + \\\n                        self.emb_bundle_id.embedding_dim + self.emb_tag.embedding_dim * 7 \n\n        self.seqnet1 = nn.LSTM(IN_UNITSQ, LSTM_UNITS, bidirectional=False, batch_first=True)\n        self.seqnet2 = nn.LSTM(IN_UNITSQA + LSTM_UNITS, LSTM_UNITS, bidirectional=False, batch_first=True)\n\n        self.linear1 = nn.Linear(LSTM_UNITS * 2 + len(self.contcols), LSTM_UNITS//2)\n        self.bn0 = nn.BatchNorm1d(num_features=len(self.contcols))\n        self.bn1 = nn.BatchNorm1d(num_features=LSTM_UNITS * 2 + len(self.contcols))\n        self.bn2 = nn.BatchNorm1d(num_features=LSTM_UNITS//2)\n\n        self.linear_out = nn.Linear(LSTM_UNITS//2, 1)\n\n\n    def forward(self, x, m = None):\n\n        ## Continuous\n        contmat  = x[:,:, self.cont_idx]\n        contmat = self.bn0(contmat.permute(0,2,1)) .permute(0,2,1)\n        contmat = contmat * self.cont_wts\n\n        content_id_prior = x[:,:,self.modcols.index('content_id')] * 3 + \\\n                            x[:,:, self.modcols.index('prior_question_had_explanation')]\n        embcatq = torch.cat([\n            self.emb_content_id(x[:,:, self.modcols.index('content_id')].long()),\n            self.emb_part(x[:,:, self.modcols.index('part')].long()), \n            self.emb_bundle_id(x[:,:, self.modcols.index('bundle_id')].long()),\n            self.emb_tag(x[:,:, self.tag_idx].long()).view(x.shape[0], x.shape[1], -1),\n            self.emb_prior(x[:,:, self.modcols.index('prior_question_had_explanation')].long() ),\n            self.emb_lpart(x[:,:, self.modcols.index('lecture_part')].long()), \n            self.emb_ltag(x[:,:, self.modcols.index('lecture_tag')].long()) , \n            self.emb_content_id_prior(  content_id_prior.long()),\n            ], 2)\n        embcatqdiff = embcatq[:,:,:self.diffsize] - embcatq[:,-1,:self.diffsize].unsqueeze(1)\n\n        # Categroical embeddings\n        embcatqa = torch.cat([\n            self.emb_cont_user_answer(x[:,:, self.modcols.index('content_user_answer')].long()),\n            self.emb_lag_time(x[:,:, self.modcols.index('lag_time_cat')].long()), \n            self.emb_elapsed_time(x[:,:,self.modcols.index('elapsed_time_cat')].long())\n            ] , 2)\n        #embcatqadiff = embcatqa - embcatqa[:,-1].unsqueeze(1)\n        embcatq = self.embedding_dropout(embcatq)\n        embcatqa = self.embedding_dropout(embcatqa)\n        embcatqdiff = self.embedding_dropout(embcatqdiff)\n\n        # Weighted sum of tags - hopefully good weights are learnt\n        xinpq = torch.cat([embcatq, embcatqdiff], 2)\n        hiddenq, _ = self.seqnet1(xinpq)\n        xinpqa = torch.cat([embcatqa, contmat, hiddenq], 2)\n        hiddenqa, _ = self.seqnet2(xinpqa)\n\n        # Take last hidden unit\n        hidden = torch.cat([hiddenqa[:,-1,:], hiddenq[:,-1,:], contmat[:, -1]], 1)\n        hidden = self.dropout( self.bn1( hidden) )\n        hidden  = F.relu(self.linear1(hidden))\n        hidden = self.dropout(self.bn2(hidden))\n        out = self.linear_out(hidden).flatten()\n\n        return out\n</code></pre>",
  "messages": [
    {
      "id": "1144148",
      "postDate": "01/08/2021 09:06:27",
      "content": "<p>Thanks to my teammate <a href=\"https://www.kaggle.com/jusco11\" target=\"_blank\">Akihiko</a> for competing with me on this, there was great learning from the community for both of us. And thanks to our hosts for a wonderful challenge. Congrats all who competed. </p>\n<p>Solution has heavily inspired by Bestfitting's <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262\" target=\"_blank\">TalkingData solution</a> which we noticed a lot of the people in this competition took part in :) </p>\n<p>We used around 35 features including the raw, in a 2 LSTM layer model, each layer single direction trained on sequence length 256, infer on 512. <br>\nBatchsize for training 2048. Hidden layer size of 512. <br>\nExample below.  </p>\n<p>First layer,</p>\n<ul>\n<li>Used features below in <code>embcatq</code>, no continuous features. </li>\n<li>No label in first layer, like in the SAINT paper.</li>\n<li>Added the difference of some of the embedding to the final embedding. This gives the model info on how similar each historical question was to the question in the sample.   </li>\n</ul>\n<p>Second layer, </p>\n<ul>\n<li>Outputs of first layer and included continuous features. </li>\n<li>Added embedding for interaction of question and chosen answer. </li>\n<li>Continuous features generated using <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> great book - this is why I would like to be able to give multiple upvotes.</li>\n<li>One important feature was the answer ratio, what percentage of students picked the same answer as the chosen answer.   </li>\n</ul>\n<p>How to handle the histories in memory was a problem, but there was plenty of space for it on the GPU, so loaded that first to numpy then to a torch tensor on GPU (history features took around 6GB); then loaded the other objects to RAM. <br>\nAttention did not work for us - looked promising on validation though - should have persisted with it.  We just took the final hidden cell from the LSTM as output.  <br>\nThe below got ~ 0.811 public, by make some changed to the architecture and bagging four models, lifted to 0.813 public, 0.816 private. </p>\n<pre><code>class LearnNet(nn.Module):\n    def __init__(self, modcols, contcols, padvals, extracols, \n                 dropout = 0.2, hidden = args.hidden):\n        super(LearnNet, self).__init__()\n\n        self.dropout = nn.Dropout(dropout)\n\n        self.modcols = modcols + extracols\n        self.contcols = contcols\n\n        self.emb_content_id = nn.Embedding(13526, 32)\n        self.emb_content_id_prior = nn.Embedding(13526*3, 32)\n        self.emb_bundle_id = nn.Embedding(13526, 32)\n        self.emb_part = nn.Embedding(9, 4)\n        self.emb_tag= nn.Embedding(190, 8)\n        self.emb_lpart = nn.Embedding(9, 4)\n        self.emb_prior = nn.Embedding(3, 2)\n        self.emb_ltag= nn.Embedding(190, 16)\n        self.emb_lag_time = nn.Embedding(301, 16)\n        self.emb_elapsed_time = nn.Embedding(301, 16)\n        self.emb_cont_user_answer = nn.Embedding(13526 * 4, 5)\n\n        self.tag_idx = torch.tensor(['tag' in i for i in self.modcols])\n        self.cont_wts = nn.Parameter( torch.ones(len(self.contcols)) )\n        self.cont_wts.requires_grad = True\n        self.cont_idx = [self.modcols.index(c) for c in self.contcols]\n\n        self.embedding_dropout = SpatialDropout(dropout)\n\n        self.diffsize = self.emb_content_id.embedding_dim + self.emb_part.embedding_dim + \\\n                        self.emb_bundle_id.embedding_dim + self.emb_tag.embedding_dim * 7 \n        IN_UNITSQ = self.diffsize * 2 + \\\n                    self.emb_lpart.embedding_dim + self.emb_ltag.embedding_dim + \\\n                        self.emb_prior.embedding_dim + self.emb_content_id_prior.embedding_dim + \\\n                        len(self.cont_idxcts)\n        IN_UNITSQA = ( self.emb_lag_time.embedding_dim + self.emb_elapsed_time.embedding_dim + \\\n                self.emb_cont_user_answer.embedding_dim) + len(self.contcols)\n        LSTM_UNITS = hidden \n        self.diffsize = self.emb_content_id.embedding_dim + self.emb_part.embedding_dim + \\\n                        self.emb_bundle_id.embedding_dim + self.emb_tag.embedding_dim * 7 \n\n        self.seqnet1 = nn.LSTM(IN_UNITSQ, LSTM_UNITS, bidirectional=False, batch_first=True)\n        self.seqnet2 = nn.LSTM(IN_UNITSQA + LSTM_UNITS, LSTM_UNITS, bidirectional=False, batch_first=True)\n\n        self.linear1 = nn.Linear(LSTM_UNITS * 2 + len(self.contcols), LSTM_UNITS//2)\n        self.bn0 = nn.BatchNorm1d(num_features=len(self.contcols))\n        self.bn1 = nn.BatchNorm1d(num_features=LSTM_UNITS * 2 + len(self.contcols))\n        self.bn2 = nn.BatchNorm1d(num_features=LSTM_UNITS//2)\n\n        self.linear_out = nn.Linear(LSTM_UNITS//2, 1)\n\n\n    def forward(self, x, m = None):\n\n        ## Continuous\n        contmat  = x[:,:, self.cont_idx]\n        contmat = self.bn0(contmat.permute(0,2,1)) .permute(0,2,1)\n        contmat = contmat * self.cont_wts\n\n        content_id_prior = x[:,:,self.modcols.index('content_id')] * 3 + \\\n                            x[:,:, self.modcols.index('prior_question_had_explanation')]\n        embcatq = torch.cat([\n            self.emb_content_id(x[:,:, self.modcols.index('content_id')].long()),\n            self.emb_part(x[:,:, self.modcols.index('part')].long()), \n            self.emb_bundle_id(x[:,:, self.modcols.index('bundle_id')].long()),\n            self.emb_tag(x[:,:, self.tag_idx].long()).view(x.shape[0], x.shape[1], -1),\n            self.emb_prior(x[:,:, self.modcols.index('prior_question_had_explanation')].long() ),\n            self.emb_lpart(x[:,:, self.modcols.index('lecture_part')].long()), \n            self.emb_ltag(x[:,:, self.modcols.index('lecture_tag')].long()) , \n            self.emb_content_id_prior(  content_id_prior.long()),\n            ], 2)\n        embcatqdiff = embcatq[:,:,:self.diffsize] - embcatq[:,-1,:self.diffsize].unsqueeze(1)\n\n        # Categroical embeddings\n        embcatqa = torch.cat([\n            self.emb_cont_user_answer(x[:,:, self.modcols.index('content_user_answer')].long()),\n            self.emb_lag_time(x[:,:, self.modcols.index('lag_time_cat')].long()), \n            self.emb_elapsed_time(x[:,:,self.modcols.index('elapsed_time_cat')].long())\n            ] , 2)\n        #embcatqadiff = embcatqa - embcatqa[:,-1].unsqueeze(1)\n        embcatq = self.embedding_dropout(embcatq)\n        embcatqa = self.embedding_dropout(embcatqa)\n        embcatqdiff = self.embedding_dropout(embcatqdiff)\n\n        # Weighted sum of tags - hopefully good weights are learnt\n        xinpq = torch.cat([embcatq, embcatqdiff], 2)\n        hiddenq, _ = self.seqnet1(xinpq)\n        xinpqa = torch.cat([embcatqa, contmat, hiddenq], 2)\n        hiddenqa, _ = self.seqnet2(xinpqa)\n\n        # Take last hidden unit\n        hidden = torch.cat([hiddenqa[:,-1,:], hiddenq[:,-1,:], contmat[:, -1]], 1)\n        hidden = self.dropout( self.bn1( hidden) )\n        hidden  = F.relu(self.linear1(hidden))\n        hidden = self.dropout(self.bn2(hidden))\n        out = self.linear_out(hidden).flatten()\n\n        return out\n</code></pre>",
      "rawMarkdown": "Thanks to my teammate [Akihiko](https://www.kaggle.com/jusco11) for competing with me on this, there was great learning from the community for both of us. And thanks to our hosts for a wonderful challenge. Congrats all who competed. \n\nSolution has heavily inspired by Bestfitting's [TalkingData solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262) which we noticed a lot of the people in this competition took part in :) \n\nWe used around 35 features including the raw, in a 2 LSTM layer model, each layer single direction trained on sequence length 256, infer on 512. \nBatchsize for training 2048. Hidden layer size of 512. \nExample below.  \n\nFirst layer,\n- Used features below in `embcatq`, no continuous features. \n- No label in first layer, like in the SAINT paper.\n- Added the difference of some of the embedding to the final embedding. This gives the model info on how similar each historical question was to the question in the sample.   \n\nSecond layer, \n- Outputs of first layer and included continuous features. \n- Added embedding for interaction of question and chosen answer. \n- Continuous features generated using @its7171 great book - this is why I would like to be able to give multiple upvotes.\n- One important feature was the answer ratio, what percentage of students picked the same answer as the chosen answer.   \n\nHow to handle the histories in memory was a problem, but there was plenty of space for it on the GPU, so loaded that first to numpy then to a torch tensor on GPU (history features took around 6GB); then loaded the other objects to RAM. \nAttention did not work for us - looked promising on validation though - should have persisted with it.  We just took the final hidden cell from the LSTM as output.  \nThe below got ~ 0.811 public, by make some changed to the architecture and bagging four models, lifted to 0.813 public, 0.816 private. \n\n\n```\nclass LearnNet(nn.Module):\n    def __init__(self, modcols, contcols, padvals, extracols, \n                 dropout = 0.2, hidden = args.hidden):\n        super(LearnNet, self).__init__()\n        \n        self.dropout = nn.Dropout(dropout)\n        \n        self.modcols = modcols + extracols\n        self.contcols = contcols\n        \n        self.emb_content_id = nn.Embedding(13526, 32)\n        self.emb_content_id_prior = nn.Embedding(13526*3, 32)\n        self.emb_bundle_id = nn.Embedding(13526, 32)\n        self.emb_part = nn.Embedding(9, 4)\n        self.emb_tag= nn.Embedding(190, 8)\n        self.emb_lpart = nn.Embedding(9, 4)\n        self.emb_prior = nn.Embedding(3, 2)\n        self.emb_ltag= nn.Embedding(190, 16)\n        self.emb_lag_time = nn.Embedding(301, 16)\n        self.emb_elapsed_time = nn.Embedding(301, 16)\n        self.emb_cont_user_answer = nn.Embedding(13526 * 4, 5)\n            \n        self.tag_idx = torch.tensor(['tag' in i for i in self.modcols])\n        self.cont_wts = nn.Parameter( torch.ones(len(self.contcols)) )\n        self.cont_wts.requires_grad = True\n        self.cont_idx = [self.modcols.index(c) for c in self.contcols]\n        \n        self.embedding_dropout = SpatialDropout(dropout)\n        \n        self.diffsize = self.emb_content_id.embedding_dim + self.emb_part.embedding_dim + \\\n                        self.emb_bundle_id.embedding_dim + self.emb_tag.embedding_dim * 7 \n        IN_UNITSQ = self.diffsize * 2 + \\\n                    self.emb_lpart.embedding_dim + self.emb_ltag.embedding_dim + \\\n                        self.emb_prior.embedding_dim + self.emb_content_id_prior.embedding_dim + \\\n                        len(self.cont_idxcts)\n        IN_UNITSQA = ( self.emb_lag_time.embedding_dim + self.emb_elapsed_time.embedding_dim + \\\n                self.emb_cont_user_answer.embedding_dim) + len(self.contcols)\n        LSTM_UNITS = hidden \n        self.diffsize = self.emb_content_id.embedding_dim + self.emb_part.embedding_dim + \\\n                        self.emb_bundle_id.embedding_dim + self.emb_tag.embedding_dim * 7 \n        \n        self.seqnet1 = nn.LSTM(IN_UNITSQ, LSTM_UNITS, bidirectional=False, batch_first=True)\n        self.seqnet2 = nn.LSTM(IN_UNITSQA + LSTM_UNITS, LSTM_UNITS, bidirectional=False, batch_first=True)\n            \n        self.linear1 = nn.Linear(LSTM_UNITS * 2 + len(self.contcols), LSTM_UNITS//2)\n        self.bn0 = nn.BatchNorm1d(num_features=len(self.contcols))\n        self.bn1 = nn.BatchNorm1d(num_features=LSTM_UNITS * 2 + len(self.contcols))\n        self.bn2 = nn.BatchNorm1d(num_features=LSTM_UNITS//2)\n        \n        self.linear_out = nn.Linear(LSTM_UNITS//2, 1)\n\n        \n    def forward(self, x, m = None):\n        \n        ## Continuous\n        contmat  = x[:,:, self.cont_idx]\n        contmat = self.bn0(contmat.permute(0,2,1)) .permute(0,2,1)\n        contmat = contmat * self.cont_wts\n        \n        content_id_prior = x[:,:,self.modcols.index('content_id')] * 3 + \\\n                            x[:,:, self.modcols.index('prior_question_had_explanation')]\n        embcatq = torch.cat([\n            self.emb_content_id(x[:,:, self.modcols.index('content_id')].long()),\n            self.emb_part(x[:,:, self.modcols.index('part')].long()), \n            self.emb_bundle_id(x[:,:, self.modcols.index('bundle_id')].long()),\n            self.emb_tag(x[:,:, self.tag_idx].long()).view(x.shape[0], x.shape[1], -1),\n            self.emb_prior(x[:,:, self.modcols.index('prior_question_had_explanation')].long() ),\n            self.emb_lpart(x[:,:, self.modcols.index('lecture_part')].long()), \n            self.emb_ltag(x[:,:, self.modcols.index('lecture_tag')].long()) , \n            self.emb_content_id_prior(  content_id_prior.long()),\n            ], 2)\n        embcatqdiff = embcatq[:,:,:self.diffsize] - embcatq[:,-1,:self.diffsize].unsqueeze(1)\n            \n        # Categroical embeddings\n        embcatqa = torch.cat([\n            self.emb_cont_user_answer(x[:,:, self.modcols.index('content_user_answer')].long()),\n            self.emb_lag_time(x[:,:, self.modcols.index('lag_time_cat')].long()), \n            self.emb_elapsed_time(x[:,:,self.modcols.index('elapsed_time_cat')].long())\n            ] , 2)\n        #embcatqadiff = embcatqa - embcatqa[:,-1].unsqueeze(1)\n        embcatq = self.embedding_dropout(embcatq)\n        embcatqa = self.embedding_dropout(embcatqa)\n        embcatqdiff = self.embedding_dropout(embcatqdiff)\n        \n        # Weighted sum of tags - hopefully good weights are learnt\n        xinpq = torch.cat([embcatq, embcatqdiff], 2)\n        hiddenq, _ = self.seqnet1(xinpq)\n        xinpqa = torch.cat([embcatqa, contmat, hiddenq], 2)\n        hiddenqa, _ = self.seqnet2(xinpqa)\n        \n        # Take last hidden unit\n        hidden = torch.cat([hiddenqa[:,-1,:], hiddenq[:,-1,:], contmat[:, -1]], 1)\n        hidden = self.dropout( self.bn1( hidden) )\n        hidden  = F.relu(self.linear1(hidden))\n        hidden = self.dropout(self.bn2(hidden))\n        out = self.linear_out(hidden).flatten()\n        \n        return out\n```",
      "votes": null
    },
    {
      "id": "1144154",
      "postDate": "01/08/2021 09:12:39",
      "content": "<p>Now this is a really classy solution! I loved this one! Thanks for sharing!</p>",
      "rawMarkdown": "Now this is a really classy solution! I loved this one! Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1144276",
      "postDate": "01/08/2021 11:05:58",
      "content": "<p>Great solution. Congrats on results and thanks for the writeup solution <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> and team</p>",
      "rawMarkdown": "Great solution. Congrats on results and thanks for the writeup solution @darraghdog and team",
      "votes": null
    },
    {
      "id": "1144957",
      "postDate": "01/08/2021 19:17:39",
      "content": "<p>Congrats and thank you for sharing. Smart solution!</p>",
      "rawMarkdown": "Congrats and thank you for sharing. Smart solution!",
      "votes": null
    },
    {
      "id": "1145478",
      "postDate": "01/09/2021 07:08:23",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!",
      "votes": null
    },
    {
      "id": "1145633",
      "postDate": "01/09/2021 09:13:03",
      "content": "<p>Hey, Can you share what was the performance if you didn't do the below? Ty! Super clean and easy to digest code!</p>\n<blockquote>\n  <p>embcatqdiff = embcatq[:, :, :self.diffsize] - embcatq[:, -1, :self.diffsize].unsqueeze(1)</p>\n</blockquote>",
      "rawMarkdown": "Hey, Can you share what was the performance if you didn't do the below? Ty! Super clean and easy to digest code!\n\n> embcatqdiff = embcatq[:, :, :self.diffsize] - embcatq[:, -1, :self.diffsize].unsqueeze(1)",
      "votes": null
    },
    {
      "id": "1146438",
      "postDate": "01/09/2021 19:18:13",
      "content": "<p>Congratulation and thanks for sharing</p>",
      "rawMarkdown": "Congratulation and thanks for sharing",
      "votes": null
    },
    {
      "id": "1146560",
      "postDate": "01/09/2021 21:19:46",
      "content": "<p>Very little, maybe ~0.001, but would have caused places in the leaderboard. Intuitively it made a lot of sense for me as I was confused how the model got information of the LSTM… when it does a forward pass on the LSTM, it did not know what the question asked at the end was (only uni-directional). </p>",
      "rawMarkdown": "Very little, maybe ~0.001, but would have caused places in the leaderboard. Intuitively it made a lot of sense for me as I was confused how the model got information of the LSTM... when it does a forward pass on the LSTM, it did not know what the question asked at the end was (only uni-directional).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1144154,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "01/08/2021 09:12:39",
      "content": "<p>Now this is a really classy solution! I loved this one! Thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1145633,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "01/09/2021 09:13:03",
          "content": "<p>Hey, Can you share what was the performance if you didn't do the below? Ty! Super clean and easy to digest code!</p>\n<blockquote>\n  <p>embcatqdiff = embcatq[:, :, :self.diffsize] - embcatq[:, -1, :self.diffsize].unsqueeze(1)</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146560,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "01/09/2021 21:19:46",
          "content": "<p>Very little, maybe ~0.001, but would have caused places in the leaderboard. Intuitively it made a lot of sense for me as I was confused how the model got information of the LSTM… when it does a forward pass on the LSTM, it did not know what the question asked at the end was (only uni-directional). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144276,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "01/08/2021 11:05:58",
      "content": "<p>Great solution. Congrats on results and thanks for the writeup solution <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> and team</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144957,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "01/08/2021 19:17:39",
      "content": "<p>Congrats and thank you for sharing. Smart solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1145478,
      "author_name": "zjjszj2",
      "author_url": "",
      "post_date": "01/09/2021 07:08:23",
      "content": "<p>thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1146438,
      "author_name": "kamalnaithani",
      "author_url": "",
      "post_date": "01/09/2021 19:18:13",
      "content": "<p>Congratulation and thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1144148": "Thanks to my teammate [Akihiko](https://www.kaggle.com/jusco11) for competing with me on this, there was great learning from the community for both of us. And thanks to our hosts for a wonderful challenge. Congrats all who competed. \n\nSolution has heavily inspired by Bestfitting's [TalkingData solution](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56262) which we noticed a lot of the people in this competition took part in :) \n\nWe used around 35 features including the raw, in a 2 LSTM layer model, each layer single direction trained on sequence length 256, infer on 512. \nBatchsize for training 2048. Hidden layer size of 512. \nExample below.  \n\nFirst layer,\n- Used features below in `embcatq`, no continuous features. \n- No label in first layer, like in the SAINT paper.\n- Added the difference of some of the embedding to the final embedding. This gives the model info on how similar each historical question was to the question in the sample.   \n\nSecond layer, \n- Outputs of first layer and included continuous features. \n- Added embedding for interaction of question and chosen answer. \n- Continuous features generated using @its7171 great book - this is why I would like to be able to give multiple upvotes.\n- One important feature was the answer ratio, what percentage of students picked the same answer as the chosen answer.   \n\nHow to handle the histories in memory was a problem, but there was plenty of space for it on the GPU, so loaded that first to numpy then to a torch tensor on GPU (history features took around 6GB); then loaded the other objects to RAM. \nAttention did not work for us - looked promising on validation though - should have persisted with it.  We just took the final hidden cell from the LSTM as output.  \nThe below got ~ 0.811 public, by make some changed to the architecture and bagging four models, lifted to 0.813 public, 0.816 private. \n\n\n```\nclass LearnNet(nn.Module):\n    def __init__(self, modcols, contcols, padvals, extracols, \n                 dropout = 0.2, hidden = args.hidden):\n        super(LearnNet, self).__init__()\n        \n        self.dropout = nn.Dropout(dropout)\n        \n        self.modcols = modcols + extracols\n        self.contcols = contcols\n        \n        self.emb_content_id = nn.Embedding(13526, 32)\n        self.emb_content_id_prior = nn.Embedding(13526*3, 32)\n        self.emb_bundle_id = nn.Embedding(13526, 32)\n        self.emb_part = nn.Embedding(9, 4)\n        self.emb_tag= nn.Embedding(190, 8)\n        self.emb_lpart = nn.Embedding(9, 4)\n        self.emb_prior = nn.Embedding(3, 2)\n        self.emb_ltag= nn.Embedding(190, 16)\n        self.emb_lag_time = nn.Embedding(301, 16)\n        self.emb_elapsed_time = nn.Embedding(301, 16)\n        self.emb_cont_user_answer = nn.Embedding(13526 * 4, 5)\n            \n        self.tag_idx = torch.tensor(['tag' in i for i in self.modcols])\n        self.cont_wts = nn.Parameter( torch.ones(len(self.contcols)) )\n        self.cont_wts.requires_grad = True\n        self.cont_idx = [self.modcols.index(c) for c in self.contcols]\n        \n        self.embedding_dropout = SpatialDropout(dropout)\n        \n        self.diffsize = self.emb_content_id.embedding_dim + self.emb_part.embedding_dim + \\\n                        self.emb_bundle_id.embedding_dim + self.emb_tag.embedding_dim * 7 \n        IN_UNITSQ = self.diffsize * 2 + \\\n                    self.emb_lpart.embedding_dim + self.emb_ltag.embedding_dim + \\\n                        self.emb_prior.embedding_dim + self.emb_content_id_prior.embedding_dim + \\\n                        len(self.cont_idxcts)\n        IN_UNITSQA = ( self.emb_lag_time.embedding_dim + self.emb_elapsed_time.embedding_dim + \\\n                self.emb_cont_user_answer.embedding_dim) + len(self.contcols)\n        LSTM_UNITS = hidden \n        self.diffsize = self.emb_content_id.embedding_dim + self.emb_part.embedding_dim + \\\n                        self.emb_bundle_id.embedding_dim + self.emb_tag.embedding_dim * 7 \n        \n        self.seqnet1 = nn.LSTM(IN_UNITSQ, LSTM_UNITS, bidirectional=False, batch_first=True)\n        self.seqnet2 = nn.LSTM(IN_UNITSQA + LSTM_UNITS, LSTM_UNITS, bidirectional=False, batch_first=True)\n            \n        self.linear1 = nn.Linear(LSTM_UNITS * 2 + len(self.contcols), LSTM_UNITS//2)\n        self.bn0 = nn.BatchNorm1d(num_features=len(self.contcols))\n        self.bn1 = nn.BatchNorm1d(num_features=LSTM_UNITS * 2 + len(self.contcols))\n        self.bn2 = nn.BatchNorm1d(num_features=LSTM_UNITS//2)\n        \n        self.linear_out = nn.Linear(LSTM_UNITS//2, 1)\n\n        \n    def forward(self, x, m = None):\n        \n        ## Continuous\n        contmat  = x[:,:, self.cont_idx]\n        contmat = self.bn0(contmat.permute(0,2,1)) .permute(0,2,1)\n        contmat = contmat * self.cont_wts\n        \n        content_id_prior = x[:,:,self.modcols.index('content_id')] * 3 + \\\n                            x[:,:, self.modcols.index('prior_question_had_explanation')]\n        embcatq = torch.cat([\n            self.emb_content_id(x[:,:, self.modcols.index('content_id')].long()),\n            self.emb_part(x[:,:, self.modcols.index('part')].long()), \n            self.emb_bundle_id(x[:,:, self.modcols.index('bundle_id')].long()),\n            self.emb_tag(x[:,:, self.tag_idx].long()).view(x.shape[0], x.shape[1], -1),\n            self.emb_prior(x[:,:, self.modcols.index('prior_question_had_explanation')].long() ),\n            self.emb_lpart(x[:,:, self.modcols.index('lecture_part')].long()), \n            self.emb_ltag(x[:,:, self.modcols.index('lecture_tag')].long()) , \n            self.emb_content_id_prior(  content_id_prior.long()),\n            ], 2)\n        embcatqdiff = embcatq[:,:,:self.diffsize] - embcatq[:,-1,:self.diffsize].unsqueeze(1)\n            \n        # Categroical embeddings\n        embcatqa = torch.cat([\n            self.emb_cont_user_answer(x[:,:, self.modcols.index('content_user_answer')].long()),\n            self.emb_lag_time(x[:,:, self.modcols.index('lag_time_cat')].long()), \n            self.emb_elapsed_time(x[:,:,self.modcols.index('elapsed_time_cat')].long())\n            ] , 2)\n        #embcatqadiff = embcatqa - embcatqa[:,-1].unsqueeze(1)\n        embcatq = self.embedding_dropout(embcatq)\n        embcatqa = self.embedding_dropout(embcatqa)\n        embcatqdiff = self.embedding_dropout(embcatqdiff)\n        \n        # Weighted sum of tags - hopefully good weights are learnt\n        xinpq = torch.cat([embcatq, embcatqdiff], 2)\n        hiddenq, _ = self.seqnet1(xinpq)\n        xinpqa = torch.cat([embcatqa, contmat, hiddenq], 2)\n        hiddenqa, _ = self.seqnet2(xinpqa)\n        \n        # Take last hidden unit\n        hidden = torch.cat([hiddenqa[:,-1,:], hiddenq[:,-1,:], contmat[:, -1]], 1)\n        hidden = self.dropout( self.bn1( hidden) )\n        hidden  = F.relu(self.linear1(hidden))\n        hidden = self.dropout(self.bn2(hidden))\n        out = self.linear_out(hidden).flatten()\n        \n        return out\n```",
    "1144154": "Now this is a really classy solution! I loved this one! Thanks for sharing!",
    "1144276": "Great solution. Congrats on results and thanks for the writeup solution @darraghdog and team",
    "1144957": "Congrats and thank you for sharing. Smart solution!",
    "1145478": "thanks for sharing!",
    "1145633": "Hey, Can you share what was the performance if you didn't do the below? Ty! Super clean and easy to digest code!\n\n> embcatqdiff = embcatq[:, :, :self.diffsize] - embcatq[:, -1, :self.diffsize].unsqueeze(1)",
    "1146438": "Congratulation and thanks for sharing",
    "1146560": "Very little, maybe ~0.001, but would have caused places in the leaderboard. Intuitively it made a lot of sense for me as I was confused how the model got information of the LSTM... when it does a forward pass on the LSTM, it did not know what the question asked at the end was (only uni-directional)."
  },
  "source": "meta"
}